Keyword research when most prompts match no keyword at all
The number everyone quotes, and the one nobody puts beside it
Semrush published its ChatGPT clickstream analysis on 7 April 2026, built on more than a billion lines of US clickstream data collected between October 2024 and February 2026. One line from it travelled everywhere: between 65% and 85% of prompts could not be matched to any traditional search keyword. If a keyword tool is your only map of demand, most of what people ask sits off the map.
The rest of that study is less convenient for the people quoting the headline. The share of prompts written in traditional search language nearly doubled in five months, from 18.9% in October 2025 to 34.9% in February 2026. Average prompt length on search-enabled queries rose from 4.7 words in early 2025 to 8.7 words in early 2026, while prompts that triggered no search fell from 24.9 words to 13.5. Read those three lines together and behaviour is converging on search, not moving away from it. People started out talking to ChatGPT in sentences. Once it began searching the web on their behalf, they started typing at it the way they type at Google.
So the standard advice, rewrite the site around long conversational questions, is aimed at a habit that is shrinking. It was never wrong exactly. It is pointed at the part of the curve heading down, and a target map built on it will date badly.
What actually broke
Google documents the mechanism in its own developer guidance on AI features: "Both AI Overviews and AI Mode may use a 'query fan-out' technique, issuing multiple related searches across subtopics and data sources." One prompt goes in. Some unknown number of searches you never see run behind it. The links that come back were selected against those searches, not against the sentence the user typed, which is the behaviour the AI Mode guide covers from Google's side.
The gap that opens is measurable and it is wide. Ahrefs ran 15,000 long-tail queries through Google, Bing, ChatGPT, Gemini, Copilot and Perplexity, then compared every citation against the ranking results for the identical prompt. Only 12% of the links cited by ChatGPT, Gemini and Copilot appeared in Google's top 10 for that same prompt. Perplexity was the outlier at 28.6%. Copilot managed 8.6%, Gemini 8.2%, ChatGPT 8.0% for in-text citations and 6.1% for the reference list underneath.
Google's own AI Overviews used to be the exception that kept classic keyword targeting honest. In July 2025 Ahrefs measured 76% of AI Overview citations coming from pages that ranked in the top 10 for the query. Re-run across 863,000 keyword SERPs and 4 million AI Overview URLs, that figure is now 38%. Roughly a third comes from positions 11 to 100 and roughly a third from pages that do not rank in the top 100 at all. Ahrefs puts the change down to fan-out directly: Google is selecting far fewer pages straight from the original SERP, because it is no longer answering the original query.
Ranking did not become less valuable. Ranking first for a term stopped being the same event as being the source of the answer to it. Those were one thing for twenty years. They are two things now, and most keyword strategies are still written as though they are one.
What keyword research turns into
If the engine decomposes the question before it retrieves anything, the unit of research is the decomposition, not the string. A cluster has to cover the sub-questions a buyer's main question breaks into, including the ones with no measurable search volume, because volume measures what people type into Google and fan-out queries are not typed by anyone.
The arithmetic of winnability changes too. Semrush's 2026 AI Visibility Index, 126 million US prompts between January and April 2026, found ChatGPT cites an average of 15 sources per response and Gemini an average of 3. Those are different games. Fifteen slots rewards breadth of coverage across a topic. Three is a podium, and on a podium you are competing with whoever the model already trusts, which is the uncomfortable finding in the piece on whether backlinks still matter.
The practical version: stop scoring a keyword by whether you can rank for it and start scoring a topic by whether you can be the most complete source on it. Same research, different exit criteria. It is the reason our Keyword Strategy clusters and scores rather than lists. A list of strings has nothing to say about coverage. Coverage answers what to write about, not how much of it to publish, which is where our guide to how much content you actually need picks up.
The modifiers you can actually observe
Fan-out is invisible. Prompt phrasing is not. Search Engine Land reported a January 2026 survey of 524 active LLM users by Stella Rising, margin of error around 4.3 points at 95% confidence, and the shape of it is useful. About 60% of people phrase the thing as a question and only 9% give a direct command. 24.5% of prompts contain the word "best". 28% mention price or a budget. 16% are explicitly location-based.
None of that is exotic, and that is the point. It is the same commercial-intent modifier set that has driven comparison and pricing pages for a decade, except now it is observed in real prompts rather than modelled. A pricing page that states real numbers, a comparison page that names competitors, a location page that actually exists: that covers a meaningful share of how people ask, and it is unglamorous enough that plenty of sites still skip it.
What you cannot see, said out loud
Google shipped generative AI performance reports in Search Console in June 2026, and they report impressions only, grouped by pages, countries, dates and devices. There is no query dimension and no click data. You can learn that a generative feature showed a link to your page. You cannot learn what was asked, and you cannot calculate a click-through rate from it. That is the design, not a rollout gap, and it sits inside the wider problem covered in measuring AI visibility.
Bing goes further with Grounding Queries in Bing Webmaster Tools, though Microsoft labels that a sample of citation activity rather than a complete log, as the Copilot guide sets out. It is currently the closest thing to real prompt data any engine hands a site owner for free. It is still not the fan-out list.
Nobody publishes the fan-out sub-queries. Not Google, not OpenAI. Any tool selling a "fan-out keyword" export is modelling what the engine probably asked, using a language model, and presenting it as data. Some of those models are good. None of them are observation. Price that output as a hypothesis worth testing, not a target list worth committing a quarter to. If we ever start selling one, hold us to this paragraph.
What we would do with a real site
Start from the questions, not the keyword file. Write out the ten questions a buyer asks between not knowing you exist and paying you, as full sentences, and only then check which have search volume. The ones that do become pages targeted the old way. The ones that do not become sections inside those pages, because sections are where fan-out finds anything.
Then check the overlap you can measure without a tool. For each priority topic, run the prompt in ChatGPT and the same prompt in Google, and list who gets cited against who ranks. When the two lists barely intersect, and on the Ahrefs numbers they usually will not, the citation list is your competitive set and the SERP is not. Most keyword strategies are still built against the SERP alone.
Then decide, in writing and once, which engine you are optimizing for. Fifteen ChatGPT slots and three Gemini slots do not reward the same page. Breadth of coverage wins the first, authority and consensus win the second, and chasing both with a single plan is how a content budget gets spent twice for one result.
Sources
Prompt matching, prompt length and search-language share come from Semrush, ChatGPT traffic analysis: insights from 17 months of clickstream data, 7 April 2026. Citations per response come from the Semrush 2026 AI Visibility Index, 26 June 2026. The fan-out description is Google's own, in AI features and your website. Citation overlap with rankings comes from Ahrefs, Only 12% of AI cited URLs rank in Google's top 10 for the original prompt and Update: 38% of AI Overview citations pull from the top 10, both by Louise Linehan and Xibeijia Guan, updated 31 May 2026. Prompt phrasing comes from Search Engine Land's report on how real people actually prompt AI. Search Console's limits are documented by Google at Generative AI performance report (Search).
Common questions
Is keyword research still worth doing for AI search?
Yes, and Semrush's own clickstream data argues for it: the share of AI prompts written in traditional search language nearly doubled between October 2025 and February 2026, from 18.9% to 34.9%. What changes is the exit criteria. You are no longer scoring a keyword by whether you can rank for it, you are scoring a topic by whether you can be the most complete source on it.
Can I see which prompts caused an AI engine to cite my page?
Almost never. Google's generative AI performance report in Search Console gives impressions grouped by pages, countries, dates and devices, with no query dimension and no clicks. Bing Webmaster Tools shows Grounding Queries, but Microsoft labels that a sample of citation activity rather than a complete log. No engine publishes the fan-out sub-queries it actually ran.
Should I write long conversational content because AI prompts are conversational?
That advice is aimed at a habit that is shrinking. Semrush measured average prompt length on search-enabled ChatGPT queries rising from 4.7 words to 8.7 words between early 2025 and early 2026, while prompts that triggered no web search fell from 24.9 words to 13.5. Prompts are converging on search-style phrasing, not drifting away from it.
Get the target map, not the keyword dump
Keyword Strategy clusters your market's questions by intent, scores each cluster against your site's real authority, and assigns every target to a page with a job. Built for coverage, because coverage is what fan-out rewards.