What we do Pricing FAQs Contact Get started
All articles

What the Research Actually Says About Getting Cited by AI

TL;DR

In July 2026 a critical survey of 45 generative engine optimisation studies was published on arXiv. It grades the evidence and most of it does not survive. The famous 40% figure describes a bigger share of an answer for a document that was already being used, not 40% more readers. Repeat the same prompt and the sources change: daily source overlap between runs sits around 0.34 to 0.42. In one test, 57.8% of ChatGPT repetitions did not even run a web search. One benchmark found body-text optimisation actually reduced the odds of being retrieved at all. The survey grades "citations predict clicks, conversions or revenue" as very low confidence. Here is what that leaves worth doing.

There is a version of this industry that talks about AI search the way people talked about meta keywords in 2003. Confident, specific, unfalsifiable. Add these phrases, get cited. Publish this file, get picked. It sells well because the alternative is admitting that a lot of it is not known yet.

So it was a genuinely useful week when, on 15 July 2026, Olivier Martinez posted a critical survey of generative engine optimisation to arXiv. It reviews 45 studies published between November 2023 and July 2026, plus the retrieval and evaluation research underneath them, and it does the thing almost nobody in the commercial space does: it grades the evidence, and says out loud which claims do not hold up.

I have read it. This is what it says, what I think it means for a business paying somebody to do this work, and where I think it should be treated with care, because it is a preprint by a single author and it deserves the same scepticism it applies to everyone else.

The first useful idea: it is a pipeline, not a ranking

Almost every sales conversation about AI visibility treats it as one thing. Are you in the answer or not. The survey's central move is to break that into seven separate stages, each of which can fail on its own.

  1. Activation. Does the assistant run a web search at all, or answer from memory?
  2. Crawling and indexing. Has your page been fetched and stored?
  3. Retrieval. Is your page pulled as a candidate for this query?
  4. Reranking and context allocation. Does it survive the shortlist, and how much room does it get?
  5. Generation and citation. Does the model use it, and does it link it?
  6. Absorption and fidelity. Does what the answer says actually match what your page says?
  7. Attention and behaviour. Does anybody read it, click it, buy anything?

The survey's blunt observation is that most published GEO work optimises stages four and five, while stages one to three and stage seven are barely studied. That is the whole industry in one sentence. The techniques being sold operate on the part of the process that is easiest to run an experiment on, which happens to be the part that only matters once you are already in the room.

It matters because of a simple bit of arithmetic the paper spells out. The probability of being cited is the probability that a search runs at all, times the probability you are retrieved given that it ran, times the probability you are cited given that you were retrieved. Three numbers multiplied together. Sell somebody an improvement to the third and leave the first two alone and you can move your metric while nothing happens to your business.

The 40% figure, corrected

If you have heard one number about GEO, it is this one. The 2023 paper by Aggarwal and colleagues at Princeton, Georgia Tech, IIT Delhi and the Allen Institute for AI, published at KDD 2024, is where "GEO can lift visibility by up to 40%" comes from. It is quoted in agency decks constantly, including by people who have plainly never opened it.

Here is what the survey says that number actually measures. Adding quotations raised a metric called position-adjusted word count from 19.3% to 27.2%. That metric is a source's position-weighted share of the generated answer. And the whole experiment ran inside a fixed context of five documents that had already been supplied to the generator.

19.3% to 27.2%
what the famous "up to 40%" GEO uplift actually measures: a larger share of the answer, for a document already placed in the model's context (Martinez, 2026, on Aggarwal et al., KDD 2024)

The survey's own words are worth repeating: this "does not mean that 40% more readers will click". It means a source that was already being used got a bigger slice of the paragraph. It says nothing about whether you would have been retrieved in the first place, and nothing about traffic.

The survey grades the general claim that "GEO increases visibility by 40%" as rejected. Not weak. Rejected.

The measurement problem, which is worse than people admit

This is the section I would put in front of anyone about to buy an AI visibility subscription.

Ask a generative engine the same question twice and you get different sources. Not slightly different. The survey cites work by Schulte and colleagues covering four engines over 45 days, finding daily source-level overlap between runs of roughly 0.34 to 0.42 on a scale where 1.0 is identical. On temperature-controlled surfaces, somewhere between 9% and 28% of decisions change across repeated runs.

Then there is the activation problem, which is starker. In one set of tests, 57.8% of ChatGPT repetitions did not activate web search at all. Over half the time, the model simply answered from what it already knew. No amount of on-page work reaches an answer that was never retrieved.

Across engines, the disagreement is worse again. Li and Sinnamon found only 26% domain overlap between Bing Chat and Perplexity. Across Google, AI Overviews and Gemini, URL-level overlap runs at 0.11 to 0.18. And 53% of the domains cited by AI Overviews do not appear in the organic top ten at all, which is the same finding, from a different direction, as the overlap work I wrote about in the state of AI search in the UK.

Put those together and you get the honest position on AI visibility scores. If the underlying sources shift by that much between two runs on the same day, a monthly score that moved by a few points has told you almost nothing about your website. It has told you about the sampling.

Optimising can make things worse

This is the finding I did not expect and the one I think is most useful commercially.

A benchmark called SAGEO Arena tested optimisation end to end, meaning it did not start with the document already in the context window. It started at retrieval, like real life. Rewriting body text in the recommended GEO style reduced average top-20 presence by about 9%, reduced top-10 presence after reranking by 16%, and reduced final citation by 6%.

Read that again. The rewrite made the page more quotable and less findable, and the second effect was larger than the first. You optimised stage five and damaged stage three.

That is a coherent mechanism, not a fluke. Retrieval systems match on relevance. Stuffing a page with quotations, statistics and authoritative-sounding framing changes what the page is about in the eyes of the retriever. If you dilute the thing the page is actually the best answer to, you can talk your way out of the shortlist.

And if everyone does it, it cancels out

The survey cites a benchmark called C-SEO Bench which tests what happens under competition rather than in isolation. Of 54 method and domain combinations, only three were significantly positive, and none of those were in question answering. Gains decline as adoption rises, which the survey describes as tending toward a zero-sum game.

This is the least surprising finding in the paper and the one the industry is least keen to discuss. Every technique that works by making your page look more citable than the next page stops working when the next page does it too. That is not a criticism of the technique. It is what competition is. But it does mean a permanent advantage is not on offer, and anyone selling you one is describing something the evidence does not support.

The confidence table

The survey ends by grading its main empirical claims. I have reproduced the grading below, with a plain-English column added by me, because this table is the single most useful thing in the paper for anybody buying this work.

ConfidenceClaimWhat that means for you
HighA document already placed in context can causally alter its rank and citationOnce you are in the shortlist, how you write matters
HighRelevance and position are major determinantsBeing genuinely the best answer, clearly stated near the top, is the reproducible lever
ModerateExtractable evidence often makes a source easier to useClear facts, figures and direct answers help, in a qualified way
LowWhite-hat interventions durably improve organic discoverabilityThe claim that GEO work gets you found in the first place is weakly supported
Very lowCitations predict clicks, conversions or revenueBeing cited is not evidence you got customers
RejectedGEO increases visibility by 40%The number in half the sales decks does not mean what they say it means

The survey's overall conclusion is that already-retrieved content can causally influence answers, and that it identified no technique with a stable, longitudinal, cross-platform causal effect on organic discoverability or on downstream clicks and conversions.

That is a careful sentence and it is worth parsing slowly. It does not say GEO does nothing. It says that across 45 studies, nobody has yet shown a technique that reliably gets you found, across platforms, over time, in a way that produces measurable business outcomes. Absence of proof is not proof of absence. But if you are the one paying, the burden sits with the person selling.

Where I would push back on the survey

I said I would apply the same scepticism, so here it is.

It is a preprint, by a single author, not yet peer reviewed at the time I am writing. Survey papers involve judgement calls about what to include and how to weight it, and a critical survey by definition sets out to critique. Some of the underlying studies it grades harshly are grading themselves harshly in the same breath, which is normal in a young field rather than scandalous.

There is also a real gap between an academic standard of evidence and a commercial one. The survey wants randomised field trials with logs and controls, which is the right bar for a scientific claim. Almost nothing in marketing clears that bar, including things that plainly work. A practitioner who says "we did this for eleven sites and eight of them started appearing" is not producing evidence at level A, but they are not producing nothing either.

And the field moves faster than the literature. Studies from late 2024 describe engines that have since been rebuilt. Some findings will not have survived contact with the current versions.

What none of that touches is the structural argument, which is the part I find most convincing: the pipeline has seven stages, the industry sells work on two of them, and the measurement is noisy enough that a lot of reported success is sampling.

Want the parts that are actually supported?

Stagg Studios does the technical foundations, structure and consistency work, checked every month. £100 a month, first month free, no lock-in.

Get startedSee pricing

What this leaves worth doing

Here is the thing. Reading all of that, you could conclude the sensible move is to do nothing. I do not think that follows, and the survey does not say it either.

What it says is that the reproducible levers sit earlier in the pipeline than the industry sells, and that the durable ones are unglamorous. So:

  1. Fix stages one to three before anything else. Can the crawlers reach you, is the page fetched, is it stored, is it retrievable for the queries you care about. This is the boring technical layer and it is the part with the highest confidence grade attached to it. It is also the part that is genuinely broken on a surprising number of small business sites.
  2. Be relevant rather than quotable. Relevance and position are graded high confidence. Quotability is graded moderate, and the SAGEO result suggests over-doing it costs you retrieval. Answer the question the page is for, clearly, near the top, and stop there.
  3. Make the facts about you agree everywhere. Not because a study proves it lifts citations, but because the fidelity research shows assistants routinely garble details, and consistent source data is the only lever you hold over that. Only 51.5% of sentences were fully supported by their cited sources in one four-engine study.
  4. Measure with repetition or do not bother. One check on one day is noise. If you or your provider are tracking AI visibility, it needs repeated runs, paraphrased prompts, more than one engine, and more than one date. Anything less is a number that will move on its own.
  5. Do not buy permanence. Competitive erosion is real. Whatever works this year works less well when everybody does it. Budget for maintenance, not for a project that finishes.

That is a less exciting list than the one in most proposals. It is also close to what good technical SEO has always been, which is the conclusion a fair reading supports and which I have argued before in the piece on SEO, AEO and GEO.

Questions people ask

Short sourced answers to the questions this research raises most often. The working is in the sections above.

Does generative engine optimisation actually work?

Partly, and less than it is sold as. The July 2026 critical survey of 45 studies grades as high confidence that a document already placed in a model's context can causally change its rank and citation within that answer, and that relevance and position are major determinants. It grades as low confidence that these interventions durably improve whether you get found in the first place, and as very low confidence that citations predict clicks, conversions or revenue. So the mechanics inside the answer are reasonably well evidenced. The business outcome is not.

Where does the "GEO lifts visibility by 40%" figure come from?

From Aggarwal and colleagues, published at KDD 2024. The survey re-examines it and finds the uplift describes a rise in position-adjusted word count from 19.3% to 27.2%, measured inside a fixed five-document context that had already been supplied to the generator. It is a larger share of the answer for a source already in use, not 40% more readers or 40% more traffic. The survey grades the general "40% more visibility" claim as rejected.

Why do AI visibility tools give different results?

Because the underlying answers genuinely change between runs, and because tools sample differently. The survey cites daily source-level overlap of roughly 0.34 to 0.42 across four engines over 45 days, 9% to 28% of decisions changing on repeated runs, and only 26% domain overlap between two major engines in an earlier study. A tool that runs 50 prompts monthly and one that runs thousands daily are measuring different things. Treat a score that moved by a few points as noise unless the method behind it is repeated, paraphrased and multi-engine.

Can optimising for AI answers make things worse?

Yes, and this is the least known finding. An end-to-end benchmark called SAGEO Arena found that body-only optimisation reduced average top-20 presence by about 9%, top-10 presence after reranking by 16%, and final citation by 6%. The mechanism is straightforward: rewriting a page to sound more citable can dilute what the page is clearly about, which is what retrieval matches on. You can make a page more quotable and less findable at the same time.

If everyone does GEO, does it stop working?

The evidence points that way. A benchmark called C-SEO Bench tested methods under competition rather than in isolation and found only three of 54 method and domain combinations significantly positive, with none positive in question answering. The survey describes gains declining as adoption increases, tending toward a zero-sum outcome. That is a reason to expect maintenance rather than a permanent gain, and a reason to be sceptical of anyone offering a lasting advantage.

So should a small business bother with any of this?

Yes, but at the right end of the pipeline. The best-evidenced levers are the early ones: being crawlable, being indexed, being retrievable, and being genuinely relevant to the question. Those are the same foundations that ordinary technical SEO has always cared about, they are frequently broken on small business sites, and fixing them gives you a chance of ranking better and of being readable to AI systems. What the evidence does not support is paying for a package of writing tricks aimed at the last two stages, sold with a percentage attached.

Is this survey the final word?

No, and it does not claim to be. It is a preprint by a single author, it was not peer reviewed at the time of writing, and survey papers involve judgement about what to include. The field also moves faster than the literature, so some studies describe engines that have since been rebuilt. What holds up regardless is the structural point: the process has seven stages, most commercial work targets two of them, and the measurement noise is large enough that a good deal of reported success is sampling.

Sources

Every figure above comes from the linked research. None of it describes a Stagg Studios client or a study I ran. Where a number is reported by the survey rather than measured by it, I have said so, and you can follow the survey's own citations to the underlying paper.

What I would do about it, and what it costs

Here is the part worth sitting with. The most expensive mistake in this field is not doing nothing. It is paying a monthly fee for work on stages four and five of a seven-stage process while stages one to three are quietly broken, and then judging the result by a score that moves on its own. That can go on for a year before anybody notices, because the report always has a number on it and the number always went somewhere.

The evidence points at the unfashionable end. Be reachable. Be indexed. Be retrievable. Be genuinely the best answer to the question your page is for, and say so near the top. Keep the facts about your business consistent everywhere they appear, because the fidelity research says assistants garble details and consistency is the only lever you hold. Then re-check it, because the engines change and so does your site.

That is what I do, as one job rather than a pile of separate products: site structure, crawlability, structured data, metadata, making sure AI crawlers can actually reach your content, and keeping it correct every month. It gives you a chance of ranking better and of being readable to the systems that answer questions about your sector. I cannot promise you a ranking or a citation, and after reading 45 studies' worth of grading, I would be more suspicious than ever of anyone who does.

£100 a month, first month free, no contract, and you deal with me rather than an account manager. Start here, or send me a note and I will tell you which stage of that pipeline your site is currently failing at.