Our first study is being run. There are no findings yet.
We would rather have this page say nothing than say something we have not measured. So instead of results, here is the question we are asking, how we are asking it, what we will publish alongside the answer, and what we will not claim even once the data is in.
The question we are asking
Narrow on purpose. A study that tries to describe “AI search” as a whole ends up describing nothing that can be checked.
- How much do answers differ between runs? Ask the same buyer question repeatedly, on the same assistant, in the same week, and count how much the set of brands named changes.
- How much do they differ between assistants? The same question on each assistant we measure, compared side by side rather than blended.
- What do they cite? Which kinds of source turn up in answers to commercial questions, and how often the same domains recur.
- How many answers does it take? How large a sample has to be before a share stops moving around — which is the number that decides what the product is allowed to show.
SparkToro asked AI tools the same recommendation questions many times and found the list of brands almost never repeated: fewer than 1 in 100 runs gave the same list.
Somebody else’s work, linked because it exists and ours does not yet.
How it is being run
The same way the product measures a client, which is the point: if the method is not good enough for a study, it is not good enough to bill an agency for.
Sampling
By default each question is asked three times on each assistant you switch on, every run, and you can ask more. A cell with fewer than three answers in its window shows a dash instead of a number.
Through the API
Each assistant is asked through its developer API with web search switched on, not through the consumer app. What a person sees in the app can differ.
Windows
Answers are grouped by week. Screens add up a rolling window (the last 28 days by default) so a figure rests on more answers; reports compare the start of the period with the end.
What is kept
Recorded on every answer: the raw text, the pages it cited, the model version that produced it, and what it cost to ask.
What gets published with it
A number on its own cannot be checked. These go out with it, or it does not go out.
The questions
The full list of prompts, verbatim, so anyone can ask them again rather than take our word for what was asked.
The counts
How many answers were collected per question per assistant, per week — not only the shares worked out from them.
The ranges
Every share with the interval around it and the number of answers behind it, on the same basis the product uses.
The model versions and dates
Which model answered and when. An assistant's behaviour changes between versions, so a figure without a version is not repeatable.
What went wrong
Questions that had to be dropped, runs that failed, anything that would change how the numbers should be read.
What we did not measure
Microsoft Copilot and Google AI Overviews / AI Mode offer no public API, so they are not measured, and no report estimates them.
What we will not claim, even with the data
Said now, while there are no results to be tempted by.
- A rank or position in any assistant
- A single 0–100 visibility score
- That a piece of work produced a change: we show what followed, with a comparison where one exists
- A forecast of the share a client will reach
- A revenue figure for visibility
- Anything about assistants we do not measure
Not a market survey
It measures how assistants answer a fixed set of questions. It says nothing about how many people ask them, or what any of it is worth in revenue.
A sample, not a census
A finite number of questions on a finite number of assistants, over a finite period. Every figure will carry the range that follows from that.
Nothing about your client
General findings do not transfer to one brand. A study is a reason to measure your own client, not a substitute for measuring them.
No claim about why
If something moves during the study we will show what moved and over which weeks. Attributing it to a reason is a different study than this one.
Until then, measure your own client
A general study would tell you how assistants behave. An audit tells you what they say about the client whose retainer is on the line, which is the only figure that settles an argument with that client.