Research

Our first study is being run. There are no findings yet.

We would rather have this page say nothing than say something we have not measured. So instead of results, here is the question we are asking, how we are asking it, what we will publish alongside the answer, and what we will not claim even once the data is in.

The question we are asking

Narrow on purpose. A study that tries to describe “AI search” as a whole ends up describing nothing that can be checked.

  • How much do answers differ between runs? Ask the same buyer question repeatedly, on the same assistant, in the same week, and count how much the set of brands named changes.
  • How much do they differ between assistants? The same question on each assistant we measure, compared side by side rather than blended.
  • What do they cite? Which kinds of source turn up in answers to commercial questions, and how often the same domains recur.
  • How many answers does it take? How large a sample has to be before a share stops moving around — which is the number that decides what the product is allowed to show.
SparkToro asked AI tools the same recommendation questions many times and found the list of brands almost never repeated: fewer than 1 in 100 runs gave the same list.
SparkToro: AIs are highly inconsistent when recommending brands
Somebody else’s work, linked because it exists and ours does not yet.

How it is being run

The same way the product measures a client, which is the point: if the method is not good enough for a study, it is not good enough to bill an agency for.

Sampling

By default each question is asked three times on each assistant you switch on, every run, and you can ask more. A cell with fewer than three answers in its window shows a dash instead of a number.

Through the API

Each assistant is asked through its developer API with web search switched on, not through the consumer app. What a person sees in the app can differ.

Windows

Answers are grouped by week. Screens add up a rolling window (the last 28 days by default) so a figure rests on more answers; reports compare the start of the period with the end.

What is kept

Recorded on every answer: the raw text, the pages it cited, the model version that produced it, and what it cost to ask.

The questions and the assistant list are fixed before the runs start, and published with the results whichever way they come out. A study whose scope is decided after looking at the data is not a study.

What gets published with it

A number on its own cannot be checked. These go out with it, or it does not go out.

  • The questions

    The full list of prompts, verbatim, so anyone can ask them again rather than take our word for what was asked.

  • The counts

    How many answers were collected per question per assistant, per week — not only the shares worked out from them.

  • The ranges

    Every share with the interval around it and the number of answers behind it, on the same basis the product uses.

  • The model versions and dates

    Which model answered and when. An assistant's behaviour changes between versions, so a figure without a version is not repeatable.

  • What went wrong

    Questions that had to be dropped, runs that failed, anything that would change how the numbers should be read.

  • What we did not measure

    Microsoft Copilot and Google AI Overviews / AI Mode offer no public API, so they are not measured, and no report estimates them.

What we will not claim, even with the data

Said now, while there are no results to be tempted by.

  • A rank or position in any assistant
  • A single 0–100 visibility score
  • That a piece of work produced a change: we show what followed, with a comparison where one exists
  • A forecast of the share a client will reach
  • A revenue figure for visibility
  • Anything about assistants we do not measure

Not a market survey

It measures how assistants answer a fixed set of questions. It says nothing about how many people ask them, or what any of it is worth in revenue.

A sample, not a census

A finite number of questions on a finite number of assistants, over a finite period. Every figure will carry the range that follows from that.

Nothing about your client

General findings do not transfer to one brand. A study is a reason to measure your own client, not a substitute for measuring them.

No claim about why

If something moves during the study we will show what moved and over which weeks. Attributing it to a reason is a different study than this one.

Until then, measure your own client

A general study would tell you how assistants behave. An audit tells you what they say about the client whose retainer is on the line, which is the only figure that settles an argument with that client.