Because the AI search space is full of black-box numbers, we open up our entire methodology here. Below are the measurement principles, definitions, and quality controls behind our reporting — everything except client data and sensitive security details.
Prompt construction — Prompts come from customer language — sales calls, support logs, search console — never keyword tools. The client owns the set, organised by intent cluster and frozen within a measurement period.
Sampling — Repeated runs per prompt per window, cold sessions only, daily collection for active programmes, and model releases annotated on every timeline. Sample size is disclosed with every figure.
Engines — ChatGPT, Gemini, Claude and Perplexity, plus Google AI Overviews and AI Mode reported separately — collapsing them hides which surface drives your presence.
Metrics — Visibility, share of voice, position, sentiment and brand safety, citation share and AI shopping visibility — each with its own protocol and disclosed sample.
Classification — Judgement metrics get human review. Prose position, entity disambiguation and the competitor set are defined with the client and applied consistently.
Disclosure — Every figure carries its prompt set, engine set, sample size and date window. Estimates are labelled as estimates. Raw runs are available to clients.
Key Takeaways
- AI answers are non-deterministic. The same prompt returns different results on different runs, so a single observation is an anecdote rather than a measurement.
- Every figure we publish or report carries its prompt count, run count, engine set and date window.
- We measure nine engines as standard, and report Google AI Overviews and AI Mode separately because they behave differently.
- We classify our metrics against the IAB's decision-grade and directional standards, and publish the protocol so the classification can be checked.
- Prompt demand figures are estimates. No AI platform publishes query volumes.
- Raw run-level data is available to clients.
Why measurement here is genuinely hard
Any measurement methodology that doesn't acknowledge these realities is overselling its accuracy.
- Non-determinism. The exact same prompt can return different answers minutes apart. This is a fundamental characteristic of generative systems rather than a bug or sampling artifact. A single observation can easily over- or understate your actual presence.
- Apparent differences are often noise. Research applying bootstrap confidence intervals to citation visibility found that many apparent differences between domains fall within the noise floor of the measurement process, and that citation rankings are unstable across repeated samples. This is why we report sample sizes and why we're cautious about small movements.
- Prompt framing changes outcomes. "Best sunscreen in India" and "which sunscreen should I buy in India" can return different brand sets. The prompt set is a methodological choice, not a neutral input.
- Engines change underneath you. Model updates shift citation patterns independently of anything a brand does. A visibility drop in the week of a major model release is not necessarily a performance change.
- Vendors disagree because methods differ. In its August 2026 framework Measuring Visibility in the AI Era, the IAB noted that more than 20 companies now sell AI visibility measurement tools, each using different methodologies that can produce different answers for the same brand or publisher. If two tools give you different numbers, at least one is measuring something else. The method matters more than the tool.
Prompt construction
Good AI visibility measurement starts with authentic inputs. Instead of relying on traditional keyword tools, we build prompts directly from customer language, sales calls, and support logs. Our prompt construction follows four strict operating rules:
- Prompts come from customers, not keyword tools. We build the initial set from client sales call recordings, support tickets, search console data and category research, then refine with the client. Keyword tools report aggregate search volume dominated by US and global queries; they are a poor proxy for what an Indian buyer types into ChatGPT.
- The client owns the set. We suggest high-impact prompts using our own models, but the client adds, edits and disables. A prompt set the client didn't choose produces numbers the client won't trust.
- Prompts are organised by intent cluster, not by keyword. Typical clusters include:
- Category discovery: "best X in India"
- Comparison: "X vs Y", "alternatives to X"
- Trust and diligence: "is X reliable", "X reviews", "problems with X"
- Buying intent: "where to buy X", "X price India"
- Local: the above, scoped to a city
- The set is frozen within a measurement period. A visibility figure measured against a prompt set that changed between periods is not a measurement, it's a chart. When we add prompts we report the previous period on both the old and new sets so the comparison stays honest.
Sampling protocol
Single-run metrics will lie to you. Here is how we enforce statistical rigor across every measurement window to ensure you are seeing real performance rather than random noise:
- Repeated sampling, never single-run. Given documented non-determinism, a single query result is an anecdote. We run each prompt multiple times per measurement window and report the distribution rather than one observation.
- Sample size is disclosed with every figure. Any visibility number we publish or report to a client carries its prompt count, run count, engine set and date window.
- Fresh sessions. Prompts are run without personalisation or conversational history, so results reflect a cold, first-time query rather than a primed one.
- Frequency. Daily collection for active client programmes, giving roughly thirty observations per tracked variant per month. Monthly reporting. We advise against weekly reporting unless a programme is actively running and the client can act on it, because at that cadence platform drift is difficult to separate from performance change.
- Model releases are annotated on every timeline. When a major model version ships, it's marked on the chart. Without this, an engine change two months ago gets misread as a content failure today.
Engine coverage
We measure across the major AI assistants, currently including ChatGPT, Gemini, Claude, Perplexity, Google AI Overviews and Google AI Mode.
We report AI Overviews and AI Mode separately. They are both Google surfaces but behave differently, and collapsing them into a single figure hides which one is actually driving your presence.
Community source monitoring
Community platforms are among the sources AI answers cite, particularly for Indian consumer categories. Reddit and YouTube appear in citations regularly, which means the discussion happening there is not only sentiment about your brand. It is material the models read.
We monitor these as sources rather than as social mentions, using each platform's official API.
What it is used for. It surfaces the language customers actually use, which feeds prompt construction. And it identifies which community threads are already being cited about your category, which produces a work queue rather than a sentiment report.
How this differs from social listening. Conventional social listening reports what people said about you. This identifies which of those conversations are appearing in AI answers, or are likely to. The overlap is partial and the purpose is different.
Limits. API access defines what we can see, and it is not the whole conversation. Private groups, direct messages and closed communities sit outside it. Platform access terms also change, sometimes at short notice, and our coverage changes with them.
What we measure
To understand your true standing in AI search, you need more than a single brand mention count. Our reporting framework tracks performance across multiple dimensions, from raw citation share to sentiment and buying-intent positioning.
Visibility
The proportion of runs in which the brand is mentioned, across the prompt set. Reported with sample size.
Share of voice
Brand mentions as a proportion of all brand mentions across the prompt set. This is the number most teams report upward, and the one most sensitive to prompt set composition, which is why the set has to be frozen and disclosed.
Position
Where the brand appears within a recommendation list. Being named third in a shortlist is materially different from being mentioned in passing, and a mention-only metric treats them identically.
Sentiment and brand safety
How the brand is characterised, including inaccuracies and risk framing. Research indicates sentiment is substantially noisier than mention frequency, so we treat sentiment as directional and flag material claims for human review rather than reporting a sentiment score as a precise figure.
Citation share
Which sources the engine cites when answering prompts in the category, and how often the brand's own domain appears among them. In our experience the brand's own site is typically a small share of citations, with the majority going to third-party sources. This is the metric that converts a score into an action plan, because it tells you which surfaces actually need work.
AI shopping visibility
For retail and D2C, whether products appear against buying-intent prompts, in what position, and which product images surface.
Levels of analysis
The same metrics are reported at several levels, because a single blended figure rarely tells a client what to do next.
By AI platform. Every metric is reported per engine as well as in aggregate. Sentiment in particular differs materially between platforms, so a blended sentiment figure can hide that one engine describes you well and another does not.
By category and product. Visibility and position are reported at category level and, for retail and D2C clients, down to individual products or SKUs. A brand can be strong in a category and absent for its highest-margin line, and a category-level figure will not show it.
By location. Reported at city level, and at finer resolution where a client's catchment justifies it. Answers for the same category can differ between locations because the sources available for those locations differ.
By language. English, Hindi and transliterated forms reported separately as well as together.
What else we analyse
Three analyses that sit alongside the core metrics.
Source analysis. For every prompt set we log and categorise the sources cited in answers, then report which sources are shaping your category's answers, how often each appears, and how often your own domain appears among them. This is the analysis that turns a score into a work queue, because it identifies the specific surfaces that need attention rather than only telling you that a gap exists.
Citation opportunities. We identify the sources your competitors appear in and you do not, prioritised by how often each source is cited in your category. See citation share.
Comparison analysis. When a prompt asks an assistant to choose between your brand and a named competitor, we record which is recommended, in what terms, and which sources that answer cited.
On this last point, a limit worth stating plainly: we report the evidence an answer drew on, not the model's reasoning. A model's internal decision process is not observable, and any vendor claiming to show you why a model chose a competitor is describing an inference. What we can show is which sources were cited, what those sources say about each brand, and what is present in a competitor's coverage and absent from yours. That is usually enough to act on, and it is a different claim from knowing the model's thinking.
How we classify and quality-check
Three of the metrics above require a judgement, not just a count. How that judgement is made is part of the methodology.
Sentiment classification
Characterisation is classified by model, with human review on any answer flagged as negative or as containing a possible factual error. We do not report a sentiment score derived purely from automated classification without that review layer, because mixed-language content, understatement and category-specific idiom are all places automated sentiment gets it wrong.
Position in prose answers
Many answers are not ranked lists. Where an answer names brands in continuous prose, we record order of first mention and flag that it is prose rather than an enumerated list, because the two are not equivalent and averaging them together would be misleading.
Entity disambiguation
Where a brand name is ambiguous, shares a name with an unrelated company, or sits inside a group structure with several entities, we define the matching rules for that client at the start of an engagement and apply them consistently. Mentions we cannot confidently attribute are recorded as unattributed rather than counted.
Competitor set
Share of voice depends entirely on who is in the comparison. The competitor set is agreed with the client, documented, and frozen alongside the prompt set. Adding a competitor mid-period changes every historical share-of-voice figure, so when a set changes we restate the prior period on both.
Decision-grade versus directional
The IAB's August 2026 framework establishes Decision-Grade versus Directional measurement standards, to help buyers assess data quality and fitness for use. We assess our metrics against those criteria and publish the protocol behind each assessment below, so the designation can be checked rather than taken on trust.
| Metric | Our assessment | Basis |
|---|---|---|
| Visibility (with disclosed sample size) | Decision-grade | Stable across sufficient repeated samples |
| Share of voice | Decision-grade within a frozen prompt set | Highly sensitive to set composition, so not comparable across differing sets |
| Citation share | Decision-grade | Cited URLs are directly observable |
| Position | Directional | Ordering varies between runs |
| Sentiment | Directional | Documented to be noisier than mention; material claims flagged for human review |
| Prompt demand estimates | Directional | Modelled, not measured |
| Small period-on-period movement | Assessed case by case | See note below |
On small movements. There is no universal threshold below which a change is noise. Whether two or three percentage points is meaningful depends on prompt count, number of repeated runs, and the historical variability of that specific prompt set.
We assess movement against the confidence interval and historical variance of the relevant set, and we do not report small changes as performance gains without sufficient evidence.
If a vendor tells you a two-point improvement is a win, ask for the sample size and the confidence interval. If they tell you any movement under a fixed number is always noise, ask how they derived the threshold.
Prompt demand estimation
No AI platform publishes prompt-level query volume. Anyone presenting AI prompt volumes as measured data is presenting a model as a measurement.
Our approach: we remodel anonymised third-party prompt data and validate it against external demand signals, computing estimates at the intent-cluster level rather than the individual prompt level. Cluster-level estimates are more robust because they don't depend on exact phrasing, which varies enormously between users.
These numbers are for prioritisation, not reporting. They help decide which clusters to work on first. They should not appear in a board deck as though they were search volume. We label them as estimates everywhere they appear, including in the product.
India-specific handling
Standard methodology treats a market as monolingual and national. Neither holds here.
Language and transliteration
Indian consumers query in English, Hindi, Hinglish and transliterated Devanagari, frequently within one session. We treat semantically equivalent prompts across these forms as a single intent cluster and report them together, while retaining the ability to segment by language. A methodology that treats "best sunscreen India" and "sabse accha sunscreen" as unrelated prompts reports an incomplete picture.
City-level segmentation
AI recommendations for the same category differ between Indian cities, because the underlying sources differ. We run location-scoped prompt variants and report by city where a client has regional concentration. National-only reporting averages this away.
Source ecosystem
In the Indian prompt sets we monitor, the observed citation ecosystem includes source categories that appear less often in Western-market equivalents: regional publications, local aggregators and listings, Quora, Reddit India, city directories and vernacular content. Our citation logging is built to capture these alongside the sources any global tool would track.
Limitations
Stated plainly, because a methodology that claims no limitations isn't one.
- We cannot measure what a specific user sees. Personalisation, location and history all affect real answers. We measure cold, unpersonalised queries, which is a proxy.
- Attribution to revenue is indirect. There is no AI equivalent of Search Console click data. AI-influenced conversions frequently arrive as direct or organic traffic. Self-reported form data and sales-call questions remain the only direct bridge.
- Prompt sets are a choice. A different reasonable set produces different numbers. We disclose ours for this reason.
- Model updates create discontinuities that we annotate but cannot control for.
- Sentiment scoring is imperfect on mixed-language content and on sarcasm, understatement and category-specific idiom.
- We cannot observe model reasoning. Comparison analysis reports the sources an answer cited and what they contain. It does not report why a model chose one brand over another, because that is not observable from outside.
- Not every engine returns citations. Where an engine answers without identifying sources, citation share cannot be computed for that engine and we say so rather than blending it in.
- Small movements are frequently noise.
Disclosure standards
We commit to the following on every published figure and every client report:
- No visibility or share-of-voice figure without its prompt set, engine set, sample size and date window.
- Estimates labelled as estimates, wherever they appear.
- Model release dates annotated on every timeline.
- Prompt set changes disclosed, with the prior period restated on both sets.
- Every statistic is linked to a primary source or to documented methodology.
- Raw data available to clients, including the underlying runs, not just the rolled-up score.
If we ever report a number to you that lacks these, ask for them.
Conclusion
In an industry built on black-box metrics and unfalsifiable scores, transparency isn't just a nice-to-have, it's the only foundation that matters. At Inner Labs, we believe you shouldn't have to take anyone's word for how AI systems perceive your brand.
By openly sharing our rigorous prompt construction, multi-run sampling protocols, citation tracking, and strict disclosure standards, we invite you to inspect the exact mechanics behind our reporting. Because when your growth strategy depends on Indian consumer queries, regional nuances, and third-party source ecosystems, you need more than a dashboard and a monthly slide.
You need data you can audit. Explore our platform capabilities or get in touch with our team to see how your brand truly stands across AI search engines today.
Sources
- IAB, Measuring Visibility in the AI Era, 3 August 2026. Establishes the metrics hierarchy, Decision-Grade versus Directional standards, and disclosure requirements. Full framework PDF
- Sielinski, R., Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement, arXiv. Citation variability, bootstrap confidence intervals, rank instability.
- Don't Measure Once: Measuring Visibility in AI Search (GEO), arXiv. Day-to-day source and brand set overlap.
- Google Search Central. Guidance on structured data and AI features.
About the author

Founder, The Inner Labs
Yohann John is the founder of The Inner Labs, an AI discovery and answer engine optimization platform for brands across India and the GCC. He works with retail, fintech, real estate and hospitality brands on how AI assistants describe and recommend them.
Team profile