Pillar Guide · Resources

The short answer: When ChatGPT answers a question, it doesn't just guess. It draws on search results, retrieved web pages, and real-time data. The links it cites serve as a diagnostic map of what shaped that response.

But here is the crucial caveat: a single citation doesn't mean one isolated web page single-handedly triggered your brand's mention. Traditional SEO teaches you to obsess over your own domain. Yet, across our client scans in Indian consumer categories, third-party platforms consistently account for the large majority of citations behind brand recommendations. Across three client engagements in two categories, the brand's own website was under 2% of cited sources. If you are only optimizing your own site, you are working on a small fraction of what the model is reading.

Yohann JohnFounder, The Inner LabsPublished 28 Aug 2026 · Updated 6 Sept 2026

Key Takeaways

  • ChatGPT does not rank your website. It draws on available sources and names the brands those sources discuss.
  • Across three client engagements in two categories, a brand's own website accounted for under 2% of the sources AI cited. Between 76% and 90% of cited sources did not mention the brand at all.
  • The source mix for Indian queries differs from Western markets: regional publications, Indian aggregators, Quora and Reddit India recur alongside global sources.
  • English and Hinglish versions of the same question can return different brands.
  • Answers can differ between Indian cities because the local sources differ.
  • Schema markup is hygiene, not a lever. llms.txt has near-zero adoption.

Why Is Your Website Not the Main Lever?

In our observed client scans across Indian categories, third-party sources account for the large majority of citations behind an answer about a brand. The full breakdown is further down this page. This runs against the SEO instinct, where your own pages are the primary asset. Here your own pages are necessary but frequently not sufficient.

When a user asks ChatGPT for the best sunscreen in India, the model doesn't scan a list of brand rankings. It retrieves live web sources and synthesizes a recommendation on the spot.

Put simply: Visibility follows conversation. If the sources the model reads aren't discussing your brand, the AI won't recommend it. The two practical consequences of this:

  • A brand with an excellent website and no third-party presence often does not appear at all.
  • A brand with a mediocre website that is well covered by aggregators, review sites and community discussion often appears repeatedly.

What Is Different About Indian Queries?

Most advice on this topic is written for US and European markets. Three things behave differently here.

The source ecosystem is not the Western one

If you follow standard Western advice on AI search, you'll be told to focus on G2, Wikipedia, and US media. In the Indian prompt sets we monitor, the observed citation mix looks different: regional news sites, Indian aggregators, listing directories, Quora, Reddit India and niche local forums recur alongside the global sources any market shares.

We logged every source cited across three client engagements: two hospitality venues and one of India's largest real estate developers. 250 non-branded prompts each, tracked on ChatGPT, Gemini, Claude and Google AI Overviews between February and September 2026.

Source typeHospitality (two engagements)Real estate developer
Third-party sites87% to 89%88.6%
Competitor websites3% to 7%8.2%
User-generated content2% to 7%2.2%
The brand's own website1.6% and 1.9%1.0%

The brand's own website came in under 2% in all three cases, across two quite different categories. And between 76% and 90% of the sources cited in those answers did not mention the brand at all.

One of those prompt sets drew on nearly 3,000 distinct domains. No single website is going to carry a category that fragmented.

There is a detail from the real estate set worth noting separately: the single most-cited domain in the entire sample was a competitor's own website. And the state RERA portal appeared at over 2% of citations, which means official regulatory records are part of what shapes answers about Indian developers.

What this is and isn't. Three engagements across two categories, published with client permission. The mix will differ in other categories and this is not a general law about Indian AI search. Named case studies and the full caveats for the hospitality data are in our hospitality guide. We will publish equivalent breakdowns for more categories as we have them.

Put simply: optimizing against a US B2B checklist means optimizing for the wrong surfaces.

Language and transliteration change the answer

Indian consumers query in English, Hindi, Hinglish and transliterated Devanagari. "Best sunscreen India" and "sabse accha sunscreen" express the same intent and can return different brand sets, because they surface from different sources.

Most brands measure only the English form and conclude they are visible. They have measured part of the conversation and reported it as the whole.

Answers vary by city

Ask the same category question from Mumbai and from Coimbatore and the recommendations frequently differ, because the sources that mention brands locally differ. National-level measurement averages this away, which is fine for a genuinely national brand and misleading for anyone with regional concentration.

What Actually Moves the Answer?

Theory is useful, but execution is what wins visibility. Based on our practical client scans, here are the five things that actually move the needle in AI search results, ordered from most impactful to foundational:

1. Presence in the sources that get cited for your category

This is the highest-leverage work and the slowest. Identify which sources ChatGPT actually cites when answering questions in your category, then earn presence in them. Aggregators, listing sites, review platforms, category publications, comparison articles you do not own.

2. Being described consistently everywhere

Models corroborate. When your company description, category, and core facts are identical across your site, LinkedIn, Crunchbase, review profiles and directories, a model becomes confident about what you are. When each source describes you slightly differently, it hedges, and hedging looks like absence.

3. Content that answers the question directly

Pages built as reference material get retrieved more than pages built as marketing. Practically: a self-contained answer in the first paragraph, question-shaped headings, one idea per section, and specifics rather than adjectives. A model extracting an answer needs a passage that stands alone without the surrounding page.

4. Community and forum presence

Reddit and Quora appear frequently in citations for consumer categories. Genuine participation over months works. Obvious astroturfing is detectable, breaches platform rules, and is a reputational risk that outlives the campaign.

5. Technical accessibility

Your important pages need to be publicly accessible, indexed by Google, and not blocked from the relevant search crawlers. This is easy to get wrong and easy to fix.

The OpenAI crawlers are not interchangeable, and treating them as one bot leads to the wrong robots.txt rules. OpenAI runs three, each controllable independently:

CrawlerWhat it governs
OAI-SearchBotWhether your content is eligible to surface in ChatGPT search features. This is the one that matters for being mentioned.
GPTBotWhether your content may be used to train OpenAI foundation models. Blocking it does not remove you from ChatGPT search.
ChatGPT-UserUser-triggered fetches during a conversation. Not the control point for search eligibility.

So if your goal is appearing in ChatGPT answers, the crawler to allow is OAI-SearchBot. Publishers who want visibility without contributing to training commonly allow OAI-SearchBot and disallow GPTBot. Changes to robots.txt take roughly 24 hours to affect search eligibility.

For Google AI Overviews and AI Mode, the foundation is normal Google indexing and crawlability, not a separate directive. Google-Extended is a distinct control governing use in Gemini and Vertex AI grounding and training, and it is not the switch for AI Overviews eligibility.

Beyond crawler access, two common technical oversights frequently undermine a site's AI visibility, which are:

  • No sitemap. Crawlers then discover pages only by following links.
  • Canonical tags all pointing at the homepage. More common than you would think, and it tells every crawler that only your homepage is worth indexing separately. Each page should carry a self-referencing canonical.

What Does Not Work as Well as People Claim?

Before investing time and resources into your AI strategy, it helps to clear out the tactics that sound good in theory but fail in practice. Here is what you can safely skip:

  • Schema markup is hygiene, not a lever. Structured data helps machines parse your pages correctly and is worth having. But Google has stated structured data is not required for AI search features, and controlled testing has found limited measurable effect on AI citation rates specifically. Add it in an afternoon and move on. Anyone selling schema as the route to AI visibility is selling you a checklist because it is easy to deliver.
  • llms.txt has near-zero adoption. No major AI provider has committed to reading it. Costs five minutes, expect nothing.
  • Publishing at volume does not substitute for being cited. Shipping sixty articles in a week produces sixty pages nobody references. Google's spam policies specifically cover scaled content abuse, and beyond the risk, undifferentiated content does not get retrieved because it does not say anything a model needs.
  • Keyword research misses the question entirely. People do not type keywords into ChatGPT, they type sentences. Keyword tools report Google volume, which is a poor proxy. Build your prompt set from what customers actually say in sales calls and support tickets.

How to Check Where You Stand?

Before optimising anything, get a baseline.

  1. Write down twenty prompts your customers would genuinely ask. Real sentences, not keywords. Include comparison and trust questions, not just discovery.
  2. Run each one several times, in a fresh session with no history. Answers vary between runs, so a single result tells you very little.
  3. Record which brands are named and in what order.
  4. Record which sources are cited. This is the diagnosis — the citation share behind your category's answers. The brands appearing instead of you are appearing because of specific sources.
  5. Repeat in Hinglish for at least five of the prompts.
  6. Repeat from two cities if you have regional concentration.
  7. Screenshot everything, dated. You will want the before picture.

Read the pattern rather than a single number. If your brand rarely appears across repeated runs of your core category questions while the same two or three competitors appear consistently, that is a meaningful gap. A single absent result is not, because answers vary between runs.

What to Expect on Timeline?

Eight to twelve weeks is a common planning window, but it is not a promise. The real timeline depends on how large the source gaps are, whether the sources shaping your category are ones you can influence, and crawl and retrieval cycles you do not control.

Two caveats worth setting with your leadership before you start. Model updates shift citation patterns independently of anything you do, so annotate your tracking with release dates or you will misread an engine change as a content failure. And small movements are frequently noise: whether a two-point change means anything depends on how many prompts you tracked and how many times you ran them.

What Is Inner Labs, and How Can We Help?

Inner Labs is an advanced AI SEO and GEO (Generative Engine Optimization) platform built to help Indian and global brands measure, understand, and actively influence how they are represented across AI systems.

Rather than stopping at basic analytics dashboards, we combine proprietary tracking technology with hands-on growth execution to bridge the gap between insight and outcome:

  • Track Real Intent: Monitor how your brand performs across high-impact prompts, factoring in regional variations, multilingual queries (English, Hindi, Hinglish), and city-level nuances.
  • Diagnose Citation Gaps: Identify precisely which third-party aggregators, review sites, and community forums are shaping AI recommendations in your category.
  • Control Brand Narrative: Track sentiment, ensure factual consistency across the web, and eliminate the "hedging" that makes models treat your brand as absent.
  • Execute & Improve: Move beyond theory with managed execution across content optimization, third-party source presence, and community engagement (such as Reddit and Quora) to secure your position as a top AI recommendation.

Frequently Asked Questions

Does ChatGPT use Google rankings?

Not directly. ChatGPT retrieves sources at the time of the question and synthesises from them. Pages that rank well often get retrieved, since ranking and retrieval share some signals, but ranking first does not guarantee being mentioned, and plenty of mentioned brands do not rank first.

Why does ChatGPT recommend my competitor and not me?

Almost always because the sources it retrieved mention them and not you. Find out which sources those are before changing anything on your own site.

Can I pay to be mentioned in ChatGPT answers?

No. There is no advertising placement inside organic AI answers today. Anyone offering guaranteed placement is describing something else.

Which AI crawlers should I allow?

For appearing in ChatGPT search features, allow OAI-SearchBot. GPTBot governs training use rather than search inclusion, so blocking it does not remove you from ChatGPT answers, and many publishers deliberately allow the first and block the second. ChatGPT-User handles user-triggered fetches and is not the search eligibility control. For Google AI Overviews, ordinary Google crawlability and indexing are the foundation. If your site has no robots.txt at all, everything is allowed by default, which is fine.

How is this different from SEO?

Yes, and it has to be measured separately. Transliterated and Hinglish prompts retrieve different sources and can return different brands. Measuring English only gives you a partial picture.

My agency says they handle AI searches. How do I check?

Ask three specific questions. Which prompts are you tracking, and how did you choose them? What is our share of voice against our top three competitors right now? Which sources is ChatGPT citing when it answers questions in our category? All three have concrete answers. Vague responses usually mean the work is not happening.

About the author

Yohann John, Founder, The Inner Labs

Yohann John

Founder, The Inner Labs

Yohann John is the founder of The Inner Labs, an AI discovery and answer engine optimization platform for brands across India and the GCC. He works with retail, fintech, real estate and hospitality brands on how AI assistants describe and recommend them.

Team profile

Begin

See how AI is shaping your brand's next decision.

Bring us your brand, competitors and category. We will show you the questions that matter and how AI currently answers them.

A 20-minute working session, not a sales deck.

Your category's must-win questions, mapped
A live view of how relevant AI models answer today
Clear gaps, competitors and next opportunities

We respond within one business day. Your details are used only to arrange the demo — see our Privacy Policy.