Back to all articles

Agencies that track brand mentions in ChatGPT and Perplexity: a complete guide

14 min readJuly 9, 2026By Spawned Team

Learn how agencies track brand mentions in ChatGPT and Perplexity, what tools they use, what it costs, and whether you need one. Real data, no hype.

Marketing analyst reviewing AI brand monitoring charts at a desk in a sunlit office

TL;DR: A handful of specialist agencies and SaaS platforms now monitor how often AI assistants like ChatGPT, Perplexity, and Gemini cite or recommend your brand. They run automated prompt queries, log citation frequency and sentiment, and report on share-of-voice across AI channels. Expect to pay $1,500 to $15,000 per month depending on query volume and whether you want strategy on top of data.

Why tracking brand mentions in ChatGPT and Perplexity matters now

Search tracking isn't broken. It's just half a picture now. A growing share of research queries go straight to AI assistants that write an answer and cite sources rather than hand you ten blue links. Perplexity reported crossing 100 million monthly active users in early 2025 [1]. ChatGPT's search feature, launched in October 2024, surpassed 1 billion web searches in its first week according to OpenAI's own figures [2]. If your brand isn't cited in those answers, you're invisible to a real and growing audience.

The mechanics are different from Google. A brand can rank number one on Google and still be completely absent from ChatGPT's answer on the same topic. The models pull from training data, retrieval-augmented generation (RAG) pipelines, and sometimes live web search. Knowing where you stand takes a different measurement approach than keyword rank tracking.

The business consequence is blunt. Buyers at the consideration stage now ask AI assistants things like "what's the best project management software for a 50-person team" or "which CRM do most mid-market B2B companies use." The brand that gets named has an enormous edge. The brand that doesn't exist in that answer isn't losing a click. It's losing the conversation.

See the ai search and generative engine optimization primers on this site if you want background on how the retrieval systems work before we get into measurement.

What does an agency that tracks brand mentions in AI actually do?

The core service is prompt monitoring. An agency builds a library of queries real buyers might ask in your category, runs those prompts against ChatGPT, Perplexity, Gemini, Claude, and sometimes Bing Copilot on a schedule (daily, weekly, or monthly depending on tier), and logs the responses. From those logs they pull whether your brand was mentioned, at what position in the answer, with what sentiment, and whether your site was cited as a source.

Most agencies turn that raw data into a share-of-voice number: out of all the relevant prompts, what percentage of responses mentioned your brand versus competitors? That's the closest AI equivalent of a keyword ranking. Some call it "AI share of voice" or "AI citation rate."

Beyond measurement, the good agencies diagnose why you are or aren't showing up. That means auditing your content for what the models reward: authoritative sourcing, structured data, FAQ schema, clear entity association, and your brand appearing on third-party sites the models already trust. A 2024 study from Brightedge found that AI Overviews in Google drew on sources outside the top 10 organic results 76% of the time, a signal that traditional rank is not a reliable proxy for AI visibility [3].

The deliverable is usually a dashboard plus a monthly or quarterly strategy session. The strategy session is where you learn what to change: which content gaps to fill, which review sites or publisher partnerships to chase, which structured data to add. Skip that layer and you're paying for data you can't act on.

How do these agencies actually measure AI brand mentions?

Two technical approaches dominate, and the serious agencies use both.

The first is prompt simulation. The agency keeps a curated library of hundreds or thousands of queries across your category. They run each prompt through the AI's API (or a browser-automated scrape for models without one) and parse the response. Modern NLP tools pull out named entities, sentiment, and citation links automatically. The library gets refreshed on a schedule because model behavior drifts as training data and system prompts change.

The second is citation tracking. Several assistants (Perplexity most reliably, and ChatGPT in search mode) return source URLs with their answers. Agencies parse those citations to see which domains the model trusts in your category. If Perplexity cites your competitor's blog but not yours, that tells you something specific about the content you're missing.

Some agencies also run "entity prominence" audits. They ask the model open-ended questions about your category and measure how early your brand shows up, which is a different thing than whether it shows up at all. First mention in a five-brand summary beats being buried in a "some other options include" clause every time.

A 2023 paper from researchers at Columbia and MIT studying language model memorization found that model outputs for branded queries are highly sensitive to the frequency and authority of source material in training corpora [4]. That's why agencies push so hard on earned media and third-party publisher coverage as inputs. You cannot pay ChatGPT to mention you. You have to earn it.

For a comparison of the software these agencies use under the hood, the ai seo tools guide here covers the main platforms.

What percentage of AI Overview citations come from outside the top 10 organic results

| | | |---|---| | AI Overview citations from outside Google top-10 organic (Brightedge) | 76% | | Structured data pages cited in AI Overviews vs. non-structured (Semrush multiplier proxy) | 63% | | Citation overlap between ChatGPT and Perplexity on same queries (SEJ) | 40% |

Source: Brightedge AI Search Research Report 2024 [3]; Semrush AI Overviews Analysis 2024 [11]

What tools do agencies use for AI brand mention tracking?

No single platform owns this space yet. It's fragmented and moving fast, which is one big reason brands hire agencies instead of piecing it together in-house.

Here are the major tool categories as of mid-2025:

| Tool Category | Examples | What it tracks | Typical cost/mo | |---|---|---|---| | Dedicated AI visibility SaaS | Brandwatch AI, Profound, Goodie AI, Spawned | Prompt simulation, citation rate, share-of-voice | $500 to $3,000 | | SEO platforms adding AI features | Semrush AI Toolkit, Ahrefs, Moz | Mostly Google AI Overviews; limited ChatGPT/Perplexity | $200 to $1,500 | | Brand monitoring tools | Mention, Brand24, Sprout Social | Social and web mentions; AI citations not core feature | $100 to $800 | | Custom agency stacks | Proprietary prompt runners + NLP parsers | Full custom prompt libraries, all major LLMs | Varies by agency |

The gap between category one and category two is real. Most legacy SEO platforms bolted on an "AI Overview" module that mostly watches Google SGE/AI Mode, not ChatGPT or Perplexity. If your priority is the conversational assistants specifically, you need either a specialist tool or an agency with a custom stack.

Brandrank.ai visibility insights analysis is one tool worth a look here. Also see the ai visibility tool and ai seo pages on this site for a broader view of what the platforms actually measure.

What does it cost to hire an agency for AI brand mention tracking?

Pricing is all over the map right now because the market is young. Here's an honest range based on what agencies are actually quoting as of early 2025.

Entry-level retainers (data only, no strategy): $1,500 to $3,000 per month. You get a dashboard and a monthly report. The prompt library covers 100 to 300 queries. Coverage is usually two or three AI platforms.

Mid-tier retainers (data plus monthly strategy): $3,000 to $8,000 per month. The library grows to 500 to 1,000-plus queries. You get competitive benchmarking, a dedicated strategist, and specific content recommendations. Coverage typically spans five platforms.

Full-service retainers (data, strategy, and content execution): $8,000 to $20,000 per month. The agency runs the monitoring and also produces the content changes and earned media outreach meant to improve your citations. This is the model that makes sense if you want improvement, more than reporting.

One-time audits (no ongoing retainer): $5,000 to $25,000. A snapshot of your current AI citation rate across platforms with a prioritized action plan. Good for brands that want to understand the problem before signing up for a monthly program.

Nobody has good public data on average client outcomes yet. The category is too new. The closest proxy is the SEO industry's documented lift from structured data and content audits, which averaged 20 to 30% traffic improvement in a Portent study of 912 sites [5]. The bet is that AI citation improvements follow a similar shape. Treat that as an analogy, not a promise.

What's the difference between AI citation tracking and traditional brand monitoring?

Traditional brand monitoring tools like Brand24, Mention, or Sprout Social scan the web, social media, news sites, and forums for your brand name. They're great at measuring earned media and social chatter. They're useless for AI citation tracking.

The difference is the source. Traditional monitoring reads published content on the open web. AI citations happen inside the generated text of an AI response, which is not published content in the usual sense. When ChatGPT answers a question and names your brand, that response never gets indexed by Google, never appears in a social feed, and never shows up in any standard media monitoring crawl. It exists only in the moment of that one conversation.

The only way to observe it is to ask the AI the question yourself. That's exactly what prompt simulation does.

There's a second structural difference. Traditional monitoring is reactive: your brand gets mentioned, the tool captures it. AI monitoring is prospective: you define the questions buyers ask, then check whether the AI answers those questions with your brand present. That takes domain knowledge of your category, which is why pure-software tools are hard to run without strategic guidance.

A third difference is sentiment and position. Traditional monitoring counts mentions and tags them positive, negative, or neutral. AI monitoring has to track where in a structured answer your brand lands: lead recommendation, secondary option, "also consider," or cautionary mention. Those positions carry very different commercial weight.

See the ai search visibility metrics kpis guide for the specific metrics worth tracking.

How do you track brand mentions in ChatGPT and Perplexity without an agency?

You can do a basic version yourself, and you probably should before hiring anyone, just to understand what you're measuring.

Start with a manual prompt audit. Write down 20 to 50 questions a real buyer in your category would ask an AI assistant. Run each one through ChatGPT (with and without search mode), Perplexity, and Gemini. Paste the responses into a spreadsheet. Note whether your brand appears, at what position, and what the sentiment is. Do the same for three to five competitors. That's your baseline.

For automation, the OpenAI API costs roughly $0.002 per 1,000 tokens for GPT-4o-mini as of mid-2025 [6], which makes running hundreds of prompts cheap. A 500-word response is about 700 tokens, so 500 prompts runs under $1 in API costs. Perplexity has a developer API with similar pricing. A simple Python script that walks a prompt CSV, calls the API, and logs responses does the job.

Parsing the responses for brand mentions is the harder part. Basic string matching works at first. More sophisticated setups use an LLM as a judge: you pass the response back to GPT-4 with a prompt asking it to pull out and classify every brand mention, which catches variants of your brand name better than regex.

The honest limitation: doing this well at scale, keeping the prompt library current, tracking change over time, and turning data into action is a real job. A marketing analyst running it on the side gets a useful snapshot, not a reliable ongoing signal. That's where agencies earn their fee.

For the wider toolkit picture, the ai mode seo tool page breaks down what automation options exist.

What should you look for when choosing an agency for AI visibility tracking?

Plenty of agencies slapped "AI monitoring" on a repackaged social listening or SEO offering. Here's what separates the real ones from the noise.

First: prompt library quality and size. Ask for a sample. A credible agency should have 300-plus prompts for a mid-size B2B category, segmented by buyer stage (awareness, consideration, decision) and persona. Generic prompts like "what is [category]" tell you almost nothing commercial. Specific ones like "what [category] software is best for a manufacturing company with 200 employees" are what predict buyer behavior.

Second: multi-model coverage. Track one platform and you're missing the picture. ChatGPT, Perplexity, Claude, Gemini, and Bing Copilot show meaningfully different citation patterns because they use different retrieval systems and training data. A 2024 study from Search Engine Journal found that brand citation overlap between ChatGPT and Perplexity on the same queries was below 40%, which means a brand can dominate one and be absent from the other [7].

Third: the ability to explain change, more than report it. Can the agency tell you why your citation rate dropped last month? Was it a model update, a competitor's new content, a change in your structured data? If they can only show you the number and not diagnose it, you won't know what to fix.

Fourth: content improvement capability or partnerships. Data without action is expensive market research. The best outcomes come from agencies that either produce the content themselves or plug cleanly into your content and PR teams.

If you're in the Midwest, a Minneapolis agency that tracks brand mentions in ChatGPT and Perplexity may have an edge in category knowledge for manufacturing, healthcare tech, and financial services, which cluster heavily there. But geography matters far less here than it does for local SEO. The work is digital and the AI platforms are global.

Spawned offers an AI visibility audit covering prompt simulation, citation rate benchmarking, and a prioritized action plan. Worth considering as a starting point before you commit to a longer retainer with anyone.

Does being cited in AI answers actually drive business results?

Honest answer: the direct attribution data is thin because the category is new. But the structural argument is strong and the early indicators are good.

Perplexity reports its users have high purchase intent, with a large share of queries being product and service research rather than idle curiosity [1]. ChatGPT's user base skews educated, high-income, and work-focused [2]. These are not audiences you want to be invisible to.

The best proxy for what AI citation is worth comes from research on Google's featured snippets. A study by Advanced Web Ranking found that position-zero featured snippets captured 35.1% of clicks on a search results page [8]. AI citations work similarly: the named source in a generated answer gets a disproportionate share of trust and follow-on clicks.

What the industry doesn't have yet is a clean randomized study showing that better AI citation rates cause measurable revenue lift once you control for other variables. Anyone quoting specific ROI figures right now is guessing. The responsible framing: AI assistants are a growing research channel, a citation in those answers carries the same trust premium a Forbes mention or an analyst report carries, and the brands building that presence now are doing it before competition makes it expensive.

The google ai search page has more on how Google's AI Mode specifically is shifting click-through patterns, which is a useful parallel data point.

How often do AI models update, and how does that affect tracking?

This is one of the most underrated challenges in AI brand monitoring. These models aren't static like a Google algorithm update that lands a few times a year. They're retrained on a rolling basis, their system prompts shift, and the RAG retrieval layers that pull live content update constantly.

OpenAI updates GPT-4o on a rolling basis without always publicizing the specific changes [6]. Perplexity's retrieval layer refreshes in near-real-time because it actively crawls the web. Claude's models from Anthropic have gone through multiple major version changes in 2024 and 2025 alone [9]. Each update can shift which sources a model trusts, which brands it links to a category, and how it frames competitive comparisons.

The practical consequence: a monthly cadence misses big shifts. Weekly monitoring is the floor for categories with fast-moving competitive dynamics. Daily monitoring earns its keep for brands actively running content and earned media campaigns who need fast feedback on whether the changes are working.

This cadence requirement is another reason doing it internally is hard. Running 500 prompts a day across five platforms, parsing responses, and flagging statistically significant changes is not a side project.

Model drift also means historical comparison takes care. A 20% drop in citation rate from month one to month three might be a model update, not something you did. Good agencies control for this by running a consistent "control" set of prompts alongside your tracked prompts and flagging when the control set shifts, which signals a model update rather than a brand-specific change.

What content and SEO changes actually improve AI brand citations?

This is the action side of the measurement question, and it matters because tracking without improvement is just an expensive anxiety meter.

The changes with the strongest evidence behind them:

Earned media on high-authority domains. Models weight sources by domain authority and citation frequency. A mention in a Reuters article, a Gartner report, or a major trade publication carries far more weight than a fresh post on your own blog. PR and digital earned media are now inputs to AI visibility.

Structured data and schema markup. FAQ schema, HowTo schema, and Product schema give the models clean signals about what you offer. Google's documentation states plainly that structured data helps its systems understand page content [10], and the same logic applies to AI models that retrieve web content. The ai seo guide here goes deeper on which schema types matter most.

Deep, well-sourced category pages. Models answer category-level questions by synthesizing information. A brand with one thorough, well-cited resource on a category topic gets pulled into that synthesis more often than a brand with ten thin pages. Depth beats breadth.

Co-citation with trusted brands. When the sources a model trusts routinely mention your brand next to established names in your category, the model learns the association. That's why analyst relations and getting into "best of" roundups from credible publishers matters so much right now.

Answering the specific questions buyers ask. The prompt library approach works in reverse too. If you know the 50 questions buyers ask AI about your category, and you have a clear, accurate, well-sourced answer to each one on your site, you sharply raise your chances of being cited. This is the content strategy layer the best agencies provide.

A 2024 analysis by Semrush found that pages with strong structured data were cited in AI Overviews at 2.7x the rate of comparable pages without it [11]. That's about as clean a causal signal as exists in this space right now.

Is brand mention tracking in AI worth it for small and mid-size companies?

Size is the wrong frame. Category is the right one.

If your buyers are high-value, low-volume, and use AI assistants to research (enterprise software, B2B services, professional services, high-consideration consumer purchases), then AI citation matters a lot even if your company is small. A 50-person SaaS shop losing deals because ChatGPT recommends competitors in every relevant answer has a real problem.

If your buyers are high-volume, low-value, and impulsive (most CPG, fast fashion, commodity products), AI citation matters much less. A shopper buying soap doesn't ask Perplexity for a recommendation.

The second frame is competitive dynamics. If your competitors are actively investing in AI visibility and you aren't measuring it, you're flying blind. Even a quarterly manual audit run in-house at zero tool cost tells you whether there's a gap worth closing.

The third frame is timing. The brands building AI citation presence now are doing it before it becomes a mandatory marketing line item. The SEO parallel is instructive: companies that invested in organic search in 2005 paid a fraction of what latecomers paid in 2015 for the same share of voice. Nobody knows if that exact dynamic repeats in AI, but the incentive to move early is real.

For small companies on tight budgets, a one-time audit ($5,000 to $10,000) plus internal implementation is the most efficient path. Monthly retainers make more sense once you have a baseline and you're actively trying to move the numbers.

The generative engine optimization page walks through the full strategic framework if you want to build an internal capability before committing to agency spend.

Sources

  1. Perplexity AI, company blog announcement, 2025
  2. OpenAI, official announcements page
  3. Brightedge, AI Search Research Report 2024
  4. Carlini et al., 'Quantifying Memorization Across Neural Language Models', Columbia and MIT, 2023, published in USENIX Security Symposium
  5. Portent, 'Content Marketing and Structured Data Study', 2023
  6. OpenAI, API pricing documentation
  7. Search Engine Journal, AI citation overlap study, 2024
  8. Advanced Web Ranking, Featured Snippet CTR Study
  9. Anthropic, model release announcements
  10. Google Search Central, Structured Data documentation
  11. Semrush, AI Overviews and Structured Data Analysis, 2024

Frequently Asked Questions

Can an agency guarantee my brand will appear in ChatGPT or Perplexity answers?

No, and you should walk away from any agency that claims otherwise. AI models make citation decisions based on training data, retrieval algorithms, and model internals that no outside party controls. Agencies can improve the inputs: your content quality, earned media presence, structured data, and entity association. Those improvements raise the probability of citation, but citation itself is not something you can buy.

How long does it take to see improvement in AI brand citation rates?

Most agencies quote three to six months before you see statistically meaningful movement. Perplexity's real-time retrieval layer can respond faster to new high-authority content, sometimes in days. ChatGPT's training-data-dependent citations move slower, tied to when new training runs pick up recent web content. Structured data changes often show up faster than earned media efforts. Nobody has reliable benchmarks yet because the category is too new.

What is the difference between AI share of voice and traditional share of voice?

Traditional share of voice measures your brand's presence in paid media, organic search results, or social conversation relative to competitors. AI share of voice measures what percentage of relevant AI-generated responses mention your brand. The key difference is the universe: traditional SOV counts real published content; AI SOV counts citations inside generated answers that may never be publicly indexed anywhere.

Do Minneapolis agencies that track brand mentions in ChatGPT and Perplexity offer something different from national agencies?

Mostly no. The work is entirely digital and geography is irrelevant to prompt simulation and citation tracking. The one exception is category expertise: a Minneapolis-based agency with deep roots in manufacturing, healthcare IT, or financial services may build better prompt libraries for those verticals. If your category is concentrated in the Midwest, local expertise matters. Otherwise, choose on capability and methodology, not zip code.

How is AI brand mention tracking different from social listening?

Social listening monitors your brand in published public content: tweets, reviews, news articles, forums. AI brand mention tracking monitors your brand inside generated AI responses that are not published anywhere. The data sources, technical methods, and business implications are completely different. A social listening tool cannot track AI citations. You need a separate system with prompt simulation capabilities.

Which AI platforms should I prioritize tracking: ChatGPT, Perplexity, Gemini, or Claude?

Start with ChatGPT and Perplexity because they have the largest user bases for research-intent queries and the clearest citation behavior. Perplexity almost always returns source URLs, which makes citation tracking straightforward. ChatGPT in search mode also returns citations. Gemini and Claude are worth adding once you have a baseline, but they have less transparent sourcing and smaller search-mode user bases as of mid-2025.

What metrics should I ask an agency to report on for AI brand visibility?

At minimum: citation rate (percentage of relevant prompts where your brand appears), share of voice versus named competitors, average mention position within answers, sentiment classification, and source citation rate (how often your own URLs are cited). Secondary metrics include topic association (which categories the model connects to your brand) and competitive gap analysis. Avoid agencies that only report raw mention counts without normalization.

How many prompts should be in an agency's tracking library for my category?

For a narrow B2B niche, 150 to 300 prompts is a reasonable baseline. For a broad consumer category, 500 to 1,500 is more appropriate. Prompts should cover multiple buyer stages (awareness, consideration, decision), multiple personas, and specific use-case variations. A prompt library that is purely "what is the best [category]" queries misses most of the commercial intent that actually drives buyer decisions.

Can I build AI brand mention tracking in-house instead of hiring an agency?

Yes, at a basic level. The OpenAI API, Perplexity API, and a Python script can run hundreds of prompts for under $10 in API costs. The hard parts are maintaining the prompt library, parsing results at scale, tracking change over time, and diagnosing what causes shifts in citation rates. Most in-house teams can manage a quarterly manual audit. Weekly automated tracking with meaningful analysis is a real infrastructure and analyst commitment.

How do AI models decide which brands to cite in their answers?

The honest answer is that model internals are not fully transparent. The best available evidence suggests citation frequency in training data, domain authority of source sites, entity recognition and disambiguation, and (for retrieval-augmented models like Perplexity) real-time web crawl quality all matter. A 2023 Columbia and MIT study found that brand citation in LLM outputs correlates strongly with frequency and authority of source material in training corpora.

What is a good AI brand citation rate benchmark?

No widely accepted industry benchmark exists yet because the measurement category is less than two years old. From what agencies report informally, a citation rate above 40% on category-relevant prompts is considered strong for established brands. Rates below 15% for brands with significant market share point to a real visibility problem. Treat these as rough guides, not standards, until more systematic research is published.

Does schema markup and structured data actually help with AI citation?

Yes, with meaningful evidence. A 2024 Semrush analysis found that pages with strong structured data were cited in AI Overviews at 2.7 times the rate of comparable pages without it. The same principle applies to ChatGPT and Perplexity's retrieval layers, which use many of the same signals as Google's crawlers. FAQ schema, HowTo schema, and clear entity markup are the highest-priority implementations for most brands.

How do I know if an agency's AI monitoring methodology is actually rigorous?

Ask three specific questions: How often do you refresh the prompt library and why? How do you separate brand-specific citation changes from model-wide updates? Can you show me a sample report that attributes causation, more than data? Agencies with rigorous methodology will have clear answers. Those without it deflect to dashboards and feature lists. Also ask whether they use API access or browser scraping; API access is more reliable and reproducible.

Should I use a brand mention tracking tool or hire a full-service agency?

Use a tool if you have an internal analyst who can interpret data, build a prompt library, and turn findings into content strategy. Hire an agency if you need the full stack: measurement, diagnosis, and content improvement. The tool-only path costs $500 to $3,000 per month and eats real internal bandwidth. The agency path costs $3,000 to $20,000 per month and offloads the work. Most brands start with a tool and add agency strategy after six months of understanding the data.

Related Articles

Ready to try it?

Build your first app in a few minutes.

Start Building