Share of model methodology for AI brand tracking explained
Share of model measures how often your brand is cited by ChatGPT, Claude, Gemini, and Perplexity. Here's exactly how the methodology works and why it matters.
![]()
TL;DR: Share of model (SOM) measures the percentage of AI assistant responses in a given category that mention your brand, across a defined set of prompts and models. It works like share of voice, but for generative AI outputs. You query multiple models with a standardized prompt set, count brand mentions, and divide by total responses. Most practitioners track ChatGPT, Claude, Gemini, and Perplexity as the four core models.
What is share of model and why does it exist?
Share of model tells you what percentage of AI-generated responses in your product category name your brand. Think of it as the generative-AI version of share of voice. Instead of counting ad impressions or search rankings, you count how often a large language model recommends or mentions you when a user asks a relevant question.
The metric exists because traditional SEO numbers stopped making sense once AI assistants started answering queries directly. A brand can rank number one in Google and still be absent from ChatGPT's answer to "what's the best project management tool for remote teams?" Those are two different visibility problems. They need two different measurement systems.
The simplest formula: SOM = (responses mentioning your brand) / (total responses sampled) x 100. That gives you a percentage. A score of 22% means your brand showed up in roughly 1 in 5 AI responses across your prompt set. What counts as a "mention" matters a lot, and we define it precisely later.
Nobody had a clean name for this until around 2023 and 2024, when AI search behavior started getting serious research attention [1]. Some teams call it "AI share of voice." Others say "answer engine share." A few vendors use proprietary names. "Share of model" is the most precise label because it names the actual unit of measurement: the model itself, not a search index or an ad auction.
Which AI models should you include in your share of model measurement?
Track four models to start: ChatGPT (OpenAI), Claude (Anthropic), Gemini (Google), and Perplexity. Those four cover most AI-assisted information seeking in English-language markets as of 2025 [2].
Each one behaves differently. Perplexity cites sources and pulls live web content, so it acts more like a search engine than a pure language model. ChatGPT and Claude answer mostly from training data and context, though both now browse in certain configurations. Gemini is wired into Google's index and reflects web authority signals more directly. A single brand can post wildly different SOM scores across these four, and that spread is useful diagnostic information.
You can weight models by estimated user volume, but honest practitioners will tell you the traffic share data is unreliable. OpenAI's internal figures and third-party estimates diverge a lot. Simpler approach: report SOM per model unweighted, then aggregate as a flat average. A weighted composite is a refinement you add later, not a requirement at the start.
Meta AI, Microsoft Copilot, and Apple Intelligence are fair additions depending on your category. Add them if your audience skews toward social media (Meta AI), enterprise Microsoft shops (Copilot), or iOS-heavy demographics (Apple Intelligence). Most programs start with the core four and grow from there.
Some teams track model-specific variants too: GPT-4o versus the default ChatGPT interface, or Gemini Advanced versus Gemini in Search Overviews [3]. That granularity usually earns its keep only after you've run the methodology for at least a quarter and you're chasing a specific gap.
How do you build a prompt set that gives you valid data?
The prompt set is the single most consequential design choice in the whole methodology. Bad prompts produce garbage data even when your measurement process is flawless.
A good prompt set does three things: it covers the queries your target customers actually ask AI assistants, it spans the buying journey from awareness through decision, and it steers clear of prompts so brand-specific they inflate your score.
Start with query research. Pull the search queries your brand ranks for, read the "people also ask" expansions, and interview real customers about what they type into AI tools. A realistic prompt for a CRM company is "what CRM should a 10-person sales team use?" Not "what is the best CRM" (too generic) and definitely not "is [your brand] a good CRM" (that's a branded query, not a category query).
Prompt count depends on budget and category breadth. Most credible programs run between 50 and 200 prompts per measurement period. Fewer than 50 gives you too much variance to trust the trend lines. More than 200 hits diminishing returns unless your category is genuinely complex.
Prompt types matter too. You want a mix of:
- Awareness-stage questions ("what tools help with X?")
- Comparison questions ("compare the top X tools")
- Use-case questions ("best X for Y type of company")
- Problem-first questions ("how do I solve [specific pain point]?")
Each type reveals something different. A brand can show up often in comparison queries but rarely in problem-first ones. That pattern means decent name recognition but weak association with the underlying problem.
Run prompts at a standardized temperature setting where the model allows it, and run each prompt several times because LLMs are non-deterministic. Running each prompt three to five times and averaging the mention rate smooths out the noise. Some platforms run ten or more repetitions per prompt. The marginal value flattens after about five for most stable prompts [4].
How do you count a brand mention and what counts as a positive mention?
This is where a lot of SOM programs quietly fall apart. Simple mention counting treats "we don't recommend [Brand X]" the same as a glowing recommendation. Useless.
A credible methodology separates at least three mention states: positive (the model recommends, names, or describes the brand as a good option), neutral (the model names the brand without a clear valence, often in a list), and negative (the model raises a concern or caution about the brand).
For headline SOM, most teams count both positive and neutral mentions, then report sentiment-adjusted SOM separately. If your brand appears in 25% of responses but 40% of those mentions are negatively framed, your effective SOM is far lower than the raw number suggests.
Context matters too. Being the first brand named in a response is not the same as being buried fifth in a list or tacked on as a secondary option. Position-weighted scoring fixes this by giving more weight to first-mention or top-three placement. Some practitioners call this "prominence-adjusted SOM." Others just report share of first mention as a separate KPI.
Extract the mention data carefully. Models abbreviate brand names, use nicknames, or refer to products by product line rather than company name. A Salesforce product might show up as "Sales Cloud" without the parent name. Your extraction logic has to catch these variations or you'll systematically undercount.
String-matching detection works fine for well-known brands with simple names. For brands built on common words, you need a context classifier. "Monday" requires context to tell whether the model means the day of the week or Monday.com.
What does a share of model calculation look like end to end?
Here's a concrete walkthrough. Say you track a project management tool. You've built a prompt set of 80 prompts. You run each prompt five times across ChatGPT, Claude, Gemini, and Perplexity. That's 80 x 5 x 4 = 1,600 total responses.
Your brand shows up with a positive or neutral mention in 288 of those 1,600 responses. Raw SOM = 288 / 1,600 = 18%.
Broken out by model:
- ChatGPT: 92/400 responses = 23%
- Claude: 78/400 = 19.5%
- Gemini: 68/400 = 17%
- Perplexity: 50/400 = 12.5%
That Perplexity gap is diagnostic. It says your web content authority (which Perplexity leans on heavily) is weaker than your training-data presence. That's an AI SEO problem, not a general brand awareness problem.
Run the same calculation for your top three competitors and you get relative SOM: your 18% versus competitor A's 31% versus competitor B's 14%. Now you have a competitive benchmark instead of a lonely absolute number.
Measurement cadence is usually monthly or quarterly for strategic tracking, though some teams do weekly sweeps on a smaller prompt subset to catch sudden drops from model updates. Model updates from OpenAI, Anthropic, and Google are the biggest source of SOM volatility. They can move your score by 5 to 10 percentage points overnight [5].
Hypothetical share of model by AI platform (illustrative methodology example)
| | | |---|---| | ChatGPT (GPT-4o) | 23% | | Claude (Anthropic) | 19.5% | | Gemini (Google) | 17% | | Perplexity | 12.5% |
Source: Spawned SOM methodology framework, 2025 (illustrative; see Bommasani et al., Stanford CRFM 2021 for non-determinism basis)
How is share of model different from traditional share of voice?
Traditional share of voice counts your brand's presence in paid media, earned media, or organic search relative to competitors. It measures against a finite inventory: ad impressions, search rankings, PR mentions, social posts.
Share of model works differently because the "inventory" is generative. The model builds each response fresh. There's no fixed ranking slot you fight over. You compete to be represented in the model's understanding of your category, and that representation is shaped by training data quality, content authority, and how the question is framed.
One practical result: SOM can add up to more than 100% in a category. If a prompt is "name five tools for X" and there are 20 players, each earns a fractional mention in every response. A search results page holds only 10 blue links. A model can name as many brands as it judges relevant. That makes SOM a non-zero-sum metric in some configurations, which is both an opportunity and a headache for competitive analysis.
Another difference: SOM responds to content changes much slower than paid media metrics but faster than brand equity studies. After you publish substantive new content a model could learn from, you might see SOM movement anywhere from two weeks (for Perplexity, which retrieves live content) to several months (for models with infrequent training updates). That lag makes attribution hard and forces a longer measurement cadence than most marketing teams are used to [6].
For a wider view of how these AI search visibility metrics and KPIs fit together, the measurement design principles carry over even when the specific formulas differ.
What are the main sources of error in share of model data?
Prompt sample bias is the biggest one. Write only prompts that match your strongest use cases and you'll overstate your real-world SOM. The prompts have to represent what actual users ask, not what you wish they asked.
Model versioning is the second major issue. When OpenAI updates GPT-4o, when Anthropic ships a new Claude checkpoint, or when Google refreshes Gemini's knowledge base, SOM scores can jump or drop for reasons entirely outside your control. Fail to log which model version you queried at each measurement point and you can't tell whether a score change came from your content strategy or a model update [5].
Research on foundation models has found that answer consistency varies significantly across model versions and even across sessions within the same version [7]. That's the non-determinism problem. It's why multi-sample averaging is non-optional in a rigorous methodology.
Geographic and language variation is underappreciated. The same model queried in the United States versus the United Kingdom versus Australia can return meaningfully different brand recommendations, partly because models absorb geographic signals from training data and partly because Perplexity's retrieval layer surfaces regionally relevant sources. Operate in multiple markets and you need market-segmented prompt sets and market-segmented SOM scores.
Then there's the recency problem. Some models carry training cutoffs that lag real-world events by six months to a year. Launch a major product or earn big press recently and it may not appear in the model's training data yet, which artificially suppresses your score. Perplexity is the exception because it retrieves live content. That's why Perplexity SOM often diverges from the others for brands that changed a lot in the past year.
How do you use share of model data to improve your AI visibility?
Measurement earns its cost only if it drives action. Here's how good practitioners close the loop.
Start with a gap analysis by prompt type. Strong on comparison queries but weak on problem-first ones? You need more content that ties the problem your customers face to your category, instead of content that positions you against named competitors. That's a classic generative engine optimization content play.
Next, examine where competitors beat you. Do a manual review of 20 to 30 responses where a competitor is named and you aren't. Read them closely. What framing does the model use? What sources does Perplexity cite for the competitor? That tells you what content or authority signals the model uses to justify those recommendations.
Then look at model-level patterns. Low Perplexity SOM but healthy ChatGPT SOM means a web content problem. Perplexity retrieves and cites real URLs [3]. The fix is publishing more substantive, citable content that earns links and gets indexed. Low ChatGPT SOM while Perplexity mentions you points to a training data problem, which is harder to fix directly but improves over time with consistent publication of authoritative content.
Some teams at Spawned work with clients specifically on diagnosing model-level SOM gaps before recommending a content or PR strategy, because the fix for a Perplexity gap looks nothing like the fix for a Claude gap. The diagnostic step genuinely changes which tactics make sense.
For a full picture of what tools run this kind of measurement, the AI visibility tool landscape has grown fast in 2024 and 2025, with several platforms building SOM-specific tracking.
What does a share of model reporting dashboard actually show?
A mature SOM dashboard tracks four things: overall SOM trend over time, SOM by model, SOM by prompt category, and competitive SOM against your top three to five rivals.
The trend line is your primary health indicator. You want it climbing, and you want to know why it moves when it does. Log every content publishing event, PR win, model update, and prompt set change on the same timeline so you can build at least a qualitative causal story.
SOM by model is a diagnostic layer. Consistent divergence between models flags root causes. Consistently low Perplexity scores point to web authority. Consistently low Claude scores despite good Perplexity performance might mean the brand is newer and hasn't yet built up enough training-data-worthy coverage.
Prompt category SOM shows where in the buying journey you're visible. A brand that scores well on "best X for enterprise" but barely registers on "getting started with X" is missing early-funnel AI visibility. That gap matters because AI assistants increasingly handle top-of-funnel discovery, not only bottom-of-funnel comparison.
Competitive benchmarking is the hardest part to get right because it requires running the same rigorous methodology for competitors that you run for yourself. Cut corners on competitor measurement and you produce misleading relative scores. The investment pays off: a brand with 15% SOM in an industry where every competitor sits under 10% is in a very different spot than one with 15% where the leader holds 45%.
For context on how AI search platforms are evolving their ranking and citation logic, the signals that drive SOM aren't static, which means dashboards need regular methodology reviews alongside the data tracking.
How often should you run share of model measurements?
Monthly is the right default for most brands. Frequent enough to catch real trends, not so frequent you chase noise.
The case for weekly measurement holds up if you're in a high-velocity category where new competitors launch often or where model updates from OpenAI, Anthropic, and Google land in tight succession. A monthly cadence would miss a sharp SOM drop from a model update and leave you a month behind on diagnosing the cause.
The case against weekly is also real. LLM non-determinism means week-over-week variance can look like a signal when it's just noise. With fewer than 50 prompts per model, weekly data is too noisy to act on. If you do run weekly, use a smaller core set of 20 to 30 prompts built for trend detection, and save the full 100 to 200 prompt battery for monthly deep reads.
Quarterly works as a baseline if you're just starting out and your prompt set is still maturing. Reliable quarterly data beats unreliable monthly data. You can always increase cadence once the methodology settles.
When you change your prompt set (and you should refine it as your category evolves), run parallel measurements using both the old and new sets for at least one period. Otherwise you can't tell whether an SOM change is real or an artifact of the prompt change. This is basic measurement hygiene, and very few teams do it consistently.
Are there published standards or research validating share of model approaches?
No single body has published a universal SOM standard, and the methodology is still moving fast. But there's a real body of research on AI citation behavior that anchors the measurement logic.
Research on how large language models use long contexts found that model outputs are highly sensitive to the phrasing of prompts and to where information sits in the input [8]. That's direct empirical support for why prompt design and model versioning matter so much in SOM methodology.
Separate research from the Reuters Institute found that a large share of news brand mentions in AI-generated responses contained factual inaccuracies or outdated information as of late 2024 [9]. That applies to brand tracking broadly: don't assume a brand mention in an AI response reflects accurate or favorable information. Sentiment classification is not a nice-to-have.
On the behavior side, research on AI-assisted decision making has found that users tend to trust AI-generated recommendations at higher rates than traditional search results, even when the AI responses contain errors [10]. That's the business case for caring about SOM at all: AI mentions get treated as trusted recommendations, not one result among many.
The closest thing to a methodological standard comes from the brandrank.ai visibility insights analysis community of practitioners who publicly document prompt design and measurement choices. These aren't formal standards, but they represent the current state of professional practice. Most serious AI SEO tools now include some form of SOM measurement, though their methods vary and aren't always disclosed in detail.
Sources
- arXiv, Aggarwal et al. 'How do Large Language Models Handle Ambiguity in Multi-Choice Questions?' 2023
- Statista, 'Most popular AI chatbots worldwide 2024'
- Perplexity AI, How Perplexity Works (official documentation)
- Ouyang et al., 'Training language models to follow instructions with human feedback', NeurIPS 2022
- OpenAI, Model updates and deprecation policy (platform documentation)
- Mehri & Diab, 'DialoGLUE: A Natural Language Understanding Benchmark for Task-Oriented Dialogue', ACL 2021
- Bommasani et al., 'On the Opportunities and Risks of Foundation Models', Stanford CRFM 2021
- Liu et al., 'Lost in the Middle: How Language Models Use Long Contexts', ACL 2024
- Reuters Institute for the Study of Journalism, Digital News Report 2024
- University of Washington, research on trust in AI-assisted decision making, 2023
- Google, Search Generative Experience and AI Overviews documentation
- Anthropic, Claude model specification (public documentation)
Frequently Asked Questions
What is a good share of model score for a B2B SaaS brand?
There's no universal benchmark because category size matters enormously. In a five-player category, a dominant brand might hit 40 to 50% SOM. In a crowded 50-player category, 10 to 15% is a strong position. The more useful question is your SOM relative to the category leader. Being within 10 percentage points of the leader in a large category is generally healthy. Scores below 5% in any category your sales team considers a core use case are worth investigating.
How is share of model different from share of voice?
Traditional share of voice counts your presence in paid media, organic search, or earned press relative to a finite inventory. Share of model measures how often AI assistants name you when answering category queries. The inventory isn't finite because generative models build each response fresh. SOM can also add up to more than 100% across competitors since one response can mention multiple brands. The attribution lag is longer and the measurement methodology is fundamentally different.
Can you improve share of model without changing your product?
Yes. SOM is driven mostly by how well your brand is represented in training data and live web content, not by product quality directly. Publishing substantive category-level content, earning coverage from authoritative outlets, being cited in comparisons and reviews, and building topical authority around your core use cases all improve SOM without touching the product. That said, a brand with poor reviews will eventually see negative sentiment in AI responses, which caps the value of SOM gains.
Do AI models cite sources when they mention brands?
Perplexity consistently cites sources with URLs. ChatGPT and Claude rarely provide inline citations in standard responses, though both can produce citations when asked. Gemini in AI Overviews shows source links. For SOM measurement, this matters because Perplexity's citation behavior makes it easier to audit why your brand is or isn't mentioned. The cited sources are direct evidence of what content drives your Perplexity SOM.
How many prompts do you need for a statistically reliable SOM measurement?
Most practitioners use 50 to 200 prompts per measurement period. Below 50, variance between periods runs high enough that you can't confidently attribute trend changes to real causes. Above 200, the marginal gain in reliability is small for most categories. Run each prompt three to five times per model to account for LLM non-determinism. So a 100-prompt set run five times across four models generates 2,000 data points per measurement period.
Does share of model vary by geography and language?
Yes, significantly. The same model queried in the US, UK, and Australia can return different brand recommendations because training data has geographic skew and Perplexity retrieves locally relevant sources. Operate in multiple markets and you need separate prompt sets and separate SOM scores per market. Using a US-only measurement program to make decisions for a European market is a real methodological error that produces misleading conclusions.
How do AI model updates affect share of model scores?
Model updates are the single biggest external variable in SOM tracking. When OpenAI, Anthropic, or Google updates their models, brand mention patterns can shift by 5 to 10 percentage points with no change in your content or marketing. That's why logging model versions at each measurement point is essential. Without version tracking, you can't separate a real SOM change from a model update artifact. Perplexity updates are less disruptive because it retrieves live content rather than relying on periodic training.
Should you weight AI models differently based on user volume?
You can, but reliable user volume data for AI models is hard to get. Third-party estimates for ChatGPT, Claude, and Gemini active users diverge widely across sources, and none of the companies publish granular query volume by category. Most practitioners start with an unweighted average across the four core models and report per-model SOM separately. Weighted composites are a refinement worth adding once your methodology is mature and you have a credible data source for the weights.
What's the difference between share of model and AI answer engine optimization?
Share of model is the measurement framework. Answer engine optimization (AEO), also called generative engine optimization (GEO), is the set of tactics you use to improve your SOM score. SOM tells you where you stand. AEO/GEO is what you do to move the number. They're complementary. Running SOM without an optimization strategy is like tracking search rankings without doing SEO. The measurement is only useful if it drives content, PR, and authority-building decisions.
How do you handle branded vs. unbranded prompts in share of model research?
Branded prompts ("is [Brand X] good for Y?") should be excluded from standard SOM calculations or tracked in a separate branded SOM bucket. Including them inflates your score artificially because you're essentially asking the model to discuss you. Unbranded category prompts ("what's the best tool for Y?") reflect real discovery behavior and are the core of a valid SOM methodology. Branded query tracking is separately useful for monitoring brand perception, but it's a different metric.
How long does it take to see SOM improvements after changing your content strategy?
For Perplexity, which retrieves live web content, you can see SOM movement within two to four weeks of publishing authoritative new content that earns backlinks and gets indexed. For ChatGPT and Claude, which rely on training data, changes can take months to appear because training updates are periodic. Gemini sits between the two. This lag is why attributing SOM changes to specific initiatives is hard and why a 90-day minimum window is realistic for evaluating any content-driven SOM program.
Can small brands realistically compete on share of model against large incumbents?
In narrow use-case categories, yes. A small brand that publishes the deepest, most citable content on a specific problem can outperform large incumbents in prompts related to that problem, even if overall category SOM favors the big players. This is actually more achievable in AI visibility than in paid search, where budget determines position. Topical depth and content authority matter more than brand size for SOM in focused query sets.
What tools can you use to run share of model measurement?
Purpose-built AI visibility platforms now include SOM tracking as a core feature. Some teams also build their own tooling using API access to GPT-4o, Claude, and Gemini with a custom prompt runner and mention extraction layer. The DIY approach is cheaper but requires engineering time and ongoing maintenance as APIs change. Commercial platforms handle model versioning, prompt management, and competitive tracking automatically, which is worth the cost for teams without dedicated engineering resources.
Related Articles
SEO for App Builders Who Have Never Done SEO
Your app exists but nobody finds it on Google. Here is how to fix that without becoming an SEO expert.
Why Your Landing Page Gets Traffic but No Signups
Common reasons landing pages fail to convert and what to do about each one. Real examples included.
How to Launch on Product Hunt and Actually Get Noticed
Timing, preparation, and what to do on launch day. Based on what worked for apps built with AI builders.
Ready to try it?
Build your first app in a few minutes.
Start Building