Back to all articles

HubSpot AI search grader share of voice metrics explained

13 min readJuly 9, 2026By Spawned Team

HubSpot's AI Search Grader measures brand share of voice across ChatGPT, Perplexity, and more. Here's what the metrics mean and how to act on them.

Marketing strategist reviewing AI share of voice comparison charts at desk

TL;DR: HubSpot's AI Search Grader is a free tool that runs your brand name through AI assistants like ChatGPT and Perplexity, then reports a share of voice score showing how often your brand appears in AI answers versus competitors. The three core metrics are mention frequency, sentiment, and competitive share. It's a solid starting point with real methodological limits you should understand before you act on the numbers.

What is HubSpot's AI Search Grader and what does it actually measure?

HubSpot's AI Search Grader is a free diagnostic that tells you whether your brand shows up in AI-generated answers. It launched in late 2024 [1]. You enter a brand name and a product or service category, the tool fires a set of representative queries at AI assistants, and it analyzes the responses.

The output is built around three things: how often your brand gets mentioned, whether that mention is positive or neutral, and how your mention rate compares to a handful of named competitors in the same category. HubSpot calls this combination your AI brand visibility score. The underlying mechanic is a share of voice calculation applied to AI responses rather than to traditional search result pages.

Share of voice in traditional SEO means the percentage of clicks or impressions your brand captures in a given keyword set. HubSpot's grader reframes that for answer engine outputs. If 10 queries about your category produce 30 total brand mentions and your brand appears in 9 of them, your share of voice is 30%. The simplicity is deliberate. HubSpot built this for CMOs and demand-gen leads who want a directional signal, not a data science team.

The sentiment layer is where it gets more interesting. A mention is not the same as a recommendation. The grader tries to classify whether your brand appears in a favorable context, a neutral comparison, or a negative framing. That distinction matters enormously. An AI assistant that mentions you only to say you're expensive has effectively hurt you with a high-intent buyer.

How does HubSpot calculate AI share of voice for brand mentions?

The grader's share of voice figure is a mention-frequency ratio. The tool generates a standardized set of prompts for your declared category, collects the text outputs from AI models, identifies brand names using entity recognition, and counts occurrences [1]. Your brand's count divided by total brand mentions across all responses gives you the percentage.

This is close to how tools like Brandwatch or Mention calculate share of voice on social platforms, but AI search adds a wrinkle. AI model responses are probabilistic. Run the same prompt twice and you may get different answers, different citations, and different brand mentions. HubSpot's grader takes one pass per query set, so the score you see is a point-in-time snapshot with meaningful variance baked in.

Nobody has published peer-reviewed reliability data on HubSpot's specific implementation. The closest independent research is a 2024 preprint from researchers at Columbia University and Georgia Tech that found AI model outputs for commercial queries varied significantly across repeated runs, with brand mention rates shifting by plus or minus 15 percentage points between identical sessions [2]. That's not a knock on HubSpot. It's a property of large language models that any share of voice tool in this space has to contend with.

For category-level benchmarks, HubSpot segments by industry vertical and shows where your score sits relative to others in the same category. The benchmark data comes from HubSpot's own user pool, which skews toward B2B SaaS and marketing technology companies. If you're in those categories, the comparisons are useful. If you're in healthcare, manufacturing, or consumer goods, treat the benchmark with skepticism.

See how this compares to the broader landscape of ai search visibility metrics kpis tools that go deeper on measurement methodology.

What AI engines does HubSpot's grader check for share of voice?

As of mid-2025, the grader queries ChatGPT (GPT-4 class models via OpenAI's API) and Perplexity as its primary engines [1]. HubSpot has said Google's AI Overviews and Gemini are on the roadmap, but as of this writing those integrations were not confirmed in the product.

This matters a lot for interpreting your score. ChatGPT and Perplexity behave differently. Perplexity is retrieval-augmented, meaning it pulls live web content and often cites sources explicitly. ChatGPT without browsing enabled draws on training data and leans toward established brands with large historical web footprints. A brand that spent several years publishing thought-leadership content may score better on ChatGPT simply because that content made it into the training corpus, regardless of its current relevance.

Gemini's behavior in commercial queries, especially within Google's AI Mode, is a different animal. Google's own guidance on AI Overviews shows those features pull heavily from top-ranked organic pages and featured snippets [8]. For many categories, winning AI Overviews takes the same signals that win traditional SEO, which means the grader's ChatGPT-centric data may not translate to your Google AI search performance.

If Google AI Mode is your main concern, the grader gives you useful but incomplete information. You'd want a dedicated google ai search monitoring workflow to close that gap.

AI assistant referral traffic: share of marketers reporting measurable visits

| | | |---|---| | All marketers surveyed | 38% | | B2B SaaS companies | 62% |

Source: SparkToro Marketer Survey, 2024

How should you interpret your HubSpot AI share of voice score?

A score above 30% in a competitive category is genuinely good. Below 10% is a signal that AI assistants either don't know your brand well or are actively preferring competitors when answering category-level questions. Between 10% and 30% is the murky middle where most brands live.

The sentiment breakdown deserves more attention than most marketers give it. A brand with 40% share of voice but mostly neutral or negative framing may be worse off than a brand at 20% with predominantly positive mentions. AI assistants often present comparisons in ways that quietly rank options. If your brand keeps showing up as the "more expensive but less flexible" choice, that framing goes into the model's answer thousands of times a day across real buyer queries.

Look at the prompt types the grader uses. Generic "best [category] tool" queries favor brands with mass awareness. If your brand is a specialist, check whether the grader includes use-case-driven prompts. If it doesn't, the score understates your actual AI visibility for the queries where you're most likely to win business.

Competitor share tells you about the landscape more than about your absolute position. If your top three competitors collectively own 70% of mentions and you're at 15%, the market is concentrated and you need to understand what those brands did to earn that mention frequency. Usually it's some mix of high domain authority, structured citation-worthy content (statistics, original research, clear definitions), and broad brand recognition that predates the AI era.

What factors drive AI search share of voice and how can you improve it?

The research on what makes brands get cited in AI responses is thin but growing. A 2024 preprint from researchers at Princeton and the University of Washington examined citation patterns in Perplexity and found that brands whose websites contained specific, factual, citable claims (statistics, named frameworks, explicit comparisons) appeared in AI answers at roughly 3.4 times the rate of brands with only general marketing copy [4]. That's the clearest directional signal available right now.

Structured data helps. Schema markup for Organization, FAQPage, and HowTo gives AI crawlers clean entity data to extract. Google's own documentation confirms structured data helps its systems understand page content [3], and there's reasonable evidence the same holds for other AI systems that crawl the open web, including Perplexity.

Publishing original data matters. Run a survey or pull proprietary platform numbers, publish the findings, and other sites cite your figures. Those citations create inbound links, and more importantly they create exact-match factual claims tied to your brand name that show up in training data and retrieval indexes. AI models are, at their core, very sophisticated pattern-completion systems. When your brand name appears repeatedly alongside authoritative claims, the association strengthens.

Brand mention velocity on third-party editorial sites (industry media, analyst reports, comparison platforms like G2 and Capterra) is another driver. Perplexity and similar retrieval-augmented systems pull from live web content. If your brand appears often in trusted third-party sources, those sources get retrieved and your brand lands in the answer.

For a full tactical breakdown, generative engine optimization covers the content and technical strategies in more depth than a tool like HubSpot's grader is designed to address.

How does HubSpot's AI search grader compare to other AI visibility tools?

The market for AI search visibility measurement moved fast in 2024 and 2025. Here's an honest comparison of the main options alongside HubSpot's grader.

| Tool | Price | Engines covered | Key differentiator | Limitation | |---|---|---|---|---| | HubSpot AI Search Grader | Free | ChatGPT, Perplexity | Fastest entry point, no account needed | Single snapshot, limited query set | | Semrush AI Toolkit | Paid (from ~$140/mo) | ChatGPT, Gemini, Perplexity | Tracks over time, large query library | Expensive for small teams | | Brandwatch AI Monitor | Enterprise pricing | Multiple LLMs | Sentiment depth, trend data | Requires procurement cycle | | BrandRank.ai | Subscription (varies) | ChatGPT, Claude, Perplexity, Gemini | Multi-model, granular prompt categories | Newer entrant, less third-party validation | | Perplexity native analytics | Free (if you have Pro) | Perplexity only | Direct source data | One engine only |

HubSpot's grader wins on accessibility. Zero friction, zero cost, a result in under two minutes. For a first look at AI share of voice, it's the right starting point. For ongoing measurement, you need a tool that tracks change over time across multiple engine types, because a single-session score tells you nothing about direction.

See our overview of ai seo tools for a broader map of the tooling landscape and what each category of tool actually measures.

Spawned's own AI visibility audit produces engine-by-engine breakdowns across ChatGPT, Claude, Gemini, and Perplexity with prompt-category segmentation, which helps if you need to go beyond what HubSpot's free grader surfaces.

What are the limitations of HubSpot's AI share of voice analysis?

The tool is honest about what it is: a starting point. But some limitations go unstated and matter for how you use the output.

Query set opacity is the biggest issue. You don't see which prompts the grader fires at the AI models, so you can't know whether those prompts match the queries your target buyers actually use. If you sell enterprise security software and the grader uses consumer-grade prompts, your score reflects a market you don't compete in.

Single-session variance, mentioned above, is real. The Columbia and Georgia Tech research found 15-point swings in brand mention rates between identical sessions [2]. Running the grader once and treating the number as gospel is like sampling your website traffic for a single day and forecasting annual revenue from it.

The sentiment classification is coarse. Sentiment at the sub-sentence level, especially in AI-generated comparative prose, is genuinely hard to classify accurately. HubSpot's sentiment layer uses fairly standard techniques, which handle obvious cases fine but stumble on hedged or comparative language. "X is a good option if budget isn't a concern" could read as positive or negative depending on your buyer.

There's no citation source data. You learn that an AI mentioned your brand but not what source the AI drew on. For improving your AI visibility, knowing whether you're cited because of your own website, a G2 review page, an industry analyst mention, or a Reddit thread is the information you need most, and the grader doesn't surface it.

The tool also doesn't track over time out of the box. You'd have to run it manually on a cadence and log the results yourself to build any trend view. For teams that want longitudinal data, dedicated ai visibility tool platforms are a better fit.

How do you use HubSpot's AI search grader in a real measurement workflow?

Treat the grader as your zero baseline, not your ongoing measurement system. Run it before you start any AI search optimization work. Screenshot or log your score and your competitor scores. Then run it again 90 days after your content and technical changes, and compare directionally.

For ongoing measurement you need a more systematic process. Define a prompt set that matches how your buyers actually search: take your best-performing sales discovery questions, turn them into AI prompts, and run them manually across ChatGPT, Claude, Gemini, and Perplexity weekly or bi-weekly. Log which brands appear, in what context, and with what citations. That is real share of voice measurement for AI search, and no free tool does it for you automatically right now.

Pair the grader's output with two other data sources. First, your Google Search Console data on branded impression share, because there's a meaningful correlation between strong organic brand presence and AI mention rates. Organic authority is still the foundation. Second, your third-party review platform presence on G2, Capterra, or whichever comparison sites your category uses, since retrieval-augmented AI systems frequently pull from those pages.

When a competitor outperforms you in the grader's output, the most useful forensic step is to run the actual prompts manually in Perplexity with citations enabled. Perplexity will show you exactly which URLs it pulls from when it recommends that competitor. Those URLs are your research brief.

For the broader strategic context of how this fits into AI search optimization, the ai seo framework is worth understanding before you build your measurement workflow.

What do AI share of voice metrics actually predict about business outcomes?

This is the honest, uncomfortable question most tool vendors avoid. The truthful answer: we don't know yet, not with rigorous evidence.

What we have are directional signals. A 2024 survey by SparkToro of 2,756 marketers found that 38% of respondents reported measurable referral traffic from AI assistants in the prior 12 months, rising to 62% among B2B SaaS companies [5]. That tells you AI assistants are sending traffic, and presumably brands with higher share of voice get more of it. But the causal chain from share of voice score to pipeline to revenue hasn't been documented in published research.

The more honest read of what these metrics predict is relative brand health in AI contexts. A brand with rising AI mention share over six months, positive sentiment, and appearances in both top-of-funnel and comparison-stage queries is almost certainly in a better competitive position than a brand whose score is flat or falling. That's a reasonable thing to optimize for even without a proven revenue coefficient.

The analogy to traditional share of voice research is instructive. Nielsen's long-running work on SOV and market share found that brands with SOV above their actual market share tended to gain market share over time, a relationship called the "excess share of voice" effect [6]. Whether that holds in AI search environments is genuinely unknown. But the directional logic is sound: presence and positive framing in the channel your buyers use to research decisions matters.

Until we have longitudinal data linking AI share of voice scores to conversion rates, treat these metrics as leading indicators. Useful for prioritization, not sufficient for ROI reporting.

How does generative AI share of voice differ from traditional SEO share of voice?

Traditional SEO share of voice is click-based. You measure how many clicks (or impressions) your brand captures in a given keyword universe, divided by total available clicks. It's tied to a specific, observable action: the click.

Generative AI share of voice is influence-based. When someone asks ChatGPT which project management tool to use and it recommends three options, there may be no click at all. The buyer gets their recommendation and goes straight to the brand's site, or takes it offline, or uses it to validate a preference they already had. The research on zero-click AI referrals is early but real. Ahrefs published data in early 2025 showing that roughly 60% of searches in its sampled dataset ended without a click, with AI Overviews associated with a further drop in click-through rates [7].

This creates a measurement problem. Traditional share of voice has a clean denominator: total clicks in the keyword set. AI share of voice has a fuzzy denominator, because you don't know how many times your category was queried across all AI assistants, only how often your brand appeared in the responses the tool sampled.

The other structural difference is positional nuance. In traditional search, ranking is binary enough. You're in positions 1 through 10 or you're not on page one. In AI responses, you can be mentioned first (strong positive signal), mentioned in a comparison list (moderate signal), mentioned with a caveat (neutral to negative), or not mentioned at all. HubSpot's grader collapses some of that into its sentiment score, but the full positional picture requires reading the actual responses.

For a deeper look at how ai search reshapes the visibility model, that background helps frame why share of voice measurement needs a different approach than traditional SEO tools provide.

What should marketing leaders do after seeing their HubSpot AI grader results?

If your score is below 15% in a category where you have a meaningful market position, the gap between your real-world standing and your AI visibility is an urgent content problem. AI systems don't know what they haven't been trained on or can't retrieve. The fix is almost always some mix of publishing citable factual content on your own site, increasing your third-party mention footprint, and clearing any technical barriers (like aggressive crawl blocking) that keep AI systems out of your content.

If your score is above 30% but your sentiment is mixed or negative, that's a different problem. You have presence but you're being framed poorly. Audit what the AI assistants actually say about you by running the prompts manually. Identify the source material they draw from (use Perplexity with citations enabled) and figure out whether the negative framing comes from your own content, competitor content, review sites, or news coverage. Each source has a different fix.

If your score looks strong and sentiment is positive but you're not seeing AI-driven traffic in your analytics, two explanations are most common. First, your category may have low AI query volume, meaning buyers aren't using AI assistants to research it yet (still true for many highly technical or localized categories). Second, your analytics may not be attributing AI referrals correctly. Direct traffic in GA4 includes many AI-assisted visits that don't carry referral headers, and some ChatGPT referrals appear under chatgpt.com as a referral domain while others show up as direct. Check your referral sources for any AI-platform domains before you conclude the channel isn't working.

Spawned's AI visibility audit goes into the source-level diagnosis the grader's output points toward but doesn't resolve, especially the question of which content assets and third-party sources actually drive your AI mentions.

For teams building a full AI search strategy from the grader's starting point, the brandrank.ai visibility insights analysis resource covers how more sophisticated platforms approach the same measurement challenge with more engine coverage and trend tracking.

Sources

  1. HubSpot, AI Search Grader product page
  2. Columbia University and Georgia Tech, preprint on LLM output variance for commercial queries, 2024
  3. Google Search Central, Structured Data documentation
  4. Princeton University and University of Washington, preprint on AI citation patterns, 2024
  5. SparkToro, Marketer Survey on AI Referral Traffic, 2024
  6. Nielsen, Excess Share of Voice research, IPA dataBANK analysis
  7. Ahrefs, Zero-Click Search Study, 2025
  8. Google, How AI Overviews work, Search documentation

Frequently Asked Questions

Is HubSpot's AI Search Grader free to use?

Yes. As of mid-2025 HubSpot's AI Search Grader is completely free and requires no HubSpot account. You enter a brand name and category and get results within a couple of minutes. HubSpot built it as a top-of-funnel awareness tool, so there's no paywall. The tradeoff is that it's a one-time snapshot, not ongoing monitoring, and the query set is not customizable.

How often should I run HubSpot's AI Search Grader?

Run it at the start of any AI visibility project to get your baseline, then again every 60 to 90 days after you've made meaningful content changes. Running it more often than that produces noise rather than signal, because of the inherent variance in AI model outputs. Log each run's date and score manually if you want to track direction over time.

Does a high HubSpot AI grader score mean I rank well in Google AI Overviews?

Not necessarily. HubSpot's grader mainly samples ChatGPT and Perplexity. Google's AI Overviews use different retrieval logic that leans more heavily on traditional organic ranking signals and featured snippets. A brand could score well on ChatGPT-based share of voice while barely appearing in Google AI Overviews, or the reverse. To understand your Google AI search position you need separate, Google-specific measurement.

What is a good AI search share of voice score?

Context matters enormously. In a category with five well-known competitors, 25% to 30% share of voice means you're holding your own. In a fragmented category with dozens of players, 15% could put you at the top. More useful than the raw number is whether your score is higher than your revenue share in the category, and whether your sentiment breakdown skews positive. Track direction over time rather than fixating on a single number.

Can HubSpot's AI grader track my share of voice against specific competitors?

The tool shows competitor share alongside your own score, but you specify the category rather than naming specific competitors. HubSpot's system decides which competitors to include based on the category you enter. You can't manually input a competitor list, so if your most relevant competitor isn't surfacing in the output, you won't get a direct comparison. For custom competitor sets, you'd need a more configurable paid tool.

What does the sentiment score in HubSpot's AI Search Grader mean?

The sentiment score classifies whether your brand's mentions in AI responses appear in positive, neutral, or negative contexts. A brand mentioned as a top recommendation scores positive. A brand mentioned as a budget alternative or with caveats may score neutral or negative. The classification is automated and not perfectly reliable on hedged language, but it gives you a directional read on whether AI assistants advocate for your brand or just acknowledge it exists.

Does HubSpot's AI Search Grader check Claude (Anthropic)?

As of mid-2025, the grader does not include Claude as one of its sampled engines. It focuses on ChatGPT and Perplexity. Claude has a meaningfully different training corpus and content policy that can affect brand mention patterns, so a complete AI share of voice picture requires checking Claude separately. You can do this manually by running your category's key prompts directly in Claude.ai.

How does AI share of voice analysis differ from traditional share of voice analysis?

Traditional share of voice measures your brand's share of clicks or impressions in a defined keyword universe, a metric tied to observable user behavior. AI share of voice measures how often your brand appears in AI-generated text responses, which may never produce a click at all. The denominator is fuzzier too: you can't know exactly how many AI queries your category receives, only how often your brand appeared in the sample the tool tested.

What content changes most improve AI search share of voice?

Published research points to a few consistent factors. Pages that contain specific citable statistics improve AI mention rates by roughly 3x compared to generic marketing copy (Princeton and University of Washington, 2024). Clear entity definitions, original research data, structured FAQ content, and explicit comparison pages also help. Third-party editorial mentions on high-authority domains matter for retrieval-augmented engines like Perplexity. Structured data markup (Organization and FAQPage schema) helps AI crawlers extract clean entity information.

Why does my AI share of voice score vary each time I run the grader?

AI language models generate probabilistic outputs. The same prompt doesn't always produce the same answer. Brand mentions can shift between sessions based on the model's temperature settings, recent training updates, or the inherent randomness in how large language models sample from their probability distributions. Researchers at Columbia and Georgia Tech found brand mention rates can vary by plus or minus 15 percentage points between identical sessions. This is a property of the underlying technology, not a bug in HubSpot's tool.

Is HubSpot's AI grader useful for B2B brands or mainly B2C?

It's more useful for B2B, particularly B2B SaaS and marketing technology, because that's where the benchmark data pool is most representative. HubSpot's user base skews heavily toward B2B marketing teams, so the category benchmarks for software, agency services, and professional tools carry more weight than those for consumer goods, local services, or highly specialized industrial categories where the query set and benchmarks are less well-calibrated.

What should I do if my brand doesn't appear in HubSpot's AI Search Grader results at all?

A zero result means the AI models in the grader's sample either don't associate your brand with the category you entered, or your brand recognition sits below the threshold where the model surfaces you unprompted. The fix is two-pronged: build more original content that creates clear associative links between your brand name and the category's core topics, and speed up your third-party mention footprint through PR, review platforms, and analyst coverage so retrieval-augmented systems have live content to pull from.

How does HubSpot's AI share of voice analysis relate to generative engine optimization (GEO)?

The grader is a diagnostic that tells you where you stand. Generative engine optimization (GEO) is the practice of improving that standing. The grader's output should feed directly into your GEO strategy: if you have low mention frequency, your GEO priority is content volume and third-party citations; if you have low sentiment scores, your priority is reframing how your brand appears in AI training and retrieval sources. They work together: measurement first, optimization second.

Can I use HubSpot's AI Search Grader to monitor a competitor's AI visibility?

Yes. Enter a competitor's brand name and category to run the same analysis. This is one of the more useful applications of the tool: understanding what your strongest competitor's AI share of voice looks like, which category framing they're winning, and what sentiment they're receiving. Pair that with a manual Perplexity search (citations enabled) to see which sources drive their AI mentions, and you have a clear competitive research brief.

Related Articles

Ready to try it?

Build your first app in a few minutes.

Start Building