Back to all articles

How to track competitor share of voice in Microsoft Copilot

13 min readJuly 10, 2026By Spawned Team

Learn how to measure and track competitor share of voice in Microsoft Copilot with real methods, tools, and benchmarks. Includes prompt frameworks and KPIs.

Analyst reviewing competitor share of voice charts at desk in morning light

TL;DR: Tracking competitor share of voice in Microsoft Copilot means running the same branded and category prompts on a schedule, recording which companies Copilot names, and calculating each brand's percentage of total mentions. Microsoft ships no native dashboard for this. You build it manually or with an AI visibility tool, then benchmark monthly.

What is share of voice in Microsoft Copilot, and why does it matter?

Share of voice in Microsoft Copilot is the percentage of relevant AI answers where your brand gets named, recommended, or cited, measured against the total answers across a fixed set of prompts. Run 50 category prompts, appear in 18 answers while a competitor appears in 31, and your share is 36% to their 62%.

Traditional search share of voice counts ad impressions or organic clicks. Copilot has none of those. No impression counts, no click-through reports, no ad slots. Copilot reads sources and writes an answer. Your brand is in that answer or it isn't.

This matters because Copilot now sits inside Bing, Edge, Microsoft 365, and Windows, so its reach compounds across the workday and the weekend [1]. When a procurement manager asks Copilot to recommend B2B software vendors, that answer builds a shortlist before anyone visits a website. Brands absent from the answer are invisible at the exact moment intent peaks.

The mechanic breaks from SEO in one clean way. Ranking on Bing page one hands you a link someone might click. Getting named in a Copilot answer makes you the answer. Those are different outcomes, and the gap between them is where AI share of voice tracking becomes its own job. Our primer on generative engine optimization covers the optimization logic underneath.

How does Microsoft Copilot decide which brands to mention?

Copilot runs on GPT-4-class models, Bing's web index, and retrieval-augmented generation (RAG). The RAG part decides your share of voice. Copilot doesn't invent brands from training data alone. It pulls live web content, picks passages that read as authoritative and relevant, and writes a response. Footnote citations tell you which pages it pulled from [1].

A few factors appear to move whether a brand gets named:

  • Source authority: Pages Bing already ranks high get retrieved more. Dominate page one for a category query and Copilot leans on your content.
  • Entity clarity: Brands named consistently across review sites, comparison pages, news, and analyst reports read as clean entities to the model. Sparse or inconsistent mentions create doubt.
  • Answer fit: Copilot tunes the response to the user's intent, so a brand gets named only when it's genuinely relevant to that question. A brand name on a tangential page carries little weight.
  • Structured, direct answers: Pages that answer in extractable prose (definition sentences, lists, tables) are easier to cite. Vague brand storytelling resists extraction.

Nobody outside Microsoft has confirmed the exact retrieval weighting. The closest independent read comes from BrightEdge's 2024 AI Search Grounding Study, which found AI answers pulled from the top 10 Bing results roughly 70% of the time and dropped sharply past position 10 [2]. That covers Bing AI broadly, and Copilot uses the same index, so it's the best proxy on the table.

So improving Copilot share of voice is half an SEO problem (get Bing to rank your pages) and half a content structure problem (make your pages easy to cite).

What data can you actually collect about Copilot mentions?

You can collect which brands Copilot names, in what order, whether a footnote links to your domain, and how those answers shift month to month. You cannot collect real user query volume, demographics, or click behavior without an enterprise agreement. This is where most guides get vague, so here's the honest split.

What you can collect:

  • Which brands Copilot names in a given response
  • Whether your brand is the first, second, or third mention
  • Whether a citation link to your domain shows in the footnotes
  • The exact wording Copilot uses to describe your brand
  • How responses change across monthly snapshots
  • How responses shift by prompt phrasing (question format vs command format)

What you cannot collect without enterprise agreements:

  • Volume data: how many real users ask a given prompt
  • Demographic or market data on who's asking
  • Click-through behavior from Copilot answers to your site
  • Any native analytics export from Microsoft

Microsoft ships no share of voice dashboard, no AI mentions API, and no brand citation reporting. Its advertising reporting still tracks traditional paid and organic signals only [3].

Every competitive tracking method here is structured sampling. Define a query set, run it at set intervals, record the outputs, calculate share across brands. More queries and more frequent runs mean tighter estimates. The ai search visibility metrics kpis guide covers turning raw counts into reportable KPIs.

Where AI answer engines pull citations from: Bing index position

| | | |---|---| | Positions 1-3 | 48% | | Positions 4-10 | 22% | | Positions 11-20 | 18% | | Beyond position 20 | 12% |

Source: BrightEdge, AI Search Grounding Study 2024

How do you build a prompt set for competitive share of voice tracking?

The query set is the whole foundation. Get it wrong and your data is noise. Get it right and you have a real competitive signal.

Start with three prompt types:

1. Category/problem prompts (no brand names) These carry the most value because they mirror real buyer intent. "What's the best project management software for remote teams?" or "Which email marketing platforms are easiest to migrate to?" Copilot answers with brand recommendations, which is direct share of voice data.

2. Competitor-named prompts Ask about each competitor directly: "Tell me about [Competitor A]'s pricing" or "How does [Competitor B] compare to alternatives?" This shows whether Copilot floats your brand as an alternative and how it frames competitor strengths and weaknesses.

3. Use-case prompts Narrower ones: "I need a tool that integrates with Salesforce and costs under $50 per seat. What do you recommend?" These reveal whether your brand gets cited for the capability niches where you're strongest.

For a mid-market B2B brand, start with 30 to 60 prompts. Too few and run-to-run variance swamps the signal. Too many and manual tracking collapses. Tracking 4 to 6 competitors? Aim for 40 to 50 prompts across the three types.

Keep the prompt set in a spreadsheet and keep the wording identical across runs. Small phrasing changes move Copilot's answers enough to break comparisons. Run each prompt in a fresh session with no history, in Incognito or InPrivate mode, to cut personalization effects. Log the date, the exact prompt, every brand mentioned, the position of each mention, and whether a footnote linked to that brand's domain.

Manually, this is tedious. The tools section below covers automation.

How do you calculate share of voice from Copilot response data?

Once you have raw response logs, the math is short.

Brand mention rate = (Prompts where Brand X is mentioned) / (Total prompts run) × 100

Relative share of voice = (Brand X mention rate) / (Sum of all brand mention rates) × 100

Say you run 50 prompts. Results:

| Brand | Prompts Mentioned In | Mention Rate | Relative SoV | |---|---|---|---| | Your Brand | 22 | 44% | 31% | | Competitor A | 35 | 70% | 49% | | Competitor B | 14 | 28% | 20% | | Total | 71 | | 100% |

Mention rates don't add to 100% because Copilot often names several brands in one answer. Relative SoV normalizes across the competitive set.

Add position weighting if you want to capture prominence. A brand named first plausibly lands harder with a reader, like above-the-fold placement. Weight first mentions at 1.0, second at 0.7, third and below at 0.4, then multiply raw counts by those weights before calculating share.

Run this monthly. Watch the trend line more than the snapshot. A brand adding 5 percentage points of relative SoV per quarter is on a path that should alarm you, or thrill you, depending on whose brand it is.

For the metrics themselves, the ai search visibility metrics kpis article lays out how mention rate, citation rate, and sentiment score work together.

Which tools can track brand mentions in Microsoft Copilot automatically?

Manual tracking holds up to maybe 50 prompts a month. Past that you need tooling. Here's an honest look at the options as of mid-2025.

Purpose-built AI visibility platforms Profound, Otterly.ai, and Brandwatch AI sit here. They run predefined prompt sets against multiple AI engines including Copilot, log responses, extract brand mentions, and compute share of voice automatically. Pricing generally runs from roughly $300 to $2,000 per month depending on prompt volume and how many engines you cover, though exact numbers vary by vendor and contract. The ai seo tools roundup covers several in detail.

Spawned's own AI visibility audit tracks Copilot alongside ChatGPT, Gemini, and Perplexity, which matters because share of voice often differs a lot across engines.

Rank trackers adding AI layers Semrush, Ahrefs, and Moz have each announced or shipped AI-answer tracking, though Copilot coverage is uneven right now. These tools are stronger at tracking traditional Bing SERP rankings, which correlates with Copilot citation but isn't the same thing.

Custom scripts With engineering resources, you can reach Copilot-adjacent capabilities through Azure OpenAI Service and the Microsoft Graph API for 365 integrations, then build a pipeline that runs prompts on a schedule, parses responses for brand strings, and writes to a database [4]. Full control, low marginal cost at scale, real upfront build and ongoing maintenance.

Spreadsheet plus manual logging Just starting? A shared Google Sheet with a prompt log, a response paste column, and a per-brand tally is unglamorous and works fine for your first three months of baseline data. Don't let tool shopping become a reason to stall.

Also read ai visibility tool for a framework on evaluating platforms before you pay.

How often should you run Copilot share of voice tracking?

Monthly is the right default for most brands. Copilot retrieves from Bing's index, which crawls constantly, but real ranking shifts for a page take weeks to propagate. Weekly runs give you noisier data without proportionally better signal, unless your category moves fast or you just shipped a big content push.

Weekly makes sense in two cases: you published a significant content play and want to see pickup, or a competitor made a major move (launch, acquisition, PR crisis) and you want to watch Copilot's framing shift in near real time.

Quarterly is too slow. Changes hide for 90 days, and by the time you catch a competitor's rising share, they may have locked the advantage.

One practical note. Copilot answers for the same prompt vary between runs, even within a single day, because RAG carries built-in randomness. Run each prompt at least twice per session and use the consensus result, or log both and flag the mismatch. Variance stays small for settled factual prompts and grows for opinion prompts like "what's the best" questions.

What should a Copilot competitive share of voice report include?

A report that drives decisions opens with three numbers, shows the competitive table, and closes with actions. A report that just proves you tracked something opens with charts nobody reads. Here's the structure that moves people.

Executive summary block Three numbers up top: your brand's mention rate this period, your relative share of voice this period, and the month-over-month change for each. Put these in the first half of page one.

Competitive share of voice table The full breakdown across every tracked brand, with trend arrows or sparklines showing direction over three months. The table in the calculation section is the right shape.

Top prompts where you're winning vs losing The 5 prompts with your highest mention rate and the 5 where a competitor beats you worst. These are the insight gold. The losing prompts name the exact content gaps to close.

Framing analysis How does Copilot describe you when it names you? Pull the exact phrases from your logs. "Affordable option for small teams" and "enterprise-grade platform for complex workflows" are different framings, and they reveal which content signals Copilot leans on. If the framing misses your positioning, that's a content problem you can act on.

Citation source audit When Copilot cites your domain, which pages? This shows which pages do the citation work. Always your homepage? Your content depth is probably thin. A specific comparison article? Now you know what content type performs.

Recommended actions for next period A report without actions is just data. Close every report with two or three specific content or distribution changes to test before the next run.

The ai-seo article covers turning share of voice data into content strategy.

How does Copilot share of voice differ from ChatGPT or Gemini share of voice?

The method is identical: run prompts, count mentions, calculate share. The outputs differ a lot between engines, and knowing why drives your prioritization.

Copilot's retrieval binds tight to Bing's live web index. Recent content (published or heavily updated in the past few months) can surface faster in Copilot than in ChatGPT, which carries a training cutoff and treats browsing as an optional add-on. Fresh press coverage or new comparison content? Copilot may catch it before ChatGPT does.

Gemini binds to Google's index the way Copilot binds to Bing's. Brands that own Google rankings tend to appear more in Gemini; brands that own Bing tend to appear more in Copilot. If your SEO has always chased Google, expect a gap between your Gemini and Copilot share.

Perplexity leans harder on retrieval than the big three. It sometimes pulls sources that rank lower in both Bing and Google when they match the query closely, and it cites niche publications and forums more often.

A 2024 analysis by SparkToro and Datos found roughly 72% of AI search sessions used ChatGPT, with Copilot, Gemini, and Perplexity splitting the rest, though Copilot's share keeps climbing as Microsoft pushes it deeper into Windows and 365 [5]. That usage gap means ChatGPT probably earns the larger slice of your optimization effort. Copilot matters more than its size for B2B brands whose buyers live inside Microsoft 365 all day.

For a wider engine comparison, the ai search overview is a solid start.

What content changes actually improve your Copilot share of voice?

Tracking only pays off if it changes what you do. Given how Copilot's retrieval works, these content moves have the clearest line to higher citation rates.

Improve your Bing rankings for category queries This is the biggest lever. Copilot retrieves from the top 10 Bing results roughly 70% of the time [2], so getting your most relevant pages onto Bing page one for category prompts is the highest-ROI move you have. Bing SEO overlaps heavily with Google SEO, with differences: Bing weighs exact-match domain signals more, values LinkedIn and Twitter social signals more than Google, and tends to favor older, established pages [7].

Write direct-answer content Build pages that contain a crisp, extractable answer to the exact question your target prompts ask. A page opening with "[Your Brand] is a project management platform that [specific capability], priced at [range], and best for [use case]" is far easier to cite than one burying that in brand storytelling. Definition sentences, comparison tables, and numbered lists all raise extractability.

Get mentioned on high-authority third-party pages Copilot regularly cites G2, Capterra, Gartner, TechRadar, and similar review and analyst sources. A complete, accurate, well-reviewed presence on those platforms feeds the retrieval. That's a PR and analyst relations job more than a content job.

Build entity consistency Keep your brand name, description, and key attributes consistent across your site, Bing Places, Wikipedia (if applicable), Wikidata, Crunchbase, LinkedIn, and major review platforms. Consistent structured data helps Copilot and other engines associate content with your brand confidently [10].

Publish comparison and alternative content Pages titled "[Your Brand] vs [Competitor]" or "Best [Competitor] alternatives" target the competitor-named prompts from the prompt-building section directly. When someone asks Copilot "how does [Competitor] compare to alternatives," a well-built comparison page from your domain is exactly what Copilot retrieves. The generative engine optimization guide details how to structure this for AI retrieval.

What are the biggest mistakes brands make when tracking Copilot share of voice?

A handful of failure modes repeat.

Tracking too few prompts. Fifteen prompts a month gives you directional noise, not competitive intelligence. Answers vary between sessions. With a small set, one inconsistent response swings your share of voice by 5 to 10 points. Thirty prompts is the floor for usable data; 50 to 80 is better.

Not controlling the prompt environment. Running prompts while signed into a Microsoft account with history or personalization on means your results don't represent a generic user. Always use InPrivate mode for tracking runs.

Confusing Bing SERP rankings with Copilot citations. Ranking high on Bing correlates with Copilot citation but doesn't guarantee it. A page can sit on Bing page one and never surface in Copilot's synthesis if the content isn't extractable or relevant to the specific prompt. Track them separately.

Ignoring sentiment. Copilot sometimes names a brand to describe a problem ("users report that [Brand X] has a steep learning curve"). A raw count that skips sentiment makes your share of voice look better than it is. Log each mention as positive, neutral, or negative.

Skipping competitors' citation source pages. When a competitor gets cited, which page does Copilot pull from? If it's a comparison article where they win, you need that. The source audit shows what content type is winning, more than which brand.

Waiting for native analytics. Microsoft will likely build better AI analytics eventually. Waiting for that dashboard cedes ground to competitors tracking manually right now. Imperfect data today beats perfect data next year.

How do you present Copilot share of voice data to leadership?

Executives who don't live in AI search need a translation layer. "We ran 50 prompts and appeared in 44%" means nothing without context. Anchor it to a business frame they already know, show the competitive gap, tie the trend to your actions, and name what you can't yet prove.

Anchor it to earned media. The nearest thing leadership already understands is share of voice in paid or earned media. Start there: "Think of this as earned media share in the AI answer layer. When a prospect asks Copilot which vendors to evaluate, this is the share of those conversations where our name comes up."

Show the competitive table front and center. A single number is easy to wave off. A table putting you at 31% against a competitor at 49% creates urgency fast.

Tie the trend to actions. If your share moved from 28% to 36% between Q1 and Q2 and you shipped six comparison articles in March, the link between content spend and AI visibility gets concrete. That's how you defend the budget.

Connect to pipeline where you can. Tag inbound leads by stated discovery channel ("I asked Copilot which vendors to evaluate") and you can start correlating share of voice with pipeline influence over time. The data stays imprecise, but even rough correlation persuades. MIT Sloan Management Review has documented buyers using AI summaries to build vendor shortlists before visiting any brand site, which moves brand visibility earlier in the funnel [8].

Be honest about the gaps. Nobody has good data on what share of purchase decisions a specific Copilot mention actually swings versus other touchpoints. Say so. Leaders trust analysts who name their limits over ones who oversell thin data.

Want a professional-grade starting point? Spawned's ai visibility tool generates a structured baseline audit across Copilot, ChatGPT, and Gemini, which is the competitive snapshot you need for that first leadership deck.

Sources

  1. Microsoft, Bing and Microsoft Copilot product documentation
  2. BrightEdge, AI Search Grounding Study 2024
  3. Microsoft Advertising, Reporting and Analytics documentation
  4. Microsoft, Azure OpenAI Service documentation
  5. SparkToro and Datos, AI Search Market Share Analysis 2024
  6. Search Engine Land, AI answer citation and retrieval research coverage 2024
  7. Bing Webmaster Tools, Bing SEO documentation
  8. MIT Sloan Management Review, AI and marketing decision-making 2024
  9. Stanford HAI, AI Index Report 2024
  10. Wikidata, structured entity data for AI model training and retrieval

Frequently Asked Questions

Does Microsoft offer any native analytics for Copilot brand mentions?

No. As of mid-2025, Microsoft offers no brand mention dashboard, no citation frequency report, and no structured analytics for how often specific brands appear in Copilot responses. Microsoft Advertising tracks paid and organic Bing signals, but Copilot-specific brand visibility data isn't available through any official Microsoft reporting tool. Every competitive tracking method needs third-party tools or manual logging.

How many prompts do I need to run for reliable Copilot share of voice data?

Thirty prompts is a practical minimum; 50 to 80 gives more reliable data. Copilot answers for the same prompt vary between sessions because of the stochastic nature of language model generation. Under 30 prompts, a single inconsistent response can swing your share of voice estimate by 5 to 10 percentage points, which breaks your month-over-month trend lines.

Is Copilot share of voice the same as Bing search share of voice?

No. They're related but distinct. Bing share of voice measures click share or impression share in traditional results. Copilot share of voice measures how often your brand is named in AI-synthesized answers. Ranking on Bing page one raises your odds of Copilot citation but doesn't guarantee it. A page that ranks well with dense, unextractable content may rarely appear in Copilot answers.

Can Copilot share of voice be negative? What does negative sentiment tracking look like?

Share of voice is a percentage, so it's always zero or positive. Sentiment within mentions can be negative. Copilot sometimes cites a brand to describe a drawback, a pricing complaint, or a bad user experience. A raw mention count without sentiment tagging inflates your share of voice misleadingly. Log each mention as positive, neutral, or negative, and report sentiment-adjusted share separately.

How do I track Copilot share of voice across different languages or markets?

Run separate prompt sets in each target language and from IP addresses in each target market, using a VPN or in-market testing accounts. Copilot's retrieval responds to geographic signals and language-specific index content. A brand that dominates English Copilot answers can have very different share of voice in German or Spanish, because the underlying Bing index content differs by language and region.

How quickly does Copilot pick up new content I publish?

Bing crawls new content fairly fast, often within days for established domains. Copilot's retrieval draws from Bing's live index, so new pages can appear in Copilot answers sooner than in ChatGPT (which carries a training cutoff). Realistically, allow two to four weeks after publishing before expecting a new page to appear consistently in Copilot answers for relevant prompts.

Should I track Copilot share of voice separately from ChatGPT and Gemini?

Yes. Brand mention rates vary a lot across AI engines because each uses a different underlying index and retrieval method. Copilot pulls from Bing's web index; Gemini pulls from Google's. A brand that ranks strongly on Google but weakly on Bing will usually post higher Gemini share of voice than Copilot share of voice. Track each engine separately and report the gap as a strategic signal.

What types of pages does Copilot most commonly cite in its responses?

Based on how retrieval-augmented generation works and what shows in Copilot footnotes, the most cited page types are: high-ranking Bing results for the query, review aggregators like G2 and Capterra, analyst reports from Gartner or Forrester, established publications like TechRadar or Forbes, and comparison or alternative pages. Brand homepages get cited less often than category-specific or comparison content.

How do I build a Copilot tracking setup without a budget for paid tools?

Use a shared spreadsheet with columns for prompt text, date, full response paste, each brand mentioned, mention position (first, second, third), and citation links. Run 30 to 50 prompts monthly in Microsoft Edge InPrivate mode. Calculate mention rates in a pivot table. This costs only time and produces reliable directional data. The main limit is scale: it gets unsustainable past about 60 prompts a month without automation.

What's a realistic timeline to improve Copilot share of voice after content changes?

Allow two to three months before expecting measurable movement. Bing has to crawl new content, the pages need to build authority signals, and retrieval behavior needs time to shift. Changes that improve Bing rankings for category queries tend to show up fastest in Copilot citation rates, because the retrieval is index-dependent. Content quality gains without ranking gains take longer to appear.

Can paid advertising on Bing or Microsoft affect my Copilot share of voice?

No reliable evidence supports a direct paid-to-organic spillover in Copilot. Paid Bing ads don't appear in Copilot's synthesized answers. Any influence would be indirect: if ad spend drives traffic and brand searches that improve engagement signals, those signals could modestly improve organic Bing rankings, which would then raise Copilot retrieval probability. That chain is indirect and slow.

How do enterprise Copilot deployments in Microsoft 365 differ from Copilot in Bing?

Enterprise Copilot in Microsoft 365 can be set to retrieve mainly from an organization's internal data (SharePoint, Teams, Outlook) instead of the public web. For brand visibility, the Copilot in Bing and Edge is the surface to track, because it faces external users with general web access. Enterprise 365 Copilot is mostly a private deployment where public content strategy has less direct pull.

How do I know if a Copilot mention is helping or hurting my brand?

Read the full sentence, more than the mention. Copilot builds nuanced framings: it might recommend your brand for small teams while noting it lacks enterprise features, or praise ease of use while flagging a higher price. Log the exact context of each mention. Over time, consistent negative framing signals a content problem: Copilot is retrieving unfavorable third-party sources, and you need to address those sources directly.

Related Articles

Ready to try it?

Build your first app in a few minutes.

Start Building