Key takeaways
- LLM citation tracking monitors when AI platforms like ChatGPT, Perplexity, Gemini, and Google AI Overviews name or link your content inside a generated answer, not just whether you rank in ten blue links.
- Citation counts move for reasons that have nothing to do with your content. ChatGPT's average citations per response dropped about 27% almost overnight after the GPT-5.3 rollout in March 2026 and never recovered, which is why single-snapshot manual checks are unreliable.
- The tooling market splits into three tiers: manual/DIY tracking (free but limited), affordable prompt trackers ($29-$250/month), and full GEO platforms that pair monitoring with content generation and crawler analytics.
- Reddit's citation share inside ChatGPT fell from roughly 6.1% in May 2026 to 3.7% in June 2026, and collapsed by about 90% in days back in September 2025 after a licensing change, proof that relying on one channel for AI visibility is risky.
- Product pages, not listicles, are now ChatGPT's most-cited content type (around 33% of citations in July 2026), so tracking needs to look at content-type performance, not just domain-level mentions.
What LLM citation tracking actually measures
When someone asks ChatGPT "what's the best project management tool for a five-person startup," the model doesn't just generate an opinion from memory. In web-search-enabled responses, it runs a handful of searches, retrieves pages, and picks a small set of sources to cite. According to Promptwatch's data on citation volume, ChatGPT cites around 5 sources per response on average, roughly half of what a classic Google results page used to show. Perplexity and Google AI Overviews are more generous, each landing near 10 sources per answer, while Microsoft Copilot is a mess, swinging from under 2 to nearly 17 sources per response within weeks as Microsoft keeps re-architecting its retrieval system.
That's the whole game, in miniature: fewer slots than traditional search, and the slots move for reasons that have nothing to do with anything you did. Citation tracking is the discipline of watching which slots you occupy, which competitors take instead, and why.
It's worth separating two things people often conflate:
- A mention is when an AI names your brand or product without a link. "Tools like Otterly help track this" is a mention.
- A citation is when the AI attributes a claim to your specific URL, usually with a footnote or inline link.
Most tracking tools report both, but citations are the harder currency because they show the model actually pulled information from your page rather than just recalling your brand name from training data.
Why manual checking doesn't work
The honest answer to "can I just ask ChatGPT myself and take notes" is: you can, but the data will lie to you. Two things make manual tracking unreliable.
First, LLM outputs are non-deterministic. Ask the same question twice and you can get different citations both times. A tool needs to run each prompt multiple times per cycle to get a real signal instead of a coin flip.
Second, and this is the part people underestimate, the platforms themselves change their citation behavior overnight. In the week before OpenAI's GPT-5.3 rollout on March 4, 2026, ChatGPT cited about 6.4 sources per search-enabled response. By late March that had settled at 4.7-4.9, a roughly 27% drop, and it hadn't recovered a month later according to Promptwatch's tracking of the citation drop. The drop hit every ChatGPT model variant at once. If you happened to run a manual audit right before that rollout and another one right after, you'd conclude your content got worse. It didn't. The platform just changed how many citation slots exist. Only continuous monitoring across a fixed prompt set lets you tell the difference between "we lost visibility" and "the whole platform cited fewer sources this week."
Reddit is the case study everyone cites, for good reason
If you want one number that explains why citation tracking needs to be ongoing, it's this: Reddit held roughly 10-14% of all ChatGPT citations in early September 2025. Within days of Reddit changing its access and licensing terms, that share collapsed to around 1% and stayed there. It never recovered in that window, per Promptwatch's domain citation data. Then, in a separate and smaller move, Reddit's share fell again from 6.11% in May 2026 to 3.71% in June 2026, a roughly 40% month-over-month decline. Even at its peak Reddit, the single most-cited domain across all of ChatGPT, held under 4% of total citations in June 2026, which tells you something else: AI citation share is far less winner-take-all than classic Google rankings. The long tail of smaller, topic-specific sites earns most citations, not a handful of giants.
Worth noting for anyone building a GEO strategy: leaning on one platform (Reddit, or any single channel) as your primary citation source is a bad bet. It can vanish in days if that platform changes its terms.
What to actually monitor
A good citation tracking setup covers a few distinct layers:
- Coverage across engines. ChatGPT, Perplexity, Gemini, Google AI Overviews, and Google AI Mode each pull from different indexes and weigh signals differently. A tool that only checks one gives you a fragment.
- Prompt sampling depth. How many times does the tool run each prompt per cycle? Once gives you a snapshot; three to five runs gives you a pattern.
- Update frequency. Weekly checks miss things that happened days ago. Daily is the realistic floor given how fast citation share moves.
- Content-type breakdown. Citations aren't evenly distributed across formats. Product pages became ChatGPT's most-cited content type in July 2026 at about 32.8% of daily citations, nearly doubling from 18% back in March, according to Promptwatch's citation-type tracking. Listicles, how-tos, and comparisons all grew too, mostly eating into generic landing-page citations rather than news, which stayed flat around 5%. A similar shift happened in Google AI Overviews, where product pages overtook listicles as the top-cited format in late July 2026.
- Crawler access. If AI crawlers can't reach your pages, none of the above matters. Comparing your own server logs against published crawler traffic mix is a sanity check worth doing.

The DIY route: free tracking methods
Before paying for anything, there are a few things you can set up yourself.
Since June 2025, ChatGPT appends utm_source=chatgpt.com to citation links, which Google Analytics 4 can pick up and segment as a distinct source. The catch is GA4 doesn't bucket this automatically; a ChatGPT click lands in "Referral," "Direct," or "(not set)" by default. You'll need to build a GA4 Exploration with a page-referrer regex filter that catches chatgpt.com, perplexity.ai, gemini.google.com, and similar domains.
Google Search Console can't show AI Overviews citations directly either. The workaround analysts use is cross-referencing keywords likely to trigger AI Overviews against GSC impression and click trends over time, then eyeballing whether a keyword's clicks dropped while impressions held steady, a signal the query moved into an AI Overview box instead of organic results.
Both of these are worth doing even if you eventually pay for a dedicated tool, because they ground the tool's reports in traffic you can independently verify.
Dedicated citation tracking tools compared
Once you're past the DIY stage, the market has three rough tiers. Cheap prompt trackers that check a handful of engines, mid-tier tools that add historical trending and source analysis, and full platforms that pair monitoring with content generation and crawler-level diagnostics.
| Tool | Starting price | Engines covered | Depth beyond monitoring |
|---|---|---|---|
| Otterly.AI | $29/mo (Lite) | 4 base, others as paid add-ons | GEO URL audits, unlimited team seats |
| Peec AI | ~$95/mo (Starter) | 6 default engines | Looker Studio export at higher tiers |
| Profound | Custom (Growth ~$399/mo per third-party reports) | Up to 9 engines | Credit-based agent workflows |
| Ahrefs Brand Radar | $199/mo per platform, $699/mo all-platform | Major engines bundled | Ties into existing Ahrefs SEO data |
| Semrush AI Visibility Toolkit | $99/mo standalone | Bundled with SEO suite | AI Search Checks inside Site Audit |
| Promptwatch | $95/mo (Essential) | 8 platforms including Meta, Grok, DeepSeek | Crawler logs, Content Agents, CMS publishing |
A few specifics worth flagging when you're comparing quotes. Otterly's advertised entry price of $29/month only covers 4 engines; getting Claude, Gemini, and Google AI Mode added on top pushes a solo user closer to $76/month, and at the Premium tier those add-ons alone can run over $700/month. Peec AI includes 6 engines by default across every brand tier and, unusually, gives unlimited team seats even on its cheapest plan, where Profound caps seats at 1 on its entry tier. Ahrefs Brand Radar is the priciest single-purpose option: full coverage (base plan plus the all-platform Brand Radar bundle) runs close to $828/month by one reviewer's calculation.


Where Promptwatch fits
Promptwatch sits in the third tier. It monitors ChatGPT, Gemini, Claude, Perplexity, Grok, DeepSeek, Meta Llama, Microsoft Copilot, and both Google AI surfaces by scraping the actual user-facing interfaces rather than relying only on APIs, which matters because what a user sees in the ChatGPT app can differ from what the API returns. Pricing starts at $95/month for Essential (50 prompts, 6,000 responses monthly, all 8 platforms), scaling to $245/month for Professional and $579/month for Business, which adds state and city-level tracking and Shopping Insights.
What separates it from most of the tools in the table above is that it doesn't stop at telling you your citation share dropped. Its AI crawler logs (Agent Analytics) show exactly which pages ChatGPTBot, ClaudeBot, PerplexityBot, and 400+ other crawlers actually visited and whether they hit errors, which is the diagnostic layer most prompt trackers skip entirely. Pair that with content gap analysis and Content Agents that draft and publish GEO-optimized pages directly to Webflow, Framer, or WordPress, and you get a loop that goes from "here's what's wrong" to "here's a published fix" without leaving the platform.

That's the practical distinction worth remembering: monitoring tools answer "was I cited?" Platforms like Promptwatch try to answer "why wasn't I cited, and what do I do about it?" If you're evaluating a full stack of options, the GEO software directory at bestgeosoftware.com is a reasonable place to browse the rest of the field, and agenticseotools.com covers tools that go further into automated execution.
Building a repeatable tracking workflow
Regardless of which tool you pick, the workflow that actually produces useful data looks roughly like this:
- Build a prompt list that mirrors real buyer questions, not just your brand name. Include category questions ("best X for Y"), comparison questions ("X vs Y"), and problem-first questions ("how do I solve Z").
- Run each prompt across every engine your audience actually uses. If your buyers are on ChatGPT and Google, don't bother tracking Copilot heavily; check your own analytics first to see where AI referral traffic actually originates.
- Check daily, not weekly. Given that a single model rollout can shift citation counts by 27% overnight, weekly checks will regularly misattribute platform-wide noise to your own performance.
- Segment citations by content type. If product pages are eating listicle citations across the board, a page that used to earn citations as a "best of" post might need a product-focused rewrite instead.
- Cross-check against your own server logs. If a crawler shows up heavily in third-party crawler traffic reports but never appears in your logs, something is blocking it, usually a robots.txt rule or an overly aggressive WAF setting.
None of this replaces judgment. A tool can tell you Reddit's citation share halved in a month; it can't tell you whether that matters for your specific category. That's still a human call, and probably the most important one in the whole process.

