Methodology
How we measure AI visibility, including what we cannot
AI answers are nondeterministic: independent research found that fewer than 1 in 100 identical runs return the same brand list, and that differences under ~5-7 percentage points are statistical noise. Most tools report a single run as if it were a ranking. We think that's astrology. Here is exactly what we do instead.
Multi-run sampling, not single snapshots
Every prompt is run multiple times per engine (1x on the free teaser, 2-3x on tracked brands). Your visibility is reported as an appearance rate, the share of sampled answers that mention you, which independent testing shows is far more stable than any “position” metric.
Confidence intervals, shown on the dashboard
Every appearance rate carries a 95% Wilson confidence interval computed from the actual sample size, displayed next to the number. When your score moves week over week, we tell you whether the change is likely real, meaning the intervals no longer overlap, or still inside the noise. We have not found another tool in this category that publishes an interval alongside its score, and if one does we will say so here rather than keep the claim.
How stable are these numbers? We measure that too
Because we run every prompt multiple times, we can measure the measurement: every multi-run report shows a run-agreement rate (how often identical runs agreed on whether you appeared) and a citation-churn rate (how much the cited sources changed between identical runs). External research finds only ~2.3% of ChatGPT citations survive three reruns (Search Engine Land) and fewer than 1-in-100 identical runs return the same brand list (SparkToro). Some vendors argue sampling once per day is enough; our position is simpler: if a number changes when you re-ask, you deserve to see by how much. A tool that samples once and shows no error bars is reporting that churn as fact.
Buyer-intent prompts, grounded in observed demand
Prompts are generated from your actual product page and phrased the way real buyers ask assistants. When you connect Google Search Console, we anchor them in the real queries your site already gets found for, so observed demand rather than modelled guesses. The prompt set is frozen for tracked brands so your trend measures the world changing, not the prompts changing.
We don't sell prompt volumes
“Prompt volume” numbers are extrapolated from small browser-extension panels covering well under 1% of AI usage, and documented discrepancies run to 245x against real search data, and roughly 91% of ChatGPT queries are asked exactly once. No AI provider publishes prompt data (a fuller debunking). We won't sell you a number nobody can measure. Where demand data matters we use real Google keyword volumes (a directional proxy at topic level) and your own Search Console data, and we label them as exactly that.
We never auto post, anywhere
Reddit removes roughly 100,000 accounts per day and runs LLM classifiers built specifically to catch automated marketing and AI-seeding. Every “auto-reply” tool is betting against that enforcement with your brand's reputation. Our earn workflows produce drafts, compliance checklists, and per subreddit rule warnings. A human reviews everything and posts from their own account, disclosed. If a thread's question isn't genuinely answered by your product, our drafts say so and skip the mention.
What the evidence says actually earns citations
We prioritize actions by published evidence, not vendor folklore: placement in the third-party “best X” lists AI engines cite (43.8% of ChatGPT source pages); classic Google top-10 rankings (~90% of Perplexity/AI Mode citations come from them); branded mentions via digital PR; and participation in the topically exact community threads engines cite (80% of AI-cited Reddit posts have fewer than 20 upvotes, so exactness beats virality). We do not recommend llms.txt (measured: no AI system reads it) or schema markup for AI citations (measured: null to negative).
Known limitations, read this
- We measure via API-grade web search (and Perplexity's API where enabled). Logged-in consumer apps personalize answers with memory and history; your customers' exact answers will differ. No tool can fully reproduce a logged in session, including the ones that claim to.
- Small sample sizes mean wide intervals. A free check (6 prompts, 1 run) is a snapshot with honest error bars, not a verdict.
- AI citation patterns are volatile. Reddit's share of ChatGPT citations once moved from ~60% to ~10% in six weeks. That volatility is precisely why we verify actions against re-measurement instead of promising fixed outcomes.
See it applied to your own brand:
Run the free check