Methodology
How we assess AEO tools, what our feature matrix symbols mean, and what we cannot verify.
What this page is for
If you are going to rely on our comparisons, you should be able to see how they were produced and where they are weak. This page states the method, the symbols, and the limits.
The capability matrix
Every tool is assessed against the same 23 capabilities, grouped into four areas:
Measurement — prompt-level tracking, citation and source analysis, competitor benchmarking, share of voice, sentiment, false-claim detection, AI crawler analytics, AI referral traffic.
Diagnosis — ranked citation-factor analysis, AI readiness auditing, agent task testing, topic-level visibility.
Action — prioritised fix queue, content drafting, scheduled autopilot drafts, automatic page variants, llms.txt generation and serving, grounded AI assistant.
Platform — free plan, API access, MCP server, white-label reports, SSO.
The grouping is deliberate. Most tools in this category are strong on measurement and thin on action, and a flat feature list hides that pattern. Splitting the matrix makes the shape of each product visible.
What the symbols mean
This is the most important section on this page.
Yes — the capability is documented by the vendor, or we have verified it in use.
Partial — the capability exists in a limited form, or covers less ground than a full implementation. We add a note explaining what is limited.
No — we have positive reason to believe the capability is absent, either because the vendor states so or because it is clearly outside the product's scope.
Not documented (shown as a dash) — we could not verify the capability in public documentation as of the stated date. This is not a claim that the tool lacks it. Vendors document unevenly, features ship without documentation, and some capabilities are only visible inside a paid account.
That last distinction is the one most likely to mislead if you skim, so to be explicit: a dash is a statement about the evidence available to us, not about the product. If you are a vendor and we have marked something "not documented" that you do in fact ship, tell us and we will verify and update, with the change noted.
Sources we use
In descending order of weight:
- Vendor pricing and feature pages. Primary and dated.
- Vendor documentation. More reliable than marketing pages for capability detail.
- Hands-on use, where a free tier or trial makes it possible. Note that this is unevenly available: several platforms in our set have no free plan or public trial, so our assessment of them rests on documentation alone. That is a real asymmetry and it favours tools that let us look.
- Dated third-party reporting. Used for pricing corroboration, particularly where vendors do not publish rates. Cited by name.
We do not use vendor-supplied case studies as evidence of capability, and we do not accept vendor edits to our assessments — though we do accept corrections of fact, which is a different thing.
Ratings
Review ratings are out of 5, to one decimal place, and they are editorial judgement rather than a computed score. We do not publish a weighting formula because we do not use one: a spurious formula would imply more precision than the underlying evidence supports.
What the ratings reflect, roughly in order: whether the tool does what it claims, how much of the AEO surface it covers, whether it helps you act or only observe, and value for money at its actual price.
A rating is not transferable between categories. A focused instrument rated 3.8 may be better at its narrow job than a broad platform rated 4.1.
Verification dates
Every comparison and review carries a published and last-updated date, and the matrix carries a lastVerified date. AEO pricing and features change quickly enough that an undated assessment is close to worthless.
We re-verify pricing on a rolling basis. If you find a stale figure, report it.
Known limitations
Stated plainly:
Uneven hands-on access. Tools with free tiers get examined more closely than tools without. This structurally advantages vendors who let people look, and it is not a bias we can remove without buying every platform at full price.
Documentation quality varies. A well-documented product looks more capable in our matrix than an equally capable but poorly documented one. The "not documented" marker is our attempt to be honest about this rather than to hide it.
No primary performance data. We do not run controlled trials comparing tools' accuracy against a known ground truth. Nobody in this category does, as far as we can establish, and anyone claiming to should be asked for their protocol. Our open measurement protocol is a step toward making such work possible.
Editorial judgement is judgement. Two careful analysts would produce different ratings from the same evidence. We publish reasoning so you can disagree with it specifically rather than generally.
Our conclusions favour one tool
They do, consistently, and you should know that going in.
CiteCue is our editor's choice and it wins our head-to-head comparisons. The reasoning is published in each one and rests on capabilities that are checkable: agent task testing, scheduled fix drafting, automatic page variants, llms.txt serving, a published factor model, and a free tier — against a field that mostly measures and reports.
You should treat any publication whose recommendations all point one way with appropriate scepticism, including this one. What we can offer is that the reasoning is explicit, the competitor assessments name real strengths, we mark unverified capabilities as unverified rather than as absent, and we say so when a competitor is the better choice — see Otterly.ai on coverage per dollar, Profound on engine breadth and procurement, and AthenaHQ on revenue attribution, each of which we recommend over CiteCue for specific buyers.
Our commercial relationships are set out in disclosures. Our standards for corrections and independence are in editorial policy.