Measurement

Compare Visibility Across AI Answer Platforms

Compare AI answer surfaces without mixing UI and API behavior, search settings, missing data or citation formats into misleading visibility rankings.

By Mohammad Alshaikhusain

Published

Sources checked

Compare AI visibility across platforms by keeping the question set and labeling rules consistent while documenting the differences you cannot control. Name the actual surface, collection method, date, locale, search configuration and available citation evidence. Report each surface's valid-answer results beside its collection failures.

The objective is a fair comparison of observed behavior under known conditions. It is not to make different systems identical or to infer that an API result reproduces every user's consumer-chat experience. A provider's name alone is not a complete measurement specification.

This guide includes a collection manifest and a fictional comparison showing how missing data can reverse the apparent story. It does not report a new benchmark of named providers.

Name the surface before the brand

A consumer chat product, a search-enabled API, a search-engine AI feature and a coding-agent session are different environments. They may expose different controls, evidence and context. Treat each as a distinct surface in the inventory.

For example, Google's documentation describes AI Overviews and AI Mode as Google Search features, while its Gemini API documentation separately describes a configurable Google Search grounding tool. A Gemini API observation should not be labeled an AI Overview merely because both involve Google technology. Google Search AI features, Gemini API grounding

Use a descriptive record such as “provider, API, search tool enabled, model identifier, date” or “consumer interface, logged-out state, displayed mode, date.” If the model identifier is not exposed, mark it unavailable. Do not infer it from a marketing label or an old collection run.

Coding-agent implementation belongs in a separate study with an environment and success criteria. Mention frequency in a chat answer does not establish whether an agent can install or use a package.

Distinguish UI and API collection

UI observation can reflect the behavior of a particular interface state. API collection can provide explicit settings and structured metadata. Neither automatically represents all users or all contexts.

Peec's documentation describes its use of UI scraping rather than API collection. Treat that as a disclosed vendor method, not proof that every logged-in, personalized or regional user sees the same answers. When comparing any two datasets, inspect the actual collection method rather than relying on their common provider label.

For UI runs, record what is observable: mode, account state where appropriate, locale, language and whether prior conversation context is present. Avoid collecting or publishing personal account information unnecessarily. Use a consistent, authorized setup.

For API runs, preserve the request configuration and relevant response metadata without exposing credentials. If a system prompt, retrieval filter or supplied product context differs between runs, it can change the task. Do not describe the outputs as directly comparable without acknowledging that difference.

Freeze the common question inventory

Use the same exact question text for the common comparison group. Keep named-brand diagnostics separate from unbranded questions. Record a version identifier so later wording changes do not silently alter the measurement.

Some surface-specific questions may still be worthwhile. For example, a coding workflow and a general buying question need different tasks. Report those as separate studies instead of forcing an artificial shared score.

When a surface cannot handle a question under the intended setup, record what happened. A refusal, tool failure, truncated response and valid omission are not the same result. Define the status policy before collecting so surprising outcomes are not discarded selectively.

The question-selection guide helps establish the inventory. The comparison cannot be stronger than the relevance and stability of that starting set.

Record search configuration and actual search evidence

Search being available does not always mean a search occurred. Google's grounding guide describes enabling a search tool and the model deciding whether search can improve the answer. Preserve the returned search and citation metadata where available rather than inferring search use solely from the request configuration. Google grounding documentation

Likewise, domain filters, search limits and localization can change the available evidence. Claude's web-search documentation exposes settings for domain controls, search limits and approximate user location. Record such settings when used; do not compare a domain-restricted run with an unrestricted one as if their retrieval opportunities were identical. Claude web-search documentation

If a UI does not expose whether or how retrieval occurred, mark that field unknown. A lack of visible citations does not establish that no external information influenced the response. Keep the report grounded in the evidence the surface actually provides.

Do not force every system to cite your domain in a discovery comparison. Supplying an allowlist containing only your site changes the question from independent selection to constrained retrieval. Such a test can be useful, but it needs a different name and interpretation.

Use a collection manifest

The manifest records comparability and exceptions before results are summarized:

FieldWhat to record
Surface IDProvider, product/interface and collection method
Question versionExact text and unbranded/seeded group
Model detailsRequested and returned identifiers where exposed
Date and repetitionCollection time, timezone and trial number
Language and localeRequested settings and observable interface state
Search configurationEnabled tool, limits, filters and relevant version
ContextSystem instructions, supplied documents and conversation history policy
Output limitsRelevant truncation limits and completion status
EvidenceRaw answer, citation metadata and available search events
StatusValid answer, refusal, truncation, tool failure or collection error

Preserve provider-specific fields in the raw record. Normalize a small common view for reporting, but do not throw away information simply because another provider lacks an equivalent field.

Document the collector version as well. A parser change can alter citation counts without any change in the answer system. If collection or extraction code changes, check representative stored responses and record the effect before joining the new data to an old trend.

Normalize citation meaning, not just URL strings

Citation formats can differ in their relationship to answer text. Some identify text spans; others accompany content blocks or lists. Preserve the original attachment and create a common record containing the source URL, associated answer context and granularity.

The current Gemini grounding guide illustrates inline URL annotations attached to text segments. Claude's web-search documentation describes citation records that include the source URL, title and cited text. These are examples of provider-specific structures; a collector should follow the API version actually used rather than assuming one permanent schema. Gemini grounding response, Claude citation fields

For a common answer-level owned-citation rate, count whether an answer contains at least one qualifying owned URL. Keep the exact-page analysis separate. Do not let a verbose citation format create an artificial advantage by counting repeated references as separate answers.

An attachment does not prove a source supports every sentence nearby. Use claim-level review for important cases and retain uncertainty where the metadata is broad. The source-tracing guide explains that assessment.

Inspect errors inside successful responses

HTTP success alone may not establish that the intended search or answer task completed. Claude's documentation explicitly describes web-search errors returned within an otherwise successful API response. Check tool results and completion state, not only the transport status. Claude search error behavior

Apply the same general collection discipline to every surface: verify the answer is complete enough for the rubric, the expected operation actually ran where required, and the response is not a partial stream or an error message mistaken for content.

Preserve retries as attempts. If a retry supplies the intended observation, connect it to the original scheduled cell while retaining the failure history. This lets the report distinguish answer behavior from collection reliability.

Do not replace a failed observation with a different question merely to fill a table. The missingness may be associated with a surface or task category and deserves investigation in its own right.

A fictional comparison with unequal completion

Assume two fictional surfaces each receive 20 unbranded scheduled questions. Surface A returns 18 valid answers and recommends the brand in six. Surface B returns 12 valid answers and recommends it in five.

MeasureSurface ASurface B
Scheduled attempts2020
Valid answers1812
Failed/not-assessable attempts28
Answers recommending the brand65
Recommendation rate among valid answers33.3%41.7%
Collection completion90%60%

It would be misleading to conclude simply that B is better for the brand. B has the higher conditional recommendation rate, but a much larger missing portion. You do not know how its unavailable answers would have behaved.

For a matched comparison, inspect questions with valid observations on both surfaces and report how that subset was selected. Suppose there are ten shared valid cases, with four recommendations on A and three on B. The matched rates are then 40% and 30%. Those values answer a narrower question and should not erase the full collection table.

This invented example does not prove a general bias direction. It shows why the denominator and common subset matter. Report both views and investigate concentrated failures before making a business decision.

Keep question groups visible

Aggregate rates can hide differences in task mix. One surface may recommend the product on general purchasing questions while another cites its documentation on implementation questions. Those are different roles in the buyer journey.

Report by the task strata you defined before collection. Avoid creating a new favorable segment after seeing the results and presenting it as the primary test. Exploratory slices can be valuable if labeled as such.

If you publish a combined index, state its weights and why they were chosen. Equal weight per surface and equal weight per valid answer are different aggregations when completion varies. Neither necessarily reflects how your buyers distribute their actual usage.

Use the visibility measurement guide to keep mentions, recommendations, citations and claim accuracy separate throughout the comparison.

Repeat and annotate changes

Repeat important questions to observe variation, while remembering that repeated runs of one prompt are not independent buyer needs. Keep the setup stable where possible and preserve each attempt.

Anthropic's evaluation guide distinguishes tasks from trials and emphasizes evaluation evidence. It supports the discipline of repeated, inspectable runs, not a universal GEO sample size or a claim that a small difference is statistically meaningful.

Annotate changes in model, interface, search tool, locale, collector or question set. A provider change can create a break in the measurement series. Keep the old and new configurations separate long enough to understand whether a trend is still comparable.

When evaluating your own content change, record these external changes alongside the intervention. The before/after measurement guide helps distinguish a useful observed improvement from an unsupported causal claim.

Turn the comparison into a decision

A useful report names the surfaces where important questions are answered accurately, where the brand is omitted, where citations are useful and where collection or interpretation remains uncertain. It then identifies the next investigation or improvement.

For example, a recurring wrong compatibility claim on one surface may justify inspecting its attached sources. A documentation citation on another may reveal a useful technical resource without implying a product recommendation. A low completion rate may require fixing collection before changing content.

Do not optimize the site around a single surprising answer. Look for recurring, relevant patterns and inspect the evidence behind them. The comparison should help allocate work, not create a universal ranking of models from a small private sample.

Can API tests stand in for consumer chat?

They can measure the configured API surface and support repeatable research. They do not automatically reproduce consumer-interface behavior. Label the surface precisely and use separate observations when the consumer experience is the decision of interest.

Should we discard answers without citations?

Not automatically. They can still provide mention, recommendation and accuracy evidence. Citation absence should remain a distinct observation. If the study specifically requires a completed search-grounded answer, define that eligibility rule before collection and report excluded cases.

Can we rank platforms from one run per question?

You can describe the observed sample, but avoid broader certainty. Variation, task mix, missingness and setup differences can affect the result. Use repeated observations and explicit limits when the decision warrants them.

Sources and review date

Sources checked September 17, 2026. Tables use fictional observations for arithmetic; no new provider benchmark was run for this article.

Explore GEO with Jam

See how Jam approaches AI visibility research and content improvements for developer-tool teams.

Explore Jam for GEO