Measurement
Evaluate Sentiment and Accuracy in AI Answers
Review AI answers with separate sentiment, recommendation and factual-accuracy labels, worked difficult cases and an explicit reviewer-disagreement process.
By Jia Chen
Published
Sources checked
Evaluate how an AI answer describes your product on separate axes: whether it recommends the product, whether its framing is favorable, and whether its factual claims are correct. A positive answer can invent a feature. A negative answer can accurately explain why the product is unsuitable. Those cases need different work.
This guide provides an annotation rubric, difficult examples and a review process. Its sample product, ExampleSync, is fictional. The example answers and product facts are teaching material, not observations about a real company or customer.
The output should be a claim ledger and a report that preserves disagreement. A single sentiment score is insufficient for deciding whether to correct information, clarify positioning, improve a product limitation or take no action.
Define the entity before judging its reputation
First confirm that the answer refers to your product. Names can overlap with ordinary words, unrelated companies or previous products. Record the expected domain, category and any disambiguating context in the observation.
Use three identity states: correct entity, wrong entity and ambiguous. Do not label a glowing paragraph about an unrelated company as positive sentiment toward yours. An ambiguous case belongs in review until the context is sufficient; the classifier should not resolve it by assuming every name match is relevant.
Also preserve the question. “Is this a good option for a one-person team?” and “Does this support enterprise self-hosting?” can reasonably produce different recommendations about the same product. Without the task, a reviewer may misread an appropriate limitation as an attack on the brand.
Use separate units for framing and facts
Sentiment usually concerns a product mention or a coherent passage. Factual accuracy concerns individual claims. Keep an answer record for context and child records for checkable assertions.
For example:
ExampleSync is a strong option for small teams, supports self-hosting, and may be expensive for high-volume workloads.
This passage includes favorable framing, a deployment claim and an unfavorable cost assessment. Review those independently. The self-hosting claim can be checked against the current contract. “Strong option” is a recommendation judgment. “Expensive” requires context and may be an opinion rather than a false fact.
A platform may define sentiment differently. Peec, for example, describes positive, neutral and negative brand portrayal separately from visibility and position. Treat that as a documented vendor convention, not proof that all sentiment scores use the same rubric. Peec metric definitions
A practical framing rubric
Apply the following rubric to the passage about the correct entity. Retain the exact wording that caused the label.
| Label | Meaning in this rubric | Illustrative language |
|---|---|---|
| Positive | A clear favorable evaluation | A good fit for a small team needing this workflow |
| Negative | A clear unfavorable evaluation | Poor fit if self-hosting is mandatory |
| Mixed | Material favorable and unfavorable evaluations | Easy to start with, but restrictive for this deployment |
| Neutral | Description without a meaningful evaluation | Provides a hosted integration service |
| Unclear | The passage cannot be judged reliably from available context | Fragmented text or unresolved entity/quotation scope |
Do not use neutral as a bucket for everything the grader cannot understand. Neutral is an interpretable description; unclear is insufficient or ambiguous evidence. Keep the two separate when reporting coverage.
Define the unit before assigning a whole-answer label. If one paragraph favors the product and another discourages it for a different requirement, preserve both passage labels. A whole-answer mixed summary can be useful, but it should link back to the actual reasons.
Negation and attribution need attention. “Some users call it unreliable, but the evidence here does not establish that” reports criticism while qualifying it. Do not assign the quoted opinion to the answer author without reading the full passage. Record who made the evaluation and whether the answer endorses it.
Keep recommendation distinct from favorable language
Use a recommendation field with yes, no, conditional and unclear. In this rubric, yes means the answer affirmatively proposes the product for the stated task. Conditional means it proposes the product only if a material condition holds, and the answer has not established that condition.
“This is a respected company” is favorable framing without necessarily recommending its product. “Choose it if managed hosting is acceptable” is a conditional recommendation. “Do not use it when self-hosting is mandatory” is a negative recommendation for that task, even if the answer praises other features.
Before aggregating, define whether your published recommendation rate includes conditional suggestions. One defensible report counts affirmative and conditional recommendations separately. Another may group them but must disclose the rule. Do not change the definition after seeing which one produces a better number.
A separate factual-accuracy rubric
Split factual passages into atomic claims and record the relevant date, plan, version, geography or deployment scope. Then check each claim against appropriate primary evidence.
| Verdict | Meaning | Required record |
|---|---|---|
| Correct within scope | Evidence supports the specific statement and its qualifications | Source, checked date and scope |
| Incorrect | Evidence contradicts a material assertion | Exact contradiction and current/historical context |
| Incomplete or misleading | A material condition is omitted or wording overstates the evidence | Missing qualification and its consequence |
| Unresolved | Available evidence is insufficient or conflicting | What was checked and what is still missing |
| Not a factual claim | Evaluation, preference or other statement without a defined factual test | Reason for exclusion from factual metrics |
An unsupported source attachment and an incorrect factual claim are different findings. A page may not support a statement that is nevertheless true according to another source. Conversely, a cited page may support a statement that is outdated. Preserve source support and factual truth in different columns.
The source-tracing guide provides the source-side method. Here, the factual verdict should depend on the best appropriate evidence, not merely whether an answer contains a link.
Work through eight difficult cases
Assume the following fictional facts about ExampleSync on the review date: it is a hosted service; Python and TypeScript clients are available; self-hosting is unavailable; SSO is available only on the Enterprise plan; public terms do not establish which competing product is fastest. These facts exist only for this exercise.
| Invented answer passage | Framing | Recommendation | Factual assessment |
|---|---|---|---|
| Choose ExampleSync for your self-hosted deployment; it is an excellent fit | Positive | Yes | Incorrect: the fictional product does not offer self-hosting |
| ExampleSync is unsuitable when self-hosting is mandatory | Negative | No | Correct within the stated requirement |
| ExampleSync offers Python and TypeScript clients | Neutral | No | Correct under the supplied facts |
| It is a good option if managed hosting is acceptable | Positive | Conditional | Evaluative fit claim; hosting condition agrees with supplied facts |
| Easy to adopt, but it is expensive | Mixed | Unclear | Evaluative assertions need a defined task/cost comparison; not established false |
| All ExampleSync plans include SSO | Neutral | No | Incorrect: fictional scope is Enterprise only |
| ExampleSync supports SSO | Neutral | No | Incomplete if the answer is evaluating a non-Enterprise plan; record the question context |
| ExampleSync is the fastest service available | Positive | Unclear | Unresolved comparative performance claim; no benchmark is supplied |
The first case is the most important trap: positive sentiment can hide a harmful recommendation. The second is another common trap: a negative answer can be accurate and useful. Correcting it would require changing the product or the buying requirement, not asking a publisher to erase a legitimate limitation.
The seventh case shows why qualifiers matter. In a general capabilities overview, “supports SSO” may be acceptable if the plan scope is explained nearby. In an answer claiming a low-tier plan satisfies an SSO requirement, it can be materially misleading. Preserve the surrounding answer rather than grading a detached phrase mechanically.
The final claim could become testable if “fastest” is defined by workload, version, environment and metric. Until then, do not convert a missing benchmark into proof that the claim is false or true.
Build a ledger that can drive a decision
Keep one row per factual claim and link it to the containing passage. Include a separate record for the passage's framing and recommendation label.
Observation and passage ID:
Question and intended task:
Entity verdict:
Exact passage:
Framing label and rationale:
Recommendation label and condition:
Claim ID and exact factual wording:
Normalized claim, date and scope:
Attached source support:
Primary factual evidence and checked date:
Accuracy verdict and rationale:
Reviewer, rubric version and unresolved questions:
Action owner, if action is justified:
A useful rationale is specific: “The answer says all plans; the primary plan matrix limits SSO to Enterprise.” An unhelpful rationale repeats the label: “Incorrect because it is wrong.” Require enough explanation that another reviewer can verify the judgment without guessing.
Semrush's brand-misinformation guide illustrates inspecting answers and their sources before pursuing corrections. The relevant lesson is the evidence-backed workflow, not an assumption that every unfavorable answer is misinformation. Semrush correction guide
Calibrate reviewers before scaling annotation
Give two reviewers the same rubric and a small independent sample. Include ordinary cases and difficult cases: accurate limitations, mixed framing, missing qualifiers, outdated facts, entity confusion and uncited claims. They should record initial labels before discussing them.
Then compare labels separately for entity, framing, recommendation and factual verdict. A team can agree strongly about tone while disagreeing about product truth. One combined agreement percentage hides that difference.
As an illustrative calculation, suppose two reviewers label eight passages. They agree on six framing labels and disagree on two. Exact framing agreement is 6/8 = 75%. That is a description of this tiny exercise, not a quality threshold or proof that the rubric is reliable across all answers.
For every disagreement, record its type:
| Disagreement type | Appropriate resolution |
|---|---|
| Different evidence | Provide the same primary material and review again |
| Different interpretation of a label | Clarify the rubric with a boundary example |
| Different claim scope | Restore the question, plan, version or missing passage |
| Genuine unresolved fact | Keep unresolved; assign further research |
| Clerical error | Correct the label while retaining the review history |
Do not merely tell the second reviewer which answer is preferred. The aim is a better rubric and evidence packet. If clarification changes a rule, version it and revisit affected records, including those initially labeled by agreement.
After calibration, retain a fresh sample that was not used to tune the instructions. Otherwise you risk demonstrating that the reviewers learned those particular examples rather than checking whether the method transfers.
Use automated grading as a reviewed assistant
Automation can propose labels and extract claims, but evaluate its output against a human-reviewed sample. Ask it to return the exact passage, selected label, supporting evidence identifier and a short rationale. A free-floating sentiment score is difficult to debug.
Anthropic's evaluation guidance discusses tasks, graders and calibration of model-based judging. The transferable principle is to test the grader against meaningful outcomes and review disagreements. It does not prescribe a universal GEO annotation sample size or certify another model's sentiment labels. Anthropic evaluation guidance
Keep a failure log. Check whether the grader systematically turns missing evidence into falsehood, mistakes a quoted opinion for the answer's position, merges two products, or rewards flattering statements. Those are actionable grading defects.
When the rubric, model or prompt changes, re-evaluate the held-out sample and document the change. A smooth dashboard trend can conceal a changed classifier. If a label cannot be reproduced from preserved evidence, exclude it from strong conclusions until reviewed.
Report tone, accuracy and review coverage together
A useful report includes the number of correct-entity passages, framing distribution, affirmative and conditional recommendations, factual claims reviewed and unresolved cases. Show counts with percentages and preserve the denominator for each.
For example, a fictional review could find twelve incorrect claims among eighty factual claims with determinate verdicts, with ten additional claims unresolved. Report both 12/80 incorrect among reviewed determinate claims and the ten unresolved cases, if that is the denominator rule you chose. Do not imply the entire answer inventory was checked when only a sample was reviewed.
The term determinate must match the data: it includes claims actually judged correct, incorrect or materially incomplete under your stated rule, and excludes unresolved and non-factual statements. If incomplete claims are included in an error numerator, say so. The measurement guide explains why silent denominator changes distort trends.
Prioritize by consequence as well as count. One invented security capability can matter more than several mildly unfavorable adjectives. A frequently repeated truthful limitation calls for positioning clarity or a product decision, not a factual correction campaign.
Choose the next action from the verdict
Use four practical routes:
- Incorrect owned fact: correct the source and dependent pages, preserving dates and scope.
- Incorrect third-party fact: prepare a narrow evidence-backed correction packet for review.
- Accurate unfavorable description: consider a better-fitting audience, clearer limitation or actual product improvement.
- Unresolved claim: research further or preserve uncertainty; do not manufacture a confident rebuttal.
A positive but incorrect claim belongs in the first two routes too. Reliability means correcting flattering inaccuracies as seriously as unflattering ones. The brand-correction procedure takes established errors through action and follow-up.
Should mixed sentiment count as negative?
Not by default in this rubric. Mixed preserves material favorable and unfavorable evaluations. You may create a separate unfavorable-framing view, but state its grouping rule and keep the original passage labels available.
Can a reviewer label a claim false without a citation?
An answer's lack of a citation does not prove falsehood. The reviewer needs appropriate contradictory evidence. If no adequate evidence is available, use unresolved rather than incorrect.
What is the most useful quality check?
Pick a surprising score and trace it back to the original question, passage, factual evidence and rubric. If another person cannot reconstruct the judgment, improve the record before optimizing the score.
Sources and checked date
References checked September 17, 2026. ExampleSync statements, annotation tables and agreement calculations are original fictional teaching examples. They are not claims about customers or measured model performance.
- Peec metric definitions: one vendor's definitions of visibility, position and sentiment.
- Semrush: correcting AI brand misinformation, dated May 8, 2026: instructional precedent for inspecting claims and sources before correction.
- Anthropic: evaluation guidance, published January 9, 2026: evaluation design and grader calibration, not a GEO-specific accuracy benchmark.
Continue reading
Explore GEO with Jam
See how Jam approaches AI visibility research and content improvements for developer-tool teams.
Explore Jam for GEO