Decision guide
Turn AI visibility measurements into decisions
Choose the right evidence, compare consistent windows, and decide when to investigate, change content, or keep observing.
9 minute read · Last reviewed: July 20, 2026
Begin with a decision, not a dashboard
Before collecting a metric, write the decision it could change. “Improve AI visibility” is too broad. A decision-ready question names the audience, topic, model or surface, time window, outcome, and action under consideration.
Example decision question
Should we update the existing comparison guide for US buyers this month, based on repeated non-mentions in the comparison-prompt cohort and observed Google demand for the same topic?
This question can be answered with inspectable evidence. It also allows “keep observing” to be a legitimate outcome.
Use an evidence ladder
1. Observation
One or more prompt runs show a mention, non-mention, rank, or citation. Useful for diagnosis; weak for generalization.
2. Repeated pattern
Comparable runs show a persistent direction across time, prompts, or models. Stronger than one answer, but still descriptive.
3. Associated movement
AI visibility and another metric move together in a documented window. This is correlation, not proof that one caused the other.
4. Incremental evidence
A comparison group, holdout, staggered rollout, or other design makes alternative explanations less likely. This can support a bounded causal claim.
5. Verified business outcome
A defined qualified event is recorded with appropriate attribution and data-quality checks. Publication also requires consent and privacy review.
Do not use level-one evidence to write a level-five claim. Most optimization work can begin with repeated descriptive evidence; public outcome claims need more.
Define the baseline before the change
A defensible baseline records:
- the hypothesis and primary decision metric;
- prompt cohort, models, language, location, filters, and run count;
- comparison dates and why the windows are comparable;
- target queries, pages, and GSC property scope;
- analytics dimensions, attribution scope, and exact key-event definition;
- known campaigns, releases, PR, seasonality, outages, or monitoring changes; and
- the threshold or evidence rule that will trigger an action.
Decision matrix
| Observed pattern | Next check | Responsible action |
|---|---|---|
| One surprising run | Answer, entity match, citations, comparable repeats | Investigate; do not generalize. |
| Repeated non-mentions on relevant prompts | Prompt validity, competitor evidence, current page coverage | Fix a diagnosed audience or evidence gap. |
| Aggregate moves, one model drives it | Per-model answers and provider change dates | Report the model-specific movement. |
| Visibility rises, qualified outcomes do not | Attribution, landing page, intent, event setup | Optimize the outcome path or revise the hypothesis. |
| No eligible runs | Date range, filters, job health | Collect or restore data; do not report zero. |
Use analytics dimensions that match the question
For visits during a measured period, inspect session-scoped source and medium. “First user source” describes the source that originally acquired the person; it can attribute a later direct, search, or AI-referral session to an older source.
Define what a qualified outcome means before reporting it. A generic event count is not automatically a lead or sale. Document event name, deduplication, currency, attribution scope, consent limitations, internal-traffic exclusions, and whether revenue is observed or modeled.
Use the GSC and AI visibility workflow to align search and prompt evidence without collapsing them.
Standards for publishing an outcome
Before publishing a study or customer result, require:
- preserved source data and a reproducible calculation;
- metric definitions, scope, dates, exclusions, and limitations;
- a clear distinction between observed, estimated, and inferred values;
- written permission for the customer name, logo, quote, data, and review copy;
- privacy review and aggregation or redaction of sensitive information;
- disclosure of Peekaboo’s commercial interest; and
- a correction owner and update date.
These requirements follow Peekaboo’s editorial and corrections policy. If the evidence cannot support the headline, narrow the headline rather than stretching the data.
Frequently asked questions
What is a good AI Visibility Score?
There is no universal score that is good across every market and prompt set. Use a stable, relevant prompt cohort; compare the same entities under the same conditions; and interpret the score with run count, model breakdown, answers, citations, and business context.
Does a higher score prove that my content strategy worked?
No. It establishes that the measured answers changed under the selected conditions. Normal output variation, model changes, competitors, external coverage, and measurement changes may also explain the movement. Stronger causal claims require a design that rules out plausible alternatives.
Which GA4 attribution dimension should I use?
Use session-scoped source and medium when evaluating the acquisition source of visits in the measured period. First-user source describes how the user was originally acquired and can mislabel a later session. State the exact dimension, attribution scope, consent limits, and event definition in every report.
When should I stop observing and make a change?
Act when the gap is repeated across enough comparable observations for the cost and reversibility of the decision, the underlying answers support a diagnosis, and the proposed action helps the intended audience even if the visibility score does not move. Higher-risk claims and expensive changes require stronger evidence.