METHOD

A method for testingwhat AI actually recommends.

PersonaSignal starts with structured shopper contexts, then changes selected test conditions—brand context, product information, model, repeated run, or Search / Grounding—to compare observable recommendation behavior.

OBSERVABLE BOUNDARYWe study answers and recommendation behavior we can observe. We do not claim access to hidden reasoning, model weights, or proprietary ranking systems.

01

RESEARCH OBJECT

Knowing a brand exists
is not the same as recommending it.

A model may describe a named brand when asked directly. That does not tell us whether the same brand will appear when a shopper asks for a solution without naming it.

OBSERVED STATE ABRAND RECOGNITION

AI can identify the supplied brand, product, or website information.

OBSERVED STATE BRECOMMENDATION

AI selects a brand or product for a defined shopper need.

SHARED STATUS LANGUAGE

These labels describe observable outcomes, not permanent brand scores.

Mentioned
The answer names the brand but does not select it as a clear recommendation.
Not Selected
Another option is selected, or the target is not chosen within the defined result set.
Unknown / Insufficient Evidence
The answer does not support a reliable classification under the recorded rules.

A status records what happened in an answer. It does not automatically explain why it happened.

02

PERSONA FIRST

Define who's asking.
Then observe what AI recommends.

We do not begin with a loose audience keyword. We build a structured shopper context that connects the person, task, full need, screening criteria, and relevant product information.

  1. 01

    PERSON

    Who is deciding?

    Life stage, experience, household, or other details relevant to the purchase.
  2. 02

    CONTEXT / TASK

    What are they trying to do?

    Market, climate, use environment, timing, and the practical task.
  3. 03

    FULL NEED

    What must the choice deliver?

    Goal, use, budget, preferences, must-haves, and exclusions.
  4. 04

    SCREENING CRITERIA

    What shapes the shortlist?

    Explicit conditions used to compare or rule out options.
  5. 05

    PRODUCT MATCHING

    What fits those criteria?

    How public product information aligns with the stated need.
  6. 06

    AI RECOMMENDATION

    What does the model select?

    The recommendation, mention, non-selection, or alternative observed.

TEST-DESIGN FRAMEWORK

This is an external test-design framework—not a claim about the model's hidden reasoning process. Structured scenarios help us compare answers; they do not represent every real shopper.

03

THREE TEST CONDITIONS

Three conditions.
Three different research questions.

The conditions are related, but they do not measure the same thing. Brand Understanding is not automatically “better performance” than a Blind / Natural result.

A

BLIND / NATURAL

Can the brand enter naturally?

The target brand is not supplied. The shopper asks from the need itself.

OBSERVE

Whether the brand enters the answer or recommendation set, and which alternatives appear.

B

BRAND UNDERSTANDING

How is the brand understood?

Brand information is supplied so the model can respond to the stated positioning and offer.

OBSERVE

Positioning, fit, constraints, uncertainty, and information gaps in the description.

C

PRODUCT FIT

Where does the product fit?

Relevant product context is supplied, then compared across defined shopper needs.

OBSERVE

Which situations fit, which do not, and which alternatives may be selected.

04

CONTROLLED COMPARISON

Hold the shopper stable.
Change selected information.

Controlled comparison helps isolate observable differences between test conditions. It does not prove why a model changed its answer.

HELD CONSTANT

SAME PERSONA+SAME PURCHASE NEED
A

NO BRAND PROMPT

Blind / Natural

Observe natural brand entry, recommendations, and alternatives.
B

BRAND INFORMATION

Brand Understanding

Observe positioning, fit, gaps, constraints, and uncertainty.
C

PRODUCT CONTEXT

Product Fit

Observe fit judgments, recommendation states, and alternatives.

COMPARE OBSERVED OUTPUTS

Compare brand entry, description, recommendation status, and alternatives to identify differences worth testing further.

05

REPEATED RUNS

One answer
can be noisy.

We repeat recorded test conditions to see which patterns persist and which results move around.

SAME TEST CONDITION

PERSONA P-07 · BLIND / NATURAL · SAME MODEL
ILLUSTRATIVE RUN RECORDS
  1. RUN 01Recommended

    Target Brand is explicitly recommended.

    ALTERNATIVE · Brand A
  2. RUN 02Mentioned

    Target Brand appears but is not clearly selected.

    ALTERNATIVE · Brand B
  3. RUN 03Recommended

    Target Brand enters the recommendation again.

    ALTERNATIVE · Brand A
  4. RUN 04Not Selected

    Another brand is selected instead.

    ALTERNATIVE · Brand C

OBSERVE

  • States that recur
  • Variation between runs
  • Alternative brands
  • Answers that remain unclear

INTERPRETATION BOUNDARY

Repeated runs support robustness and pattern comparison. They do not, on their own, establish statistical significance or represent the wider market.

06

MODEL COMPARISON

The same question
can surface different answers.

With the shopper, need, and test mode recorded, we can compare observable outputs across GPT, Gemini, or other selected models.

HELD CONSTANT

SAME PERSONA+SAME NEED+SAME TEST MODE
MODEL A

GPT

ILLUSTRATIVE OUTPUT

RESULT
MAIN ALTERNATIVE
Brand A
DESCRIPTION NOTE
The answer emphasizes ingredient fit and screening criteria.
MODEL B

Gemini

ILLUSTRATIVE OUTPUT

RESULT
Mentioned
MAIN ALTERNATIVE
Brand B
DESCRIPTION NOTE
The answer emphasizes use context and public product information.

COMPAREObserve model-specific differences in recommendation state, alternatives, and description. This is not a provider ranking.

The research question and agreed Scope determine whether one model or a model comparison is useful. Differences are not reduced to an “intelligence level.”

07

SEARCH / GROUNDING

When search is available,
sources become observable too.

Some projects examine what changes when a model can retrieve public information. We can record whether a source or brand appears and compare grounded with non-grounded output.

SAME PERSONA + SAME NEED

ILLUSTRATIVE COMPARISON
A

NON-GROUNDED

Without public retrieval

RECOMMENDATION
Record recommendation state
MENTION
Record whether the brand appears
ALTERNATIVE
Record the selected alternative
SOURCE OBSERVATION
Not available in this condition
B

SEARCH / GROUNDED

With public retrieval

RECOMMENDATION
Record recommendation state
MENTION
Record whether the brand appears
ALTERNATIVE
Record the selected alternative
SOURCE OBSERVATION
Record public sources that appear

SOURCE TYPES OBSERVED

  • 01Official
  • 02Retail
  • 03Editorial
  • 04Community

COMPARE OBSERVED OUTPUTS

Compare recommendation, mention, alternatives, and source observations across recorded conditions.

SOURCE BOUNDARY

A source appearing in an answer does not prove that it caused the recommendation. The research does not claim access to hidden source weights.

08

HUMAN REVIEW & TRACEABILITY

Automation can scale runs.
Human review makes the record defensible.

Each finding is mapped back to the shopper context, prompt condition, model, run record, recommendation status, and available source observation.

  1. 01

    RAW RUN

    Model answer

  2. 02

    STRUCTURED RECORD

    Recorded fields

  3. 03

    CLASSIFICATION

    Status label

  4. 04

    SOURCE / CONTEXT CHECK

    Evidence check

  5. 05

    HUMAN REVIEW

    Reviewed record

  6. 06

    REPORT FINDING

    Written finding

TRACEABLE STRUCTURED RECORD

Every finding has a route back to the test.

These illustrative fields use the same logic as the Sample Report's Raw Data structure.

RUN IDRUN-024
PERSONA IDP-07
TEST MODEBLIND / NATURAL
MODELGPT
GROUNDING STATENON-GROUNDED
RECOMMENDATION RESULTMENTIONED
MAIN ALTERNATIVEBRAND A
SOURCE RECORDNOT APPLICABLE
REVIEW STATUSREVIEWED

HUMAN-REVIEW BOUNDARY

Human review checks classification, context, and evidence records. It does not turn judgment into model certainty or decide what AI should answer.

09

EVIDENCE BOUNDARIES

What the research supports.
What it does not prove.

A useful research record is explicit about its boundaries. Findings are tied to the tested model, version, time, shopper context, prompt, and condition.

SUPPORTED BY THIS RESEARCH

Reviewable observations

Recommendation behavior
Whether the target was Recommended, Mentioned, Not Selected, or could not be classified.
Recurring patterns
Which outcomes persist or move across repeated recorded runs.
Competitor substitution
Which alternatives appear when the target is not selected.
Model differences
How observable outputs differ across selected models and conditions.
Source observation
Which public sources appear when Search / Grounding is available.
Questions worth reviewing
Public information that appears missing, ambiguous, inconsistent, or worth testing.

NOT AUTOMATICALLY PROVEN

Claims outside the evidence

Market representativeness
Structured shopper scenarios are not a statistically representative consumer sample by default.
Hidden reasoning
Observable outputs do not reveal internal chain of thought, weights, or ranking logic.
Causal attribution
Controlled differences do not prove why an output changed.
Source effect
A source appearing does not prove that it caused a recommendation.
Future behavior
No test guarantees how a model will recommend later.
Business outcomes
No result guarantees ranking, traffic, recommendation, sales, or an optimization effect.

METHOD SUMMARY

Not one AI answer.
A record you can compare and review.

PersonaSignal turns recommendation questions into structured evidence through shopper contexts, controlled conditions, repeated runs, model comparison, source observation, and human review.

The record remains bounded by the Persona, model, test condition, time, and public information available.