Prompt sensitivity

Same question, reworded: up to 5 different brands reached #1

Prompt Sensitivity Report
Prompt Sensitivity Report
Retail · athletic footwear & sportswearPrompt sensitivity series · 01
750
Responses
4
Topics
40
Prompts
4
Providers
Section 01

Research question and sample

Research question

When the same underlying topic is expressed through slightly different wordings, does the set of brands an agent cites stay the same, and does their order stay the same?

Design

The dataset covers 4 topics. Each topic is expressed through 10 near-identical paraphrases meant to elicit the same answer. Each paraphrase was issued to 4 AI providers across multiple runs, yielding 750 responses. Because the 10 paraphrases per topic are treated as the same intent, any difference across them is attributed to prompt sensitivity.

Sample scope · confirmed from the data
4 topics × 10 paraphrases40 prompts
Providers4
Responses750

No figures beyond those directly supported by the data are used.

Method and definitions
Parent-brand rollup

Sub-brands folded into parent (Jordan → Nike; Offline by Aerie → Aerie).

Noise removal

Pure retailers excluded (Dick’s, REI, Target). Decathlon retained but flagged as borderline.

Membership stability — “who is cited”

Mean Jaccard overlap between the brand sets of two paraphrases, across all 45 pairs per topic. 1.0 = identical set.

Rank stability — “the order”

Each paraphrase’s responses are aggregated into one ranking via reciprocal-rank points (a brand at position p earns 1/p per response). Measured by distinct #1 brands, top-3 set overlap, and mean Spearman rank correlation.

Rephrasing effect — section 05

The drop in response-level overlap from same-wording pairs (model noise floor) to different-wording pairs, controlled within provider.

Section 02 · question 1

Does the cited brand set change between paraphrases?

For each topic: distinct brands seen, the stable core (brands present in all 10 paraphrases), and mean / worst-pair membership overlap.

Mean overlapWorst-pair overlap
Jaccard · 0 → 1
Running shoes
7 core / 16 distinct
0.75 / 0.57
Casual sneakers
11 core / 24 distinct
0.73 / 0.58
Gym & fitness
9 core / 45 distinct
0.60 / 0.43
Sportswear value
12 core / 44 distinct
0.61 / 0.41
Figure 1. Mean brand-set overlap per topic (bars) vs the worst single pair of paraphrases (dots).
Reading

Membership is largely stable. Every topic keeps a solid core (7–12 brands) present in all 10 paraphrases, and mean overlap is high (0.60–0.75). Instability lives at the edges: narrow topics barely move; broad topics surface long tails of one-off brands that flip in and out with wording.

Answer to Q1

Yes, slightly. Who is cited shifts at the periphery when phrasing changes, but the central recommended set stays anchored. The narrower the topic, the more stable the set.

Section 03 · question 2

Does the order of the brands change?

Order is far less stable than membership. Distinct #1 brands across paraphrases, mean top-3 set overlap, and mean rank correlation.

Top-3 overlapRank correlationWorst pair
Distinct #1 brands →
Running shoes
0.62 · 0.90
Casual sneakers
0.90 · 0.76
Gym & fitness
0.54 · 0.54
Sportswear value
0.27 · 0.61
Figures 2–3. Order stability per topic (top-3 set overlap and Spearman rank correlation, with the worst pair marked) and the number of distinct brands that ever reach #1 across the 10 paraphrases. More filled dots = a more contested, wording-sensitive leader.
Reading

The picture splits by leader strength. In topics dominated by a single brand, order holds — that brand keeps #1 across all or nearly all paraphrases and rank volatility is close to zero. In fragmented topics with no dominant brand, order is fragile: the #1 slot rotates among several brands (up to 5 different leaders in the most contested topic) and top-3 overlap falls as low as 0.27, so the leading trio is almost fully reshuffled by wording alone. A high top-3 overlap can still hide an unstable leader — the same few brands may recur while which one leads keeps changing.

Answer to Q2

Yes, and more than membership does. Rephrasing rarely changes who is in the pool but routinely changes who leads — except where a single brand dominates the category, where even the order stays stable.

Section 04

Breakdown by provider

Prompt sensitivity is not uniform across providers. Each metric is averaged over the four topics.

MembershipTop-3 overlapRank corr.
Higher = steadier · avg distinct #1 →
OpenAI
0.65 · 0.47 · 0.57
2.75
Gemini
0.63 · 0.51 · 0.53
2.75
Perplexity
0.61 · 0.62 · 0.59
2.50
Google AI Mode
0.60 · 0.51 · 0.47
3.00
Figure 4. Provider stability across three metrics (higher = steadier under rephrasing). Bars scaled 0 → 0.8.
Provider × topic
Deeper purple = more stable setGreen = order preserved · red = order reshuffled
Running shoes
Casual sneakers
Gym & fitness
Sportswear value
OpenAI
0.72
0.70
0.62
0.54
0.69
0.52
0.75
0.32
Gemini
0.72
0.69
0.52
0.58
0.66
0.54
0.48
0.44
Perplexity
0.78
0.62
0.51
0.54
0.67
0.71
0.47
0.51
Google AI Mode
0.78
0.65
0.49
0.47
0.69
0.53
0.44
0.21
Figures 5–6. Membership overlap and rank correlation by provider × topic. Every provider is weakest on the broad value topic.
Reading

Perplexity is the most order-stable provider (rank correlation 0.59, top-3 overlap 0.62, fewest distinct leaders). OpenAI is the most membership-stable (0.65) with solid order stability (0.57). Gemini sits close behind, and Google AI Mode is the least stable of the four on rank correlation (0.47) and produces the most distinct leaders. The differences are real but moderate: all four providers cluster within about 0.05 on membership and 0.12 on rank correlation.

Cross-cutting pattern

The same topic effect appears inside every provider: all four are most stable on the narrow running-shoes topic and least stable on the broad value topic. Topic breadth dominates provider choice — the spread across topics within any single provider is larger than the spread across providers within any single topic.

Section 05

How much of the change is really rephrasing?

A key check: agent outputs vary run-to-run even with identical wording. To isolate the rephrasing effect, we compare the overlap of two responses to the same wording (the model’s noise floor) against two responses to different wordings, controlled within provider.

Same wording — noise floorDifferent wording
Response-level overlap · 0.30 → 0.70
Running shoes
0.66 → 0.62 +0.04
Casual sneakers
0.61 → 0.52 +0.09
Gym & fitness
0.57 → 0.48 +0.09
Sportswear value
0.48 → 0.38 +0.09
Figure 7. Same-wording overlap (noise floor) vs different-wording overlap. The small gap between the two is the true rephrasing effect.
Reading

Most of the observed brand-set variance is the model’s own stochasticity, not rephrasing. Even with identical wording, two runs overlap only 0.48–0.66. Different wording drops this by just ~0.04–0.09 more. This qualifies the working assumption that all variance is prompt sensitivity: the honest rephrasing effect is the drop (baseline − treatment), which is modest — roughly 8 points of Jaccard on average.

~8 pts
Average Jaccard drop attributable to rephrasing
Section 06

Conclusions

01
Membership is sticky; order is not

Rephrasing rarely changes which brands appear (mean overlap 0.60–0.75, stable core in every topic) but routinely changes their ranking, especially the top-3 and the #1 slot.

02
Topic breadth is the main driver

Narrow topics with a clear category leader are highly stable in both membership and order. Broad, fragmented topics are where wording bites — the most contested topic had as many as 5 different brands reach #1 across paraphrases.

03
A dominant brand anchors the order

When one brand clearly owns a category it stays #1 regardless of phrasing. Order instability is a symptom of a contested top of the ranking, not of rephrasing per se.

04
Provider differences are modest

Perplexity and OpenAI are the steadiest on rank, Google AI Mode the most volatile, yet all four cluster within a narrow band (about 0.12 on rank correlation), so provider choice nudges stability at the margin rather than driving it.

05
Much of the variance is model noise

Same-wording runs already disagree substantially; the marginal rephrasing effect is modest once that noise floor is accounted for.

06 · practical implication
Report the leader as a distribution

A single-phrasing measurement of brand rank is unreliable, especially for broad topics on less stable providers. Robust brand-visibility tracking should average across several paraphrases and multiple runs, and report the leader as a distribution, not a single winner.

All figures derive solely from the uploaded dataset (750 responses; 4 topics × 10 paraphrases; 4 providers). Cleaning: parent-brand rollup; pure retailers removed; Decathlon flagged. These are stated in Section 01 and can be adjusted.

Appendix · FAQ

Applying this to AI visibility work

Questions that come up when teams take these results into prompt coverage, citation tracking and content planning. Every answer below is bounded by this dataset.

Is this the same as tracking keyword rankings?+

No. A keyword is one string; a topic is many. The same intent phrased 10 ways returns a brand set that overlaps 0.60–0.75 but a leader that can change up to five times. A single phrasing is one draw from a distribution, not a rank.

How many paraphrases do we need per topic?+

More than one, and more than one run each. This study used 10 paraphrases per topic across multiple runs per provider. Two runs of identical wording already overlap only 0.48–0.66, so a single prompt read once carries most of that variance and none of the correction.

Our brand appears on some phrasings and not others. Is that a problem?+

It means the brand sits outside the stable core — the 7–12 brands present in all 10 paraphrases of a topic. Core membership is the durable position; peripheral membership depends on wording. Broad topics carry long tails of one-off brands: Gym & fitness surfaced 45 distinct brands against a core of 9.

Which provider should we optimize for?+

Not a useful frame at this level. All four providers cluster within about 0.05 on membership and 0.12 on rank correlation, and each one is weakest on the same broad topic. Provider gaps are real inside a topic, though — OpenAI holds order on Gym & fitness (0.75) and is the weakest of the four on Sportswear value (0.32) — so read providers per topic rather than as a league table.

We went from #3 to #1 this month. Is that real?+

Not on the strength of one phrasing. In the most contested topic here, five different brands reach #1 across 10 paraphrases of the same intent and top-3 overlap falls to 0.27. Report the leader as a distribution — how often the brand leads across paraphrases and runs — rather than as a position.

How do we separate a real change from model noise?+

Establish the floor first. Two responses to identical wording overlap 0.48–0.66; different wording removes only 0.04–0.09 more. Movement inside that band is not yet a signal. Hold the prompt set fixed, repeat it, and compare distributions between periods.

What does this change about content work?+

Where one brand owns the category, order does not move with phrasing — the work is being cited at all, not out-phrasing the leader. Where the top is contested, order is already volatile, and that volatility is the headroom: broad topics carried the largest brand pools and the weakest leader stability in this sample.