Skip to content

Citelift / guides

Guide · updated 17 September 2026 · 10 min read

Repeating buying questions in ChatGPT and Claude: a 48-answer pilot

A small consumer-browser study of visible source links across 12 buying questions, with repeat-overlap results, exact prompts and downloadable data.

We asked 12 buying questions twice in ChatGPT and twice in Claude, retaining 48 answers on 9 September 2026. None of the 24 within-platform repeat pairs had exactly the same set of visible text-link URLs. This is a small observation about repeatability, not proof that either platform is inaccurate or that a particular content strategy works.

Update, 14 September 2026: A second 48-answer repeatability wave using the same questions is appended below with its own methods, data and audit links. It is a separate dated observation, not a controlled trend or a pooled result with the original pilot.

The questions covered apparel, home, beauty, supplements, pet and food. They asked what a shopper should compare or measure, without naming a brand. The exact prompts, observed links and per-question results are available to inspect.

Dated findings at a glance

CollectionConsumer interfaceRetained answersComplete repeat pairsIdentical URL setsPairs with no shared URLMean pair overlap
9 Sep 2026ChatGPT24 / 2412 / 120 / 122 / 1216.0%
9 Sep 2026Claude24 / 2412 / 120 / 124 / 1210.5%
14 Sep 2026ChatGPT24 / 2412 / 120 / 122 / 1223.2%
14 Sep 2026Claude24 / 2412 / 120 / 123 / 1216.3%

Each percentage is the unweighted mean of 12 within-interface, within-question Jaccard overlaps. The rows are separate dated convenience samples. They are not pooled, and differences between dates or interfaces are not controlled comparisons.

Mean visible-URL overlap for each dated consumer-interface wave. Four separate bars show 16.0% for ChatGPT and 10.5% for Claude on 9 September, then 23.2% for ChatGPT and 16.3% for Claude on 14 September.

Source and credit: Citelift funded and published the work. Siddesh Patil was the single observer; Codex assisted with offline analysis. The chart and table were calculated from the public per-question overlap files, not from the private answer text. Download the standalone evidence bundle, ZIP, four-row aggregate CSV or chart as SVG.

The deterministic ZIP is 37,700 bytes with SHA-256 bc87336c01378ba00831383ab575f33ef2956bf026fc3a2164f12a3904d12611.

Reproduce the aggregate

The public reproduction script uses only Python's standard library and the two public run, visible-link and overlap tables. From the application directory, run:

python3 public/resources/consumer-citation-repeatability-study/reproduce.py --root public --check

It independently counts retained answers and answer–URL occurrences, deduplicates URL and host fields, and recalculates pair counts and the mean Jaccard overlap. --check fails if the result differs from the downloadable aggregate. Run without --check to print the recomputed CSV. Raw answer text, excluded attempts and collapsed citation-group destinations are unavailable publicly, so this reproduces the published calculations rather than independently verifying browser collection.

If you do not have the repository, download and extract the standalone evidence bundle. Its README gives one command that runs against the included public-safe source tables.

What we observed

Measure ChatGPT Claude
Retained answers 24 24
Answers with visible text links 24 24
Answer–URL occurrences 143 202
Distinct normalized URLs across answers 118 182
Distinct source hosts, excluding www. 56 163
Repeated question pairs 12 12
Pairs with identical observed URL sets 0 0
Pairs with no shared observed URLs 2 4
Mean within-platform URL-set overlap 16.0% 10.5%

An answer–URL occurrence counts each normalized URL once per answer. The same URL in another answer counts again. “Source hosts” preserve subdomains other than www.; this is not a count of independent businesses.

Overlap uses the Jaccard measure: shared URLs divided by the combined set of unique URLs. For example, if two answers each show six URLs and share two, the combined set has ten URLs and overlap is 20%. The averages above give each of the twelve question pairs equal weight. They do not mean that 16.0% or 10.5% of all sources were retained.

Both interfaces can collapse additional citations behind a group label. We did not expand those groups or include image-only links. These numbers describe visible text-link destinations, not every citation or every source consulted. Differences in interface presentation make the two columns unsuitable as a platform quality ranking.

What this changes about a manual visibility check

A single answer is a dated observation. In this sample, repeating the exact question in a fresh conversation changed the visible URL set every time. A practical measurement record should therefore retain the question, platform, date and source URLs, and include repeats before drawing conclusions about persistent visibility.

That is an interpretation of this pilot, not a statistically representative estimate. A missing URL in the second answer also does not establish that the page became inaccessible or lost a ranking. This design cannot identify why a link appeared or disappeared.

Our free manual checker prepares buying questions and analyzes answers you paste. It separates a brand-name match from a link to your domain and lets you download the result. It does not query an AI service, verify the linked claims or infer a visibility score from this study.

Watch one retained comparison

Watch the 70-second walkthrough or read the plain-text transcript. It interprets one retained pair from this historical 9 September pilot; it does not introduce a new collection wave or turn the result into a benchmark.

Use the pilot as a measurement checklist

The result supports a cautious store-level routine:

  1. Fix the buyer question, market, language, platform and search setting before the first run.
  2. Use a fresh conversation for each attempt and run at least one repeat before treating a visible link as persistent.
  3. Save the full answer privately where platform terms permit, then record brand mentions and visible text-link URLs as separate fields.
  4. Open each claimed store URL before acting on it. A visible citation does not prove that the destination supports the nearby statement; the product-fact HTML audit explains why source and rendered reviews can differ.
  5. Compare like with like at the next interval. Do not change the prompt and call the new answer a trend.

For a reusable store-level log, use how to measure AI visibility. For Google's own AI surfaces, keep Search Console eligibility and reporting separate from consumer-chat observations; the Google AI Overviews and AI Mode guide covers that distinction.

One inspectable product-page example

The first ChatGPT answer to the summer-overshirt question included a Sunspel linen overshirt product page. Opening that page on the collection date showed its linen composition, care instructions and garment measurements. This illustrates a merchant product page appearing among the observed links for an informational buying question.

It does not establish that merchant pages are preferred overall, that the store uses Shopify, or that the link drove a sale. We did not systematically classify all source types or audit every destination and nearby claim. Treat the downloadable URLs as observations to investigate, not endorsements.

How we collected the answers

Each question was submitted in a fresh ChatGPT temporary chat and a fresh Claude incognito chat. ChatGPT displayed High effort but did not expose the underlying model; Claude displayed Opus 5 and High. Every prompt requested US English, web search and cited sources. Search was requested, but retrieval execution was not independently audited.

We ran two passes, interleaving platforms by category. This was not randomized. Existing paid consumer subscriptions were used; incremental LLM API spending was $0. Consumer-product results must not be presented as API results. Account effects, IP location and platform changes were not controlled.

A browser restart interrupted collection. Five confirmed first-pass ChatGPT answers were lost before archiving and the same questions were resubmitted. One additional Claude submission before that restart could not be confirmed. Those attempts are outside the 48 retained answers. Missing archives were replaced without selecting answers based on favorable findings.

Offline analysis removed tracking parameters and fragments, excluded platform help/privacy links and blank/image-only anchors, and deduplicated each answer's URLs. Other query parameters, paths and HTTP/HTTPS variants remain distinct. Run metadata includes capture times and hashes of the private archives.

Limits and disclosure

This convenience sample contains only two questions per category and two answers per question per platform, all collected on one day. Repeats are not independent shoppers. No uncertainty interval, market-wide citation rate, sales effect or causal uplift is estimated. The separate planned 36-store Shopify cohort was not completed and is not the denominator here.

Citelift funded and published the pilot and sells content and visibility software. One Citelift-operated observer collected the data with Codex assistance; there was no independent double coding. The study did not test Citelift-generated articles or merchant outcomes.

Raw answer text remains private to avoid reproducing full model and third-party passages. The public derived data supports checking the calculations, but does not independently verify the complete browser collection. Archive hashes identify retained files; they are not a substitute for access to those files. A future repeat may produce different results.

Download and inspect

For definitions and a repeatable store-level log, see how to measure AI visibility.

Second repeatability wave: 14 September 2026

On 14 September 2026 we repeated the same 12 US-English shopper questions twice in each consumer interface, producing 48 of 48 planned retained answers with no unavailable canonical attempts. ChatGPT used fresh Temporary chats with Unpersonalized visibly selected, High effort and web search selected in the composer; the underlying model name was not exposed. Claude used fresh incognito conversations displaying Opus 5 and High, and every retained answer visibly reported web search. Existing consumer subscriptions were used, with no paid API calls.

Fourteen ChatGPT submissions matched the frozen prompt bytes exactly. Ten preserved the exact question and suffix but used one separating space instead of two newlines; the run metadata labels them whitespace_normalized rather than rewriting them as exact. All 24 Claude canonical prompts matched exactly. Seven earlier answered Claude attempts were excluded for protocol compliance before analysis and without regard to their citation outcomes: six materially paraphrased both repeats of three questions, and one reused context after a reload instead of starting fresh. Exact or fresh-conversation replacements became the canonical runs. ChatGPT had no personalized or non-Temporary submissions.

Measure — 14 September wave ChatGPT Claude
Planned attempts 24 24
Retained answers 24 24
Unavailable attempts 0 0
Answers with visible text links 24 24
Answer–URL occurrences 145 189
Distinct normalized URLs 116 166
Distinct source hosts, excluding www. 55 147
Complete repeat pairs 12 / 12 12 / 12
Pairs with identical observed URL sets 0 0
Pairs with no shared observed URLs 2 3
Mean within-engine URL-set overlap 23.2% 16.3%

The counting and Jaccard definitions match the historical pilot above. In this second dated wave, none of the 24 complete within-engine pairs had identical visible URL sets. Two ChatGPT pairs and three Claude pairs shared no observed URL. These are consumer-interface observations for this collection date. They do not establish that repeatability improved or declined from 9 September, and they do not identify a causal effect from any Citelift or merchant content change.

The 9 and 14 September values are not a controlled trend. ChatGPT did not expose its underlying model, and model, retrieval and interface behavior can change without a visible label. Claude displayed Opus 5 and High, but matching labels do not establish an unchanged system. Temporary and incognito conversations reduce saved-chat and profile context; they do not control account, IP, location or platform experimentation. Accepted whitespace-only prompt deviations remain visible. Repeats are not independent shoppers, and the two engines should not be pooled into one rate.

What the source spot check covered

The occurrence-level source-class and support fields remain unreviewed in the derived link table: 0 of 145 ChatGPT occurrences and 0 of 189 Claude occurrences have embedded assessments. A separate deterministic eight-occurrence audit reviewed the first qualifying link in repeat 1 for four fixed questions in each engine. That is 4/145 ChatGPT and 4/189 Claude occurrences, or 8/334 overall. In the fixed sample, all four ChatGPT sources were classified as editorial or reference publishers and all four supported the nearby claim tested. Claude's sample contained two editorial or reference publishers, one merchant or retailer and one unknown; its support results were one supported, one partially supported, one unsupported and one unverifiable.

This spot check is not representative, not an error rate and not a full-answer review. Six of the eight selected labels displayed a collapsed +1 or +2 group, and those additional destinations were not reviewed. The remaining occurrences are unreviewed; they are not “other,” “supported” or “unsupported.” Publisher ownership also does not establish source quality or independence.

Data, disclosure and use

The frozen questions and normalization rules match the 9 September repeatability protocol. Inspect the frozen question panel, run metadata, visible links, per-question overlap, aggregate summary, method note and source-audit sidecar. Raw answer text and the seven excluded attempts remain private.

This wave measures visible links in answers to unbranded consumer shopping questions and the stated audit coverage. It does not estimate market share, search volume, referral traffic, sales, unreviewed citation accuracy or causal uplift. These shopping questions did not ask merchants to select a SaaS tool, so their Citelift mention and owned-domain-link observations must not be treated as a measure of Citelift or broader SaaS visibility. Use the wave to decide how many repeats and how much source review a monitoring routine needs, not to rank the two platforms.

Questions.

Did repeated questions produce the same visible links?

In the original 9 September pilot, no pair had an identical visible URL set. We asked 12 questions twice on each platform, producing 24 within-platform pairs. The appended 14 September wave also had no identical pair among its 24 within-platform pairs. Neither small sample is a universal benchmark.

Does this compare citation accuracy?

No. The original pilot counts observed text-link URLs and their overlap without a systematic destination audit. The 14 September wave adds a fixed eight-occurrence spot check, not a full-answer or representative accuracy review. Collapsed citation groups were not expanded.

Did this research use paid API calls?

No. Both waves used existing ChatGPT and Claude browser subscriptions and analyzed the saved links offline. Incremental LLM API spending was zero; subscription and hosting costs still exist.

, founder of Citelift. Citelift writes and publishes product-linked articles on your Shopify blog and checks whether AI assistants name your store.

Citelift is listed on the Shopify App Store: Citelift on the Shopify App Store.

Run the check after reading Repeating buying questions in ChatGPT and Claude: a 48-answer pilot