Skip to content

Citelift / guides

Guide · updated 18 September 2026 · 9 min read

Product facts on Shopify storefronts: what an HTML audit misses

A 17-page product-information pilot across six retail categories, with source-text counts, browser corrections, methods and downloadable observations.

A simple HTML extraction can miss product facts that a browser displays. We selected 18 storefronts across six categories and found an eligible product page for 17 using a fixed homepage-link rule. In the extracted text, 16 pages contained product composition information, 13 contained relevant numerical selection details, and 12 contained use, care or storage instructions. Browser review then recovered care instructions on three pages where the initial extraction had not found them.

This is a small product-information pilot conducted on 9 September 2026. It does not measure what an AI assistant retrieves or recommends. A missing field in our extraction is not evidence that an entire store lacks that information.

Download the page-level observations or inspect the frozen selection protocol.

When the product has selectable options, continue with the variant price and availability audit; a parent-page extraction cannot establish every variant's current offer.

What the source text contained

The denominator below is the 17 selected product pages, not all products on those stores or a representative sample of Shopify merchants.

Check Present in extracted source text What qualified
Composition 16 of 17 Product-specific material or ingredient information; not necessarily a complete ingredient list
Numerical selection information 13 of 17 A relevant size, dimension, weight, volume or serving amount
Use, care or storage 12 of 17 Instructions for the selected product; manufacturing descriptions did not count
All three checks 11 of 17 At least one qualifying item for every check above

These are manual observations of a particular extraction. They are not completeness scores, rankings or compliance assessments. For example, identifying whey protein does not establish a complete ingredient list for every flavor. A model’s height linked to a clothing size provides some selection information; it is not a full garment measurement table.

The method retained text within the HTML main element when available, otherwise body text, excluding scripts, styles, templates and noscript content. That choice can miss information outside main, information loaded later, and text inside images. Source HTML can also contain hidden panels, navigation and recommendations; we manually checked that counted facts applied to the selected product. Debug fields, prices, review counts and other products’ details did not qualify.

Three browser checks changed the interpretation

On these selected pages, the initial extraction did not find product care instructions. Opening the product in a browser and inspecting its details did:

The observation is that the two review methods exposed different information. It does not isolate whether JavaScript, document structure, interaction or an extraction limitation caused each difference. Nor does it establish how Google, ChatGPT, Claude or another crawler processes those pages.

This is why a content audit should preserve both the original observation and its correction. Replacing “not found in the extraction” with “the brand omitted it” would overstate the evidence.

What a useful product fact sheet contains

Before writing a buying guide, record the facts a reader needs for the actual decision:

Category Composition Numerical selection detail Use or care
Apparel Fabric and lining Relevant garment measurements Washing and drying
Home Main materials Dimensions and compatibility Cleaning or installation
Beauty Product ingredient list Package quantity Label application instructions
Supplements Product and variant ingredients Label serving size and quantity Label directions; do not invent medical advice
Pet Materials or ingredients Fit or portion information Care and use limitations
Food Ingredients Package or serving quantity Preparation or storage

Keep the source URL, selected variant and date beside each fact. If the page links to a chart or manual, open it. If a fact appears only in an image, verify it visually and record that location. Confirm conflicting or unavailable information with the merchant before using it in copy.

Our free product-description scorer checks text you paste using deterministic rules. It does not fetch the whole storefront, read product-label images or verify factual accuracy. Use its prompts to guide review, then supply missing verified facts yourself. The content-tool evaluation worksheet helps compare drafts and publishing workflows using the same evidence.

Watch a claim check from source sheet to correction

Watch the 75-second walkthrough or read the plain-text transcript. The Alpine Daypack shown is a fictional demonstration, not a customer or merchant result. It shows why a 100-point structure score can coexist with unsupported claims and why the source comparison is a separate required step.

Worked example: the scorer cannot tell 92 cm from 82 cm

This synthetic fixture, completed 14 September and rerun 18 September 2026, is permission-safe and reproducible. Harbor Linen Apron is not a store, customer or product. Download the completed audit row and blank worksheet rows, use the exact before and after scorer inputs, inspect the machine-readable executed result, and review the public proof map.

The source sheet says the apron is 82 cm long. The deliberately incorrect description says 92 cm long. Both versions are visible plain text; there is no hidden accordion or JavaScript ambiguity in this fixture.

Pass Source or rendered observation Decision
Source fact Catalog field: length_cm = 82; checked 14 September 2026 Treat 82 cm as the allowed value for this exercise
Visible rendered fact “measures 92 cm long by 68 cm wide” Discrepancy: stop publication and check the selected variant
Correction “measures 82 cm long by 68 cm wide” Matches the source sheet; record who approved the change
Deterministic scorer before 100/100; it finds four measured values, including 92 cm It proves the text contains specifications, not that 92 is true
Deterministic scorer after 100/100; it finds four measured values, now including 82 cm The unchanged score is expected; the fact check happened outside this tool

The downloadable inputs are 100 words each and differ only at 92 cm versus 82 cm. We executed both against the checked-in scorer on 18 September 2026. This is the actual compact result:

before: score 100; 10/10 checks passed; failed checks: none
after:  score 100; 10/10 checks passed; failed checks: none

The scorer uses ten equally weighted writing checks. It recognizes measured-value patterns, buyer wording, question words, material wording, a trade-off, sentence length, tone and a number. It has no source sheet to compare against, so a plausible but wrong measurement can pass every check. This is a negative-capability test of one deterministic fixture, not a writer benchmark or evidence that other drafts are accurate.

Complete the same worksheet without Citelift:

  1. Copy one product and selected variant into the source columns. Record the URL or admin field, date and value.
  2. Open the customer-visible page. Expand its details and linked charts, then record the exact visible value separately.
  3. Compare the two cells. Use match, discrepancy or unresolved; do not turn “not found” into “false.”
  4. For a discrepancy, pause the draft. Confirm which source is current with the merchant before changing public copy.
  5. Record the before text, after text, reviewer and date. Re-open the page after saving to verify the visible result.

For the exact demonstration, download the text file, copy only the 100-word BEFORE INPUT into the scorer, save the result, and repeat with AFTER INPUT. The expected result above is also written into the downloadable worksheet. A different result means the scorer or parsing behavior has changed and this dated example should be rechecked before it is cited. The 75-second walkthrough shows the same source-first method with a separate fictional daypack. Citelift proof and progress explains where this narrow proof sits in the wider product evidence.

Use the scorer after this evidence pass for clarity prompts. It can flag that the copy lacks specifications or material wording; it cannot validate dimensions, prices, ingredients, stock, care instructions, certification or suitability. For a publishing workflow that can hold a draft while a person resolves a conflict, continue to the editorial-control walkthrough.

Turn the observation into a page audit

Use two passes and keep their results separate:

  1. Source pass: save the selected product URL, variant and date; inspect the HTML or extracted main text; mark each required fact as found, not found or ambiguous and quote only a short locating phrase in private notes.
  2. Rendered pass: open the same product in a browser; expand details, size, ingredient and care controls; inspect linked charts or manuals; record whether the browser confirms, corrects or leaves the source observation unresolved.
  3. Conflict pass: when the page, label image, linked file and catalog field disagree, stop. Ask the merchant which source is current before using the fact in public copy.
  4. Article pass: carry approved facts into a content brief with the source URL and checked date beside each assertion. The six fictional blog examples show where composition, selection and use facts belong across categories.

This workflow does not produce a percentage completeness score. It produces a list of usable facts, unresolved gaps and corrections. That is the evidence a writer needs.

Selection and collection

We used the first three store-domain-signal candidates per category in an existing roster: apparel, home, beauty, supplements, pet and food. Each homepage contained a Shopify store-domain signal at collection. This is evidence of a Shopify relationship, not a certification of every part of the storefront stack or the brand’s ownership, revenue or market size.

For each homepage, the collector selected the first distinct eligible physical-product hyperlink in document order. Its implementation recognized same-host HTTPS /products/ links. Gift cards and clearly unrelated merchandise were skipped and logged. Bundles were eligible; selection was not random and did not attempt to choose a brand’s best-known product.

The fixed rule produced 17 selected pages. Fly By Jing had no eligible hyperlink in the collected initial homepage HTML and was retained as unavailable, without replacement. That does not mean it has no product pages. Parachute’s selected bundle page exposed inspiration text, but its component choices did not populate during browser review; its footer linked to fabric and care resources. We do not treat that incomplete bundle review as a whole-store information gap.

Public product requests were bounded and checked against the collected robots.txt. We archived response dates, URLs and hashes, then manually coded the three checks and reviewed selected pages in the browser. The CSV preserves the source-text observations and a separate browser-follow-up note. Image labels, linked manuals, alternate variants and retailer destinations were not exhaustively transcribed. There is no final whole-page completeness score.

Limits and disclosure

Citelift funded and publishes this research and sells Shopify content software. The brands were not customers or endorsers in this study. One Citelift-operated observer used Codex assistance for collection and analysis; there was no independent double coding.

The sample is small, deliberately ordered and drawn from an existing convenience roster. It is unsuitable for estimating the prevalence of information gaps across DTC or Shopify stores. Pages, variants and regional rendering can change. Browser access used an India-based connection; storefront presentation may differ elsewhere.

No medical, product-safety, certification or performance claims were validated. No sales, citations, rankings or causal uplift were measured. The study did not test Citelift-generated content. Incremental LLM API spending was $0; no paid generation was needed for collection or counting.

Raw HTML and browser text remain private. The downloadable observations are paraphrased evidence notes, source URLs and archive hashes, not full copies of the stores’ pages. Hashes identify the retained archives; they do not independently validate the collection. If a source or observation needs correction, contact support with the selected URL and detail.

Questions.

Does an HTML audit prove a product page lacks buying information?

No. In this pilot, browser review recovered care instructions on three pages where a simple extraction of the initial HTML did not find them. Check rendered details and linked resources before diagnosing a gap.

Does the study show which stores AI assistants recommend?

No. It examines product-information availability, not assistant retrieval, citations, rankings, sales or product quality.

How many stores were included?

We selected 18 stores from a pre-existing convenience roster across six categories. The fixed homepage-link rule found a product page for 17. The remaining store was recorded as unavailable without replacement.

, founder of Citelift. Citelift writes and publishes product-linked articles on your Shopify blog and checks whether AI assistants name your store.

Citelift is listed on the Shopify App Store: Citelift on the Shopify App Store.

Run the check after reading Product facts on Shopify storefronts: what an HTML audit misses