ResearchFindingsPublicationsPeopleContactProtocol on GitHub ↗
Research

Four waves. One protocol.

Recommendation algorithms quietly shape billions of purchase decisions, yet almost no one audits them from the outside. Our program makes the algorithmic shelf of online retail measurable, comparable, and accountable — platform by platform, year by year.

Waves

Two complete, two planned.

Shopify sampled under category quotas; BigCommerce taken as a census with the list fixed before the crawl. Seasonal and longitudinal waves are pre-registered.

01
Surveyed
300
Analyzed
221
74% of the list yielded data
Wave 1 · complete

Shopify

300 stores surveyed, 221 analyzed. Price steering, popularity bias, brand concentration, representation, accessibility.

Blocks · with prices806 · 621
Collection windowAug 5–6, 2026
Designquota sample · 6 categories
View findings
02
Census list
1,394
Analyzed
1,014
73% of the list yielded data
Wave 2 · census complete

BigCommerce

Cross-platform replication on 1,394 candidate stores, 1,014 analyzed. First results confirm the Shopify pattern.

Blocks · with prices3,093 · 2,648
Collection windowAug 11–14, 2026
Designcensus · list fixed before crawl
View findings
03
Nov 2026 · planned

Seasonal slice

Black Friday snapshot: do recommendation systems collapse onto a narrow set of doorbusters?

Protocolv2.1 · unchanged
Time slicesingle day
04
Feb 2027 · planned

Longitudinal wave

Same stores, six months later. Pre-registered design. Does the steering gradient persist?

Designpre-registered
Time points2
Method

From the outside, the way a shopper meets the shelf.

Black-box. Reproducible. Open. Five steps, each written down before the first store was crawled.

01Step 01

Sample

Stores are drawn from public catalogs under category quotas (Wave 1) or as a fixed census list (Wave 2): an active storefront, a catalog above a minimum size, English language. The list is fixed before collection and never trimmed to a round number.

Stores analyzed1,235
02Step 02

Seed

Three seed products per store: A — popular (top of the best-selling sort), B — unpopular (the tail, a different product line), C — median-priced (a third line).

Seeds per store3
03Step 03

Capture

A crawler with a clean profile and an openly research-identifying user-agent opens each seed’s product page, locates recommendation blocks by heading patterns with a structural fallback, takes screenshots and extracts title, price, brand and alt-text status. No login, request rate below an ordinary shopper.

Coverage · W296%
04Step 04

Measure

Per block: price delta and upsell share, popularity overlap, brand concentration, representation index, alt-text status. Means are winsorized; significance is tested with mixed-effects models that respect the store as a cluster.

Metrics per block6
05Step 05

Publish

Anonymized datasets — category and catalog size, no URLs — are deposited on Zenodo with DOIs alongside each publication. The protocol is versioned on GitHub under CC BY 4.0. Capture losses are reported, never silently dropped.

LicenseCC BY 4.0
Anatomy

Anatomy of a recommendation block.

The block is the unit of analysis. Every metric in the protocol is computed on one of these four layers.

A1

Product page (seed)

Three seeds per store — popular, unpopular, median-priced. The page a shopper actually lands on, opened without login.

A2

Recommendation block

Located by heading patterns with a structural fallback. The unit of analysis: every metric is computed here, the store is the cluster.

A3

Price layer

Mean price of the recommended items against the seed. Winsorized to [−100%, +200%]; the gradient by seed-price band is the central finding.

A4

Image & alt-text layer

Who appears on the cards (fashion and beauty, coded against the published codebook) and whether a screen reader gets anything at all.

How we operate

Three rules we don't break.

01Rule 01

No platform cooperation

We audit from the outside — the way a shopper meets the shelf. No API access, no partnerships, no permission required. Any researcher can rerun the measurement.

02Rule 02

Open protocol

Sampling rules, seed selection, capture procedure, metric definitions and the visual-representation codebook are published on GitHub under CC BY 4.0.

03Rule 03

Open data

Anonymized datasets are deposited with DOIs (Zenodo) alongside each publication. Capture losses are reported, never silently dropped.

Zenodo DOIs — September 2026

AVBR
Open science

Rerun the shelf yourself.

The protocol, codebooks and datasets are open. Take the measurement to a platform we have not reached yet — or write to the lab.

Protocol v2.1 · CC BY 4.0Datasets with DOIs · ZenodoNo platform cooperation