The scoreboard for AI-native retail just went public. It doesn't rank the retailers you'd expect.
The first public scoreboard for AI-native retail discovery just landed — and it ranks the wrong retailers, on purpose. That's the story.
TL;DR
- Digital Commerce 360 and ReFiBuy launched the AI Commerce Rankings on 9 July 2026, a quarterly benchmark scoring the Top 1000 U. S. online retailers on readiness for agentic shopping.
- The inaugural leaders — Online Labels, Nixon, Everlane, Brooklinen — sit outside the top 100 of the traditional web-sales ranking. Online Labels is by revenue. It's #1 here.
- The methodology has four inputs: Bot Friendliness, AI Source Traffic, Diversity of AI Sources, 90-Day Momentum. Scale doesn't feature. Catalog machine-readability does.
- DC360 and ReFiBuy separately estimate $180B in U. S. online sales were influenced by agentic AI in 2025 — roughly 15% of the $1.2T Top 1000 total. That's the demand-side justification for the index existing at all.
- This is the third public benchmark for AI-retail visibility to launch in five months, after Parcel Perform's AI Visibility Index (Feb 2026) and Yotpo's LLM Brand Benchmark (Dec 2025). The competitive frame is now real. That, more than any single ranking, is the story.
What actually happened
At around 08:00 ET on 9 July, Digital Commerce 360 — the incumbent trade publication for U. S. ecommerce measurement — announced a partnership with ReFiBuy, a two-year-old firm that coined the phrase agentic commerce optimization and has been quietly selling readiness audits to Top 1000 merchants since late 2025. The output is a new quarterly index, embedded inside DC360's flagship 2026 Top 1000 Report, that scores every ranked retailer on how prepared their catalog is to be found, parsed, and recommended by AI shopping agents — ChatGPT's shopping surface, Gemini, Perplexity, and the agent stacks being built on top of them.
The methodology sits on four legs:
-
Bot Friendliness — can AI agents actually access and interpret the retailer's product data? This is a
robots.txt,llms.txt, schema.org, feed-quality and API-surface question. - AI Source Traffic — what proportion of site traffic already arrives from AI discovery platforms?
- Diversity of AI Sources — is the retailer visible across multiple engines, or dependent on one?
- 90-Day Momentum — is the AI-sourced traffic growing or decaying quarter-over-quarter?
Scale is conspicuously absent. So is brand equity. So is ad spend. The result, published in the DC360 Top 1000 database and covered by ReFiBuy on its ai-rank microsite, is a ranking that shares almost none of its top ten with the traditional web-sales list. Amazon and Walmart still appear near the top — they are inescapable — but the retailers immediately below them are, by DC360's own framing, "entirely absent from the top 10" of the conventional ranking.
The first-cycle leaders are worth naming precisely, because they carry the argument:
- Online Labels — a specialty labels-and-stickers merchant, roughly in the Top 1000 by revenue.
- Nixon — watches and accessories, mid-tier DTC.
- Everlane — apparel, DTC-native.
- Brooklinen — home textiles, DTC-native.
These are not category-killers. They are catalog-clean, SKU-narrative-rich, structurally legible merchants — and they happen, right now, to be legible to a machine reader in a way that most $1B+ retailers are not.
The framework — why this ranking exists, and what it is really measuring
Here is the thing worth understanding, and it is not obvious from the press release.
For twenty-five years the substrate of retail discovery has been the same: humans type queries into Google, Google surfaces links, links land on product pages, product pages convert. The entire discipline of ecommerce — SEO, PLA, product-feed hygiene, review acquisition, PPC budgeting — is a set of adaptations to that substrate. The Top 1000 ranking measures success on that substrate.
The substrate is changing. Not in some 2029 sense; in a 2025 sense. DC360's own estimate — $180B of 2025 U. S. online sales "influenced" by agentic AI — is aggressive and worth reading with a raised eyebrow1. But the direction is not in serious dispute among people who look at the data. ChatGPT's shopping surface, Perplexity's shopping mode, Google's AI Overviews, and the emerging class of dedicated shopping agents (Rabbit, Nova, the Anthropic and OpenAI computer-use models) are collectively re-routing a growing share of purchase-intent queries away from the classical funnel.
The AI Commerce Rankings are not, then, a ranking of retailers. They are a ranking of catalog machine-readability — and the argument being made, quietly, is that catalog machine-readability is about to be as strategically important as SEO was in 2005. Retailers who won the first substrate — through scale, through brand, through Google authority — do not automatically win the second one. In fact, quite often they lose it, because their catalogs are large, messy, inconsistently attributed, and encumbered by legacy commerce platforms that were never built to hand structured product knowledge to a third-party agent.
Online Labels wins the inaugural index not because it is a great retailer. It wins because it has a clean, narrow, well-attributed catalog that an LLM can read without tripping over itself. That is either a profound insight about the next decade of retail, or a temporary artefact of an early scoring methodology. This piece's editorial view is that it is meaningfully both — but more the former than the latter.
Where this fits in a crowded field
This is the point where the analysis has to slow down, because the "first-of-its-kind" claim in the DC360 release is doing more work than it can carry. It is not the first public benchmark in this space. It is the third in seven months, and the differences between them matter.
| Benchmark | Publisher | Launched | Measures | Update cadence |
|---|---|---|---|---|
| LLM Brand Benchmark | Yotpo | Dec 2025 | Brand-level visibility across OpenAI + Gemini across 5 dimensions (GEO visibility, citations, answer quality, structured readiness, trust) | Ad hoc |
| AI Visibility Index | Parcel Perform | Feb 2026 | Brand-level visibility, LLM ranking, brand-trust sentiment across ChatGPT, Gemini, Perplexity | Weekly |
| AI Commerce Rankings | Digital Commerce 360 + ReFiBuy | Jul 2026 | Retailer-level agentic readiness — bot friendliness, AI source traffic, source diversity, momentum | Quarterly |
Yotpo's benchmark is a brand study — Nike, Patagonia, Adidas — and its methodology (an n8n pipeline running structured prompts through two LLMs) is a marketing-side view. Parcel Perform's index is explicitly described by its founder Dr Arne Jeroschewski as "the Nielsen ratings for AI Commerce," and its cadence — weekly, free, no registration — is aggressive. It is closer to a ratings-service model.
DC360/ReFiBuy is doing something different and, if it holds, more consequential. It is a retailer-side benchmark, distributed through the ranking that C-suite retail executives already read, and it treats agentic readiness as a strategic KPI rather than a marketing metric. The distribution is the moat: DC360 is the incumbent list. Being ranked poorly there — or well — reaches the boardroom. Being ranked well on a startup's weekly index reaches the CMO's dashboard, which is a smaller room.
The competitive question, therefore, is not "which of these is best?" It is: which of these becomes the number a CEO cites in a 2027 earnings call? DC360's distribution advantage makes it the front-runner. Its narrower methodology — retailer readiness, not brand sentiment — makes it more defensible against gaming. That is the bet worth watching.
The quieter story — what changes on the ground
Take the framework seriously and there is a concrete change coming to retail engineering roadmaps. It looks like this:
Immediate (0–90 days). Catalog hygiene stops being a project owned by SEO and starts being a project owned by architecture. That means schema.org/Product completeness at SKU level, resolvable canonical URLs, structured attribute coverage (materials, dimensions, compatibility, use-cases), and — increasingly — dedicated llms.txt and product-feed endpoints designed for agent consumption. If you are a retailer with an Adobe Commerce or Salesforce Commerce Cloud stack running product data through legacy PIM, you have discovery work to do before you have build work to do.
Medium (3–12 months). The AI Source Traffic and Diversity of AI Sources metrics push retailers into a measurement problem they have not solved. Server logs need to be re-instrumented to identify agent-driven sessions distinctly from human sessions. Vendor tooling here is thin. ReFiBuy, Parcel Perform, Yotpo, and a handful of newer entrants are competing to become the measurement layer, and none of them yet has definitive traffic-attribution methodology. Expect vendor churn.
Longer (12–24 months). If AI-sourced traffic continues to compound, the underlying commerce stack is remodelled around agent consumption. That means product APIs designed for agent read patterns rather than human browsing, price and inventory endpoints that respond to natural-language queries, and — inevitably — a commercial layer where retailers pay agent platforms for placement, or agents charge retailers for verified inventory feeds. The economic model of AI commerce is unwritten. The infrastructure that determines who wins it is being written now.
None of that is inevitable. All of it is directional.
Where the argument gets weaker
The Signal Score is 9, not 10, and the missing point is worth naming.
The traffic estimates are load-bearing and lightly evidenced. DC360's $180B agentic-influenced sales figure is presented without a methodology appendix in the public materials so far. "Influenced by" is a soft attribution — it does not mean the sale was closed inside an agent; it can mean the shopper consulted an LLM at some point in the journey. Under a strict definition (agent-initiated, agent-completed transaction), the number is meaningfully smaller. Probably an order of magnitude smaller, in mid-2026. This does not invalidate the index — the direction is what matters — but readers should not quote the $180B number as if it is a bureau statistic.
Methodology transparency is thin at launch. ReFiBuy has published high-level component names, not weights, not sampling detail, not error bars. Parcel Perform, whose weekly index has an eight-industry footprint, publishes its prompt libraries openly and Yotpo publishes its rubric. DC360/ReFiBuy has not yet done either at the same depth. That will need to change if this is to become the reference benchmark.
The leaders may not survive the second quarter. Bot friendliness is easy to game once retailers know it is being measured. Ninety-day momentum is inherently unstable at the top of a young ranking. Expect the inaugural leaderboard to reshuffle materially in the Q4 2026 refresh. That is normal. But it should temper the "Online Labels beats Amazon" framing, which is a subset of the truth.
One benchmark is not a market. Three benchmarks is not a market either. The category is pre-consolidation. Any of the three current entrants could win, none could win and a fourth could emerge, or the LLM platforms themselves could publish first-party inventories that make all three redundant. That last scenario — OpenAI or Google publishing an inventory of "shoppable-ready" retailers — is the one that would end this category most decisively.
What this means for the people who actually have to act
For retailers in the Top 1000. Get the AI Commerce Ranking score before the Q4 refresh. If you are outside the top 200 by score, the intervention is boring and unglamorous: PIM cleanup, structured-data audit, llms.txt publishing, schema.org completeness at SKU level, and instrumenting server logs for agent traffic. The vendors selling this work (ReFiBuy for audits; Yotpo, Bloomreach, Constructor.io on the search-and-discovery side) are converging on a similar playbook. This is not exotic engineering. It is deferred hygiene.
For retailers outside the Top 1000, and mid-market DTC brands. The story is more interesting. The inaugural leaders — Online Labels, Nixon, Everlane, Brooklinen — are your reference class, not Amazon's. Catalog cleanliness at your scale is an achievable competitive edge. The window is 12–18 months before the majors close the gap through spend.
For commerce-platform vendors. Shopify, BigCommerce, Adobe Commerce, Salesforce Commerce Cloud, Commercetools — the pressure to ship first-party agentic-readiness tooling is now on your roadmap whether it was there yesterday or not. Shopify's early lead on robots.txt for agents and its structured shopping feed positions it best. Adobe and Salesforce enterprise stacks have the most technical debt to close.
For consumers. Nothing to do. This is a supply-side story. But if in twelve months your ChatGPT shopping session starts surfacing brands you have never heard of at the expense of the incumbents, this benchmark — and the engineering work it is provoking — is a decent part of why.
For policy and competition observers. Watch the platform layer. If agent providers begin monetising placement inside shopping surfaces, the antitrust framing for AI commerce arrives fast, and the AI Commerce Rankings become an exhibit in the argument about whether the ranking reflects merit or paid access. It is not there yet. It will get there.
Uncertainty ledger
- Whether $180B of "AI-influenced" 2025 online sales survives independent replication. High uncertainty on the number, low uncertainty on the direction.
- Whether DC360's quarterly cadence is fast enough to remain the reference benchmark against Parcel Perform's weekly cadence. Cadence usually wins.
- Whether OpenAI, Google, or Anthropic publish first-party retailer inventories in the next 12 months. If they do, third-party benchmarks compress.
- Whether the inaugural top ten survive one refresh cycle. Prediction: 4–6 of the current top ten will not be in the top ten in Q4 2026.
- Whether the underlying methodology will be published in sufficient detail to be audited. If it isn't, the benchmark degrades into marketing.
Bottom line
For twenty-five years, being a big retailer meant being findable by Google. It is starting to mean being legible to a machine that reads your catalog and decides, on the shopper's behalf, whether to recommend you. Digital Commerce 360 and ReFiBuy have just published the first scoreboard for that second substrate, and — pointedly — it does not rank the retailers who won the first one. Whether Online Labels stays at number one is beside the point. The point is that the scoreboard now exists inside the report the retail C-suite already reads, and that is how a niche measurement problem becomes a boardroom KPI. Retailers should get their score before Q4. Vendors should ship tooling before Q1. Everyone else can watch the leaderboard reshuffle and treat it, correctly, as an early index on which retailers understood the substrate change first.
Sources
- Digital Commerce 360, 2026 Top 1000 Report — AI Commerce Rankings Overview (Tier 1, primary)
- Digital Commerce 360, Digital Commerce 360 and ReFiBuy Launch First-of-its-Kind AI Commerce Rankings, 9 July 2026 (Tier 1, primary press release; syndicated via USA Today)
- Digital Commerce 360, Ecommerce Trends: What shoppers are using AI to buy, 26 Feb 2026 (Tier 1, methodology preview)
-
ReFiBuy, Articles /
ai-rankmethodology page (Tier 2, vendor primary) - Parcel Perform, AI Visibility Index launch coverage, 18 Feb 2026 (Tier 2, competing benchmark)
- Yotpo, LLM Brand Benchmark: AI Visibility Report, Dec 2025 (Tier 2, competing methodology)
- Ranksignal.ai, Retail industry AI-search visibility benchmarks, Mar 2026 update (Tier 3, contextual)