Skip to content

Data and method

Every answer the Reach Agent gives is produced from the same graph. This page describes what is in it and how a question becomes an answer.

The two vocabularies

The platform holds two bodies of language and one map between them.

Product vocabulary is how sellers describe what they sell. Over 360 million product listings from more than a million independent online stores across many platforms, resolved into 29.6 million product clusters. A cluster is a semantic neighborhood: every way the market writes the same product, grouped together. That is what lets "Merino Wool DK Weight Yarn" and "soft baby blanket yarn" land in the same place.

Buyer vocabulary is how people search. 101.4 million analyzed keywords with search volume, cost per click, and competition, plus 70.5 million natural-language questions of the kind people ask out loud rather than type.

The intersection is the map between them: which searches should reach which products. That mapping is the thing the platform exists to compute, and it is what makes a ghost product findable. Counts here are as of July 21, 2026.

The other layers

Layer What it carries
Price history Observed listing prices over time, as dated change events per product
Weekly dimensions How demand moves by season, audience, occasion, and use case
Product attributes Material, type, brand, and other facets extracted per product, which is what your catalog groups are built from
Store registry The independent stores behind the listings

How a question becomes an answer

The Reach Agent classifies your question, dispatches the skills that can answer it, and each one queries the graph directly. Nothing is answered from a summary written earlier or from a model's recollection. A skill that cannot reach the data it needs reports that instead of estimating.

Findings are then synthesized into one answer, and the evidence footer is attached from the actual queries that ran, not from what the answer claims to have used.

Freshness

The layers refresh on their own cadences. Rather than publishing a single date that would be wrong for most of the data most of the time, every answer carries the snapshot date of each source it used, and dated figures state their vintage in place: "counts as of July 16, 2026".

When a build date cannot be established, the field is omitted. A date is never inferred to fill the slot. See Read an answer.

Your catalog

Your catalog is yours. It is matched against the graph so your products can be located in it, and it is never merged into the market data other accounts read. Answers about your products are computed from your catalog alone.