Methodology
Version 0.1 · Series begins April 2026 · Updated monthly
What the index measures
The GroceryChop Price Index tracks the price of a fixed basket of everyday grocery items — observed on actual store shelves, not estimated from surveys — across major U.S. grocery chains. It answers a simple question every month: did the same groceries, at the same stores, get more or less expensive?
Where the prices come from
GroceryChop observes millions of in-store prices daily as part of its price-comparison product. Every observation is classified by price basis: shelf prices (what a shopper pays in the store, from roughly 29 grocery banners) or marketplace prices (delivery-platform listings, which typically include platform markups, from 60+ additional banners). The two bases are never mixed.
Two series: shelf and marketplace
The headline index is shelf prices only. Marketplace (delivery-platform) prices are published as a separate, clearly labeled series in the dataset. Why include them at all? Because a platform markup that is roughly constant cancels out in our matched-sample month-over-month calculation — the same item at the same store is compared to itself, so the markup divides out of the ratio. Marketplace price changes are therefore valid signal, and they extend coverage to metros and chains the shelf panel doesn't reach. Marketplace price levels, however, are delivery prices — higher than shelf — and are labeled as such wherever they appear. Never compare a marketplace level to a shelf level; that difference is the platform markup, not geography or inflation.
The basket: specified items, not loose categories
Like the Bureau of Labor Statistics, we price tightly specified items — for example “whole milk, one gallon” or “large white grade-A eggs, dozen” — rather than averaging everything in a category. Each of the basket's 19 items carries strict matching rules (product title, package-size window, and unit checks), and every observed price is rescaled to the item's reference size so a 3-lb value pack and a per-pound listing land on the same comparable series. Specialty, organic, and premium variants are excluded so the series tracks the staple, not the assortment.
Weighting
Category weights come from the BLS Consumer Price Index relative importance tables (December 2025, 2024 expenditure weights) — public, external figures reflecting how American households actually spend. Weights are never derived from how much data we happen to collect, which would skew the index toward whichever chains we observe most.
Month-over-month change: matched sample
The monthly change compares only the same item at the same store across both months, aggregated with a trimmed Jevons (geometric mean of price ratios) — the same elementary formula BLS uses. This matters because our store panel grows over time: a naive comparison of monthly averages would measure changes in which stores we observe, not changes in prices. The matched-sample design guarantees the index moves only when prices move. Monthly changes are chained into the index level (April 2026 = 100, re-based at a methodology change — see below).
What we refuse to publish
- Thin cells. Every published figure must clear minimum thresholds (hundreds of observations, multiple stores, at least three distinct banners). Cells below the floor show “—” — withheld, never estimated.
- Retailer comparisons. We publish aggregates only — never “Chain A vs. Chain B” pricing, and never individual product price lists.
- Year-over-year figures before they exist. The series begins April 2026; annual comparisons begin when 13 months of data exist (spring 2027), not before.
- An “all groceries” number carried by one category. The headline change requires at least two categories with sufficient matched data.
Known limitations
- Convenience panel. Coverage follows where GroceryChop users shop, not a designed statistical sample. Metro pages exist only where density genuinely supports them.
- Provisional early months. April–May 2026 reflect a smaller matched panel and may be revised in interpretation; the panel is stable from June 2026 onward.
- No seasonal adjustment. Our figures are raw observed changes; the official CPI is seasonally adjusted. Compare with that in mind.
- Price levels are indicative; changes are the precise figure. A market's basket cost depends on which items and stores cleared our sample thresholds there, and that composition differs between markets and between the two price bases. So comparing two basket costs — city against city, or shelf against delivery — mixes genuine price differences with coverage differences. Matched-sample month-over-month changes have no such problem: they compare the same item at the same store, so composition cancels. This is why we publish no “delivery premium” figure; doing it honestly requires matching items across the two bases, which is planned but not yet built.
- National dispersion includes geography. The national price-spread figure includes genuine regional differences; metro dispersion is the within-market number.
How this differs from the official CPI
The BLS CPI food-at-home index is the authoritative measure of U.S. grocery inflation, built from a designed sample and seasonally adjusted. Our index is complementary, not competing: it observes orders of magnitude more prices, publishes days earlier, and reaches metro detail official data doesn't — at the cost of a convenience panel and no seasonal adjustment. Where the two diverge, the divergence itself is usually the interesting story.
Revisions and versioning
Each published figure carries a methodology version. If the method materially changes, the full series is recomputed under the new method and the change is documented here — we do not silently mix methodologies within a series.
What an “observation” counts, and what changed
From August 2026 onward, an observation is a distinct price: the first time we see a given product at a given store in the month, plus every time that price changes. Reading the same unchanged price again on a later day does not count again. Our minimum-sample floors apply to that number, and each figure also carries n_readings — how many times those prices were read.
Earlier months counted readings. For April through July 2026, n_observations is a count of readings, and those figures are frozen: we do not restate published numbers. The difference matters and is not uniform. Measured across our corpus over the 30 days to 20 September 2026, readings exceeded distinct prices by 4.98× overall — and by 21.3× for one chain that supplied about half of July’s national shelf sample. A cell at the 400 floor under the old counting therefore rested on roughly 80 distinct prices, and fewer where one frequently-read retailer dominated. We are publishing both numbers from now on so the distinction is visible rather than implied.
The same change applies to the price itself, from August 2026. An item’s price in a cell is the median of the prices behind it — and from August that median is taken over distinct prices, one per store per price level. Earlier months took it over every reading, so a price re-read three hundred times counted three hundred times toward the median. The measured effect on August is large and uneven: the national marketplace headline is about 21% higher under the new median than it would have been under the old one (beverages +66%, eggs +62%), while the shelf headline moves −1.9%. Marketplace prices are where the same listings are read over and over, so that is where the old method pulled hardest.
Because of that, August 2026 cannot be compared with July 2026. A month-over-month across that boundary would measure the change of method, not a change in prices. We therefore publish no month-over-month and no year-over-year for August, and the index series restarts at 100 rather than being chained through. The break is visible in the data on purpose: mom_pct is empty and index_value is 100. Comparisons from September 2026 onward are on a consistent basis.
2026-09-20 — August 2026 published, with two caveats. First, the counting and median changes above: August is the first month on distinct prices, and it is not comparable with July. Second, price collection was materially degraded from 27 July to 26 August 2026 — 26 of those 31 days — during which the depth of our per-store coverage fell sharply (for one large chain, per-store depth fell about 83%). August’s sample is therefore composed differently from July’s: fewer prices per store, and a different mix of stores within each metro. We publish it because the floors it clears are real ones — 48,350 distinct marketplace prices across 2,656 stores, and 20,448 shelf prices across 1,451 stores — but a reader comparing metros within August should know that the panel behind them moved during the month.
Corrections
Methodology changes are applied to the full series, as above. Data errors are corrected individually: the affected cells are withdrawn, never silently overwritten, and every correction is listed below with the date, the cells affected, and the reason. The withdrawn values remain available in the corrections dataset.
2026-09-18 — Indianapolis, July 2026, marketplace prices withdrawn. The July sample for Indianapolis marketplace prices included two stores outside the metro — a Jewel-Osco in Oak Park, Illinois (212 readings) and a Safeway in Kaneohe, Hawaii (1 reading) — whose records had been assigned Indianapolis-area ZIP codes by a defect in how store ZIP codes were recorded. The beverages figure ($9.43) met our minimum sample size only because of those readings, and the overall marketplace figures ($4.72 and $4.91) included them — about a quarter of their readings, lifting the figure by an estimated $0.08–0.09. Because July cannot be recomputed exactly under the rules it was published with, we have withdrawn these four figures rather than publish estimates. Indianapolis shelf-price figures for July are unaffected and remain published. The ZIP-code defect was fixed on 2026-09-18.
2026-09-20 — Marketplace figures withdrawn for Raleigh (July 2026), Phoenix (May–July 2026), Jacksonville (July 2026) and New York (July 2026). Some marketplace prices are collected through a shared shopping service that, when it cannot place a search at the shopper's location, answers from a single default store. Several of those stores were recorded under the ZIP code of whoever searched last, so their prices were counted in metros they are not in: a New Jersey supermarket supplied 97% of Raleigh's July marketplace readings, and a Massachusetts store appeared in the Phoenix and New York samples. Separately, one Jacksonville store's record carried different prices for the same product on the same day, so its readings cannot be attributed to that store. We withdrew every figure that met our minimum sample only because of those readings, every figure whose value they changed by 5% or more, and those where they made up 30% or more of the sample — 39 figures in all, listed individually in the corrections dataset. Shelf-price figures, and every other metro, are unaffected. The ZIP-recording defect was fixed on 2026-09-18; the next monthly release is on hold until checks that keep such prices out of the index are in place.