Industrias13 min read

Competitor Dispensary Prices: How to Design the Dataset Without Crossing the Regulator

Benjamin Arjona
Benjamin Arjona

September 25, 2026

Competitor Dispensary Prices: How to Design the Dataset Without Crossing the Regulator

The question always comes up in the same meeting. The data team proposes tracking competitors' prices, someone from management asks whether that is legal in the state, and the project is put on hold until somebody checks with a lawyer. Meanwhile, the dispensary across the street changes the price of its ounce and nobody records it.

The answer is not in a legal opinion. It is in how the dataset is designed: which sources go in, which fields are captured, what is deliberately left out, how a store that appears on three platforms is unified, and how often each menu is queried. Those five decisions, taken well, put the project on solid ground. Taken badly, no legal framework fixes it.

We did not write this guide from assumptions. In June 2026 we collected the public menus of Dutchie, Leafly and Weedmaps across the whole of the United States and built a proprietary AUTOScraping dataset with 8,424 store rows, 6,076 unique dispensaries in 30 states and 6,001,359 SKUs with a list price. Every statement that follows about which field exists, on which platform, how complete it is and what it takes to use it comes from that dataset.

What state regulations cover, and what they do not

State cannabis regulation is dense, but almost all of it points at two things: how you sell and what you publish. Age verification, product claims, where an advert may appear, how a discount is communicated. All of that reaches your sales operation and your communications to the consumer.

Reading another dispensary's public menu does not fall into that category. Consumers do it every day when they compare before leaving home, and the managers of your own stores do it before setting Friday's price. The menu is published precisely for that.

From there comes the operating line worth fixing at the design stage, the one that settles most of the doubts: the dataset feeds the internal dashboard, not the public website. Using a competitor's price to decide your own is market intelligence. Publishing a price comparison on your website is advertising, and cannabis advertising has its own state-by-state rules that your marketing team and your lawyer already handle. The boundary is easy to respect: the data is there to decide, not to communicate.

With that boundary clear, the rest is design. And design comes down to five rules.

1. Public sources only, and public sources are enough

The rule is binary, which is why it works: if seeing a piece of data requires logging in, it does not enter the project. Nothing behind a login, no seed-to-sale systems, no third-party point-of-sale systems. What it takes to compete is on the menus any consumer opens in a browser.

The next question is whether those public sources cover the market. The answer, measured across the whole of the United States, is yes, but only if the three main ones are used together:

SourceStore rows collectedUnique dispensaries coveredStores found only on this sourcePriced SKUs
Weedmaps3,8533,6852,2773.08 million
Dutchie3,2723,0551,6632.16 million
Leafly1,2991,2215020.76 million

Of the 6,076 unique dispensaries, 4,442 (73%) publish a menu on a single platform, and only 251 are on all three. Weedmaps contributes the most exclusive stores; Leafly the fewest, but its 502 exclusive stores are nowhere else. A series built on a single platform is not a series of the market: it is a series of that platform, and it will have gaps exactly where the competitor that matters most to you is.

The dataset serves that purpose from day one: with address and coordinates at 100%, the dispensaries within 10, 25 or 40 kilometres of each of your own stores can be listed, and the source each one appears on can be seen.

2. Only the fields the analysis needs, and knowing which ones exist

A well-designed competitive pricing dataset has few fields. The fields that serve for decisions come complete in public menus. The ones that do not come are not invented.

This is the real completeness of each field across the 6 million SKUs collected, and the place each one deserves in the design:

FieldDutchieLeaflyWeedmapsConsolidatedPlace in the dataset
List price100%100%100%100%In. It is the central column
Category100%100%100%100%In, with taxonomy mapping across platforms
Address and coordinates99% to 100%100%100%100%In. Defines the competitive radius
Weight (published field)89%64%100%92%In, with enrichment from the product name: reaches 94%
Brand96%95%47%71%In, with a coverage note: Weedmaps omits it on more than half of its SKUs
THC / CBD82%74%69%74%In as secondary segmentation, with a note
Active discount14% of SKUsNot exposed16% of SKUs13%In as a field separate from list price. Only has value if captured daily
Store ratingNot exposed92%97%62%Optional. It belongs to the store, not the product
StockDoes not existDoes not existDoes not existDoes not existOut. No platform publishes it
Licence numberDoes not existDoes not existDoes not existDoes not existOut. No source exposes it
Personal dataNot capturedNot capturedNot capturedNot capturedOut by design

Two readings come out of that table, and neither is the one that people who stall the project tend to expect.

  1. The four fields that support any pricing decision (price, category, location and weight) come complete or almost complete in the public source. There is no need to go looking for anything anywhere else. With price and weight normalised, 94% of SKUs are expressed in dollars per gram, which is the only unit in which an eighth, a quarter and an ounce can be compared.
  2. Two of the fields people usually ask for do not exist. None of the three platforms publishes stock, and none exposes a licence number. A vendor promising competitor inventory or licence cross-references from public menus is promising something the source does not have. Knowing that in advance avoids buying hot air, and also avoids someone going where they should not in order to get that data.

3. What is left out on purpose

The design is defined as much by what it includes as by what it decides not to include. What is left out, even where it would be technically possible:

  • Any field that identifies a person. It is not captured, not even "just in case". A dataset with no personal data cannot have a personal-data incident, and that peace of mind is worth more than any extra column.
  • Any content behind a login. The menus needed are public; there is no reason to go further.
  • Seed-to-sale systems and third-party points of sale. They are not public sources and they add nothing the menu does not already have: list price and active promotion are already published at 100% and 13% respectively.
  • Your own non-public prices, if someone offers you an "industry benchmark" in exchange for adding them to a shared pool. Your series is yours. Monitoring what is public has nothing to do with handing over what is private, and it is worth keeping the two conversations separate.

4. One outlet, one row

This is the problem nobody anticipates and that distorts every count. The same dispensary usually publishes a menu on more than one platform. Of the 8,424 store rows collected, 2,348 (27.9%) were the same outlet repeated on another source. Without unifying them, an analyst counting competitors in their radius inflates them by more than a quarter, and every price from those stores is counted two or three times in any median.

Since no platform exposes a licence number, unification is done with three signals within the same jurisdiction: coordinates within 150 metres, identical normalised address, and identical normalised name plus city. A store with no coordinates is joined only by address or by name. With that method, the 8,424 rows are reduced to 6,076 unique dispensaries.

The result changes concrete readings. Michigan has 582 dispensaries with a public menu, not the 838 retail licences the state reports, because the dataset counts what the consumer can see, not what the regulator has on file. And the typical size of a store, measured in published SKUs, is 707 (median), with the largest 10% above 2,147 and a maximum of 14,751. Without unification those numbers mean nothing, because a store with 1,000 SKUs present on three platforms shows up as three stores with 3,000 SKUs.

5. Pace, volume and geography

Query volume is not a field in the dataset, but it is a design decision with consequences. A daily capture of the forty menus in a radius is irrelevant traffic for any platform. Thousands of requests per minute are something else, and at that point what matters is not any framework: it is that you are degrading a third party's service. Reasonable frequency, controlled pace.

To size it with numbers from the dataset: a full menu weighs around 2.5 MB, so the entire United States universe, 6,076 stores, comes to roughly 15 GB per full run. With 25 parallel processes and about 10 seconds per store, a national pass is projected at under an hour. It is a design estimate, not a timed run, but it makes clear that a daily refresh of the whole country is feasible, and that a daily refresh of a forty-store radius is trivial.

The daily pace is not a luxury. The only field in the dataset that loses value if it is not captured every day is the active discount: the 13% of SKUs on promotion we saw in June 2026 is a snapshot of that day. The series of that 13%, captured every day, is what says which competitor discounts, how deeply, in which category and on which day of the week.

Geography is the other point that does not fail with an error message. US dispensary menus show different things depending on where they are viewed from, and a server in another state does not see the same thing as a customer standing in Grand Rapids. The result is not a visible error: it is data that loads perfectly into the dashboard and describes another market. Capturing from the right geography is the difference between seeing the menu your customer sees and seeing a different one.

Which decisions a dataset designed this way enables

With those five rules applied, the dataset ends up small, public and complete. And that is enough to change four decisions that most dispensaries today make on intuition:

DecisionWhich field feeds itWhat the dataset showsWhat is left out
Which format level to fight atPrice and weight normalised to dollars per gram (94% of SKUs)Nationally, a gram of flower averages $8.94 and an ounce $48.10, that is, $1.72 per gram: five times less for the same flower. Median flower price per SKU runs from $12 in Oregon to $59 in New JerseyNothing. Price and weight are public
When and how much to discountActive discount, separate from list price, captured daily13% of SKUs with an active promotion on any given day; the sector's discount rate is already 26% of value soldLeafly does not expose promotions: the discount series is built from Dutchie and Weedmaps
How much inventory to buyFormat and category mix per competitorWhen competitors push ounces and half ounces to the front of the menu, it is rarely a bet on the volume buyer: it is inventory that is not turning. On Dutchie, pre-rolls are already the category with the most listings (22%)Stock: no platform publishes it. The oversupply signal is read in price and format, not in inventory
When to worryTime series of median price by category and radiusA price falling 3% per quarter is a maturing market; one falling 3%, then 5%, then 9%, is a market coming apart. The difference is not in today's price but in the seriesNothing, except time: the menu from six months ago exists in no archive if nobody kept it

The four decisions have something in common: none is solved by looking at a competitor's menu once. All four need a continuous, normalised series with no gaps. And that is the point that makes everything above urgent: a competitor's menu from six months ago cannot be bought, rebuilt or recovered. Either it was captured on the day it was live, or that data is lost for good.

Conclusion

Monitoring competitor dispensary prices is not a legal problem waiting for a ruling. It is a design problem, solved before the first line of code is written. The data that serves to compete (price, category, location, weight, promotion) is published on menus anyone can read, comes complete or almost complete on the three main platforms, and requires touching nothing private. The data people fear capturing (persons, logins, internal systems) is not needed, and the data they tend to ask for without knowing it does not exist (stock, licences) is in no public source.

What does demand work is what almost nobody anticipates: covering the three platforms because 73% of stores are on only one; unifying the 27.9% of duplicate rows so the same competitor is not counted three times; normalising weight and taxonomies so that prices are comparable; capturing from the right geography; and doing it every day, because today's discount leaves no trace tomorrow.

A dataset designed with those five rules is small, public, complete and defensible. And it is the only one that makes it possible to answer, with data rather than intuition, which format to compete in, when to discount, how much to buy and when to worry.

Start with your competitive radius

If you work with data at a dispensary or a chain, we can build you a map of your competitive radius from the dataset: which dispensaries with a public menu there are within 10, 25 or 40 kilometres of each of your stores, which platform each one publishes on, how many SKUs it has and which fields come complete for your market. It is the input for designing your dataset with the five rules in this article before capturing a single price.

Request your radius map and tell us which state you operate in. The answer is data, not a slide deck.

Sources

Benjamin Arjona

Written by

Benjamin Arjona

Hace más de 10 años que trabajo con datos web. Si hay algo que aprendí es esto: las empresas que ganan no son las que tienen más información, son las que la tienen primero. Soy co-founder de AUTOScraping, la empresa que armamos con Francisco Battan y Cesar Farhat desde Santiago del Estero. Hoy trabajamos con compañías en USA, Europa y LATAM, y cada día estoy más convencido de que construir desde acá es una ventaja.

Stay in the loop

Web scraping tips, industry news and use cases — weekly, no spam.

Share this article

Did you find it useful?

Related Articles

More from the same category