The question always comes up in the same meeting. The data team proposes tracking competitors' prices, someone from management asks whether that is legal in the state, and the project is put on hold until somebody checks with a lawyer. Meanwhile, the dispensary across the street changes the price of its ounce and nobody records it.
The answer is not in a legal opinion. It is in how the dataset is designed: which sources go in, which fields are captured, what is deliberately left out, how a store that appears on three platforms is unified, and how often each menu is queried. Those five decisions, taken well, put the project on solid ground. Taken badly, no legal framework fixes it.
We did not write this guide from assumptions. In June 2026 we collected the public menus of Dutchie, Leafly and Weedmaps across the whole of the United States and built a proprietary AUTOScraping dataset with 8,424 store rows, 6,076 unique dispensaries in 30 states and 6,001,359 SKUs with a list price. Every statement that follows about which field exists, on which platform, how complete it is and what it takes to use it comes from that dataset.
What state regulations cover, and what they do not
State cannabis regulation is dense, but almost all of it points at two things: how you sell and what you publish. Age verification, product claims, where an advert may appear, how a discount is communicated. All of that reaches your sales operation and your communications to the consumer.
Reading another dispensary's public menu does not fall into that category. Consumers do it every day when they compare before leaving home, and the managers of your own stores do it before setting Friday's price. The menu is published precisely for that.
From there comes the operating line worth fixing at the design stage, the one that settles most of the doubts: the dataset feeds the internal dashboard, not the public website. Using a competitor's price to decide your own is market intelligence. Publishing a price comparison on your website is advertising, and cannabis advertising has its own state-by-state rules that your marketing team and your lawyer already handle. The boundary is easy to respect: the data is there to decide, not to communicate.
With that boundary clear, the rest is design. And design comes down to five rules.
1. Public sources only, and public sources are enough
The rule is binary, which is why it works: if seeing a piece of data requires logging in, it does not enter the project. Nothing behind a login, no seed-to-sale systems, no third-party point-of-sale systems. What it takes to compete is on the menus any consumer opens in a browser.
The next question is whether those public sources cover the market. The answer, measured across the whole of the United States, is yes, but only if the three main ones are used together:
| Source | Store rows collected | Unique dispensaries covered | Stores found only on this source | Priced SKUs |
|---|---|---|---|---|
| Weedmaps | 3,853 | 3,685 | 2,277 | 3.08 million |
| Dutchie | 3,272 | 3,055 | 1,663 | 2.16 million |
| Leafly | 1,299 | 1,221 | 502 | 0.76 million |
Of the 6,076 unique dispensaries, 4,442 (73%) publish a menu on a single platform, and only 251 are on all three. Weedmaps contributes the most exclusive stores; Leafly the fewest, but its 502 exclusive stores are nowhere else. A series built on a single platform is not a series of the market: it is a series of that platform, and it will have gaps exactly where the competitor that matters most to you is.
The dataset serves that purpose from day one: with address and coordinates at 100%, the dispensaries within 10, 25 or 40 kilometres of each of your own stores can be listed, and the source each one appears on can be seen.
2. Only the fields the analysis needs, and knowing which ones exist
A well-designed competitive pricing dataset has few fields. The fields that serve for decisions come complete in public menus. The ones that do not come are not invented.
This is the real completeness of each field across the 6 million SKUs collected, and the place each one deserves in the design:
| Field | Dutchie | Leafly | Weedmaps | Consolidated | Place in the dataset |
|---|---|---|---|---|---|
| List price | 100% | 100% | 100% | 100% | In. It is the central column |
| Category | 100% | 100% | 100% | 100% | In, with taxonomy mapping across platforms |
| Address and coordinates | 99% to 100% | 100% | 100% | 100% | In. Defines the competitive radius |
| Weight (published field) | 89% | 64% | 100% | 92% | In, with enrichment from the product name: reaches 94% |
| Brand | 96% | 95% | 47% | 71% | In, with a coverage note: Weedmaps omits it on more than half of its SKUs |
| THC / CBD | 82% | 74% | 69% | 74% | In as secondary segmentation, with a note |
| Active discount | 14% of SKUs | Not exposed | 16% of SKUs | 13% | In as a field separate from list price. Only has value if captured daily |
| Store rating | Not exposed | 92% | 97% | 62% | Optional. It belongs to the store, not the product |
| Stock | Does not exist | Does not exist | Does not exist | Does not exist | Out. No platform publishes it |
| Licence number | Does not exist | Does not exist | Does not exist | Does not exist | Out. No source exposes it |
| Personal data | Not captured | Not captured | Not captured | Not captured | Out by design |
Two readings come out of that table, and neither is the one that people who stall the project tend to expect.
- The four fields that support any pricing decision (price, category, location and weight) come complete or almost complete in the public source. There is no need to go looking for anything anywhere else. With price and weight normalised, 94% of SKUs are expressed in dollars per gram, which is the only unit in which an eighth, a quarter and an ounce can be compared.
- Two of the fields people usually ask for do not exist. None of the three platforms publishes stock, and none exposes a licence number. A vendor promising competitor inventory or licence cross-references from public menus is promising something the source does not have. Knowing that in advance avoids buying hot air, and also avoids someone going where they should not in order to get that data.
3. What is left out on purpose
The design is defined as much by what it includes as by what it decides not to include. What is left out, even where it would be technically possible:
- Any field that identifies a person. It is not captured, not even "just in case". A dataset with no personal data cannot have a personal-data incident, and that peace of mind is worth more than any extra column.
- Any content behind a login. The menus needed are public; there is no reason to go further.
- Seed-to-sale systems and third-party points of sale. They are not public sources and they add nothing the menu does not already have: list price and active promotion are already published at 100% and 13% respectively.
- Your own non-public prices, if someone offers you an "industry benchmark" in exchange for adding them to a shared pool. Your series is yours. Monitoring what is public has nothing to do with handing over what is private, and it is worth keeping the two conversations separate.
4. One outlet, one row
This is the problem nobody anticipates and that distorts every count. The same dispensary usually publishes a menu on more than one platform. Of the 8,424 store rows collected, 2,348 (27.9%) were the same outlet repeated on another source. Without unifying them, an analyst counting competitors in their radius inflates them by more than a quarter, and every price from those stores is counted two or three times in any median.
Since no platform exposes a licence number, unification is done with three signals within the same jurisdiction: coordinates within 150 metres, identical normalised address, and identical normalised name plus city. A store with no coordinates is joined only by address or by name. With that method, the 8,424 rows are reduced to 6,076 unique dispensaries.
The result changes concrete readings. Michigan has 582 dispensaries with a public menu, not the 838 retail licences the state reports, because the dataset counts what the consumer can see, not what the regulator has on file. And the typical size of a store, measured in published SKUs, is 707 (median), with the largest 10% above 2,147 and a maximum of 14,751. Without unification those numbers mean nothing, because a store with 1,000 SKUs present on three platforms shows up as three stores with 3,000 SKUs.
5. Pace, volume and geography
Query volume is not a field in the dataset, but it is a design decision with consequences. A daily capture of the forty menus in a radius is irrelevant traffic for any platform. Thousands of requests per minute are something else, and at that point what matters is not any framework: it is that you are degrading a third party's service. Reasonable frequency, controlled pace.
To size it with numbers from the dataset: a full menu weighs around 2.5 MB, so the entire United States universe, 6,076 stores, comes to roughly 15 GB per full run. With 25 parallel processes and about 10 seconds per store, a national pass is projected at under an hour. It is a design estimate, not a timed run, but it makes clear that a daily refresh of the whole country is feasible, and that a daily refresh of a forty-store radius is trivial.
The daily pace is not a luxury. The only field in the dataset that loses value if it is not captured every day is the active discount: the 13% of SKUs on promotion we saw in June 2026 is a snapshot of that day. The series of that 13%, captured every day, is what says which competitor discounts, how deeply, in which category and on which day of the week.
Geography is the other point that does not fail with an error message. US dispensary menus show different things depending on where they are viewed from, and a server in another state does not see the same thing as a customer standing in Grand Rapids. The result is not a visible error: it is data that loads perfectly into the dashboard and describes another market. Capturing from the right geography is the difference between seeing the menu your customer sees and seeing a different one.
Which decisions a dataset designed this way enables
With those five rules applied, the dataset ends up small, public and complete. And that is enough to change four decisions that most dispensaries today make on intuition:
| Decision | Which field feeds it | What the dataset shows | What is left out |
|---|---|---|---|
| Which format level to fight at | Price and weight normalised to dollars per gram (94% of SKUs) | Nationally, a gram of flower averages $8.94 and an ounce $48.10, that is, $1.72 per gram: five times less for the same flower. Median flower price per SKU runs from $12 in Oregon to $59 in New Jersey | Nothing. Price and weight are public |
| When and how much to discount | Active discount, separate from list price, captured daily | 13% of SKUs with an active promotion on any given day; the sector's discount rate is already 26% of value sold | Leafly does not expose promotions: the discount series is built from Dutchie and Weedmaps |
| How much inventory to buy | Format and category mix per competitor | When competitors push ounces and half ounces to the front of the menu, it is rarely a bet on the volume buyer: it is inventory that is not turning. On Dutchie, pre-rolls are already the category with the most listings (22%) | Stock: no platform publishes it. The oversupply signal is read in price and format, not in inventory |
| When to worry | Time series of median price by category and radius | A price falling 3% per quarter is a maturing market; one falling 3%, then 5%, then 9%, is a market coming apart. The difference is not in today's price but in the series | Nothing, except time: the menu from six months ago exists in no archive if nobody kept it |
The four decisions have something in common: none is solved by looking at a competitor's menu once. All four need a continuous, normalised series with no gaps. And that is the point that makes everything above urgent: a competitor's menu from six months ago cannot be bought, rebuilt or recovered. Either it was captured on the day it was live, or that data is lost for good.
Conclusion
Monitoring competitor dispensary prices is not a legal problem waiting for a ruling. It is a design problem, solved before the first line of code is written. The data that serves to compete (price, category, location, weight, promotion) is published on menus anyone can read, comes complete or almost complete on the three main platforms, and requires touching nothing private. The data people fear capturing (persons, logins, internal systems) is not needed, and the data they tend to ask for without knowing it does not exist (stock, licences) is in no public source.
What does demand work is what almost nobody anticipates: covering the three platforms because 73% of stores are on only one; unifying the 27.9% of duplicate rows so the same competitor is not counted three times; normalising weight and taxonomies so that prices are comparable; capturing from the right geography; and doing it every day, because today's discount leaves no trace tomorrow.
A dataset designed with those five rules is small, public, complete and defensible. And it is the only one that makes it possible to answer, with data rather than intuition, which format to compete in, when to discount, how much to buy and when to worry.
Start with your competitive radius
If you work with data at a dispensary or a chain, we can build you a map of your competitive radius from the dataset: which dispensaries with a public menu there are within 10, 25 or 40 kilometres of each of your stores, which platform each one publishes on, how many SKUs it has and which fields come complete for your market. It is the input for designing your dataset with the five rules in this article before capturing a single price.
Request your radius map and tell us which state you operate in. The answer is data, not a slide deck.




