Industrias12 min read

Cannabis Data in the United States: What Exists and What Decisions It Supports

Benjamin Arjona
Benjamin Arjona

September 25, 2026

Cannabis Data in the United States: What Exists and What Decisions It Supports

The U.S. cannabis market is no longer growing by opening new stores. It is growing by competing for every customer on the shelf. In June 2026, three key figures made that clear: only 13% of 6 million products had an active promotion, flower prices ranged from $12.00 in Oregon to $60.00 in Utah, and 73% of dispensaries limited their visibility by publishing their menu on a single platform.

Translated into boardroom terms: growth by expansion is over, and what remains is winning share. That requires knowing what the store across the street charges, what it is promoting, and which categories are moving, at a frequency no quarterly report can deliver.

That information exists, but it is fragmented and scattered across dozens of state portals and the dynamic menus of thousands of dispensaries, with different formats, criteria, and cadences, where today's prices change tomorrow and are lost for good.

To answer that with numbers rather than assumptions, at AUTOScraping we built a proprietary dataset from the public menus of Dutchie, Leafly, and Weedmaps in the United States, consolidating 8,424 store records and validating every field. The result is an exclusive asset with no intermediaries: a proprietary dataset of 6,076 unique dispensaries across 43 states and the District of Columbia, along with 6 million SKUs with real prices. It is not a generic report or a public download; it is the foundation behind every decision that follows.

This guide walks through the cannabis data sources available in the U.S., which decision each one supports, and, above all, which problems only surface once the data is on the table.

The four types of sources

Source typeWhat it containsWhich decision it supportsAccessHow we handle it
Dispensary menusList price, category, brand, weight, declared potency, promotionsPricing, assortment, competitive positioningPublic web, requires extractionProprietary dataset: 6 million SKUs from 6,076 dispensaries, deduplicated and normalized, with recurring refresh
State open dataLicenses, aggregate sales, average prices, lab resultsWhere to enter, where to exit, how to size a marketOfficial portals, sometimes downloadableOn-demand extraction: direct download where CSV is available, dashboard and PDF scraping where it is not, and matching to menus by name and coordinates
Certificates of analysisPotency and contaminants per batchRegulatory exposure and supplier controlPDFs on brand and lab websitesOn-demand extraction: PDF parsing and unit normalization to a common scale, scoped to the brands in the client's catalog
Seed-to-sale systemsProduct movement from cultivation to saleNot applicable from the outsideRestricted to licenseesNot extracted. If the client is a licensee, their own data is integrated with the other three layers

The first two support most analytics projects. The third is specialized and becomes critical when the rules change. The fourth is often misunderstood, and it is worth explaining why. But the practical difference between the first and the rest is that the other three can be queried tomorrow; today's menus cannot.

Dispensary menus: the most important source, and the one nobody archives

Every dispensary publishes its full menu online, with products, brands, prices at each weight tier, active promotions, and declared potency. It does so because it needs consumers to compare before deciding where to buy.

Those menus live on a handful of platforms, including Dutchie, Leafly, Weedmaps, and Jane, plus dispensary-owned websites. The first three account for the 6 million SKUs in the dataset: Weedmaps contributes 51%, Dutchie 36%, and Leafly 13%. Anyone building a competitive map from a single platform sees less than half the market. And anyone stacking listings from several platforms without deduplicating them inflates it.

Which decision it supports. It is the only source that contains the actual price at which a product is selling in a specific store. Without it, prices are set against an average that describes no real market.

Our dataset shows this at the SKU level: the median price of a flower product on Dutchie and Leafly ranges from $12.00 in Oregon and $15.00 in Michigan to $59.00 in New Jersey and $60.00 in Utah, with California at $42.00 and Florida at $50.00. A national average of that series lands in the middle and is useless at either extreme. And the price level is not uniform across platforms either: the median list price is $24.99 on Dutchie and Weedmaps, and $29.00 on Leafly.

What makes it urgent. It is the only source that disappears. The menu updates and the previous state is not stored anywhere. There is no official repository, no vendor that sells it, and no way to rebuild it.

State regulators' open data portals

Several states publish more information than the industry uses, but the format determines how usable it is. Massachusetts delivers ready-to-consume files. Washington, Oregon, and New York publish queryable datasets on open data portals. California and Colorado offer dashboards built for viewing, not for processing. Michigan publishes documents.

That difference defines the strategy: where there is a CSV you download, where there is a dashboard you check whether it exposes the underlying data, and where there is a PDF you extract.

These are the seven largest or most representative programs. The rest fall within the same range, so a state's absence here does not mean it publishes nothing.

StateWhat it publishesFormatCadence
MassachusettsLicenses, sales by establishment, price per ounce, cultivation, lab testingCSV and JSONWeekly to monthly
WashingtonLiquor and Cannabis Board Marijuana DashboardOpen data portalPeriodic
OregonHarvest, prices, sales, licensesOpen data portalMonthly
New YorkActive OCM licenses, cannabinoid hempOpen data portalPeriodic
CaliforniaLicenses, harvests, sales, and unit priceFour DCC dashboardsMonthly
ColoradoQuarterly Market Update, regulatory activityDashboard and PDFQuarterly
MichiganCannabis Regulatory Agency reportsPDF and DOCXMonthly

Which decision it supports. These sources answer market-sizing and market entry or exit questions: how many active licenses there are, how much the state sold, how the average price moved. This is the layer for deciding which state gets the next store, and which one to pull out of.

Its limit. It arrives aggregated and delayed, from weeks to a quarter depending on the state. It is useful for understanding a market, not for operating within it. And it does not tell you what our dataset does: Michigan has fewer stores than California but more SKUs in total, at 1,871 products per dispensary, while Florida averages 297. An assortment benchmark that does not distinguish between those two markets is comparing different things.

Certificates of analysis: potency data scattered across thousands of PDFs

Every batch that reaches the shelf has a lab analysis behind it covering potency, cannabinoid profile, and contaminant screening. Those certificates are usually published, often through a code on the packaging.

It is a source that is rich and inconvenient in equal measure. The certificates live in PDFs hosted on dozens of brand and lab websites, each with its own layout, and each lab reports in different units: some in percentages, others in milligrams per gram, others in milligrams per package. Without normalizing to a common unit there is no possible comparison, neither between products nor against any regulatory threshold.

It pays to know what the menu covers before going to the certificate. In our dataset, declared potency (THC and CBD) is present in 74% of SKUs, with Dutchie at 81.5%, Leafly at 74.1%, and Weedmaps at 68.8%. That is enough to read market trends; it is not enough to audit a full catalog. For that, you need the certificate.

Which decision it supports. Quantifying exposure when a rule changes. When a potency threshold is modified, the question that reaches the executive team is how many SKUs in the catalog fall out of compliance and how much revenue they represent. Without normalized certificates, that answer is an estimate. With them, it is a calculation.

Seed-to-sale systems: why they are not a public source

Metrc and equivalent systems record the movement of every plant and every product from cultivation to sale. But they are not public: Metrc operates a managed-access API for integrations, designed so licensees can connect their systems to the state program, not for open queries.

What is public is what some regulators derive from those systems: the Colorado and Oregon dashboards are fed by Metrc, but what reaches the public is aggregated, not the transaction-level record. In practice, seed-to-sale is not part of the set of sources you can work with from the outside, and it does not need to be.

Which question each source answers

QuestionAppropriate sourceWhat you need to know first
How many real competitors do we have in each state?Menus, deduplicated across platformsWithout deduplication, the count is 28% higher
Do our prices track the current market or last quarter's?MenusPrice fill rate is 100%; compare by state, not against the national average
Are our brands gaining or losing shelf space?MenusBrand fill rate is 71% consolidated, 47% on Weedmaps
When and how deeply do competitors discount?Menus, daily seriesIt is a snapshot of the 13% on promotion; the pattern requires history
Which competitor products sell out?Menus, daily seriesStock does not exist as a field; it is inferred from absence
Which state makes sense for the next opening?State open data, cross-referenced with menusThe match is by name and coordinates, not by license number
Is the market expanding or consolidating?State open dataArrives with a lag of weeks to a quarter
What is our exposure if a potency threshold changes?Certificates of analysisDeclared THC on menus covers 74%

In short: state open data answers questions about the market, and menus answer questions about the competition. Executive decisions almost always need both layers, and the second is the one nobody else has.

Three ways to get the data, and only one that gives you an edge

There are three ways to reach the information in this guide. All three are legitimate and all three have their place. But only one produces something your competitors do not have, and it is worth being clear about that before allocating budget.

Download what is already published. This is the right route for everything a regulator delivers as CSV, and the most common mistake is building a pipeline to obtain something Massachusetts already publishes as a download. What you get is the market baseline: how many licenses, how much the state sold, how the average price moved. It is necessary information and, by definition, the same information anyone who knows where the portal is already has.

Buy an industry report. Clear methodology, broad coverage, no internal work. The limit is exclusivity: it is the same data your competitors have, with the vendor's scope rather than yours. It works as a baseline, not as an advantage.

Build your own series. This is the only route that produces information nobody else has, because the scope is your own: your real competitors (which are almost never "everyone in the state" but the six within fifteen miles), the categories that matter to the business, SKU-level granularity, and the frequency the decision demands. And it is the only one where the value grows over time. A menu captured today is worth little; the same menu captured every day for a year is each competitor's discount calendar, the list of products that sold out and came back, and the price history no purchased report will ever include, because nobody was saving it.

Conclusion

Cannabis in the U.S. is an industry where information is abundant and poorly distributed. Knowing what each source contains, in what format, and with what lag avoids two expensive mistakes: launching an extraction project for something already published as CSV, and assuming a state report will answer a question that only a competitor's menu can resolve.

Three practical criteria follow from this. Always check whether the data is already available for download before building anything. Use open data to size the market and menus to decide where to compete. And start the historical menu series as early as possible, because it is the only data on this list that cannot be recovered later: state reports will still be available next year; today's menu will not.

This map gets updated because the sources change too. At AUTOScraping we track those changes closely and turn them into usable data. You can learn more about that process at Data Factory.

If you want to see what your market looks like under these same criteria, get in touch.


Sources

Keep reading

Benjamin Arjona

Written by

Benjamin Arjona

Hace más de 10 años que trabajo con datos web. Si hay algo que aprendí es esto: las empresas que ganan no son las que tienen más información, son las que la tienen primero. Soy co-founder de AUTOScraping, la empresa que armamos con Francisco Battan y Cesar Farhat desde Santiago del Estero. Hoy trabajamos con compañías en USA, Europa y LATAM, y cada día estoy más convencido de que construir desde acá es una ventaja.

Stay in the loop

Web scraping tips, industry news and use cases — weekly, no spam.

Share this article

Did you find it useful?

Related Articles

More from the same category