Find a buyer

Methods from the OECD and Statistics Canada

Data valuation: three methods and real deal prices

Data valuation puts a dollar figure on a dataset, either to sell or license it or to count it as an asset. Here are the three accepted methods, why the usual cost-plus math breaks for data, and the real prices AI buyers have paid.

Analyst at two monitors of charts with a printed report on the desk

The methods

Three data valuation methods

These come from national accounting rules (SNA 2008) as summarized by the OECD. Each data valuation method answers a different question, so many valuations use more than one.

1. Market-based

Value the data at the price comparable data has traded for. It's the most direct method, but the OECD notes only a small share of data trades on standard terms, so good comparables are rare.

2. Cost-based

Add up what it cost to collect, clean and store the data. Statistics Canada used this approach to estimate CAD 9 billion to 14 billion of investment in data in 2018. It sets a floor, not a sale price.

3. Income-based

Estimate the future income the data will earn, such as license fees, and discount it to today's value. The OECD lists this net present value method as the other accepted fallback.

Why cost-plus fails

Data can be copied at almost no cost, so the OECD says it can't be priced at marginal cost plus a markup. Sellers price by the value to each buyer, which can differ widely for the same dataset.

Anchors

Real prices AI buyers have paid

Comparable deals are the market-based method in practice. These figures come from filings or are labeled as reported.

DataPriceSource
Platform posts (Reddit)$203.0M across January 2024 deals, 2–3 yearsReddit S-1 (filing)
Academic content (Taylor & Francis to Microsoft)$10M+ initial fee plus recurring paymentsInforma (filing)
Business records (HVAC companies)$150,000The Information (reported)
Nonfiction books (HarperCollins)$5,000 per title, $2,500 to the authorMusic Ally (reported)
Photos (Photobucket)$0.05 to $1 per photo; more per videoPetaPixel (reported)

What moves the price

Factors buyers weigh in data valuation

FileYield lists the factors it sees in negotiations: exclusivity, permitted use, freshness, quality, provenance and the cost of reproducing the dataset. Each one changes what a buyer will pay for the same rows, which is why data valuation starts with the buyer.

  • Exclusivity. A buyer pays more to keep data from competitors.
  • Permitted use. Training, display or both, and for how long.
  • Freshness. Ongoing feeds can earn recurring fees, as Reddit's do.
  • Provenance and consent. Clear rights cut the buyer's legal risk.
  • Replacement cost. Data that's hard to recreate is worth more.

To see where datasets like yours sell, compare the AI data marketplace options or the full list of places to sell data to AI companies.

Person browsing dataset listings on a laptop

Questions

Common questions

How much is my data worth?

It depends on who buys it and on what terms. Because data can be copied for almost nothing, buyers pay for what it's worth to them, not what it cost you. Exclusivity, freshness, provenance and permitted use all move the price. Comparable deals, like those on our AI data licensing deals page, are the best anchor.

What is the most common data valuation method?

When a market price exists, compare to it. Most data isn't traded on standard terms, so the OECD notes statisticians usually fall back on the cost of producing the data, or on the income it's expected to earn, discounted to today.

Why can't I price data at cost plus a margin?

The OECD points out that data can be copied at near-zero marginal cost, so cost-plus pricing doesn't work. Information goods are usually priced by their value to each buyer, which varies widely.