Find a buyer

Get your data ready to sell

Data anonymization before you sell your data

Data anonymization removes names, IDs and other personal details from a dataset before it leaves your company. These specialists de-identify, structure and curate data so AI buyers can license it with less risk.

Lawyer and executive reviewing a printed agreement

The specialists

14 data anonymization and prep specialists

Each one helps with data anonymization or another step that makes a dataset sellable: removing PII, parsing documents into clean records, or curating what's worth selling. Prices are shown where the company publishes them.

Limina logo
LiminaDe-identification
Companies removing PII from text, PDFs, images or audio
Free to start75 calls/day free; Batch plan quoted
Tonic.ai logo
Tonic.aiDe-identification
Teams that need realistic data without the real people in it
Pay as you goper 1,000 words (Textual)
Skyflow logo
SkyflowDe-identification
Companies tokenizing PII while keeping records linkable
Not publishedcontact sales
Datavant logo
DatavantDe-identification
Health data owners licensing records under HIPAA
IQVIA Privacy Analytics logo
IQVIA Privacy AnalyticsDe-identification
Health datasets that need a HIPAA Expert Determination
Not publishedquote; fit check in ~5 business days
Protecto logo
ProtectoDe-identification
Teams masking sensitive data before it reaches AI tools
BigID logo
BigIDDe-identification
Companies that need to find their sensitive data first
MOSTLY AI logo
MOSTLY AISynthetic data
Companies that want to share a synthetic copy of tables
Free tierusage-based enterprise; open-source SDK
Gretel (NVIDIA) logo
Gretel (NVIDIA)Synthetic data
Teams generating private synthetic tables
Not publishednow part of NVIDIA
Unstructured logo
UnstructuredDocument parsing
Owners of PDF, email and document archives
$0.015/pageafter 10,000 free pages
LlamaIndex (LlamaParse) logo
LlamaIndex (LlamaParse)Document parsing
Teams parsing messy multi-format documents
From $0/moStarter $50/mo, Pro $500/mo
Reducto logo
ReductoDocument parsing
Archives with tables, charts, scans or handwriting
$10/1,000 pagesParse, pay as you go
DatologyAI logo
DatologyAIData curation
Owners of large text corpora selling a curated set
Encord logo
EncordData curation
Image, video, audio and medical-imaging collections
Not publishedplans listed without prices

Methods

Four ways to make data safe to sell

The right data anonymization method depends on the data. Health records follow HIPAA's rules, while documents and logs usually need PII found and redacted first.

  1. Redact or maskFind names, emails, account numbers and faces, then remove them or swap in fake values. Limina and Tonic.ai do this across text, files and audio.
  2. TokenizeReplace each identifier with a consistent token so records still link up. Skyflow and Datavant work this way.
  3. Certify (health data)Under HIPAA, remove the 18 Safe Harbor identifiers or get an Expert Determination. IQVIA Privacy Analytics issues those opinions.
  4. Go syntheticSell generated records that keep the patterns but contain no real people. MOSTLY AI and NVIDIA's Gretel tools do this for tables.

What buyers ask

What buyers check before licensing

Buyers want proof of what was removed and how. micro1's data partnerships agree on the approved data scope, security standards and anonymization requirements before anything is shared.

HHS says health data de-identified by Safe Harbor or Expert Determination is no longer protected health information, so it can be used and shared. That makes the method you choose part of the product you sell.

  • A list of the fields removed or replaced.
  • The method used, and who certified it for health data.
  • Where the data came from and what consent covers it.

Then compare the companies that buy data or read how AI data licensing terms work.

Best for: Companies with operational data: SOPs, CRM records, project histories

  • Pays for operational data
  • Scope and redaction agreed first
  • $500M gross run rate (reported)

micro1 is a data lab that makes expert human data, RL environments and evaluations for frontier AI labs. Its data partnerships program licenses companies' operational data, such as SOPs, knowledge bases, CRM data, project histories and QA processes. The data scope, anonymization and internal approvals are agreed before anything is shared. TechCrunch reported in August 2026 that micro1 reached a $500 million gross run rate.

De-identification and privacy

FileYield logo

FileYield

Free to list Readiness check

Best for: Companies documenting PII status, structure and provenance before listing

  • Free to list
  • No file upload needed
  • Buyer outreach
  • 2,566 data types

FileYield's listing flow has a readiness step: document the data's quality, structure, provenance, PII status and the terms a buyer would need. Its homepage says you can flag PII and compliance requirements early, and that sellers and buyers remain responsible for legal review. Listing is free and the files stay with you.

Limina logo

Limina

Free to start De-identification

Best for: Companies removing PII from text, PDFs, images or audio

  • 50+ entity types, 52 languages
  • Formerly Private AI

Finds and removes names, contact details, health and payment identifiers in text, PDFs, images and audio before a dataset leaves the building. It can run inside the seller's own infrastructure, and its Batch plan is aimed at one-time dataset monetization.

Tonic.ai logo

Tonic.ai

Pay as you go De-identification

Best for: Teams that need realistic data without the real people in it

  • SOC 2 Type II, HIPAA
  • 100+ PB processed

Tonic Textual finds names, emails, addresses and account numbers in free text, files and audio, then redacts them or swaps in realistic fake values so the data still reads naturally. Tonic Structural does the same for database tables, so a seller can share realistic records without the real people in them.

Skyflow logo

Skyflow

Not published De-identification

Best for: Companies tokenizing PII while keeping records linkable

  • SOC 2 Type II, PCI DSS L1
  • Values kept in your vault

Replaces personal and business-sensitive values in training data with tokens, keeping the links between records intact so the cleaned data is still usable. The real values stay locked in a separate vault the seller controls.

Datavant logo

Datavant

Not published De-identification

Best for: Health data owners licensing records under HIPAA

  • Patient tokens link datasets
  • 2,000+ tokenized datasets

For health data: strips or tokenizes patient identifiers so records can be shared or licensed under HIPAA, and arranges Expert Determination reviews. Its tokens let a buyer link the same patient across datasets without seeing who the patient is.

Best for: Health datasets that need a HIPAA Expert Determination

  • Expert Determination
  • 200+ customers since 2007

Measures re-identification risk and issues a HIPAA Expert Determination opinion, so a health dataset can be sold with documented evidence that it is de-identified. It also anonymizes medical images (DICOM), clinical documents and free text.

Protecto logo

Protecto

Not published De-identification

Best for: Teams masking sensitive data before it reaches AI tools

  • 200+ sensitive data types
  • SOC 2, HIPAA BAA

Detects and masks personal, health and business-confidential data with tokens that keep the meaning of the text, so it can be passed to AI models or partners. Its main focus is data flowing into LLMs and agents, not one-off dataset delivery.

BigID logo

BigID

Not published De-identification

Best for: Companies that need to find their sensitive data first

  • Data discovery and labeling
  • Founded 2016

Scans a company's databases, cloud storage, SaaS apps and file shares to find and label which data holds personal or sensitive information. That inventory is the first step in deciding what can be sold and what has to be cleaned first.

Synthetic data

MOSTLY AI logo

MOSTLY AI

Free tier Synthetic data

Best for: Companies that want to share a synthetic copy of tables

  • Open-source SDK
  • Brand now part of Syntho

Trains on a seller's real tables and generates synthetic records that keep the statistical patterns without exposing real people. A company can then share or license the synthetic version in place of the raw data.

Best for: Teams generating private synthetic tables

  • NVIDIA NeMo Safe Synthesizer
  • Acquired 2025 (reported)

Gretel's site now forwards to NVIDIA, which offers NeMo Safe Synthesizer: it learns from a sensitive table, detects and replaces names and addresses, and outputs synthetic records with optional differential privacy. A seller can offer the synthetic copy instead of the real records.

Document parsing and structure

Unstructured logo

Unstructured

$0.015/page Document parsing

Best for: Owners of PDF, email and document archives

  • 65+ file types
  • SOC 2, HIPAA

Turns piles of PDFs, slides, emails, scans and spreadsheets into clean, structured JSON with tables and reading order preserved. This is the format AI buyers expect for document collections.

Best for: Teams parsing messy multi-format documents

  • LlamaParse
  • SOC 2 Type II

LlamaParse reads messy, multi-format documents and turns them into structured, machine-readable output. A seller uses it to make a document archive usable for AI training or retrieval.

Reducto logo

Reducto

$10/1,000 pages Document parsing

Best for: Archives with tables, charts, scans or handwriting

  • Citations to source page
  • 1B+ pages processed

Parses PDFs, scans, spreadsheets and slides, including tables, charts and handwriting, into structured JSON with citations back to the source page. It turns a document archive into clean records a buyer can load.

Curation and quality

Best for: Owners of large text corpora selling a curated set

  • Quality and topic scoring
  • $46M Series A

Cleans and filters large training corpora (removing badly formatted, empty, short and test-contaminated data), scores them by quality and topic, and mixes sources into training-ready sets. A seller with a large text corpus can use it to deliver a curated version instead of a raw dump.

Encord logo

Encord

Not published Data curation

Best for: Image, video, audio and medical-imaging collections

  • De-duplication, outliers
  • 5+ PB on platform

Organizes, searches, de-duplicates and labels large image, video, audio, document and medical-imaging collections, with outlier detection and quality checks. A seller uses it to turn raw footage or files into a labeled, documented dataset.

Questions

Common questions

What is the difference between anonymization and de-identification?

De-identification removes or replaces identifiers, and is the term US health law uses. Anonymization usually means the data can't reasonably be traced back to a person at all. Synthetic data goes further: it creates new records that keep the patterns but contain no real people.

What does HIPAA require to de-identify health data?

HHS recognizes two methods. Safe Harbor removes 18 listed identifiers, from names to device serial numbers, for the person and their relatives, employers and household members. Expert Determination has a qualified expert certify that the risk of re-identification is very small. Data de-identified either way is no longer protected health information.

Do AI data buyers require anonymized data?

Buyers set their own rules, but the scope is usually agreed up front. micro1, for example, defines the approved data, anonymization and redaction requirements before a company shares anything. FileYield's listing asks sellers to state PII status.

How much does data anonymization cost?

Few specialists publish prices. Unstructured charges $0.015 a page after 10,000 free pages, Reducto $10 per 1,000 pages, and LlamaIndex starts free with paid plans from $50 a month. Limina and MOSTLY AI have free tiers. The rest quote each project.