How much is company data worth?
The short answer
Published AI data programs list amounts from about $10,000 to more than $1 million per partnership. Handshake AI currently publishes a range reaching $4 million, and named companies have reported offers or payments in the hundreds of thousands of dollars.
Those numbers do not tell you what your company will get.
Assess the opportunity against four factors:
- What the records show about how work gets done.
- How well those records connect decisions, actions, and outcomes.
- Whether buyers want that kind of data right now.
- What data and usage rights you can actually license.
Buyer programs screen for company size, operating history, data quality, specialized expertise, and buyer fit. They do not publish acceptance rates, so program payout ranges cannot tell you your odds of receiving an offer.
This guide compares published amounts, shows what current buyers seek, and explains how to assess your data, avoid deal blockers, and judge an offer. For the full seller-side path, read how to sell your company's data to AI companies.
What are companies getting paid?
Public numbers fall into several categories. A program range, an individual offer, a completed transaction, and a pending bid do not mean the same thing.
| Source | Amount | Type | What it means | Checked |
|---|---|---|---|---|
| Scale AI | $10K–$1M+ | Program range | Scale calls this illustrative value per data partnership | Sep 23, 2026 |
| Prism by Labelbox | About $50K to $1M+ | Program range | Smaller targeted datasets can start around $50K; enterprise partnerships can reach $1M+ | Sep 23, 2026 |
| Handshake AI | $100K–$4M | Program range | Current published payout range | Sep 23, 2026 |
| micro1 | $500K+ / $1M+ | Program tiers | Published tiers for different partnership profiles | Sep 23, 2026 |
| Warmly | Up to $300K | Reported offer | Mercor made the offer. Warmly declined it. | Sep 23, 2026 |
| cielo24 | Hundreds of thousands | Reported transaction | Roughly 13 years of operating history monetized during wind-down | Sep 21, 2026 |
| Spirit Airlines | $10M selected bid | Selected bid | Google was selected in a bankruptcy auction. Court approval remained pending before the Sep. 30 hearing. | Sep 23, 2026 |
Two cautions matter.
First, the program figures are published by the organizations running the programs. They show what those organizations say qualifying partnerships can be worth. They do not show how many applicants receive an offer, what the median deal looks like, or where most completed deals fall within the range.
Second, the named examples are very different transactions.
Warmly received an offer and declined it. cielo24 monetized its history while winding down. Spirit involves a large enterprise archive, bankruptcy proceedings, competitive bidding, and significant privacy review.
Averaging those numbers would produce a precise-looking number with almost no useful meaning.
When comparing deals, ask:
What exactly was the buyer getting?
Which records? Which rights? One-time or ongoing access? Exclusive or non-exclusive? What use was permitted?
Those questions matter more than whether two companies look similar from the outside.
Which kinds of company data do buyers want right now?
Demand changes quickly, and no single buyer represents the whole market.
But current programs give us useful signals.
| Type of data | Current signal | What we are seeing |
|---|---|---|
| Cross-system operating workflows | Strong | Prism currently labels cross-system interaction data as urgent. Scale and micro1 both emphasize real operational workflows. |
| Support history with resolutions | Strong | Prism lists longitudinal customer support as high demand. Scale and micro1 explicitly seek support and customer-operations records. |
| Engineering and code-change history | Strong | Prism lists code-generation and edit traces as high demand. Scale and micro1 both identify engineering workflows and project history. |
| Expert decisions in specialized fields | Strong in some domains | Prism currently labels financial and scientific reasoning traces as urgent. Scale emphasizes domain depth; micro1 explicitly looks for decision-making patterns and specialized expertise. |
| Sales and CRM workflows | Active interest | Scale and micro1 both identify CRM, sales processes, and revenue workflows as useful operational data. |
| General email and documents without linked outcomes | Less clear on their own | Email and documents appear in current programs, but buyers repeatedly emphasize workflow context, relationships, decisions, and outcomes rather than raw volume alone. |
These are signals, not permanent market rankings.
A dataset can be unusual and useful but still have no active buyer today. Conversely, something that attracted little interest six months ago can become much more relevant as AI systems and buyer priorities change.
What makes company data valuable?
Think of company data as a record of work, not a pile of files.
Take one support ticket.
By itself, it tells you that a customer had a problem.
Now imagine your systems let someone follow the entire chain:
Customer problem → investigation → internal discussion → engineering action → resolution
That history shows much more.
It captures what happened, what information people had, how they reasoned about it, what action they took, and whether the action worked.
That kind of connected operational context is exactly what several current programs say they are looking for.
The same pattern appears throughout a business:
Sales Lead → qualification → objections → proposal → won/lost
Engineering Bug → investigation → code change → test → release
Operations Exception → investigation → approval → action → outcome
Support Issue → diagnosis → escalation → fix → resolution
An individual Slack message, Jira ticket, CRM record, or document tells only part of the story. The history that connects them shows decisions, actions, and results.
An HVAC company
Imagine an HVAC business with 12 years of service records.
One archive consists mainly of old invoice PDFs.
Another contains:
- equipment and fault codes
- technician observations
- suspected causes
- diagnostic steps
- parts replaced
- follow-up visits
- whether the repair solved the problem
Both archives contain the same number of records.
But the second archive captures years of real diagnostic judgment and outcomes. It also looks much more like the operational data current programs describe as useful.
Two software companies
Both have operated for ten years.
Company A has ten years of Slack messages and documents.
Company B can connect customer problems to internal discussions, engineering tickets, code changes, releases, and final customer outcomes.
On paper, both have "ten years of company data."
The underlying assets are very different.
Score your own data
Use this checklist to identify strengths to test and gaps to close.
Rate each question 0, 1, or 2.
| Question | 0 | 1 | 2 |
|---|---|---|---|
| Decisions. Do the records show people making real judgment calls? | Rarely | Sometimes | Often, and in detail |
| Context. Can you see why they decided what they did? | No | Partly | Yes, with surrounding facts and discussion |
| Actions. Can you see what happened next? | No | In some systems | Yes, linked to the decision |
| Outcomes. Can you tell whether it worked? | No | Sometimes | Yes, outcomes are recorded |
| Hard to recreate. Would someone need to operate a company like yours for years to reproduce this history? | No, it is public or inexpensive to collect | Somewhat | Yes |
How to read the score
This is a DataDeals screening heuristic, not a valuation model and not evidence that a buyer will make an offer.
0–4: The archive currently shows relatively few of the characteristics buyers emphasize. Before approaching buyers, define the asset more carefully and see whether better-connected records exist elsewhere in the company.
5–7: Investigate the useful signals, then check for gaps in context, outcomes, connectivity, and uniqueness.
8–10: The data shows many of the characteristics current programs describe as attractive. The next questions are whether there is active buyer demand and whether the relevant records can actually be licensed.
Do not turn the score into a dollar multiplier.
Its job is to tell you where to investigate next.
Can you estimate a rough dollar range from the company profile alone?
No. Company size and industry do not identify the asset, buyer demand, or licensing rights. Public deals do not support a reliable price table by company profile.
There are too few disclosed completed transactions, and the public program ranges do not tell us how actual deals are distributed within them.
A relatively small company can hold highly specialized history that a buyer wants. A much larger company can have millions of generic or poorly connected records and receive no offer.
Current programs publish ranges spanning tens of thousands to millions of dollars. Reported offers and transactions document six-figure amounts, but none sets a price for your specific asset.
The actual asset still has to be defined.
Things that can strengthen an opportunity include:
- connected records across systems
- visible decisions and outcomes
- specialized domain knowledge
- difficult-to-recreate operating history
- clean chronology and metadata
- clear provenance and usable rights
- active buyer demand
- more than one interested buyer
- ongoing access that the buyer values
Things that can weaken it include:
- generic material
- disconnected records
- missing history
- poor exports or metadata
- unclear rights
- substantial privacy or contractual restrictions
- no current buyer for that type of data
What stops or shrinks a deal?
Rights and privacy
This is where many opportunities become more complicated.
Owning a system does not give a company unrestricted rights to license every record inside it.
Common issues include:
Employee communications. Review Slack, email, meeting transcripts, and chat for personal information, employment-law duties, privacy rules, internal policies, and other restrictions.
Customer data. CRM records and support tickets often contain names, contact information, confidential business information, and data governed by customer agreements.
Regulated data. Health, financial, consumer, and international personal data can bring additional legal requirements.
Contractor and vendor work. Your right to reuse code, documents, designs, or other material created by third parties depends on the underlying agreements.
Third-party confidential material. Check partner information, NDA material, licensed content, and third-party intellectual property before including them in a license.
Separate records you can license from those requiring exclusion, de-identification, aggregation, permission, or legal review.
The rights review determines which records and uses the license can cover.
Data quality
Data quality determines how much work a buyer must do before the archive is useful.
Problems include:
- missing timestamps
- inconsistent exports
- broken links between systems
- incomplete histories
- duplicate records
- poor metadata
- unclear authorship or provenance
The more work required to make the asset understandable and usable, the harder the transaction may become.
Costs
Preparation costs depend on the asset, privacy work, and deal structure. Budget for the actual transaction, not a generic rate.
They can include:
- Legal review: contracts, privacy, intellectual property, employment issues, and the eventual license agreement.
- Data preparation: exporting, cleaning, joining, filtering, de-identifying, and formatting the records.
- Internal staff time: especially from IT, security, engineering, legal, or operations.
- Outside specialists: where privacy, security, or technical work requires them.
- Intermediary or advisory fees: if another party sources buyers, structures the opportunity, or helps negotiate the deal.
Some buyers absorb substantial preparation work themselves. Prism, for example, says it handles ingestion, normalization, and PII scrubbing within its program. That should not be assumed for every transaction.
No buyer yet
Rarity does not create demand. Check for active buyers before investing in extensive data preparation.
Someone offered you $250K. Is that good?
Judge the $250K against the data, rights, access, exclusivity, and liability the buyer receives. Ask:
- What exact data is included?
- Is access one-time or ongoing?
- What can the buyer do with the data?
- Is the license exclusive? For how long?
- Are future records included?
- Does the buyer have sublicensing or resale rights?
- Who carries the legal risk if something in the dataset creates a claim?
- Could another buyer want the same asset?
$250K for a narrow, one-time, non-exclusive license is a very different deal from $250K for permanent exclusivity plus all future data.
Compare the deal, not just the headline number.
Red flags worth examining
- The data scope is vague, such as "all operational data."
- A one-time payment buys permanent rights for any future use.
- The buyer can freely resell or sublicense the asset with no additional economics for you.
- Exclusivity has no clear end date.
- Future data is included automatically at no additional price.
- Broad indemnity language puts most of the legal risk on the seller.
- The process gives you no realistic time to obtain legal advice.
- A buyer pressures you to commit before you understand whether other buyers may be interested.
Public deal data does not support a standard percentage premium for exclusivity.
Price exclusivity, future access, broader permitted use, and sublicensing as additional rights you are giving the buyer.
How does a company-data deal work?
The details vary, but a seller will usually need to work through these questions.
1. Define the asset
Describe what the records actually show.
This is weak:
Ten years of company data.
This is useful:
Eight years of customer-support cases linked to internal investigations, engineering actions, and final resolution outcomes.
2. Check the rights
Work out what can be included, what must be removed or changed, and where legal review is needed.
3. Check buyer fit
Find out whether anyone is currently buying this type of information.
4. Scope what can be shown
Be prepared to show a sample, schema, workflow, or limited subset before agreeing final terms.
That does not mean handing over the entire raw archive before there is an agreement.
5. Create market feedback
Where practical, test the opportunity with more than one potential buyer.
One buyer's view is one data point.
6. Negotiate the rights and economics
Price is only one term.
Negotiate scope, exclusivity, permitted use, future access, sublicensing, liability, retention, deletion, and payment structure alongside price.
7. Prepare and deliver the agreed data
Export, clean, de-identify, package, and transfer the data according to the final agreement.
How long does it take?
Public deal reports do not establish a market-wide timeline.
Prism currently advertises roughly 21 days to first wire for its own licensing process. That is useful as one program-specific reference, but it should not be treated as the typical time for the market.
A straightforward opportunity handled through an established program can move much faster than a transaction requiring complicated legal review, extensive de-identification, multiple counterparties, or court approval.
Get a starting estimate
The DataDeals calculator gives you an initial benchmark using basic company information such as company size, operating history, and location.
It then separately looks at the types of data you may have, current buyer fit, relevant market evidence, and what still needs to be investigated.
It takes about two minutes.
You do not upload your company's raw data.
A calculator result is a starting benchmark, not a buyer valuation.
The strongest pricing signal comes when a real buyer has reviewed a defined opportunity and puts money behind an offer.
About 2 minutes · No raw data upload
Frequently asked questions
Is there an average price for company data?
No meaningful market average is available.
Published numbers mix advertised program ranges, individual offers, completed transactions, and pending bids. Combining them into an average would misstate what companies actually receive.
Is company data priced per gigabyte?
No. There is no standard per-gigabyte rate for company-data licenses.
Volume is one input. Current programs also assess workflow context, quality, decision-making, domain expertise, uniqueness, buyer demand, and rights.
Does more data mean more money?
No. More records do not guarantee a higher price.
A huge archive of generic or disconnected records can be less useful than a smaller dataset that clearly captures decisions, actions, and outcomes.
Can a small company have valuable data?
Yes. A smaller company can hold specialized operating history that is difficult to recreate.
But individual programs have their own eligibility rules. Handshake currently says it is looking for companies with at least 20 full-time employees and three years in operation. Scale lists 40+ employees as its preferred starting point, while micro1 says its Enterprise Data Partnership looks for operationally mature companies with 30+ employees.
Those are program rules, not universal market minimums.
Is old company data valuable?
Yes. Historical records can have value when they capture usable operating context.
cielo24 reportedly received hundreds of thousands of dollars for roughly 13 years of historical Slack, Jira, email, and Google Drive data during its wind-down.
That is evidence that historical operating records can be monetized. It does not establish a standard price for old data.
Is it legal to license company data to AI companies?
A company can license data it has the right to share, subject to applicable privacy, contract, and other restrictions.
Employee information, customer contracts, regulated data, third-party intellectual property, and confidentiality obligations determine which records and uses require special care.
Have qualified counsel review the actual dataset and agreement before committing to a transaction.
Do we need employee consent to license Slack or email?
Consent requirements depend on applicable law, employee location, company policies, contractual obligations, the records being shared, and how they will be processed. Check these before sharing Slack or email data.
Current programs commonly describe de-identification or removal of personal information as part of their processes.
Who buys company data?
AI companies can acquire operational data directly or through data partners and intermediaries.
Public company-data programs currently include Scale AI, Prism by Labelbox, Handshake AI, and micro1. Mercor made the reported offer to Warmly.
The active buyer set can change quickly.
What do intermediaries charge?
Intermediary fees vary by agreement. Confirm these terms before engaging one:
- Does the seller pay you?
- Does the buyer pay you?
- Can you receive fees from both?
- Is the fee based on the transaction value?
- Is there a cap?
- Are buyer-paid referral fees disclosed?
- Does a buyer-paid fee reduce what the seller owes?
DataDeals' standard seller-side pricing is a 10% success fee on cash proceeds, capped at $100,000. There is no upfront seller fee or monthly retainer, and if no transaction closes, there is no success fee. Any buyer or referral fee paid separately to DataDeals is disclosed and credited dollar-for-dollar against the seller's success fee.
How long does a deal take?
There is no published market-wide timeline.
Prism advertises around 21 days to first wire for its own program. Diligence, privacy review, negotiation, data preparation, and approvals extend more complicated deals.
Does exclusivity raise the price?
Exclusivity limits your ability to license the asset elsewhere. Price that restriction as a separate deal term.
Public deal data does not support a standard percentage premium.
Is a calculator result a valuation?
No.
It is a benchmark based on limited inputs.
A real offer from a buyer who has reviewed the defined opportunity is a much stronger signal of what the market will pay.
About DataDeals
DataDeals is an independent seller-side research and advisory business founded and operated by Ola Rask. We help companies identify, evaluate, package, and license proprietary operating data to AI buyers.
Our standard seller-side pricing has no upfront fee and no monthly retainer. If a transaction closes, DataDeals charges 10% of cash proceeds, capped at $100,000. If no transaction closes, the success fee is $0.
Any buyer or referral fees paid separately to DataDeals are disclosed and credited dollar-for-dollar against the seller success fee so we do not collect twice on the same transaction. Read how DataDeals works to see the review and seller-side process.
Sources
DataDeals tracks buyer programs, reported offers and transactions, public filings, independent reporting, and research on data valuation.
Market-sensitive sources are rechecked before publication and periodically as the market changes. The evidence tables show when the underlying market information was last checked.
- Scale AI Data Partnerships
- Prism by Labelbox Enterprise
- Prism Data Categories
- Handshake AI Data Partnerships
- micro1 Data Partnerships
- The Information reporting on Warmly
- Forbes reporting on cielo24
- Spirit Airlines bankruptcy sale reporting and docket references
- Deloitte research on valuing data assets
- Fraunhofer ISST enterprise data valuation framework
