Backed by Y Combinator

Turn your data into revenue

Ooak Data licenses the operational and business data that defines how companies actually work, anonymizes it, and turns it into training material for frontier AI models. You get paid for something you already have.

  • Earn up to $300K
  • 2 to 4 weeks
  • No engineering work required

Estimator

What is my data worth?

Connected data is worth more than big data. Every tool you include adds another thread to the story, and the value is in the threads that join up.

What is my data worth?

Answer a few questions for a preliminary estimate.

Tools you used 0 selected

From your systems to the model

The data that teaches models how work actually gets done

Every tool your company ran on, anonymized by Ooak into one consistent dataset.

  • Emails
  • Slack & Teams
  • Documents
  • Tickets
  • Code & PRs
  • CRM
  • Support
  • Spreadsheets
  • Intelligence
  • Efficiency
  • Velocity
  • Reliability
  • Autonomy
  • Judgment
Ooak Data Anonymized

Your systems

What models gain

Why now

The AI labs need new sources of data

  1. 01

    The easy supply is exhausted

    Web text, books, code and forums have all been scraped and trained on.

  2. 02

    Models now need complexity

    Longer, multi-step, multi-tool tasks take richer datasets, not more short-form text.

A single workflow, across multiple systems

  1. Customer email
  2. Support ticket
  3. Slack thread
  4. Jira issue
  5. Pull request
  6. CRM update

Read our research →

Your part

Licensing your company data

We run the whole process, from the first call to the signed agreement. The technical work is ours.

Book a call and tell us which tools you are ready to license.

  1. 1

    Chat with our team

    We map your workflows, tools and data footprint, and give you an estimate of what your data is worth. Fifteen minutes.

  2. 2

    Decide what is in scope

    You pick the systems and date ranges. Anything you exclude is annexed to the contract before a single record moves.

  3. 3

    We evaluate your data

    Hand in hand, we count what is actually there: volume, coverage and how deeply your tools connect. Under an hour of your time, and not a line of engineering work.

  4. 4

    We anonymize your data

    Our in-house engine strips all PII. Names, companies and addresses are replaced consistently across every source, and faces, logos and signatures are blacked out.

  5. 5

    Get paid

    We pay fast, and your raw archive is deleted within 30 days - a contract obligation, not a policy.

Our core technology

Anonymization is not a step in our process. It is what we build.

We built the engine, we run it, and no third party ever touches your raw archive. One entity dictionary is applied across every connected source, so a name replaced in an email is the same replacement in the ticket, the ledger and the commit. Faces, logos and signatures are detected and blacked out, page by page.

Your source

#launch-hydra-serum
Camille Roux 09:12

@Théo we are 4,000 units short for the Vantelle FR drop on the 18th. Milan plant confirmed 11,000 of 15,000.

Entity dictionary

  • Camille RouxSofia Klein
  • ThéoMarc
  • Vantelle FRRetailer 3
  • MilanCity 2

One dictionary, every source. Same input, same output.

Anonymized

#launch-product-a
Sofia Klein 09:12

@Marc we are 4,000 units short for the Retailer 3 drop on the 18th. City 2 plant confirmed 11,000 of 15,000.

The workflow survives. The identities do not. That consistency across sources is what makes an anonymized dataset usable, and it is the hard part - we master it.

What a dataset looks like

Three datasets, three shapes

Three of the datasets we have licensed. Different sectors, different sizes, different stacks.

  • Edtech

    People
    400+
    History
    6 years
    Tools
    15

    Google Workspace, MySQL, BigQuery, MongoDB, GitLab, ClickUp, AWS, Azure, GCP

  • Construction tech

    People
    200+
    History
    4 years
    Tools
    7

    Google Workspace, Slack, Jira, Asana, Airtable, GitHub, HubSpot

  • Fleet logistics

    People
    15+
    History
    4 years
    Tools
    7

    Google Workspace, Slack, WhatsApp, Jira, Notion, Supabase, GitHub

FAQ

Questions founders ask

  • In most cases, yes. Your company owns its operational records, and the contract licenses them for training only, with your exclusions annexed. We check jurisdiction and requirements during scoping. Counsel reviews everything before signature.

  • The buyer receives a cleaned dataset under a code name. Company names, people, addresses and logos are replaced or blacked out consistently across every source. We have recorded zero re-identification incidents to date. You review a sample before delivery.

  • Only Ooak. The raw archive never leaves Ooak's own pipeline and is never handed to another company. Our QA team works on the anonymized output, not the raw files. The buyer sees the code-named dataset and nothing else.

  • Whoever holds signing authority for the entity at that moment: a director, a liquidator or an administrator. We have closed deals in each of those situations. We ask for the authority document up front so the contract is not challenged later. Payment goes to the entity, not to individuals.

  • It is deleted within 30 days of delivery. That is a contract obligation, not a policy, and you receive written confirmation. Only the anonymized dataset is retained for the licence term.

  • Deals to date range from $20K, the smallest closed, to several hundred thousand dollars. The number depends on years of history, headcount, the number of tools and how deeply they connect to each other. We never quote a number before counting. The estimator gives a range; the audit gives a figure.

  • About one hour in total. One 15-minute intro call, one scoping pass by call or async, and either a native export from each tool or read-only access. No engineer is needed on your side. We do the extraction ourselves.

  • We acquire the operational, strategy, and business data that defines how companies actually work. This includes documents and files (reports, contracts, presentations), project management data (Jira, Notion, Trello), CRM and sales data, financial records, product analytics, internal wikis, and code repositories. We are interested in the full context, not just isolated files, but the relationships between tools, teams, and workflows.

Next step

Ready to find out what your data is worth?

We will walk you through the process, answer your questions, and give you an honest valuation. No commitment required.

Book a call Email us directly