Domain data, cleaned and verified
Runix Data cleans, structures and builds training and evaluation data in six domains: code, finance, cybersecurity, legal, embodied AI and AI for Science, with code as the focus. Every record ships with its source, its licence and the checks it passed.
Built on Runix Pipeline, the data tooling its engagements run on.
// one accepted record, pretty-printed { "id": "code-task-000142", "source": { "repo": "example-org/example-parser", "base_commit": "3f9c2e1", "licence": "MIT" }, "dedup": { "cluster": "c-5521", "kept": "1 of 3" }, "checks": { "tests_with_reference_patch": "pass", "tests_without_patch": "fail", "secrets_scan": "clean", "statement_matches_tests": true }, "verdict": "accepted", "dropped_reason": null } // and one that did not make it { "id": "code-task-000143", "verdict": "dropped", "dropped_reason": "tests pass without the patch" }
An illustrative record. The repository, commit and IDs are placeholders.
Six domains, six definitions of clean
The stages are the same everywhere: ingest, clean, structure, mask, report. What counts as a duplicate, a valid record or a leak changes with the field, and that is where each domain gets its own rules.
Repositories Focus
Code
Repository and task data for code models and coding agents, from training examples to evaluation tasks that have to run.
- Licence and provenance recorded per file, and code whose licence does not allow the use left out
- Secrets, credentials and personal data scrubbed before anything leaves the pipeline
- Deduplication across forks, vendored dependencies and generated or minified code
- Task statements that describe what the tests actually check
Code · verification
Every task is run, not read
A coding task only teaches or measures something if its tests can tell a right answer from a wrong one. Each task is executed in a clean container: the reference solution must make its tests pass, and the untouched repository must make the same tests fail.
Filings · transactions
Finance
Filings, statements, market data and transaction records, for models that have to read numbers as carefully as an analyst does.
- Units, currencies and fiscal periods normalised before any two figures are compared
- Restated figures kept apart from the originals they replace
- One entity across tickers, legal names and registry identifiers
- Account numbers and personal data masked, failing closed
Advisories · logs
Cybersecurity
Advisories, vulnerability records, logs and threat reports, for models that triage, detect and explain.
- The same vulnerability, reported by several sources, merged into one record
- Affected version ranges normalised so they can be compared and queried
- Indicators defanged, so no live payload or working link ships in a dataset
- Labels for detection tasks, each with the rule that assigned it
Contracts · case law
Legal
Contracts, case law and regulation, for models that have to cite what they rely on.
- Clauses segmented so each can be retrieved, compared and labelled on its own
- Citations normalised, so the same case is always the same case
- Every document tagged with its jurisdiction and the date it took effect
- Privileged material and personal data redacted before processing continues
Trajectories · sensors
Embodied AI
Robot trajectories, teleoperation sessions and multi-sensor recordings, for policies that learn from demonstration.
- Cameras, joint states and force readings aligned on one clock
- Calibration kept with every episode, not in a separate file nobody ships
- Long recordings segmented into episodes, with failed or unsafe ones filtered out
- Action spaces normalised across robots and controllers
Antibodies · proteins
AI for Science
Antibody and protein records, where a bad merge changes the answer rather than the formatting. Our work here is limited to biological data.
- Sequence formats converted and validated, not just parsed
- One numbering scheme applied across sources that use competing ones
- Accession numbers reconciled where two databases disagree
- Names, aliases and identifiers normalised to one entity
Not listed
Another domain
The stages carry over from one field to the next; the rules do not. Tell us what the data is and what the model has to do with it, and we will say plainly whether we have the judgement for it.
Describe your data →Domain judgement on shared tooling
Runix Pipeline runs the stages the same way on every engagement. Runix Data is the layer of judgement on top of it: the rules that decide, field by field, what survives and what is dropped.
You receive
Runix Data
Runix Pipeline
Your sources
What every delivery carries
The data is half of a delivery. The other half is what lets you check it without taking our word for it.
Provenance per record
Where each record came from, when it was collected and what was done to it, so a bad answer downstream can be traced back to the record that caused it.
A licence per record
The licence or permission each record was used under. Material whose terms do not allow your use is left out, not flagged and shipped anyway.
A quality report
Coverage, duplication, extraction confidence and the checks each record passed, plus what was dropped and why.
Personal data masked
Identified and masked before it reaches a training set. Detection fails closed: a record we cannot clear is held back, not passed through.
Evaluation kept apart
Evaluation data is split from training data by source rather than by row, so near-duplicates cannot sit on both sides of the split and inflate a score.
The schema, written down
Every field defined and every known gap stated, so the next team to use the data does not have to reverse-engineer it.
How early access works
Three steps, each with a person on the other end: Runix Data is scoped per engagement, not bought off a shelf.
01Send a sample and the task
A slice of the real data, not a description of it, and what the model has to do with the result: train on it, be evaluated on it, or both. We reply within one business day.
02Get a scoped plan and a quote
The rules we would apply, the checks each record has to pass, what we think is not worth doing, and the price, all agreed before any work starts.
03Receive the data and the report
Batch or continuous, as files, a database or an endpoint, whichever your training and evaluation jobs actually consume. Every delivery comes with its quality report.
Common questions
Which domains do you work in?
Code, finance, cybersecurity, legal, embodied AI and AI for Science, where our work is limited to biological data: antibody and protein records. Code is our focus. If your field is not on the list, tell us what the data is and we will say plainly whether we can do it well.
Do you clean our data, or build new data?
Both. We clean and structure data you provide or have the rights to use, and we build task data, such as verified coding tasks, to a specification agreed with you in writing before work starts.
What happens to the data we send?
It is processed only to do the work you asked for. It is not used to train models, ours or anyone else's, and it is not sold. The details are in our Privacy Policy and on our Security page.
How is it priced?
Per project or by volume, quoted before any work starts. You see the scope and the number together, so there is nothing to reconcile afterwards.
How does Runix Data relate to Runix Pipeline?
Runix Pipeline is the tooling underneath: the stages every engagement runs through, from ingest to delivery. Runix Data is the service built on it, with the domain rules and judgement calls that differ from one field to the next.
Send us a sample of the hard part
A slice of the real data and what the model has to do with it. We reply within one business day, and the scoped plan that follows includes the parts we think are not worth doing.
Request early access