For AI Labs & Robotics

The judgment layer your physical AI is missing.

You have the motor data and the pixels. What you don't have is why an experienced worker made the call they made — verified against the rule and the result.

Four ways to source physical-world data

Teleop logs

Motor and visual — no semantic judgment.

Simulation

No real stakes, wrong distribution.

Scraped text

Pretraining, not physical-world RLHF.

OkAI

Verified human judgment at the moment it mattered — checked against regulation, confirmed by outcome.

the one that doesn't exist elsewhere

Not annotation — creation

Annotation labels data that already exists. OkAI creates the first record of a decision that would otherwise have vanished — the instant a worker says it out loud, because we pay them to.

What you license

Three layers, one corpus.

Licensed corpus Available now

Verified gold traces, occupation-coded, outcome-linked.

Eval-as-a-service Rolling out

Benchmark your models on real occupational judgment.

Specialist models In development

Distilled on high-DQM traces, per occupation.

licensed sampleStop Work
S

Situation

The crew is ready to start decking on the second floor, about 20 feet up. The perimeter guardrail isn't up yet and no one is tied off, and we're behind schedule. Do we start now, or hold until fall protection is in place?

B

Background

Grounded in OSHA 29 CFR 1926.501(b)(1) — Fall Protection

A

Assessment

The risk reads as: Stop Work.

High confidence
R

Recommendation

OkAI's call

Stop Work

Not worth the risk. Make it safe first.

Licensed, never sold

b2e9a4c7…c680O*NET 47-2061.00ConstructionOSHA_CONSTRUCTIONrevocable

The shape of a preference pair

Situation What the worker described.
Rejected The verdict OkAI recommended.
Chosen The verdict the worker overrode to.
Reasoning Why — in the worker's own words.

A preference pair no simulation can generate — the situation, the rejected verdict, the chosen verdict, and the human reasoning behind the correction. Sourced from real stakes. In this pilot, 1 of 4 logged decisions was an override.

Auditable by construction

Provenance-hashed

Every sample carries a cryptographic hash tying it to its origin.

Occupation-coded

SOC-coded, so you can slice the corpus by the work it came from.

Outcome-verified

Each verdict is checked against what actually happened on the floor.

The proof, coming next

We're publishing an empirical validation: models trained on high-DQM gold traces against synthetic-data and generic-RLHF baselines — including a frontier-model-with-retrieval arm — on our occupational benchmark. This is the number that prices the corpus. Not a claim, a result.

Get the data class that doesn't exist anywhere else.