Find what your benchmark
doesn’t know to test.
Playtest.ai is building a competitive human red-team network for world models and interactive agents—designed to discover novel, reproducible failures that static benchmarks miss.
Verified failures become regression tests. The hardest cases become targeted training data.The next generation of world models will need humans
exceptionally good at breaking them.
Not a generic crowd. A capability network whose evaluators earn reputation by finding failures that are novel, severe, reproducible, and diagnostically useful.
Use play not only to train world models, but to continuously adversarially evaluate and improve them. Playtest.ai finds what the lab did not know to test.
A failure is only valuable
if the lab can understand and reproduce it.
Playtest.ai aligns the evaluator’s observation and exact actions with simulator state, resulting transitions, and session evidence.
Discovery finds the weakness. State makes it diagnostic.
Seeing something strange
is not the same as proving a failure.
A useful adversarial case needs the interaction, underlying state, expected consequence, observed divergence, and a path back to reproduction.
Observation without exact intervention
Ot → Ot+1Large-scale, but exact input and hidden state are often missing.
Scale from your own policy
Cheap and abundant, but sampled from behaviour the model already produces.
Reproducible interaction evidence
St + Ot + At ↛ St+1Human discovery, qualified state, exact actions, divergence, and verification.
Reports describe what happened. Regression artifacts make it actionable.
One transition.
Fully aligned.
Not another folder of video.
A queryable record of what the model saw, what the human did, and what the world actually became.
Same world.
Different action.
Where supported by the environment, Playtest.ai can restore a simulator checkpoint and execute alternate interventions from the same underlying state.
Plausible futures
can still be wrong.
If changing the action barely changes the predicted future, the model may be relying on visual momentum rather than action-conditioned dynamics.
Your policy is not
a human distribution.
Synthetic rollouts are efficient, scalable, and repeatable. Humans contribute the behavioural variation a model team’s own policy may not produce.
The action tells you what happened.
The human can tell you why.
Optional intent capture adds a research signal where the collection calls for it. It is not assumed to be universally useful.
Worlds have
memory.
The rendered frame may not expose all the state required to explain an outcome minutes later. Long-horizon collections can preserve the interactions and hidden variables between cause and consequence.
Games are not just content.
They are experimental systems.
Game engines expose something most real-world video cannot: direct access to the simulator underneath the pixels.
Game environments do not replace real-world robotics data.
They provide controllable environments for studying action, dynamics, planning, and behaviour with strong ground truth.
Every verified failure
becomes an asset.
The network searches beyond the standard suite. Strong discoveries are verified, attributed, packaged as regression cases, and used to define the next targeted collection.
“The model follows the visual prior instead of the action.”
“An object disappears after leaving and returning.”
“Contact physics look plausible but produce the wrong outcome.”
“Geometry drifts during a long interaction.”
“The model cannot recover after a failed action.”
“A successful benchmark strategy collapses under a novel perturbation.”
A benchmark that learns
every time the model fails.
Challenge
Define the checkpoint, task family, model claim, and adversarial skill tracks.
Discover
Specialist evaluators search for novel failures outside the known suite.
Verify
Reproduce the strongest cases and attach synchronized interaction evidence.
Regress
Promote accepted failures into replayable tests for future checkpoints.
Improve
Turn hard cases into targeted data, retrain, and run the campaign again.
Don’t buy the story.
Run the ablation.
Same model. Same eval. Progressively richer supervision.
One adversarial learning loop.
Several research directions.
Bad ground truth is worse
than no ground truth.
If a datum fails the agreed gate, it does not ship.
The data should come with an answer to:
“Can we actually train on this?”
Collections can define environment rights, affirmative training rights, permitted use, sublicensing, participant consent, consent versioning, pseudonymous IDs, PII separation, and build provenance.
Rights are negotiated per collection.
We are not competing with
your internal harness.
Your suite should test every failure you already know. Playtest.ai supplies the moving frontier: people who find the next one.
That is the scarce layer.
Give us a checkpoint.
We’ll try to break it.
Start with one playable task family and one or two checkpoints. We assemble the right evaluators, hunt for unknown failures, verify the strongest cases, and return a reusable regression package.
Bring us
one model claim.
We turn a playable model into a live adversarial campaign, then return the evidence your team can use.
ONE MODEL
A checkpoint—or two checkpoints you need to compare.
ONE TASK FAMILY
A controllable environment where the claim should hold.
ONE CLAIM
The capability your existing benchmark says is working.
Technical questions,
plain answers.
Both, over time. Each campaign is an adversarial evaluation; verified discoveries accumulate into a living regression benchmark.