ADVERSARIAL EVALUATION FOR INTERACTIVE AIPLAYTEST.AI LABS

Find what your benchmark
doesn’t know to test.

Playtest.ai is building a competitive human red-team network for world models and interactive agents—designed to discover novel, reproducible failures that static benchmarks miss.

Verified failures become regression tests. The hardest cases become targeted training data.
State persistenceAction adherenceGeometryPhysicsLong-horizon memory
ADVERSARIAL EVALUATION CAMPAIGN
CHECKPOINT WM-042 LIVE
VERIFIED FAILURE / 017NOVEL

The world forgets an object after the evaluator leaves and returns.

A state-persistence failure discovered outside the static suite, then reproduced from the same qualified origin state.

place object leave scene return later object missing
NOVELTY 0.96×SEVERITY 0.84×REPRODUCIBILITY 1.00×DIAGNOSTIC VALUE 0.92
EVALUATOR 147PERSISTENCE SPECIALIST
#03
State persistence97
Action adherence91
Geometry86
Physics82
Long-horizon memory94

Reputation follows verified skill—not volume of reports.

01DISCOVER
02VERIFY
03REGRESSION
04TARGETED DATA
05RETEST
A COMPETITIVE HUMAN RED TEAM

The next generation of world models will need humans
exceptionally good at breaking them.

Not a generic crowd. A capability network whose evaluators earn reputation by finding failures that are novel, severe, reproducible, and diagnostically useful.

01

SPECIALIZE

Evaluators develop measurable skill profiles across persistence, action adherence, geometry, physics, and memory.

02

COMPETE

Campaigns reward verified discoveries—not noisy report volume or benchmark familiarity.

03

ACCUMULATE

Every accepted failure expands the living benchmark and makes the next checkpoint harder to fool.

Use play not only to train world models, but to continuously adversarially evaluate and improve them. Playtest.ai finds what the lab did not know to test.

THE EVIDENCE BEHIND THE FAILURE

A failure is only valuable
if the lab can understand and reproduce it.

Playtest.ai aligns the evaluator’s observation and exact actions with simulator state, resulting transitions, and session evidence.

Discovery finds the weakness. State makes it diagnostic.

RENDERED OBSERVATIONLOCKED
RENDERED OBSERVATIONUNLOCKED
RENDERED OBSERVATIONSCRIPTED
RENDERED OBSERVATIONWAITING ON EVENT
DEPENDING ON THE ENVIRONMENT INTEGRATIONtransformsvelocityphysicscollisionobject stateinventoryabilitiescooldownsquest stateNPC stategame eventscheckpointsRNG state
A REPORT IS NOT A REGRESSION TEST

Seeing something strange
is not the same as proving a failure.

A useful adversarial case needs the interaction, underlying state, expected consequence, observed divergence, and a path back to reproduction.

01 / PASSIVE VIDEO

Observation without exact intervention

Ot → Ot+1

Large-scale, but exact input and hidden state are often missing.

02 / POLICY ROLLOUTS

Scale from your own policy

Cheap and abundant, but sampled from behaviour the model already produces.

03 / VERIFIED FAILURE

Reproducible interaction evidence

St + Ot + At ↛ St+1

Human discovery, qualified state, exact actions, divergence, and verification.

Reports describe what happened. Regression artifacts make it actionable.

THE DATA OBJECT

One transition.
Fully aligned.

Not another folder of video.

A queryable record of what the model saw, what the human did, and what the world actually became.

trajectory.schemaALIGNED TRANSITION
(
01observation_t,02state_t,03action_t,04state_t+1,05task,06outcome,07human_intent
)
episode_idtrajectory_idbranch_idparent_branch_idsimulation_tickRNG_seedbuild_shaconsent_version
CONTROLLED INTERVENTIONS

Same world.
Different action.

Where supported by the environment, Playtest.ai can restore a simulator checkpoint and execute alternate interventions from the same underlying state.

MATCHED STATE BRANCH SET
branch_set_04f2
CHECKPOINTSt
sim_state: identicalrng_state: identicalobservation: identical
INTERVENTION A₁Use keySₜ₊₁¹DOOR OPENS
INTERVENTION A₂Attack doorSₜ₊₁²FAILURE
INTERVENTION A₃Walk awaySₜ₊₁³STATE PERSISTS
HELD CONSTANT
  • physics
  • object state
  • NPC state
  • inventory
  • timers
  • RNG
Only the intervention changes.

Where environment integration allows, restore fidelity can be measured and divergent restores excluded before delivery.

DID THE MODEL LEARN THE ACTION?

Plausible futures
can still be wrong.

If changing the action barely changes the predicted future, the model may be relying on visual momentum rather than action-conditioned dynamics.

IDENTICAL STARTING STATE
St / frame_0148
ACTION PROMPTOPEN DOOR
generated future A
ACTION PROMPTWALK AWAY
generated future B
Do the predicted futures respond to the intervention?
01action controllability
02intervention sensitivity
03next-state prediction
04branch consistency
05causal consistency
06persistent-state evaluation
07planning
HUMAN BEHAVIOUR

Your policy is not
a human distribution.

Synthetic rollouts are efficient, scalable, and repeatable. Humans contribute the behavioural variation a model team’s own policy may not produce.

SYNTHETIC / POLICY ROLLOUTSEFFICIENT · SCALABLE · REPEATABLE

High-volume samples from the model team’s policy distribution.

HUMAN TRAJECTORIESMESSY · EXPLORATORY · ADAPTIVE

Mistakes, hesitation, misuse, exploration, recovery, abandonment, and unexpected strategies.

We collect both.World models need dynamics.Agents need behaviour.
LANGUAGE SUPERVISION

The action tells you what happened.
The human can tell you why.

Optional intent capture adds a research signal where the collection calls for it. It is not assumed to be universally useful.

SYSTEM

“You tried this door four times. What were you expecting?”

PLAYER

“I thought I had to break it. The key looked decorative.”

goalbeliefintended actionperceived affordanceconfusionrecovery strategyabandonment reason
Ask rather than infer.
MORE THAN SHORT CLIPS

Worlds have
memory.

The rendered frame may not expose all the state required to explain an outcome minutes later. Long-horizon collections can preserve the interactions and hidden variables between cause and consequence.

object permanencepersistent inventorydelayed consequenceshierarchical taskslong-horizon planningtool userecoverysubgoal tracking
trajectory.long_horizonPERSISTENT STATE
01Pick up keyinventory.key = true
02Leave roomdoor.closed = true
03Complete another objectivequest.branch = 04
04Return laterinventory.key = true
05Open original doordoor.open = true
CONTROLLABLE WORLDS

Games are not just content.
They are experimental systems.

Game engines expose something most real-world video cannot: direct access to the simulator underneath the pixels.

Game environments do not replace real-world robotics data.
They provide controllable environments for studying action, dynamics, planning, and behaviour with strong ground truth.

ENVIRONMENTCONTROL SURFACE
  • restore state
  • replay rare events
  • modify one intervention
  • expose hidden variables
  • vary initial conditions
  • alter difficulty
  • repeat experiments
  • hold out complete environments
FROM UNKNOWN FAILURE TO KNOWN TEST

Every verified failure
becomes an asset.

The network searches beyond the standard suite. Strong discoveries are verified, attributed, packaged as regression cases, and used to define the next targeted collection.

01

The model follows the visual prior instead of the action.

02

An object disappears after leaving and returning.

03

Contact physics look plausible but produce the wrong outcome.

04

Geometry drifts during a long interaction.

05

The model cannot recover after a failed action.

06

A successful benchmark strategy collapses under a novel perturbation.

Evaluators discover it.We make it reproducible.
THE ADVERSARIAL LEARNING LOOP

A benchmark that learns
every time the model fails.

01

Challenge

Define the checkpoint, task family, model claim, and adversarial skill tracks.

02

Discover

Specialist evaluators search for novel failures outside the known suite.

03

Verify

Reproduce the strongest cases and attach synchronized interaction evidence.

04

Regress

Promote accepted failures into replayable tests for future checkpoints.

05

Improve

Turn hard cases into targeted data, retrain, and run the campaign again.

MEASURE THE SIGNAL

Don’t buy the story.
Run the ablation.

Same model. Same eval. Progressively richer supervision.

AObservation only
B+ exact actions
C+ privileged state
D+ counterfactual interventions
E+ human language / intent
WORLD-MODEL METRICS
next-state prediction
action controllability
object-state consistency
persistent-state accuracy
branch consistency
long-horizon drift
AGENT METRICS
task success
recovery success
action prediction
sample efficiency
held-out task generalization
held-out environment generalization
BUILT FOR INTERACTIVE MODELS

One adversarial learning loop.
Several research directions.

01

WORLD MODELS

Action-conditioned dynamics and state prediction.

02

INTERACTIVE VIDEO

Learning how user input should change generated futures.

03

GENERALIST AGENTS

Human demonstrations, recovery, tools, and long-horizon behaviour.

04

GAME AGENTS

Cross-environment behaviour and planning.

05

EMBODIED RESEARCH

Controlled virtual perception-action research before or alongside physical data.

QUALITY GATES

Bad ground truth is worse
than no ground truth.

01

SYNCHRONIZATION

Actions aligned with observations.

02

TELEMETRY

Required fields present.

03

STATE VALIDITY

Required simulator state captured.

04

RESTORE FIDELITY

Checkpoint replay within agreed tolerance.

05

SPLITS

Train / validation / held-out separation.

06

PROVENANCE

Environment, build, rights, and collection metadata linked.

If a datum fails the agreed gate, it does not ship.

RIGHTS-CLEARED BY DESIGN

The data should come with an answer to:
“Can we actually train on this?”

Collections can define environment rights, affirmative training rights, permitted use, sublicensing, participant consent, consent versioning, pseudonymous IDs, PII separation, and build provenance.

Rights are negotiated per collection.

Discuss a collection’s rights structure
COLLECTION MANIFESTEXPLICIT PROVENANCE

environment_rightsdefined

training_rightsdefined

permitted_usedefined

sublicensingdefined

participant_consentdefined

consent_versionversioned

pseudonymous_idsdefined

pii_separationdefined

build_provenancelinked

THE SCARCE LAYER

We are not competing with
your internal harness.

Your suite should test every failure you already know. Playtest.ai supplies the moving frontier: people who find the next one.

01novel failure discovery
02specialist evaluator reputation
03adversarial human strategies
04reproducible interaction evidence
05cross-checkpoint comparison
06regression-ready artifacts
07targeted failure data

That is the scarce layer.

ADVERSARIAL EVALUATION SPRINT

Give us a checkpoint.
We’ll try to break it.

Start with one playable task family and one or two checkpoints. We assemble the right evaluators, hunt for unknown failures, verify the strongest cases, and return a reusable regression package.

Design a sprint
EXAMPLE SPRINT STRUCTURECUSTOMIZED PER MODEL

011 playable task family

021–2 model checkpoints

03curated specialist evaluators

04five adversarial skill tracks

05novelty + severity + reproducibility scoring

06top failures independently verified

07comparative failure report

08regression package + targeted data brief

Illustrative scope only. The final campaign follows the model, environment, and claims under evaluation.
FOR WORLD MODEL & AGENT TEAMS

Bring us
one model claim.

We turn a playable model into a live adversarial campaign, then return the evidence your team can use.

01

ONE MODEL

A checkpoint—or two checkpoints you need to compare.

02

ONE TASK FAMILY

A controllable environment where the claim should hold.

03

ONE CLAIM

The capability your existing benchmark says is working.

We’ll find the test you’re missing.

hello@playtest.ai
FAQ

Technical questions,
plain answers.

Both, over time. Each campaign is an adversarial evaluation; verified discoveries accumulate into a living regression benchmark.