Training environments that grow with every rollout.

930 generates training environments for computer-use models from natural language: tasks, diagnostic grading, and reward signals. Every failure becomes training data. New tasks target the weaknesses each rollout reveals.

What we believe

Your model learns only what its environments teach.

The limit is not compute or architecture. The environments your model trains in decide what it can learn. Most teams still hand-build their own eval and training setups, and they discard failures instead of learning from them.

We're building infrastructure where environments are programs, grading names each broken check, and every session adds its environments and traces back to the shared universe.

The loop

Each cycle turns rollouts into new tasks.

930 generates environments from natural language. It also reads your rollouts, finds where your model fails, and generates new tasks for those weaknesses. Across teams, each cycle adds tasks where models fail most.

  1. 01

    Evaluate

    Run tasks with criterion-level grading. Each failure names the broken check and shows expected against actual.

  2. 02

    Find blind spots

    Review rollouts across sessions. Group failures by cause and show what remains untested.

  3. 03

    Generate targeted tasks

    Create new environments where the model fails, so the next round tests what broke.

  4. 04

    Train & repeat

    Export training data, fine-tune, and re-evaluate. Each cycle adds to the shared universe.

Technical bets

Four bets on how agent training should work.

Each shapes the platform, and we test each in production.

Environments are programs

You describe an environment in plain language. 930 compiles it to code and validates it at runtime. Each one runs as a state machine with event handlers, world generators, and grading logic, deterministically from a seed. You can fork and combine them.

Grading that shows what broke

Composable rubric functions grade each task. Every criterion returns a grade and written feedback: "cell (row_1, amount) = 450, expected 500" instead of "73%". The same results are RL reward signals, with one auxiliary reward per criterion.

Every session is saved

930 stores every agent action with full state snapshots. You can replay any moment, fork from any failure, and reseed a session to reproduce it. One session produces eval grades, RL reward signals, and SFT training data, with no separate pipelines.

A shared universe that compounds

Every team's training generates environments the whole platform can use. Failures show blind spots, and 930 turns them into new tasks for the shared universe.

See it in action

Models train in real environments.

Criterion-level grading across CRM workflows, spreadsheet checks, and spatial layouts.

Cross-app workflow

Model onboards a client across CRM, spreadsheet, and task list in one session.

Find the right contact, record the contract in revenue, and capture follow-ups in todos. 930 grades each step against real UI state.

Example prompt

Onboard Quigley-Block as a new client with a $126000 contract. Update their CRM status and deal value, add the contract as revenue in Excel, and create three onboarding …

Spreadsheet reasoning

Model reconciles multiple sheets without changing source data.

Cross-reference tabs, infer what is missing or inconsistent, and write only the derived output. Grading checks correctness and that originals stay intact.

Example prompt

Reconcile invoices against payments: 1. Switch to the 'Invoices' sheet and review all invoice IDs 2. Switch to the 'Payments' sheet and note which invoice IDs have payme…

Spatial layout

Model arranges furniture under explicit spatial constraints.

In a floor-plan editor with movable pieces, the model satisfies relationships like "sofa opposite windows" and "lamp symmetry" while avoiding collisions. Grading checks positions, symmetry, and collisions.

Example prompt

Arrange the living room: place the sofa against the wall opposite the windows, put the coffee table in front of the sofa, place a lamp on each side of the sofa, move the…

We're building the training infrastructure for computer-use models.

If you work on agent capabilities, model evaluation, or synthetic training data, we should talk. We're looking for research partners who want to train agents on harder work.

hello@930labs.co @930_labs

9°30'W · Cabo da Roca The westernmost point of continental Europe, where the known world once ended and the frontier began.