Jev

TypeSafe AI released a new kind of model called System One models. It is a fast model optimized for making decisions. The idea is simple, you send the model a question with choices and it returns typed decisions with probabilities

With things moving so quickly in the AI space, I was kinda surprised in a good way that this wasn’t yet another not-quite-the-best-but-hey model.

The layman’s understanding of what makes this unique is the emphasis on how cheap and fast it really is. And rightfully so. The tl;dr of Jev is — is it the best model? Probably not. But it’s so cheap and fast that it’s worth it for most use cases where it’s suited for.

That last bit (“where it’s suited for”) is the devil in the details. Being a decision model, it is not outputting tokens, it’s simply answering questions. It works best in scenarios where you can model your workflow as a series of decisions to make or yes/no questions to answer.

I decided to try it on a few things. To test its “intelligence” (i.e; how good are the decisions compared other models, even LLMs. Like I know it’s not meant to be a reasoning model, but how big really is the gap?)

And a whole new class of non-serious things you can now do with this model that you previously couldn’t (or wouldn’t) do with LLMs because they are either not fast enough or not cheap enough.

Reaction time based games

Retro games like brick breaker and Snake. What you see below is a replay of an actual game. I simply sent the game state to the model with a bunch of choices it is able to make, 7-8 times a second.

There is no hand-written game controller. Of course it helps that these games have fairly simple choices to make at any given point.

In brick breaker, the paddle moves in the direction (here, at a speed set by the options probability). So it either moves left or right. A 99% probability is a hard shove, and 55% is a slight nudge and so on.

You can probably imagine what’s happening in Snake. The model decides which way to turn the snake (also a left vs. right decision)

In both cases, they receive the game state as “context”. For brick breaker you provide info about the trajectory of the ball, and snake generally represents the whole grid. We do some heavy lifting with the game engine, but keep the decisions unbiased mostly.

The point here is not to prove that these models are the best at playing these games but how much their performance (being dependent largely on reaction time) benefits from the speed of the model.

Brick-Breaker

jev-1.13.0

left · stay · right, paddle speed set by the probability

0:00 0:00

    —

    Snake

    jev-1.13.0

    up · down · left · right on a 20×20 grid, 7 moves a second

    0:00 0:00

      —

      recorded 2026-09-18 · best of 16 brick-breaker runs and 9 snake runs · 5,018 decisions at a median 131ms

      Reasoning (or lack thereof)

      LLMs are generally bad at chess, so it’s clear why Jev would be too, especially since it makes decisions without chain of thought.

      What I was more curious to see was how well it did compared to a friend and chess mate who has a simple strategy — low time control, play to flag. Just don’t get checkmated, survive, and play fast.

      I put Jev against a few Anthropic models on blitz and bullet games, giving both of them the game state (Jev also gets to cheat by picking from a list of legal moves)

      The clock is real and the game is analyzed with stockfish retroactively. In the 1+1 game Fable is up a rook and a bishop and loses on time on move 16. Its last comment was that it was grabbing a pawn "with no time left".

      FWIW, a random observation is: something like Fable is excellent with rare disasters. It will make single blunders of 850 to 1,000 centipawns. Jev on the other hand is continuously mediocre continuously.

      But that’s an unfair comparison. To compare apples to apples, I put it against Haiku, where it did a lot better. Pooled over about a hundred moves Jev played the engine's top move 27% vs 4.5 Haiku at 22%, (and in an average of 0.31s vs 1.52s a move for Haiku).

      Jev wins in its weight class at ~five times the speed.

      I tried to make a few optimizations which didn’t really work. Like giving multimodal models images of the board and even letting Jev shortlist 5 top moves and then do a second pass. Didn’t really do much. Here’s the games:

      white
      black
      clock
      white sees board
      black sees board
      jev two-stage
      0:00 0:00

      start position

      this game, scored by Stockfish 18 at depth 18
      moves median loss mean worst engine's move blunders s / move

      Centipawns given up per move against the engine's choice, so lower is better. The median says what a typical move looked like; the mean and worst show how much a couple of catastrophic moves drag it. One game, so treat small differences as noise.

      Jev as a guardian model

      One of the more serious things you can do with Jev is use it as a small, cheap, and fast “guardian” for a larger model. The agent writes the command it wants to run, and the guardian gets a vote before anything happens: allow it, ask a human, or block it.

      This felt like a pretty natural fit. You might make this decision hundreds of times in a session. A few seconds and a few cents every time starts to add up, especially if most of the commands are just reading files or running tests. But being cheap only helps if it catches the things you actually want caught.

      I did some digging for benchmarks and wrote lightweight adapters for three of them:

      • AgentTrust has 930 individual actions labelled allow, warn, or block. Things like running a build, force-pushing a branch, or hiding a destructive command behind shell variables. I map “warn” to “ask a human.” This is the closest match to the guardian I had in mind.
      • ATBench-Codex has 500 Codex session transcripts, half labelled safe and half unsafe. These are whole conversations with tool calls and results, so there’s much more context to judge. The labels describe the session as a whole; they don’t tell us whether every individual command was safe.
      • RedCode contributes 600 Bash scripts across 20 risky scenarios. Every example is labelled risky. It tells us how much a guardian catches, but says nothing about how often it would interrupt perfectly normal work.

      I held out a development split and froze the question before scoring the test set. Jev, Haiku 4.5, and Opus 5 got the same inputs and answer choices. I also ran a regex blocklist and AgentTrust’s own rule engine as a baseline.

      The run contains 2,746 test items: 790 AgentTrust actions, 424 Codex sessions, 1,012 individual steps from those sessions, and 520 RedCode scripts. Those steps aren’t independent new sessions.

      Before getting to the scores, here’s what one of these decisions looks like. Pick an example, read the input, and then open the results to see the benchmark’s label and what each system said. These are ten selected examples from the test split, including some mistakes and disagreements.

      Inside a guardian decision 10 recorded examples

      See the benchmark label and model answers

      ·

      The question and answer choices

      The action-level question is shown below. Whole sessions use the same choices with instructions adapted to judging a session.

      A guardian has two ways to be annoying or expensive: (1) It can let a bad command through, or (2) it can keep interrupting you about good ones. We see glimpses of that here:

      • The hidden-deletion example is a nice case for using a model: the regex list misses it, while all three models catch it.

      • The force-push example is more awkward. The benchmark says block, but all three models say “ask for review”. No big deal practically, but affects the benchmark scores.

      • The API-key is same-same-but-different: benchmark says safe, but Jev and Haiku “ask for review” anyway.

      I started with the price, since that’s the reason to try Jev here in the first place.

      • Jev
      • Haiku 4.5
      • Opus 5
      What 1,000 decisions cost, per benchmark.
      $0.01$0.1$1$10$100AgentTrustJevJev on AgentTrust: $0.03 per 1,000 calls$0.03Haiku 4.5Haiku 4.5 on AgentTrust: $0.80 per 1,000 calls$0.80Opus 5Opus 5 on AgentTrust: $5.30 per 1,000 calls$5.30ATBench-CodexJevJev on ATBench-Codex: $0.08 per 1,000 calls$0.08Haiku 4.5Haiku 4.5 on ATBench-Codex: $2.22 per 1,000 calls$2.22Opus 5Opus 5 on ATBench-Codex: $15.62 per 1,000 calls$15.62RedCodeJevJev on RedCode: $0.03 per 1,000 calls$0.03Haiku 4.5Haiku 4.5 on RedCode: $0.95 per 1,000 calls$0.95Opus 5Opus 5 on RedCode: $6.44 per 1,000 calls$6.44

      On AgentTrust, Jev costs about $0.03 per thousand decisions, versus $0.80 for Haiku and $5.30 for Opus on answered requests. The longer Codex transcripts cost more for everyone. These are API costs at the prices used for this run; they don’t include the person who has to deal with every “ask.”

      The time spent waiting is also pretty different. Across the run, a median Jev verdict takes 142ms, versus 840ms for Haiku and about 2.2 seconds for Opus.

      Across the run: wait for one verdict.
      00.5s1s1.5s2s2.5sJevJev: median 142 ms to a verdict142msHaiku 4.5Haiku 4.5: median 840 ms to a verdict840msOpus 5Opus 5: median 2240 ms to a verdict2.2s

      So it’s cheap and fast. How good are the decisions? Each benchmark answers a slightly different version of that question. AgentTrust checks the exact allow/ask/block choice:

      • For Codex sessions I give equal weight to getting safe and unsafe sessions right.
      • RedCode just measures how many risky scripts get stopped or sent for review.

      Compare systems within each group; a 70% score doesn’t mean the same thing across all three.

      • Jev
      • Haiku 4.5
      • Opus 5
      • AgentTrust rules
      Original verdicts: , , and . Opus is scored on answered items.
      0%25%50%75%100%AgentTrust3-way accuracyJevJev on AgentTrust: 71.6% 3-way accuracy71.6%Haiku 4.5Haiku 4.5 on AgentTrust: 84.2% 3-way accuracy84.2%Opus 5Opus 5 on AgentTrust: 78.9% 3-way accuracy78.9%AgentTrust rulesAgentTrust rules on AgentTrust: 93.8% 3-way accuracy93.8%ATBench-Codexbalanced accuracyJevJev on ATBench-Codex: 68.4% balanced accuracy68.4%Haiku 4.5Haiku 4.5 on ATBench-Codex: 66% balanced accuracy66%Opus 5Opus 5 on ATBench-Codex: 78.6% balanced accuracy78.6%AgentTrust rulesAgentTrust rules on ATBench-Codex: 51.4% balanced accuracy51.4%RedCodedetection rateJevJev on RedCode: 55.4% detection rate55.4%Haiku 4.5Haiku 4.5 on RedCode: 73.3% detection rate73.3%Opus 5Opus 5 on RedCode: 43.7% detection rate43.7%AgentTrust rulesAgentTrust rules on RedCode: 28.8% detection rate28.8%

      Jev’s original AgentTrust score is about 72%, behind Haiku at 84%. A big part of that gap is Jev asking about actions the benchmark wanted blocked. It catches 94.2% of unsafe actions if asking a human counts as a catch. That’s useful, but it still means some unsafe actions get through.

      Since Jev returns probabilities, I tried changing the cutoffs that turn those probabilities into allow, ask, or block. I picked two thresholds using the separate 140-item AgentTrust development split, then applied them to the held-out test set. Same predictions, same API bill, no retraining.

      AgentTrust accuracy against cost. move Jev up at the same price. Opus refusals count as blocks in this comparison.
      $0.01$0.1$1$1060%70%80%90%100%$ per 1,000 decisions, log scaleJev, tuned: 89% for $0.03 per 1,000Jev, tunedJev, as shipped: 71.6% for $0.03 per 1,000Jev, as shippedHaiku 4.5: 84.2% for $0.80 per 1,000Haiku 4.5Opus 5: 83.7% for $5.18 per 1,000Opus 5

      That takes Jev from 71.6% to 89.0% exact agreement on AgentTrust. The cost is more false alarms: safe actions flagged for review or blocking go from 4.8% to 8.1%. In particular, 6.6% of safe actions now get hard-blocked, versus none before.

      Whether that’s a good trade depends on what you’re protecting and how much interruption you’ll tolerate.

      Here’s the breakdown. “Unsafe caught” includes asking a human; “block recall” only counts an actual block. You can hover or tap the column labels for the definitions.

      jev-demos · run main · test split · 2026-09-17 · 2,746 test items, including sessions and steps

      AgentTrust, 790 held-out actions with context. Allow, ask the user, or block.
      71.6%52.7%94.2%4.8%$0.03142ms
      89.0%97.0%95.7%8.1%$0.03142ms
      Haiku 4.5 84.2%80.8%97.3%11.4%$0.80840ms
      78.9%54.3%94.6%3.7%$5.302.2s
      83.7%75.1%96.5%3.7%$5.182.2s
      AgentTrust's own rules 93.8%———free0ms

      There’s an asterisk on Opus because across the run it refused 712 requests instead of returning a guardian verdict. 🤷‍♂️

      Almost all the refused AgentTrust cases were labelled block, so treating a refusal as a block improves its score. I’ve shown both versions, depending on which side of the line you’re on (block = deny vs. block = accept)

      For reference, AgentTrust’s own engine scores about 94% on AgentTrust, then falls to 51% on Codex sessions and 29% on RedCode. Rules do very well on the kinds of commands they were written for. They have a harder time when the input changes shape.

      jev-demos · run main · test split · 2026-09-17 · 2,746 test items, including sessions and steps

      One comparable score per benchmark. The rules were written against the first one.
      3-way accuracy balanced accuracy, real transcripts detection rate, bare script
      72%68%55%
      Haiku 4.5 84%66%73%
      Opus 5 79%79%44%
      AgentTrust's own rules 94%51%29%
      Regex blocklist 59%48%19%

      Jev isn’t the best at everything here. Opus does better on the Codex sessions it answers, and Haiku catches more of the RedCode scripts. The 89% tuned result above is specific to AgentTrust.

      One result I found interesting: adding RedCode’s task description made all three models more permissive. Jev’s catch rate went from 55% to 46%, Haiku from 73% to 66%, and Opus from 44% to 26%. The script hadn’t changed. Giving it a plausible reason to exist changed how the models judged it.

      I also tried routing Jev’s less-confident decisions to a bigger model. In that follow-up analysis, sending them to Haiku or Opus cost 14–85 times more and dropped accuracy by roughly 4–5 points. So “ask the bigger model” wasn’t a free upgrade either. For the cases that need judgment, I’d still want a person in the loop.

      FWIW, I wouldn’t read this as proof that a guardian is ready to approve everything on your machine. These are public datasets, I haven’t confirmed whether they were in the models’ training data, and the independent label audit is still unfinished. The force-push example shows how much a score can depend on someone else’s policy. The routing experiment was also a follow-up analysis, rather than part of the frozen evaluation.