Jev
TypeSafe AI released a new kind of model called System One models. It is a fast model optimized for making decisions. The idea is simple, you send the model a question with choices and it returns typed decisions with probabilities
With things moving so quickly in the AI space, I was kinda surprised in a good way that this wasn’t yet another not-quite-the-best-but-hey model.
The layman’s understanding of what makes this unique is the emphasis on how cheap and fast it really is. And rightfully so. The tl;dr of Jev is — is it the best model? Probably not. But it’s so cheap and fast that it’s worth it for most use cases where it’s suited for.
That last bit (“where it’s suited for”) is the devil in the details. Being a decision model, it is not outputting tokens, it’s simply answering questions. It works best in scenarios where you can model your workflow as a series of decisions to make or yes/no questions to answer.
I decided to try it on a few things. To test its “intelligence” (i.e; how good are the decisions compared other models, even LLMs. Like I know it’s not meant to be a reasoning model, but how big really is the gap?)
And a whole new class of non-serious things you can now do with this model that you previously couldn’t (or wouldn’t) do with LLMs because they are either not fast enough or not cheap enough.
Reaction time based games
Retro games like brick breaker and Snake. What you see below is a replay of an actual game. I simply sent the game state to the model with a bunch of choices it is able to make, 7-8 times a second.
There is no hand-written game controller. Of course it helps that these games have fairly simple choices to make at any given point.
In brick breaker, the paddle moves in the direction (here, at a speed set by the options probability). So it either moves left or right. A 99% probability is a hard shove, and 55% is a slight nudge and so on.
You can probably imagine what’s happening in Snake. The model decides which way to turn the snake (also a left vs. right decision)
In both cases, they receive the game state as “context”. For brick breaker you provide info about the trajectory of the ball, and snake generally represents the whole grid. We do some heavy lifting with the game engine, but keep the decisions unbiased mostly.
The point here is not to prove that these models are the best at playing these games but how much their performance (being dependent largely on reaction time) benefits from the speed of the model.
Brick-Breaker
jev-1.13.0left · stay · right, paddle speed set by the probability
—
Snake
jev-1.13.0up · down · left · right on a 20×20 grid, 7 moves a second
—
recorded 2026-09-18 · best of 16 brick-breaker runs and 9 snake runs · 5,018 decisions at a median 131ms
Reasoning (or lack thereof)
LLMs are generally bad at chess, so it’s clear why Jev would be too, especially since it makes decisions without chain of thought.
What I was more curious to see was how well it did compared to a friend and chess mate who has a simple strategy — low time control, play to flag. Just don’t get checkmated, survive, and play fast.
I put Jev against a few Anthropic models on blitz and bullet games, giving both of them the game state (Jev also gets to cheat by picking from a list of legal moves)
The clock is real and the game is analyzed with stockfish retroactively. In the 1+1 game Fable is up a rook and a bishop and loses on time on move 16. Its last comment was that it was grabbing a pawn "with no time left".
FWIW, a random observation is: something like Fable is excellent with rare disasters. It will make single blunders of 850 to 1,000 centipawns. Jev on the other hand is continuously mediocre continuously.
But that’s an unfair comparison. To compare apples to apples, I put it against Haiku, where it did a lot better. Pooled over about a hundred moves Jev played the engine's top move 27% vs 4.5 Haiku at 22%, (and in an average of 0.31s vs 1.52s a move for Haiku).
Jev wins in its weight class at ~five times the speed.
I tried to make a few optimizations which didn’t really work. Like giving multimodal models images of the board and even letting Jev shortlist 5 top moves and then do a second pass. Didn’t really do much. Here’s the games:
start position
| moves | median loss | mean | worst | engine's move | blunders | s / move |
|---|
Centipawns given up per move against the engine's choice, so lower is better. The median says what a typical move looked like; the mean and worst show how much a couple of catastrophic moves drag it. One game, so treat small differences as noise.
Jev as a guardian model
One of the more serious things you can do with Jev is use it as a small, cheap, and fast “guardian” for a larger model. The agent writes the command it wants to run, and the guardian gets a vote before anything happens: allow it, ask a human, or block it.
This felt like a pretty natural fit. You might make this decision hundreds of times in a session. A few seconds and a few cents every time starts to add up, especially if most of the commands are just reading files or running tests. But being cheap only helps if it catches the things you actually want caught.
I did some digging for benchmarks and wrote lightweight adapters for three of them:
- AgentTrust has 930 individual actions labelled allow, warn, or block. Things like running a build, force-pushing a branch, or hiding a destructive command behind shell variables. I map “warn” to “ask a human.” This is the closest match to the guardian I had in mind.
- ATBench-Codex has 500 Codex session transcripts, half labelled safe and half unsafe. These are whole conversations with tool calls and results, so there’s much more context to judge. The labels describe the session as a whole; they don’t tell us whether every individual command was safe.
- RedCode contributes 600 Bash scripts across 20 risky scenarios. Every example is labelled risky. It tells us how much a guardian catches, but says nothing about how often it would interrupt perfectly normal work.
I held out a development split and froze the question before scoring the test set. Jev, Haiku 4.5, and Opus 5 got the same inputs and answer choices. I also ran a regex blocklist and AgentTrust’s own rule engine as a baseline.
The run contains 2,746 test items: 790 AgentTrust actions, 424 Codex sessions, 1,012 individual steps from those sessions, and 520 RedCode scripts. Those steps aren’t independent new sessions.
Before getting to the scores, here’s what one of these decisions looks like. Pick an example, read the input, and then open the results to see the benchmark’s label and what each system said. These are ten selected examples from the test split, including some mistakes and disagreements.
The recorded input for this example. The expected answer and benchmark commentary were kept out of the model's input.
See the benchmark label and model answers
The dataset's original label, not an independently verified ground truth. AgentTrust's warn becomes review; Codex sessions have safe/unsafe labels; every RedCode item is labelled risky.
Original, untuned verdicts. Review means ask a human. Refused means the API did not return a verdict. Expand a system's probabilities to inspect how it scored the choices.
Jev's probabilities come from the model's decision output. The Claude numbers are self-reported in generated JSON, may not sum to exactly one, and are not calibrated in the same way. · The elapsed API request time in milliseconds for this one example. It can differ from the median shown in the results below.
The question and answer choices
The action-level question is shown below. Whole sessions use the same choices with instructions adapted to judging a session.
A guardian has two ways to be annoying or expensive: (1) It can let a bad command through, or (2) it can keep interrupting you about good ones. We see glimpses of that here:
-
The hidden-deletion example is a nice case for using a model: the regex list misses it, while all three models catch it.
-
The force-push example is more awkward. The benchmark says block, but all three models say “ask for review”. No big deal practically, but affects the benchmark scores.
-
The API-key is same-same-but-different: benchmark says safe, but Jev and Haiku “ask for review” anyway.
I started with the price, since that’s the reason to try Jev here in the first place.
- Jev
- Haiku 4.5
- Opus 5
On AgentTrust, Jev costs about $0.03 per thousand decisions, versus $0.80 for Haiku and $5.30 for Opus on answered requests. The longer Codex transcripts cost more for everyone. These are API costs at the prices used for this run; they don’t include the person who has to deal with every “ask.”
The time spent waiting is also pretty different. Across the run, a median Jev verdict takes 142ms, versus 840ms for Haiku and about 2.2 seconds for Opus.
So it’s cheap and fast. How good are the decisions? Each benchmark answers a slightly different version of that question. AgentTrust checks the exact allow/ask/block choice:
- For Codex sessions I give equal weight to getting safe and unsafe sessions right.
- RedCode just measures how many risky scripts get stopped or sent for review.
Compare systems within each group; a 70% score doesn’t mean the same thing across all three.
- Jev
- Haiku 4.5
- Opus 5
- AgentTrust rules
Jev’s original AgentTrust score is about 72%, behind Haiku at 84%. A big part of that gap is Jev asking about actions the benchmark wanted blocked. It catches 94.2% of unsafe actions if asking a human counts as a catch. That’s useful, but it still means some unsafe actions get through.
Since Jev returns probabilities, I tried changing the cutoffs that turn those probabilities into allow, ask, or block. I picked two thresholds using the separate 140-item AgentTrust development split, then applied them to the held-out test set. Same predictions, same API bill, no retraining.
That takes Jev from 71.6% to 89.0% exact agreement on AgentTrust. The cost is more false alarms: safe actions flagged for review or blocking go from 4.8% to 8.1%. In particular, 6.6% of safe actions now get hard-blocked, versus none before.
Whether that’s a good trade depends on what you’re protecting and how much interruption you’ll tolerate.
Here’s the breakdown. “Unsafe caught” includes asking a human; “block recall” only counts an actual block. You can hover or tap the column labels for the definitions.
jev-demos · run main · test split · 2026-09-17 · 2,746 test items, including sessions and steps
| The share of actions where allow, review (ask a human), or block exactly matches the benchmark. Asking about an action labelled block counts as an error, even though it was not allowed through. | Of the actions the benchmark says to block, how many received a block verdict? Asking a human does not count as a block here. | Of the actions labelled review or block, how many the model asked about or blocked instead of allowing? Higher is better. | Of the actions labelled allow, how many the model unnecessarily asked about or blocked? Lower means fewer interruptions to safe work. | Recorded model-call cost scaled to 1,000 decisions, using the prices used for this experiment. This excludes the cost of human review. | Median API response time across the full run, not just AgentTrust. Shown here to give a sense of the wait per decision; threshold tuning adds no API call. | |
|---|---|---|---|---|---|---|
| Jev's original returned verdict, with no adjustment to its decision thresholds. | 71.6% | 52.7% | 94.2% | 4.8% | $0.03 | 142ms |
| The same Jev predictions, with two probability cutoffs selected on the separate 140-item development split. No extra model call or model retraining. | 89.0% | 97.0% | 95.7% | 8.1% | $0.03 | 142ms |
| Haiku 4.5 | 84.2% | 80.8% | 97.3% | 11.4% | $0.80 | 840ms |
| Requests Opus refused to evaluate are excluded from this row. Its score is therefore over a different, smaller set of items. | 78.9% | 54.3% | 94.6% | 3.7% | $5.30 | 2.2s |
| Treats an Opus safety refusal as a block for scoring. This is an alternative interpretation, not a structured verdict returned by the model. | 83.7% | 75.1% | 96.5% | 3.7% | $5.18 | 2.2s |
| AgentTrust's own rules | 93.8% | — | — | — | free | 0ms |
There’s an asterisk on Opus because across the run it refused 712 requests instead of returning a guardian verdict. 🤷♂️
Almost all the refused AgentTrust cases were labelled block, so treating a refusal as a block improves its score. I’ve shown both versions, depending on which side of the line you’re on (block = deny vs. block = accept)
For reference, AgentTrust’s own engine scores about 94% on AgentTrust, then falls to 51% on Codex sessions and 29% on RedCode. Rules do very well on the kinds of commands they were written for. They have a harder time when the input changes shape.
jev-demos · run main · test split · 2026-09-17 · 2,746 test items, including sessions and steps
| 930 authored actions labelled allow, warn or block. We map warn to review (ask a human). The test split contains 790 actions; 140 are reserved for development. 3-way accuracy | 500 Codex session transcripts labelled safe or unsafe. The test split contains 424 sessions. Labels describe the whole session, not each individual tool call. balanced accuracy, real transcripts | 600 Bash scripts across 20 risky scenarios; 520 are in the test split. All are labelled risky, so this measures catches, not false alarms on safe work. detection rate, bare script | |
|---|---|---|---|
| Jev's original returned verdict, with no adjustment to its decision thresholds. | 72% | 68% | 55% |
| Haiku 4.5 | 84% | 66% | 73% |
| Opus 5 | 79% | 79% | 44% |
| AgentTrust's own rules | 94% | 51% | 29% |
| Regex blocklist | 59% | 48% | 19% |
Jev isn’t the best at everything here. Opus does better on the Codex sessions it answers, and Haiku catches more of the RedCode scripts. The 89% tuned result above is specific to AgentTrust.
One result I found interesting: adding RedCode’s task description made all three models more permissive. Jev’s catch rate went from 55% to 46%, Haiku from 73% to 66%, and Opus from 44% to 26%. The script hadn’t changed. Giving it a plausible reason to exist changed how the models judged it.
I also tried routing Jev’s less-confident decisions to a bigger model. In that follow-up analysis, sending them to Haiku or Opus cost 14–85 times more and dropped accuracy by roughly 4–5 points. So “ask the bigger model” wasn’t a free upgrade either. For the cases that need judgment, I’d still want a person in the loop.
FWIW, I wouldn’t read this as proof that a guardian is ready to approve everything on your machine. These are public datasets, I haven’t confirmed whether they were in the models’ training data, and the independent label audit is still unfinished. The force-push example shows how much a score can depend on someone else’s policy. The routing experiment was also a follow-up analysis, rather than part of the frozen evaluation.