I’ve spent the past few days looking at Jev, TypeSafe’s new AI model, and trying to work out where it fits.
The idea I keep coming back to is this: perhaps the next generation of AI agents doesn’t need one model that does everything. Perhaps different parts of the job need different kinds of intelligence. ⚡
An agent has to understand a goal, make judgments along the way, and eventually do something. Those aren’t always the same kind of task. Writing a thoughtful response and deciding which of a few allowed actions fits the current situation ask different things of a model.
Jev made me think about that separation more carefully.
What Jev actually does
TypeSafe calls Jev its first “System One” model. In its documentation, the interface is built around a state and questions about that state. Instead of generating a written answer, Jev returns values that software can use directly.
There are three question types:
- Choice: choose from options you define.
- Score: assess something against levels you specify.
- Noul: return a probability that the answer to a yes/no question is yes.
Choice and Score return probability distributions and a separate confidence value. Noul returns a probability, without that separate confidence field.
You can ask several questions in one call. According to TypeSafe, they are evaluated in parallel and independently against the same state. That independence matters: one question cannot quietly depend on another answer from the same call. You have to handle that dependency in code or a later request.
I find this interesting because sometimes the thing an agent needs next is quite small. Which route fits this request? Does this message ask for a refund? Is there enough information to continue?
It doesn’t necessarily need another paragraph.
My Jev AI analogy: brain, nervous system, controller, body
This is how I’m currently picturing it:
🧠 The LLM is the brain. It interprets the goal, plans, and revisits the plan when something unexpected happens.
⚡ Jev is the fast nervous system. It evaluates clearly defined questions and returns structured judgments that the rest of the system can use.
🧩 The code is the controller. It combines those judgments with rules, permissions, and checks. It decides which action is allowed, or when the LLM or a person needs to take over.
🖱️ Browsers, APIs, mouse, and keyboard are the body. More precisely, the tools that operate them carry out the actions.
I put that idea into this animated infographic:

My proposed architecture, not a biological model or a benchmark result. “Screen” means information extracted into text or structured data before it reaches Jev. Confidence is available for Choice and Score; a high value never replaces permission to act.
Of course, the analogy isn’t perfect. A brain is part of a nervous system, and an LLM isn’t a literal brain. I’m using the comparison to separate responsibilities, not to suggest that this is how human cognition works.
Most importantly: Jev does not move the mouse. It returns judgments. The controller and tools do something with them.
What this could look like in an agent
Imagine an assistant handling a message about a possible duplicate charge. This is a design example, not a system I’m presenting as tested.
The application gathers the message, relevant transaction records, and the refund policy. An LLM could help interpret a complicated request. Jev could then evaluate narrower questions: does the message request a refund? Does the supplied evidence suggest a duplicate? Which support route fits?
The controller still has work to do. It checks the actual records, the account permissions, and whether approval is required. It might open a review screen, request missing information, or ask a person to confirm. The API performs the action only after those conditions are met.
A confident model answer would not be permission to move money.
There is also a practical limit to the browser part of my illustration. Jev currently accepts text input, including JSON objects and arrays of text. It does not directly accept images, audio, or video. A browser agent would need to extract useful page information first, or use another component to interpret a screenshot.
That extra work belongs in the architecture, too.
Where I would be careful
TypeSafe makes substantial speed claims in its launch announcement. Those are the company’s results, not independent measurements for the architecture I’m describing. TypeSafe also says its headline workflow gains are likely at the higher end of real-world gains.
I would want to measure the whole loop: collecting state, getting judgments, applying checks, executing an action, and reading the result. A fast model call doesn’t tell me how long all of that takes.
I would also keep type safety separate from correctness. Choosing an allowed option prevents an invented option. It doesn’t prevent choosing the wrong allowed option.
The same caution applies to confidence. TypeSafe’s confidence documentation explains that the value is derived from the probability distribution. It isn’t a second opinion, and it isn’t a guarantee that the action is safe. Thresholds need testing on the actual task.
For me, that leaves an important role for ordinary code. If a rule can be checked exactly, I would check it exactly. I wouldn’t ask a model to guess whether an amount exceeds a configured limit.
Why I’m interested beyond automation
This connects with what I’m exploring in AI personality engineering. A conversational system’s behavior depends on more than the model writing its replies. Memory rules, routing, and permissions also shape what it does.
I can imagine a companion keeping its conversational voice in an LLM while a separate decision layer helps route narrow questions. That is a possible application, not evidence that Jev improves personality or makes a companion safer. The boundaries would still need to be designed and tested.
It’s one of the connections I want to explore through my writing on conversational AI and human–AI interaction.
My next step would be a small comparison: one workflow using an LLM for the judgments, and another using Jev with the same inputs and permissions. I’d look at wrong decisions and unnecessary escalations alongside speed.
For now, the nervous-system analogy helps me think about who does what. I don’t yet know how far it will hold up in practice.
What do you think? Does it work for you, or does it hide something important?


