Why Jev, a System 1 model built to make decisions instead of writing text, could change how we design AI agents, RAG chatbots and business automation
By the engineering team at PiTangent Analytics & Technology Solutions
What if the biggest problem with Agentic AI isn’t that AI isn’t smart enough? What if the real problem is that we keep asking generative AI to make too many small decisions?
Most conversations about AI agents are about capability. Which model reasons better. Which one writes better code. Which one can plan ten steps ahead. These are fair questions, and the progress on them has been real.
But when you put an agent into production, something else starts to matter. The agent doesn’t spend most of its time writing essays or solving hard problems. It spends its time making small calls. Is this request in scope? Which agent should take it? Do I need to search the knowledge base? Is this answer safe to send? Should I try again?
In most systems built today, each of those small calls goes to a large language model. Every one adds a few seconds, a little cost and a small chance of an odd answer. On their own, these are minor. Stacked across a multi-step workflow and thousands of daily requests, they become the main reason agentic systems feel slow, expensive and hard to trust.
This is why a new kind of model called Jev caught our attention. Jev doesn’t generate text at all. It makes decisions. And our view, after building agentic systems for clients, is simple:
Jev could be the missing piece in the Agentic AI puzzle. Not because it is smarter than today’s generative models, but because it may be the right tool for a large share of the work we have been forcing those models to do.
A typical agent today works something like this. A user request comes in. An LLM reads it and decides what kind of request it is. The same LLM, or another one, picks a tool. The tool returns data. An LLM reads the data and decides whether it is enough. An LLM writes the answer. Sometimes another LLM checks that answer before it goes out.
In other words, it is LLM all the way down.
This design is popular for good reasons. LLMs are flexible. You can change behaviour by changing a prompt. You don’t need training data. For a prototype, it is the fastest way to get something working.
Latency stacks up. Each model call waits for the one before it. If a workflow makes seven calls and each takes two seconds, the user waits fourteen seconds. Most of that time goes on routing and checking, not on the answer the user actually wants. (These numbers are a hypothetical illustration, not a measurement.)
Cost compounds. A request that looks cheap becomes expensive once you count every classification, routing and validation call behind it. Multiply that by volume, then multiply again by the number of agents in the system.
Errors multiply. Suppose each step in an eight-step workflow is right 97% of the time. That sounds fine. But the chance that all eight steps go right is only about 78%. Roughly one request in five takes at least one wrong turn somewhere. (Again, a hypothetical figure used to show the maths.)
Control gets harder. When an LLM decides where to route a request, it answers in text. You have to parse that text, handle the times it adds an explanation you didn’t ask for, and guess how sure it was. When something goes wrong, you end up reading prompts and outputs to work out why.
None of these problems come from the model being weak. They come from using one kind of intelligence for every kind of task.
Look at the questions a production agent asks itself during a single request:
Now look at the shape of the answers. Yes or no. One option from a short list. A score from one to five. None of them needs a paragraph. These are decision problems, not writing problems.
The psychologist Daniel Kahneman described two modes of human thinking. System 1 is fast and automatic. It is how you recognise a friend’s face or sense that the car ahead is about to change lanes. System 2 is slow and deliberate. It is how you work through a tax problem or plan a trip. Most of what we do in a day runs on System 1. System 2 is expensive, so the brain saves it for when it is really needed.
Modern LLMs, and reasoning models in particular, sit much closer to System 2. They work step by step, one token at a time. That is exactly what you want for summarising a contract or debugging code. It is far more than you need to decide whether an email is a complaint or a sales enquiry.
A hospital gives a useful picture. A triage nurse sees every patient who walks in and makes quick, practised calls: urgent or not, which department, can they wait. Specialists see only the patients who need them. No hospital sends every arrival straight to a senior consultant. It would be slow and costly, and it would waste the consultant’s time on cases a nurse handles well.
Most agentic systems today send every patient to the consultant.
Jev comes from TypeSafe AI. It went into early access on September 15, 2026, the same day the company came out of stealth with a $40 million seed round led by DCVC. TypeSafe describes it as the first of a new class of “System One models,” a direct nod to Kahneman
The idea is easy to explain. You give Jev some program state along with a set of typed questions, and it answers all of them in one parallel pass, returning structured values with calibrated probabilities instead of generated text.
The state can be messy: a customer email, a log excerpt, a JSON record. The questions are defined by you, at the time of the request. TypeSafe’s documentation describes a few question types: pick one option from a list, give a score on a rubric, or answer yes or no. For every question, Jev returns its answer, the probability it gives each possible answer, and an overall confidence value.
Here is a simple example based on TypeSafe’s published material:
State: “My Stripe account won’t connect for 3 days…”
Question: is_urgent (yes/no) “Is this urgent?”
Response: is_urgent -> yes, probability 0.999
There is no sentence to parse, no stray explanation and no formatting to clean up. The output goes straight into an if statement.
It helps to be clear about what Jev is not.
It is not a smaller LLM. A small LLM still writes text, one token at a time. According to TypeSafe, Jev doesn’t generate text at all.
It is not the same as JSON mode. Many LLMs can be told to reply in a fixed format. But they still generate that reply token by token, with the format enforced on top. You get an answer, but usually no reliable sense of how sure the model was. Jev’s output is, by design, a probability spread across the answers you defined.
It is not a traditional classifier. Classic classifiers are trained for one task on labelled data. TypeSafe positions Jev as general-purpose: you write the questions and options at request time, with no task-specific training.
It is not a replacement for generative AI. Jev cannot write a reply to a customer, summarise a document or explain its reasoning in words. It is built to decide, not to create.
TypeSafe makes bold performance claims. Its website advertises results 193.6 times faster and 444.6 times cheaper, measured on its own System One workflows, and lists input pricing of $42 per billion tokens. It describes response times of roughly 70 to 500 milliseconds, against seconds for large generative models. The company says Jev is built on a new architecture, a new sampler and a new training method it calls Reinforcement Learning for Calibrated Decisions (RLCD).
This article is meant to help people make architecture decisions, so we want to separate facts from claims.
What is established is the interface. Jev takes state and typed questions and returns typed answers with probabilities. That design has real consequences however the benchmarks turn out.
What is not yet established is how well it performs outside TypeSafe’s own tests. The architecture has not been disclosed. RLCD has not been published as a paper. The speed and cost figures come from TypeSafe’s own workflows. One independent write-up of TypeSafe’s four-workflow benchmark puts Jev at around 68% accuracy, close to mid-tier LLMs, while being much cheaper and faster. Reviewers have also noted that the benchmark uses the average of two frontier models as its “correct” answer, which builds in a bias toward those models.
A few limits matter too. Because Jev doesn’t produce free text, it cannot make up sentences or facts the way LLMs sometimes do. But it can still pick the wrong option. That is a classification error, and it is still an error. The research we reviewed also notes that answers to separate questions are not guaranteed to agree with each other logically, and that TypeSafe’s own documentation lists weak spots such as arithmetic and dates.
So the honest position is this. Jev’s design is interesting on its own merits and its published numbers are promising, but teams should test it on their own data before trusting it with anything important.
Any one of these on its own is nice to have. Together, they change what is practical.
Think about what happens when a decision costs a tiny fraction of a cent and comes back in a fraction of a second. You stop rationing decisions. Today, many teams skip checks because each one adds a second of waiting and another line on the API bill. With a fast, cheap decision layer, you could check every step, validate every output and score every retrieved document without the user noticing.
The question shifts from “can we afford to check this?” to “what else should we be checking?”
Confidence matters just as much. When an LLM routes a request, you usually get a label and no reliable signal of doubt. When a model returns a calibrated probability, that number becomes a control. TypeSafe’s pitch is that your software sets the thresholds for when Jev acts on its own and when it asks for review. In practice that might mean acting automatically above 0.9, asking a stronger model between 0.6 and 0.9, and sending anything below 0.6 to a person.
One point is easy to miss. Calibration and accuracy are not the same thing. A well-calibrated model is not always right. It is honest about how often it is right. If it says 80%, it should be correct about 80% of the time. For automation, that honesty can be more useful than a small gain in raw accuracy, because you can design around the model’s uncertainty. Whether Jev actually achieves good calibration still needs independent testing.
Then there is consistency. A model that returns probabilities over a fixed set of answers, without free-form sampling, should in principle give the same answer to the same input. That makes behaviour easier to test, easier to monitor and easier to explain to an auditor. We treat this as a design implication worth verifying, not a guarantee.
If most of an agent’s work is decision-making, a different architecture starts to make sense. Instead of putting a generative model at the centre of everything, you put a fast decision layer at the front door and call generative models only when they are needed.
The generative model is still there, still doing the hard and valuable work. It just isn’t being asked to do everything.
Here is how this layer could affect each part of orchestration.
Intelligent routing. Every agentic system has a router, whether it is explicit or buried in a prompt. Making routing an explicit Jev call turns it into a typed decision with a confidence value. You can log it, test it and change the options without rewriting a prompt.
Agent selection. In multi-agent systems, choosing the right specialist is often done by a “supervisor” LLM. That supervisor adds a full generation step before any real work begins. A single choice question over the available agents could do the same job in a fraction of the time.
Tool selection. “Which tool next?” is a textbook choice question. The answer maps directly to a function call. There is nothing to parse, and no risk of the model naming a tool that doesn’t exist, because it can only pick from the tools you listed.
Model selection. Not every generative step needs the most expensive model. A System 1 layer could judge how hard a request is and send simple ones to a small model and hard ones to a strong one. This is one of the most direct ways a decision layer could reduce spending, because it controls where the biggest costs go.
Workflow branching. Real workflows branch constantly: refund or replacement, new customer or existing, domestic or international. Each branch point is a decision. Running them all through one fast layer puts the workflow’s logic in one visible place instead of scattering it across prompts.
Input classification. Intent, language, sentiment, urgency, product line. Because Jev answers many questions in one pass, you could ask all of these at once and get a full profile of a request before anything else runs. TypeSafe says adding more questions barely changes latency. If that holds, it encourages asking many narrow questions instead of one broad one, which is usually better design anyway.
Context filtering. Long contexts are slow and costly, and irrelevant context can confuse a generative model. A decision layer could score which documents, messages or records are relevant before anything reaches the LLM. The generative model then works with less, better material.
Output validation. Before an agent sends a reply or takes an action, something should ask: does this answer the question, does it follow policy, does it include what it must? Each of those is a yes/no or score question. Today, adding an LLM-based checker can double the cost of a response. A fast, cheap checker makes validation something you can run on every output.
Guardrails. Guardrails are decisions about what is allowed. Is this request asking for something we don’t do? Does this output contain personal data? Is this action within the user’s permissions? Because Jev’s answers are limited to the options you define, the guardrail can’t be talked into producing some other kind of output. Its judgment can still be wrong, though, so it should sit alongside hard rules in code, not replace them.
Confidence-based escalation. This is where calibrated probabilities matter most. A system that knows when it is unsure can hand off to a stronger model or a person at the right moments. It is the difference between an agent that fails quietly and one that asks for help.
Retry and fallback decisions. When a tool call fails or a result looks thin, the system must decide: retry, try another source, use a fallback, or stop. Handling this with a fast decision call keeps error handling cheap, which matters because failures tend to pile up exactly when systems are under load.
Cost-aware orchestration. Once decisions are cheap and explicit, cost itself can become an input. The system can ask whether a request justifies a premium model, or whether a cached or template answer would do.
Latency-sensitive workflows. Some decisions must happen in milliseconds: live chat handoffs, voice agents, fraud checks at the point of payment, real-time alerts. Generative models are often too slow for these loops. A System 1 model could bring AI judgment into places where it was previously ruled out on speed alone.
Deterministic business decisions. Businesses need decisions that are repeatable and explainable. “The model chose option B with 0.94 confidence” is easier to record, audit and defend than a paragraph of generated reasoning. It is not a full explanation, but it is a clear, testable record.
A company runs an assistant with four specialist agents: orders, billing, technical support and account management. A customer writes: “I was charged twice for my last order and now the app won’t let me log in.”
Without a System 1 layer: A supervisor LLM reads the message and reasons about which agent should handle it. Because there are two problems, it has to notice both, and it may not. Say it picks billing. The billing agent’s LLM decides which tool to call, calls the payments API, reads the result, decides whether a refund is allowed and writes a reply. The login issue either gets missed or needs the supervisor to reason again. A checker LLM reviews the final text. That is six to eight generative calls running one after another, each needing its output parsed, and the most important decision (are there two issues here?) depends on the supervisor reading carefully.
With a System 1 layer: One Jev call asks several questions at once. Is there a billing issue? Is there a login issue? How urgent is this? Does this look like a duplicate charge? Both issue flags come back yes, so the billing and technical agents start in parallel. Inside each agent, tool choice is a simple choice question. Whether the refund can go ahead is a Jev judgment combined with a hard rule in code on the amount. A generative model is called once to write a single reply covering both problems. A final Jev check confirms the reply addresses both issues before it is sent. If confidence on the refund decision is low, the case goes to a person. One or two generative calls, and every decision is explicit, logged and able to run in parallel.
At PiTangent, we design and build agentic AI systems for clients: support assistants, knowledge chatbots, sales and operations automation, and document-heavy workflows. When you build agentic systems in the real world, the hard part isn’t always making the model smarter. Often, it’s deciding when not to use the smartest model.
Some patterns keep coming back.
We have written prompts that say “Answer with only YES or NO,” and then written extra code for the times the model replied “Yes.” with a full stop, or “Yes, because…” followed by a paragraph. That parsing code is small, but it is the kind of thing that breaks quietly in production.
We have caught ourselves adding a small LLM call to decide whether to make a big LLM call. It works. It is also a sign that the tool doesn’t quite fit the job.
We have watched latency creep up as workflows grew. Each new agent or check made sense on its own. Together, they made the system feel slow, and the extra seconds were mostly spent on routing and validation, not on the part the user cared about.
We have had to explain inference bills to clients, and trace which calls were adding real value and which were simply deciding what to do next.
We have tuned prompts and temperature settings for consistency, only to see behaviour shift when the model version changed underneath us.
And on every project, we have balanced the same four forces: intelligence, accuracy, cost and speed. Improving one usually cost us another.
Looking back, if a model like Jev had been available earlier, we think our architectures would have looked different. Many routing, classification and validation steps could have gone to a lightweight decision layer. Generative models could have been saved for the steps that truly needed them: writing responses, summarising, reasoning through unusual cases. Workflows would likely have been faster, with fewer unnecessary calls and lower API bills for our clients. They would have been more stable, because the most frequent decisions would come from a component designed for consistency. And they would have been easier to watch. A dashboard of typed decisions and confidence scores is far easier to monitor than a log full of generated text.
Most of all, there would have been a clean line between deciding and generating. That kind of separation is good engineering in any system. With AI it has been hard to achieve, because the same model was doing both jobs.
We are not saying Jev would have solved every problem. We are saying it addresses a gap we kept running into, and that gap is about architecture, not about writing better prompts.
The economics of AI agents are mostly about call counts. Here is a simple, hypothetical example to show the shape of the change. No real prices or benchmark figures are involved.
Imagine a support agent that handles 100,000 requests a month. For each request, the current design makes six decision calls to an LLM (intent, scope, routing, retrieval choice, sufficiency check and output check) plus one generation call to write the answer. That is 700,000 LLM calls a month.
Now move the decisions to a System 1 layer. Because Jev answers several questions in one pass, the first four decisions could become a single call. The sufficiency and output checks might be one or two more. The generation call stays. And some requests, such as order status checks, could be answered from a template once they are correctly classified, so they never reach the LLM at all.
The result is around 100,000 generative calls or fewer, plus a few hundred thousand very cheap decision calls. The expensive calls fall sharply. The cheap calls rise. Whether this saves money in practice depends on real prices and real accuracy, but the direction is clear: expensive intelligence gets used only where it earns its cost.
There is an interesting twist here. Jev takes its name from the economist William Stanley Jevons. Jevons is known for the observation that when something becomes cheaper to use, people often end up using more of it, not less. We suspect the same could happen with AI decisions. The biggest benefit of cheap decisions may not be a smaller bill. It may be that teams can finally afford to check everything they should have been checking all along. And systems that check more are systems you can trust more.
It would be easy to frame all this as System 1 against generative AI. That framing is wrong.
A busy restaurant kitchen has an expediter at the pass. The expediter reads every order, decides which station gets it, checks each plate before it goes out and sends back anything that isn’t right. The expediter doesn’t cook. The chefs don’t manage the flow of orders. The kitchen works because each role sticks to what it does best.
That is the relationship we see between Jev and generative models. A simple way to put it:
Jev answers “What should the system do?”
GenAI answers “What should the system say or create?”
This is a way of thinking, not a rigid rule. In real systems the two will work in loops. A generative model might produce a plan, and Jev might choose which step to run next. Jev might classify a request, a generative model might draft the reply, and Jev might check that the reply meets policy. A generative model could even help engineers draft the question schemas that Jev answers.
Some early demos shared by developers show a two-speed pattern: a generative model sets goals every few seconds while a System 1 model makes quick, moment-to-moment choices in a much faster loop. That mirrors how people work. We plan slowly and act quickly, and each mode supports the other.
Retrieval-augmented generation (RAG) is the most common pattern in enterprise AI today. It is also one of the clearest examples of hidden decision-making.
A good RAG chatbot has to make many choices before it writes a single word. What is the user actually asking? Does this even need retrieval, or is it a greeting or a follow-up? Which knowledge source is relevant: the HR handbook, the product docs or the contracts library? Which metadata filters apply, such as region, product version or date? Are the retrieved passages really relevant? Is there enough to answer, or should it search again with a different query? Should it answer at all, or admit it doesn’t know, or pass the user to a person?
Many RAG systems handle these choices poorly or skip them. They retrieve on every message and pass whatever comes back to the LLM, hoping the model sorts it out. That is where a lot of confident wrong answers come from. The model was handed weak context and asked to write something anyway.
A fast decision layer could take on most of these choices, leaving the generative model to do what it does best: turn good context into a clear, well-written answer.
An employee asks: “Can I carry over unused leave into next year if I’m based in the UK office?”
Without a System 1 layer: The message goes straight to retrieval. The system searches all documents using the raw question and pulls the top results, which include the India leave policy, a general benefits FAQ and a UK policy from two years ago. Everything goes to the LLM, which is asked to answer. The model may blend the policies into a plausible but wrong answer. Or the team adds one more LLM call to judge whether the context is good enough and another to review the reply. Several generative calls, each adding seconds, and the key question (which policy applies here?) is left to the model’s reading of a large, mixed context.
With a System 1 layer: One Jev call answers several questions at once: intent (policy question), retrieval needed (yes), source (HR policies), region (UK), sensitivity (standard). Retrieval runs with the right filters. A second Jev call scores each passage for relevance and asks whether the context is enough to answer, with a confidence value. If confidence is high, the LLM writes the answer from a small, clean context. If it is low, the system says it couldn’t find a clear answer and offers to connect the employee with HR. One generation call, a couple of cheap decision calls, and a clear point where the system chooses not to guess.
The real shift is that “should I answer?” becomes an explicit, logged decision instead of something the generative model decides by accident.
The same idea reaches well beyond agents and chatbots.
Traditional automation is deterministic but brittle. A rule like “if the subject contains ‘refund’, send it to billing” is fast, cheap and predictable, right up until a customer writes “I want my money back” and the rule misses it. Teams end up maintaining long, fragile lists of rules that never quite cover real life.
Generative AI solved the flexibility problem but brought new ones. It understands messy language, but it is slower, costs more per decision and can behave unpredictably. For high-volume automation, running an LLM on every single item is often too slow or too expensive.
A System 1 model could sit in the middle. It understands messy input the way an AI model does, but answers in the fixed, typed way automation needs, at speeds and prices much closer to ordinary software. That combination opens up a lot of everyday business processes.
In customer support, it could sort and route tickets by intent, urgency and product before any person or agent sees them. In fraud and risk triage, it could score transactions or applications in real time and pass only the unclear ones to deeper review. In document processing, it could decide what kind of document has arrived, whether it is complete and which workflow it belongs to. In CRM automation, it could tag leads, judge buying intent from emails and pick the right follow-up sequence. In IT operations, it could sort alerts by severity and likely cause, cutting through the noise during an incident. In compliance, it could flag records or messages that need a closer look. In data pipelines, it could label and filter large volumes of records, a job that is often too costly for generative models at scale. And for API and email routing, it could decide where each incoming request belongs based on what it means, not which keywords it happens to contain.
A finance team receives thousands of supplier invoices a month by email, in many different formats.
Without a System 1 layer: One option is rules and templates. They work for the biggest suppliers and fail on everyone else, sending a large share of invoices to manual handling. The other option is to send every invoice to an LLM to read, classify, check for problems and decide the next step. That copes with variety, but it means one or more generative calls per invoice, slower processing during month-end peaks, and text outputs that must be parsed before the finance system can use them.
With a System 1 layer: Each incoming email goes through one Jev call with a set of questions. Is this an invoice, a credit note or something else? Is there a purchase order number? Does the supplier look like a known vendor? Is anything unusual, such as changed bank details or an amount far outside the normal range? Every answer comes with a confidence value. Clean, high-confidence invoices go straight to standard extraction and matching. Anything flagged, like changed bank details, goes to a person right away. Only the truly unusual cases, such as a disputed invoice with a long explanatory email, go to a generative model, which summarises the issue for the finance team. The arithmetic (totals, tax, matching amounts) stays in normal code, where it belongs.
The difference is not only fewer model calls. It is a clearer decision boundary. Routine items move automatically, risky items go to people, and generative AI is used where reading and summarising genuinely help.
Put these ideas together and a layered picture starts to form. Each layer handles what it is good at and passes work upward only when it has to
Code handles anything that must be exact: calculations, permissions, hard business limits. System 1 models handle fast judgment on messy input. Generative and reasoning models handle language and complex thinking. People handle the cases where the stakes are high or the system isn’t confident.
This is not a new idea in organisations. Most well-run teams already work this way. Clear procedures cover routine work, experienced staff make quick calls, specialists take the complex problems, and leaders make the high-stakes decisions. What is new is that AI may now have a proper layer for the second of those roles.
For leaders evaluating Agentic AI, RAG or intelligent automation, a few practical steps follow from this.
Start by auditing your AI calls. List every LLM call your system makes and mark each one as either “deciding” or “generating.” In our experience, teams are often surprised by how many of their calls are decisions in disguise. That list is your map of where a System 1 layer could help.
Begin with decisions that are high in volume and low in risk. Ticket tagging, routing and relevance scoring are good candidates. They are frequent enough to show real differences in speed and cost, and a mistake there is easy to catch and fix.
Build your evaluation on your own traffic. Vendor benchmarks, including TypeSafe’s, are a starting point, not proof. Take a few hundred real examples, label them properly, and compare a System 1 approach with your current LLM approach on accuracy, calibration, latency and cost.
Design your questions with care. A System 1 model can only choose from the options you give it. If the right answer isn’t on the list, it will still pick something. Keep questions narrow, make the options clear, and include an “other” or “none of these” option where it makes sense. Enforce logical rules between related answers in your own code.
Use confidence as a control, and watch it over time. Set thresholds for acting, escalating and asking a person. Then track how confidence scores move. A rise in middling scores often means new kinds of input are arriving that your questions weren’t designed for.
Keep hard rules in code. Maths, permissions and legal limits should never depend on a model’s judgment, however fast or cheap it is.
Finally, think about the operational side. Jev is currently in early access through a hosted API. Before relying on it, look at data handling, availability and how dependent you want to be on a single provider. A sensible approach is to build your decision layer as a clean interface, so the model behind it (Jev, a constrained LLM or something else) can change without redesigning the whole system.
For the past few years, progress in AI has mostly meant bigger, more capable models, and the default architecture has been one powerful model doing everything. That approach gave us working prototypes quickly. It is much less suited to systems that must be fast, affordable and predictable at scale.
What comes next may look less like a single genius and more like a well-run team. Code for what must be exact. Fast, cheap judgment for the hundreds of small decisions every workflow needs. Powerful generative models for the work that truly calls for language and reasoning. People for the moments that matter most. Multiple forms of intelligence, each doing what it is best suited for.
Whether Jev becomes the model that defines this middle layer is still an open question, and independent testing will answer it. But the need for the layer is real. We have felt its absence in every agentic system we have built.
The future of Agentic AI may not depend only on making generative models more capable. It may depend just as much on a fast, inexpensive decision layer that knows when, where and how to use generative intelligence, and when not to use it at all.