Table of Contents
Most AI agents make decisions and produce outputs by burning through tokens and making a multitude of tiny decisions. They can answer questions like: Should this email go to billing or support, or is this request safe to run? However, digging deeper into the infrastructure of these AI agents reveals that their decision-making process is inefficient and can be costly for businesses. A new kind of model called Jev is built to fix this exact dilemma.
What is Jev AI?
Jev is a proprietary model from TypeSafe AI, a San Francisco-based company founded in 2024. Unlike large language models (LLMs), it does not generate natural-language text. A user can send it a “state.” Simply put, a state is a snapshot of the situation it needs to judge. It could be a customer’s email, a log of what an agent just did, or a game screen. Then, the user can ask Jev about their state, like “Did the task succeed?” or “Is this a billing issue?” Jev replies with structured answers based on the provided formats. It can reply with a choice from the list you supplied, a score, or a yes/no, each with an associated probability.
System 1 vs. System 2 Models
The names “System 1” and “System 2” come from Daniel Kahneman and his framing of human thinking. System 1 is fast and intuitive, while System 2 is slow and deliberate. Existing LLMs, with their chain-of-thought and multi-second reasoning traces, sit firmly in the System 2 category.
System 1 thinking is driven by the part of your brain that recognizes a friend’s face, reads a stop sign, or answers “2+2” without any conscious calculation. You do not need to weigh options in System 1 thinking because you just know.
System 2 is slow, deliberate, and effortful thinking. This system kicks in when you do long division, plan a road trip, or decide whether to take a new job. It requires focus, and you can feel yourself “thinking it through.”
| Aspect | System 2 (Claude, GPT agents) | System 1 (Jev) |
| Output | Free-form text, code, plans | Typed decisions plus probabilities |
| Strength | Open-ended reasoning, writing, debugging | Bounded, repeatable judgments |
| Speed and Cost | Seconds, priced per token | Built to be fast and cheap |
| Failure Mode | Can hallucinate or drift off-format | Can only answer the question you wrote |
Neither system is better, as they do different jobs. The most interesting part is combining the two. For now, Jev AI is the only System 1 AI on the market.
How Jev Fits into a Claude or GPT Agent
When integrating System 1 and System 2, there is a clear division of labor. The LLM handles the System 2 work of writing and reasoning, while Jev handles the System 1 work of making decisions intuitively, quickly and many times. Here is how that could look inside an agent.
1. Router
Model-routing middleware can use Jev to evaluate a request and select the most appropriate model based on defined criteria. Simple tasks can be routed to faster, lower-cost models. On the other hand, complex requests can be escalated to stronger models. This ensures that Claude or GPT agents are only invoked when deep reasoning is required.
Jev returns a classification and a confidence score in milliseconds, and your middleware routes based on that answer:
- Simple, well-defined requests go straight to a fast, low-cost path, sometimes skipping the LLM entirely if the answer is a lookup or a template response.
- Ambiguous or complex requests get escalated to Claude or GPT, where the cost of deeper reasoning is justified.
- Borderline cases, where Jev’s confidence is low, can be flagged by default for a stronger model or human review, so uncertainty fails safe rather than being silently misrouted.
For example, a customer support agent might use Jev to sort incoming messages into “password reset” (handled by a scripted flow, no LLM needed), “billing dispute” (routed to a mid-tier model), or “contract negotiation” (escalated to Claude with full reasoning). As a result, the most expensive model runs only when the task actually needs it, rather than processing every request end-to-end regardless of difficulty.
2. Guardrail and Verifier
Imagine a claims-processing agent reviewing a customer submission. The AI agent reviews several records and attempts to submit a claims update. However, the underlying system returns an error message that states that the update has failed. However, on the customer side, the agent has told them that the claim was successfully updated.
Before the responses are sent, Jev can review the evidence and notice a mismatch. Once Jev sees the discrepancy, it will flag the response and unsupported and assign a high probability that the agent’s conclusion is incorrect. Instead of allowing a potentially costly mistake to reach the customer, the organization catches the error automatically. In this role, Jev acts like a quality-control reviewer, checking the work of a more powerful but less predictable AI.
3. Confidence-Based Escalation
Imagine a healthcare benefits assistant is trying to determine whether treatment is covered by an insurance plan. When a member asks if their upcoming procedure will be covered by their plan, Jev can review the policy and provide a concise answer and a confidence score.
Jev AI determines the confidence score using three decision paths. It has three levels of confidence: high, medium, or low. Below is an explanation of what each means:
- High confidence: Jev answers directly, and the agent relays the answer to the member without needing a human. A clearly listed, unambiguous benefit falls here.
- Medium confidence: The language of the policy is vague, or the case has unusual details. Instead of guessing, the system asks the member a clarifying question (dates of service, in-network vs. out-of-network provider, prior authorization on file) and re-checks with the added context.
- Low confidence: The case involves a genuinely ambiguous clause or conflicting language of policy. It is routed to a benefits specialist for manual review, with Jev’s reasoning attached, so the human is not starting from zero.
The more uncertainty it has, the more review a human must do. Not every decision deserves the same level of scrutiny. Jev helps determine when the system can ask autonomously, when it should ask for more information, and when expert review is warranted. This allows organizations to automate routine cases while maintaining oversight where uncertainty is highest.
4. Plumbing and Integration
LangChain provides access to Jev through a TypeSafeClassifier integration. Additionally, TypeSafe offers an agent skill that can be installed in Claude Code and other coding agents. Jev operates as a hosted model accessible through TypeSafe’s HTTP API, making it easy to incorporate as:
- A tool call within an agent workflow.
- Middleware in an orchestration layer.
- A classification or verification step in existing AI systems.
This enables Jev to fit naturally into modern agent architectures with minimal integration overhead.
Where System 1 Models Are Useful
Because of its different thinking process, the uses for system 1 models also differ from those of system 2 models.
Support and Ticket Triage: Deciding between billing versus technical support without paying for a full language model.
Moderation Pipelines: System 1 models can run many yes/no policy checks in parallel on the same input.
Agent Tool Selection: Pick which tool or sub-agent runs next, with a confidence score attached.
Workflow Branching: True/false checks inside automations, where a malformed output would break something downstream.
Extraction by Selection: Jev copies rather than generating original content, which avoids a model subtly altering a number or address.
Real-time control: Community demos have shown Jev controlling Minecraft, Subway Surfers, drones, and driving sims, though the same coverage stresses these aren’t proof it has solved robotics or game AI.
Conclusion
The agent era has meant using one big model for every task. Jev points to a different design: a slow, thoughtful System 2 model for reasoning and writing, and a fast System 1 model for the dozens of small judgment calls around it. Even if Jev itself does not hold up, the architecture is worth adopting. Start small by putting one classifier in front of your agent. Then measure it against your current approach and expand from there.
If you are looking to integrate a system like Jev AI, reach out to AppsChopper to discuss the possibilities today.







