LLMs Can Generate Paragraphs. But What If You Just Need a Decision? Meet Jev
There’s a new AI model that’s been getting people talking, and you might have heard of it. It’s called Jev.
But what exactly is Jev, and what makes it different from the general-purpose LLMs we use every day?
Let’s take a look at what Jev is all about.
1. Meet Jev
What Is Jev?
Jev is an AI model from TypeSafe AI designed to make structured decisions rather than generate long-form text.
To understand why it exists, let's start with a simple problem.
Imagine you're building an application that receives 100,000 customer messages. For every message, your application needs to decide:
- Is this a billing problem?
- Is this a technical problem?
- Is this an account problem?
You could send every message to a general-purpose LLM like GPT and ask it to classify each message. But you're using a model capable of writing essays and generating code just to select one of three options.
Jev is designed for this narrower kind of task: making decisions based on the information provided.
Three Things Jev Can Answer
Jev provides three types of decisions called Choice, Score, and Noul.
1. Choice
Which option is correct?
- Input: "My payment failed, and I was charged twice."
- Options:
billing,technical,account - Example output:
billing
This is useful when you need to classify information into a predefined set of options.
2. Score
How well does this fit a defined scale?
- Input: "The production service is down for every customer."
- Scale:
0 = minor,1 = moderate,2 = severe - Example output:
2.91
This is useful when you need to evaluate something against a predefined scale.
Note that Score returns a continuous value along your scale, not a whole number. An output of 2.91 sits very close to "severe", so your application can round it or apply its own thresholds.
3. Noul
How likely is a particular statement to be true?
- Input: "The production service is down."
- Question: "Does this incident require immediate escalation?"
- Example output:
0.95, meaning a 95% estimated probability of a positive answer, assuming the output represents a calibrated probability.
This is useful when you need a structured answer to a specific question rather than a long explanation.
Take Jev for a Spin
The fastest way to try it is through the Jev Playground.
Go to https://jevplayground.com/.
It's an independent playground where you can experiment with Jev.
Choice
Here, you can see some information in the state textbox, along with a question and a set of choices based on that information.
In this example, the state describes a billing problem, and the question asks which team should handle the customer's request.
When we run it through Jev, it selects Billing, which is the correct choice.
Score
This checks the severity of the issue.
Noul (Boolean)
This checks whether the issue is financial or not.
2. Working with Jev
These are the basics. Let's go through some more details now.
Let's look at how to work with Jev more effectively. We'll cover the types of input it accepts, how to structure a request, how it evaluates multiple questions, and how to design Choice questions that avoid forced or misleading answers.
Understanding State
In Jev, state is the information the model uses to evaluate your questions.
A state can be provided in three formats:
- String: Plain text, such as a customer message or incident report.
- JSON object: Structured information containing fields such as an order ID, customer status, or error message.
- Array: A collection of context items that provide information for a decision.
For example, a customer support application might provide this state:
{
"customer_message": "My payment failed, but I was charged twice.",
"account_status": "active",
"previous_contact": "No previous complaints"
}
Jev can use this information to answer questions about the customer's issue, account, or next steps.
However, Jev is a text-based model. It does not directly accept images, audio, video, or other binary inputs as state. If your application receives these formats, you need to extract relevant information into text or structured data before sending it to Jev.
Understanding Request Anatomy
A Jev request consists of three main components:
model: The Jev model you want to use.state: The information the model should evaluate.questions: A map of named questions you want answered.
Each question has its own identifier and definition. The identifier helps your application locate the corresponding answer in the response.
Here's an example:
{
"model": "jev-latest",
"state": {
"customer_message": "My payment failed, but I was charged twice."
},
"questions": {
"issue_type": {
"type": "choice",
"instructions": "What is the primary issue?",
"criteria": {
"billing": "Payment failures, duplicate charges, or refunds",
"technical": "Application errors or technical failures",
"account": "Account access or account management issues",
"other": "The issue does not fit any of the other categories"
}
},
"needs_follow_up": {
"type": "noul",
"instructions": "Does this customer need a follow-up from support?"
}
}
}
In this example, the request contains two questions about the same customer message.
The first uses Choice to classify the issue. The second uses Noul to evaluate whether follow-up is needed.
Notice that each question has its own type and instructions. Choice uses a criteria field to define the available options. Score uses criteria to define an ordered scale, while Noul can evaluate a yes-or-no question without requiring criteria.
This structure lets you define multiple decisions in one request rather than making a separate request for every question.
Parallel, Independent Evaluation
Jev can evaluate multiple questions in the same request. However, each question is evaluated independently against the shared state.
This means one question does not receive another question's answer as additional context.
Consider an AI agent that needs to process a customer support ticket. You might want it to:
- Identify the issue type.
- Determine its severity.
- Decide whether it needs human intervention.
You can ask all three questions in one request. But the severity question cannot use the answer from the issue classification question, and the escalation question cannot inspect the severity answer.
Each question evaluates the original state independently.
Designing Better Choice Questions
Choice is useful when you need to select one answer from a predefined set of options.
However, the quality of the result depends partly on how you define those options.
Consider this example:
{
"type": "choice",
"instructions": "What is the primary issue?",
"criteria": {
"billing": "Payment failures or duplicate charges",
"technical": "Software bugs or application errors",
"account": "Login or account access problems"
}
}
What happens if a customer asks about a shipping delay?
None of these options fits the issue. Without a suitable fallback, the model may still select one of the available options, producing a misleading classification.
Include an "other" or "none" option
When your predefined choices might not cover every possible input, include a fallback option.
For example:
{
"type": "choice",
"instructions": "What is the primary issue?",
"criteria": {
"billing": "Payment failures or duplicate charges",
"technical": "Software bugs or application errors",
"account": "Login or account access problems",
"other": "The issue does not fit any of the categories above"
}
}
Now, the application has an option for issues that fall outside the expected categories.
This is especially useful when processing real-world data, where customer messages and other inputs may not always fit neatly into predefined classifications.
Make the options distinct
A fallback option is only one part of good question design. The other options should also be clearly distinguishable.
For example, if both billing and account include subscription problems, Jev may have difficulty distinguishing between them.
Define each option in terms of what it covers, and avoid unnecessary overlap. If two options represent different outcomes, their descriptions should make that difference clear.
Remember that Choice selects from the options you provide. It does not automatically expand the list when an unexpected case appears.
3. Try out a practical demo
Getting Started
You can visit the Jev Console, sign up for an account, and get an API key.
You can also try out this project:
https://github.com/RijulTP/jev-lab
Clone the repository and run make run to try the demo mode without an API key.
If you have an API key, run make run-live and add your key to the .env file.
Playground: Running Your First Example
You'll now be on the Playground page.
Select the Login Outage preset.
Here, you can see the state and the types of questions being asked.
The first question uses Choice to determine which team should handle the issue.
The second question uses Score to report the severity of the issue.
When you press Run:
Jev determines the appropriate team to handle the issue and its severity.
Fan Out: Answering Multiple Questions at Once
Let's explore the next page, Fan Out.
Normally, you'd ask an AI one question at a time. For example, you'd first ask which team should handle an issue, then ask how urgent it is. With eight questions, you'd need to make eight separate calls.
Jev's fan-out feature lets you ask all eight questions in a single call. It evaluates the same message against every question at once, with little effect on response time as you add more questions.
On this page, we'll try it with a support ticket and eight questions covering topics such as the issue, its urgency, the customer's frustration, and the product area involved.
Consider this support ticket:
"Our API started returning 500 errors, and we can't process orders."
To handle this ticket, you need answers to eight questions, including which team should handle it, how urgent it is, and whether the message is trying to trick the system.
There are three ways to get these answers:
- Way 1: One question at a time. Send the ticket, ask a question, and wait for the answer. Repeat this process eight times.
- Way 2: Run eight separate checks in parallel. Send eight requests at once. This gets you the answers faster, but you still make eight separate calls.
- Way 3: Ask all eight questions in one call. Send the ticket once, along with all eight questions, and receive the answers together.
Jev's fan-out feature uses the third approach, evaluating all eight questions against the same ticket in a single call.
It can also reduce costs because each call must include the ticket text for Jev to evaluate.
Here's the comparison:
- Eight separate calls: The ticket is sent eight times, costing around 2,852 units.
- One combined call: The ticket is sent only once, costing around 829 units.
Combining the questions reduces the amount of text processed and the overall cost.
Confidence Gates: Using Confidence Scores
Let's move to the next page, Confidence Gates.
Jev doesn't just give you an answer. It also exposes how strongly it favours that answer, and the way it does this depends on the question type.
You can use this score to decide what happens next:
- High confidence: The system handles the request automatically (green: Automatic).
- Medium confidence: A person reviews the answer before taking action (orange: Review).
- Low confidence, risky requests, or trick messages: The request goes directly to a person (red: Human).
Imagine a support bot working the night shift. A customer sends a message, the AI reads it, chooses an answer, and takes action. Most of the time, it gets things right.
But one night, a customer asks a question about visas. The AI can't find a suitable team, so it confidently routes the request to Sales. The customer waits until morning for someone to notice the mistake.
The problem isn't that the AI stopped working. It did what it was designed to do: choose the most likely answer. The real problem is that the system trusted the answer without checking how reliable it was.
That's what this page is about: using confidence scores to decide when to trust an answer and when to ask a human to step in.
Here's how each question type expresses certainty:
- Choice questions (Which team?) Each option gets a probability, and the response includes a separate confidence value. If one option has a much higher probability than the others, confidence is high. If the probabilities are spread across several options, confidence is low.
- Score questions (How urgent is it?) Jev uses the same idea across a range of possible scores, and also returns a separate confidence value.
- Noul (Boolean) questions (Is this a trick message?) Noul doesn't return a separate confidence field, because its output is already a probability. A value of 0.99 indicates a strong Yes, 0.01 indicates a strong No, and 0.5 indicates uncertainty.
These numbers don't guarantee that an answer is correct. They show how strongly Jev favours one answer over the alternatives.
This allows your software to handle answers differently based on their confidence scores instead of treating every answer as equally reliable.
The Three Levels
A support system can use confidence scores to sort requests into three groups:
- Automatic (green): Jev is confident, and the risk is low. The system handles the request on its own, such as routing a ticket or sending a standard reply.
- Review (orange): Jev is less certain. A person checks the answer before the system takes action.
- Human (red): The request is risky, confusing, or contains a trick message. It's sent directly to a person for handling.
You decide the confidence thresholds for each group based on the risk of getting an answer wrong.
For example, a simple balance inquiry might be handled automatically with moderate confidence. However, approving a ₹10 lakh refund may require very high confidence and a human review.
The same AI can handle both tasks. The difference is how much confidence you require before allowing it to take action.
Exploring the Results
Let's move to this page, which contains 30 questions.
Jev processes all 30 questions and categorizes the results based on their confidence levels.
You can see the numbers at the top of the page, which show how the results are distributed across the different categories.
This page lets you see how Jev categorizes the results.
You can adjust the categories using the sliders.
Try lowering the safety slider to around 0.10. You'll notice that even ordinary support tickets start moving into the Human category. This happens because the threshold is now so low that even a small safety concern sends the ticket to a person.
You can explore more thresholds using the dropdown below.
Support Desk Chat: Putting Everything Together
So far, we've explored each part separately.
- Playground: Ask Jev a question and get an answer.
- Fan Out: Ask multiple questions in one call.
- Confidence Gates: Use confidence scores to decide whether to act automatically, ask a person to review, or hand over the request to a human.
Now, let's bring everything together on the Support Desk Chat page.
This page simulates an online store's support desk. It handles orders, refunds, technical issues, account access, product questions, and complaints.
When a customer sends a message, Jev evaluates it, your rules decide what happens next, and the system sends a reply. This brings everything we've explored so far into one working example.
How It Works: Three Simple Steps
Every customer message goes through three steps, each handled by a different part of the system.
1. Jev evaluates the message
Jev answers six questions in a single call. These cover the customer's intent, urgency, frustration, possible attempts to override instructions, safety concerns, and whether they want to speak to a human.
Jev returns the answers along with their confidence scores. It doesn't decide what action to take.
2. Your rules decide what happens next
The system checks a list of rules in order and follows the first rule that matches. For example:
- If the message tries to manipulate the system, refuse the request.
- If someone is in danger, send the case to a human urgently.
- If the customer asks for a human, hand over the conversation.
- If Jev is unsure, ask the customer for more information.
- If the message is unrelated to support, send an out-of-scope reply.
- If the customer is very angry or their issue is blocked, send the case for priority handling.
- Otherwise, use the standard reply for that type of request.
3. The system sends a prepared reply
The system uses a pre-written response based on the decision. Jev doesn't write the message that the customer sees.
Why Keep These Steps Separate?
Each part has a clear responsibility. Jev evaluates the message, your code decides what to do, and a prepared template sends the reply.
This makes the system easier to understand and check. If something goes wrong, you can see whether the problem came from Jev's evaluation, the rules, or the reply.
The What Just Happened panel shows how the system handled each message.
Trying It Out
I've asked a question on this page, and you can see the replies below.
4. Conclusion
This is an emerging approach, with companies like OpenAI developing a Decisions API and Cloudflare developing Clef.
It offers an alternative to using general-purpose LLMs for decision-making. Instead, you can integrate a specialized model like Jev into your workflow to make decisions more efficiently.