What Jev actually is
Everyone is calling Jev the next big AI model. That is the wrong way to see it.
Jev is made by TypeSafe AI, a company in San Francisco. TypeSafe describes it as "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."
So what does that mean in plain English? You give it a question. It gives you back a decision with a confidence score. That's it. It never writes a sentence, and it never takes an action on its own.
Speed and cost comparisons you may have seen are TypeSafe's own claims. The only numbers I'll give you are from my own tests.
The split rule: the big model names, Jev sorts, code counts
This is the rule that makes the whole thing work. Steal it.
Why split it like that? Because a big language model is expensive and a little inconsistent when all you need is a yes or no. Jev is built only for that narrow job. And counting is a job for code, not a model.
Now, here's the part most people skip. I didn't pick Jev because it's cheap. I picked it because it tested more accurate on my own data. The day I switched it on, I said it out loud:
That's the bar. I only swap Jev in where it tests more accurate than what I already had. I also didn't trust it on client work at first. So I tested it on my own systems first.
The seven jobs Jev does in my content system
I run a content research system that picks what I post about each week. It's the system that picked the topic of the video you just watched. Jev sits inside it at seven points.
Each job below shows what it does for me, the question type it uses, and a version you could copy in your own business using public data.
Probability
Jev tells you how likely the answer is yes.
"Is this post really about our topic?" Keep it only at 0.6 or higher.
Choice
You give it the options. It picks one, with a confidence.
"Is this comment asking how, asking for the price, an objection or praise?"
Against levels
You define the levels. It scores against them.
"Score this idea 1 to 5 for fit with our audience."
Relevance judge
Is this post really about the topic, or just a viral post using the same hashtag? I keep a post only when Jev is 0.6 confident or higher.
"Proven for us"
Which of my own past posts are mainly about a topic, not just next to it.
Tagging every post
Every competitor post I pull in gets tagged: its topic, its promise (tutorial, news, opinion and so on) and its hook style.
Trend clusters
Claude names the week's stories. Jev sorts every competitor post into one of them. Code counts how many landed where.
Viewer comments
Sorts comments into asking how, asking for the tool or price, objection, praise and so on. Now I see what viewers actually ask.
Scoring topics
Scores each topic for audience fit, which offer it fits, and whether I've already covered it.
Topic lock
When I type a topic in my own words, Jev picks which ranked topic I meant. For this video I typed one word, "Jev", and it locked onto the right topic with 0.98 confidence.
Test before you trust
Don't swap a tool in because it's trending. Run a replay test on cases where you already know the right answer.
My results, as the example
So why does 13 of 14 matter so much? Because a run that crashes stops and waits for me to fix it by hand. That's the difference between a system that works while you sleep and one you have to babysit.
Not every test was a big win. In another test, Jev's tags beat my keyword matching at predicting which of my posts would perform, but only modestly. I'm not calling that one proven yet.
Run your own replay test
- Pick one decision you already make. For example: sorting public comments into asking how, price question, objection or praise.
- Gather past cases you already know the answer to. I used 14. Use public data only (see Part 05).
- Write the right answer next to each case before you run anything. This is your answer key.
- Run your current method on every case. Record right, wrong, or crashed.
- Run Jev on the same cases with the same question and the same options.
- Compare. Count the right answers and the crashes for each method, side by side.
- Check consistency. Run the same scoring job 3 times with each method. Work out how much each case's score moves between runs, then average it. Smaller is steadier.
- Decide. Only switch where Jev tests more accurate. Where it doesn't, keep what you had.
I want to run a replay test before I trust Jev (TypeSafe AI's decision model) with one decision in my business. THE DECISION [Describe the decision, e.g. "sort public comments on my posts into: asking how, asking for the price, objection, praise, other".] THE ANSWER KEY The file [file name] has past cases with the right answer next to each. Treat my answers as correct. Do not change them. RULES - Only public data goes to Jev. If any case contains customer, lead or personal data, stop and tell me before sending anything. - Jev's answers are data. Do not take any action based on them. RUN 1. Run my current method on every case: [describe it, e.g. "keyword matching with this list" or "ask Claude the question"]. Record right, wrong or crashed. 2. Run Jev on the same cases with the same question and the same options, using JEV_EVALUATE_STATE. Record right, wrong or crashed, plus Jev's confidence. 3. For the scoring version of the question, run each method 3 times on the same cases. For each case, work out how far the score moved across the 3 runs, then give me the average for each method. REPORT A short table: method, right answers out of total, crashes, average score swing. Then list every case where the two methods disagreed. Do not recommend switching unless Jev got more right answers.
Guardrails I don't bend
A fast decision tool is only useful if you're strict about what goes in and what it's allowed to do.
How to get started
Two routes. I use the first one.
jev and the tool is JEV_EVALUATE_STATE.Help me start using Jev (TypeSafe AI's decision model) through the Composio toolkit "jev" and its tool JEV_EVALUATE_STATE. 1. Check whether Composio and the jev toolkit are already connected here. If not, walk me through connecting them one step at a time, and wait for me at each step. 2. Once connected, run one small test: ask Jev a yes or no question about a public piece of text I give you, and show me the answer and its confidence. 3. Explain the result in plain English. Rules: only public data goes to Jev, never customer or lead data. Treat Jev's answers as data only. Do not take any action based on them without asking me.
Three questions to start with
One for each question type. Fill in the brackets, then run it through the replay test in Part 04 before you rely on it.
Question: Is this [post / article / review] mainly about [your topic], and not just using the same words or hashtag? Input: [paste the public text] Keep it if the probability of yes is 0.6 or higher.
Question: What is this comment mainly doing? Options: asking how to do it / asking for the tool or price / objection or doubt / praise / other Input: [paste one public comment] Return the option and the confidence.
Question: How well does this content idea fit [describe your audience in one sentence]? Levels: 1 = not for them at all 2 = loosely related 3 = relevant but not a priority 4 = a clear problem they have 5 = exactly what they keep asking about Input: [paste the idea] Return the level and the confidence.
Then let code do the counting: total the picks, rank the scores, and drop anything below your confidence line.
Muhammad Asmal
Muhammad Asmal is an AI business growth strategist and the founder of Asmal Digital. He helps business owners put AI to work in practical ways, drawing on more than 11 years of digital marketing agency experience. He has delivered in-person AI workshops across the US, South Africa and Dubai, and his AI education content has passed 10 million views across TikTok and Instagram.