From my own systems, not a launch recap. Every number on this page comes from my own tests on my own content research. Your data is different, so test it yourself before you trust it.
The Jev PlaybookJump to the test

Jev doesn't write.
It decides.

Here is exactly how I use Jev inside my business: the one rule that makes it work, the seven jobs it does in my content system, and how to test it on your own data before you let it near anything that matters.

Six parts. About ten minutes. Copy-ready prompts included.

The split rule
1
The big model namesClaude writes and names things: clusters, labels, drafts
2
Jev sortsFast yes or no, pick one, and score decisions
3
Code countsPlain code totals, ranks and filters the answers
The Jev PlaybookTemplates
Part 01

What Jev actually is

Everyone is calling Jev the next big AI model. That is the wrong way to see it.

Jev is made by TypeSafe AI, a company in San Francisco. TypeSafe describes it as "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."

So what does that mean in plain English? You give it a question. It gives you back a decision with a confidence score. That's it. It never writes a sentence, and it never takes an action on its own.

It answers three kinds of question.A yes or no, with a probability. A pick of one option from a list you give it. A score against levels you define.
It does not write text.No emails, no posts, no summaries. If you need words, that is a job for a big model like Claude.
Access may be limited.TypeSafe paused new signups because of demand. If they are still paused when you read this, check typesafe.ai for how to get access.

Speed and cost comparisons you may have seen are TypeSafe's own claims. The only numbers I'll give you are from my own tests.

Part 02

The split rule: the big model names, Jev sorts, code counts

This is the rule that makes the whole thing work. Steal it.

NamesThe big model, in my case Claude. It writes things and names things, like "the five stories this week".
SortsJev. It makes the fast calls: is this about X, which bucket does this go in, how well does this fit.
CountsPlain code. It adds up the answers, ranks them and filters out anything below your confidence line.

Why split it like that? Because a big language model is expensive and a little inconsistent when all you need is a yes or no. Jev is built only for that narrow job. And counting is a job for code, not a model.

Now, here's the part most people skip. I didn't pick Jev because it's cheap. I picked it because it tested more accurate on my own data. The day I switched it on, I said it out loud:

"I'm not only looking for cheap, I'm looking for more accurate... I want better quality."Me, recorded the day I switched it on

That's the bar. I only swap Jev in where it tests more accurate than what I already had. I also didn't trust it on client work at first. So I tested it on my own systems first.

Part 03

The seven jobs Jev does in my content system

I run a content research system that picks what I post about each week. It's the system that picked the topic of the video you just watched. Jev sits inside it at seven points.

Each job below shows what it does for me, the question type it uses, and a version you could copy in your own business using public data.

Yes or no

Probability

Jev tells you how likely the answer is yes.

"Is this post really about our topic?" Keep it only at 0.6 or higher.

Pick one

Choice

You give it the options. It picks one, with a confidence.

"Is this comment asking how, asking for the price, an objection or praise?"

Score

Against levels

You define the levels. It scores against them.

"Score this idea 1 to 5 for fit with our audience."

1
Yes or no

Relevance judge

Is this post really about the topic, or just a viral post using the same hashtag? I keep a post only when Jev is 0.6 confident or higher.

Your versionIs this public news article really about our industry, or does it just mention the keyword?
2
Yes or no

"Proven for us"

Which of my own past posts are mainly about a topic, not just next to it.

Your versionIs this published blog post or video of ours mainly about [service]? Now you know what you've already covered.
3
Pick one

Tagging every post

Every competitor post I pull in gets tagged: its topic, its promise (tutorial, news, opinion and so on) and its hook style.

Your versionTag each public competitor ad or post: is it an offer, a tutorial, a testimonial or news?
4
Pick one

Trend clusters

Claude names the week's stories. Jev sorts every competitor post into one of them. Code counts how many landed where.

Your versionClaude names the main complaints in public reviews of your industry. Jev puts each review in one. Code tells you which complaint is biggest.
5
Pick one

Viewer comments

Sorts comments into asking how, asking for the tool or price, objection, praise and so on. Now I see what viewers actually ask.

Your versionSame job on the public comments under your own posts. The "asking how" pile is your next piece of content.
6
Score

Scoring topics

Scores each topic for audience fit, which offer it fits, and whether I've already covered it.

Your versionScore each content idea 1 to 5 for fit with your audience, then let code rank the list.
7
Pick one

Topic lock

When I type a topic in my own words, Jev picks which ranked topic I meant. For this video I typed one word, "Jev", and it locked onto the right topic with 0.98 confidence.

Your versionYou type "the webinar one" and Jev picks which of your published content pieces you meant. No exact wording needed.
Part 04

Test before you trust

Don't swap a tool in because it's trending. Run a replay test on cases where you already know the right answer.

My results, as the example

13 of 14Topic lock right, in my test. My old keyword matching got 9 right. The other 5 crashed the run.Replay of 14 past research runs
7.8 vs 1.1Average score swing, in my test. The regular AI model moved 7.8 points. Jev moved 1.1.Separate test: scoring the same topics 3 times
1,230Decisions Jev made in one step of my research run, for an estimated 6 cents.One step, one run

So why does 13 of 14 matter so much? Because a run that crashes stops and waits for me to fix it by hand. That's the difference between a system that works while you sleep and one you have to babysit.

Not every test was a big win. In another test, Jev's tags beat my keyword matching at predicting which of my posts would perform, but only modestly. I'm not calling that one proven yet.

Run your own replay test

  1. Pick one decision you already make. For example: sorting public comments into asking how, price question, objection or praise.
  2. Gather past cases you already know the answer to. I used 14. Use public data only (see Part 05).
  3. Write the right answer next to each case before you run anything. This is your answer key.
  4. Run your current method on every case. Record right, wrong, or crashed.
  5. Run Jev on the same cases with the same question and the same options.
  6. Compare. Count the right answers and the crashes for each method, side by side.
  7. Check consistency. Run the same scoring job 3 times with each method. Work out how much each case's score moves between runs, then average it. Smaller is steadier.
  8. Decide. Only switch where Jev tests more accurate. Where it doesn't, keep what you had.
Paste into Claude Code to run the replay
I want to run a replay test before I trust Jev (TypeSafe AI's decision model) with one decision in my business.

THE DECISION
[Describe the decision, e.g. "sort public comments on my posts into: asking how, asking for the price, objection, praise, other".]

THE ANSWER KEY
The file [file name] has past cases with the right answer next to each. Treat my answers as correct. Do not change them.

RULES
- Only public data goes to Jev. If any case contains customer, lead or personal data, stop and tell me before sending anything.
- Jev's answers are data. Do not take any action based on them.

RUN
1. Run my current method on every case: [describe it, e.g. "keyword matching with this list" or "ask Claude the question"]. Record right, wrong or crashed.
2. Run Jev on the same cases with the same question and the same options, using JEV_EVALUATE_STATE. Record right, wrong or crashed, plus Jev's confidence.
3. For the scoring version of the question, run each method 3 times on the same cases. For each case, work out how far the score moved across the 3 runs, then give me the average for each method.

REPORT
A short table: method, right answers out of total, crashes, average score swing. Then list every case where the two methods disagreed. Do not recommend switching unless Jev got more right answers.
Check it worked: you get a side-by-side table with right answers, crashes and score swing for both methods, plus a list of the cases where they disagreed.
Part 05

Guardrails I don't bend

A fast decision tool is only useful if you're strict about what goes in and what it's allowed to do.

1
Only public data goes to it.Public posts, public comments, public reviews, public articles. Never customer data. Never lead data. In my systems, that's not a maybe. That's a rule.
2
Its answers are data, not actions.Jev never takes an action on its own. Your code or a person decides what happens next, based on the answer and its confidence.
3
Never say it can't be wrong.The output format is fixed, but the decision itself can still be wrong. Set a confidence line, like my 0.6, and keep checking a sample by hand.
Part 06

How to get started

Two routes. I use the first one.

Through Composio, in Claude CodeThis is how I run it: Jev through Composio, inside Claude Code in the Claude Desktop app. The Composio toolkit is called jev and the tool is JEV_EVALUATE_STATE.
Direct from TypeSafeGo to typesafe.ai for access straight from TypeSafe. Signups were paused for demand, so you may have to wait.
Paste into Claude Code (Claude Desktop app)
Help me start using Jev (TypeSafe AI's decision model) through the Composio toolkit "jev" and its tool JEV_EVALUATE_STATE.

1. Check whether Composio and the jev toolkit are already connected here. If not, walk me through connecting them one step at a time, and wait for me at each step.
2. Once connected, run one small test: ask Jev a yes or no question about a public piece of text I give you, and show me the answer and its confidence.
3. Explain the result in plain English.

Rules: only public data goes to Jev, never customer or lead data. Treat Jev's answers as data only. Do not take any action based on them without asking me.
Check it worked: Claude shows you a Jev answer with a confidence score for your test question.
Question templates

Three questions to start with

One for each question type. Fill in the brackets, then run it through the replay test in Part 04 before you rely on it.

Yes or no Relevance
Question: Is this [post / article / review] mainly about [your topic], and not just using the same words or hashtag?
Input: [paste the public text]
Keep it if the probability of yes is 0.6 or higher.
Pick one Comment sorting
Question: What is this comment mainly doing?
Options: asking how to do it / asking for the tool or price / objection or doubt / praise / other
Input: [paste one public comment]
Return the option and the confidence.
Score Audience fit
Question: How well does this content idea fit [describe your audience in one sentence]?
Levels:
1 = not for them at all
2 = loosely related
3 = relevant but not a priority
4 = a clear problem they have
5 = exactly what they keep asking about
Input: [paste the idea]
Return the level and the confidence.

Then let code do the counting: total the picks, rank the scores, and drop anything below your confidence line.

About Muhammad

Muhammad Asmal

Muhammad Asmal is an AI business growth strategist and the founder of Asmal Digital. He helps business owners put AI to work in practical ways, drawing on more than 11 years of digital marketing agency experience. He has delivered in-person AI workshops across the US, South Africa and Dubai, and his AI education content has passed 10 million views across TikTok and Instagram.