What can you build with Jev?

What can you build with Jev?

Practical uses for fast, cheap decision models

Sarim Malik

6 min read

Jev is a new model built to evaluate questions with predefined possible answers. It returns answers in a few hundred milliseconds, in a structured format, with a probability for each possible answer.

Practically, Jev helps you make a decision. You provide it context, questions, and possible answers, and it answers each question using probabilities, which you can then use to make a decision.

We spent the past two weeks building with Jev’s API. In this post, we’ll show practical examples of what you can build with it.

What can you ask Jev?

You can ask Jev three kinds of questions. Every figure in this post calls the Jev API live, and the time and cost in each corner come from your request.

Choice picks the best option from a list you provide, and returns a probability for every option.

Which emoji best represents “Q4 revenue report”?

    Choosing an emoji
    Live
    Using Choice to find the emoji that best represents the title

    Score rates something on a scale you define, like 1 to 5. The score can land between levels, like 4.7.

    How well does the “📊” emoji represent the title “Q4 revenue report”?

    Live
    Using Score to rate how well an emoji represents the title

    Check (Jev calls it Noul, and OpenAI calls it Predicate) returns the probability that a statement is true. Your code decides how sure is sure enough.

    The “📅” emoji is a suitable match for the title “Office lunch menu”.

     
    Live
    Using Check to test whether an emoji represents the title

    Examples

    Prioritizing notifications

    We all get tons of notifications every day from different products, whether natively on your device or via email.

    Knowing which ones need your attention now and which can wait is really critical.

    So instead of doing it manually yourself or passing this on to a reasoning model, which can take many seconds and has real cost per notification, we asked Jev two questions: Does this notification need the recipient to act? And is it time-sensitive?

    Live0/8 requests0 priority
    80%
    Prioritizing incoming notifications with Jev

    A notification is marked priority when both answers come back at 80% or higher. That threshold is our choice. Change it in the figure’s bottom corner, then replay the inbox to sort it against the new cutoff.

    Click the priority pill on any notification to see Jev's answers.

    This example isn't perfect: we're not passing any context about the user.

    Take the first notification, for example, Maya's message in #team about lunch at noon. Our two questions are written for work, so they treat lunch plans as something that can wait. So it's really important how you specify the criteria to the model, because your criteria determine how the model answers your questions.

    Secondly, the model is only going to make a decision as good as the state or context you provide, which is true for any language model. So in another situation, if we provided more context about the user, perhaps its decisions would have been different.

    Searching by meaning

    When you search a web page in your browser, it runs a literal search: it only finds the exact characters you typed.

    Semantic search matches what you mean instead, so it can find an answer that shares none of your words.

    Jev can do semantic search. In the figure below, we ingested every sentence from Wikipedia’s article about platypuses and made it searchable.

    Ask anything about the platypus, or click refresh for another question.

    In this demo, we first split Wikipedia’s article on the platypus into 211 sentences. During each search request, we ask two questions:

    • A Choice question ranks all 211 sentences, using their IDs as the options. Choice accepts up to 255 options, so the whole article fits in one request.
    • A Check (Noul) question asks whether any sentence answers the question, where true means at least one sentence states or directly implies the answer.

    We need both because Choice always ranks some sentence first, even when none answers. Ask "How fast can a platypus swim?" and Choice picks a swimming sentence at 85%, but the article never gives a speed. The Check returns 5%, below our 35% cutoff, so the figure dims the matches.

    One question rarely covers everything you need to know, so it often pays to ask several in one request: here, Choice finds where the answer is, and Check tells you whether there is one.

    Each search takes about 250 ms and costs 0.05¢. If you want to dig further into this use case, check out this guide.

    Stay in touch

    Join our mailing list to receive future posts straight in your inbox.

    Checking spelling

    Think of a Grammarly-style editor that checks your spelling as you type. Each time you finish a word and press space, we send Jev that word and the text around it. Once you type the next word, we check it again, since some mistakes only show up in what comes after.

    The request sends two Checks in one call, the same pattern as the semantic search example above:

    • Non-word: "Is this word, as written, a misspelled non-word?" This catches recieved and publsh. Names, jargon, and regional spellings count as correct.
    • Wrong word: "Is this word spelled correctly but wrong for this sentence?" This catches form for from and loose for lose.

    A typo gets a red squiggly line, and a wrong word gets a blue one.

    LiveTotal cost: 0¢
    Checking spelling as you type

    We underline a typo when Jev is at least 80% sure, and a wrong word at 50%. The wrong-word bar is lower on purpose: in our tests, correct words never scored above 37%, so 50% still leaves a safe margin. Lowering it from 80% to 50% caught 8 of 12 wrong words instead of 5.

    It still misses some easily confused pairs, like then and than or your and you’re. Those scored between 20% and 46%, below our bar.

    If you fix a word yourself, its underline disappears, and Jev checks it again once you pause or press space.

    Because the rules are written in plain language, you could also make the spellcheck programmable: tell it what to ignore, like startup names, acronyms, or slang, and it skips those words.

    Finding settings

    So far, each example asks Jev one round of questions. Here, Jev walks down the menus of iPhone Settings one level at a time, and each answer decides which question we ask next. Sorting something into a tree like this is called hierarchical classification.

    Search the phone for a setting you want to change, or click refresh for another one.

    9:41

    Settings

    Apple Account
    Airplane Mode
    Wi-Fi
    Bluetooth
    Cellular
    Personal Hotspot
    Battery
    General
    Accessibility
    Action Button
    Apple Intelligence & Siri
    Camera
    Control Center
    Display & Brightness
    Home Screen & App Library
    Search
    StandBy
    Wallpaper
    Notifications
    Sounds & Haptics
    Focus
    Screen Time
    Face ID & Passcode
    Emergency SOS
    Privacy & Security
    Game Center
    Wallet & Apple Pay
    Apps
    Settings
    Live
    Finding a setting in iPhone Settings

    Rewriting emails

    Jev also works well next to a larger model. Here, each email comes with a few things its writer cares about. Jev scores the draft on each one, and a language model rewrites it until every score reaches 4 out of 5, or until it has tried three times.

    Click play to improve the draft, or refresh for another email and its criteria.

    Sam, The launch deck was due yesterday and I still don't have it. This is the second time this month this has happened, and honestly I'm tired of chasing you for it. The whole team is waiting on your slides before we can do the dry run, and now we're going to have to push it. I need the deck today, no excuses. Jordan

    • Sounds friendly
    • Makes one clear ask
    • Doesn't blame anyone
    • Explains why it matters
    Live
    Rewriting an email until it meets the writer's criteria

    What’s next

    To recap, all of these use cases could have been done with structured outputs from previous LLMs. None of this is fundamentally new.

    What's different about Jev is that it's incredibly fast and cheap. TypeSafe measures it at 40–200x faster than frontier models, and 445x cheaper. You've probably noticed this yourself using the figures in this post.

    It's faster because it doesn't write. An LLM produces its answer one token at a time, and your code then parses it. Jev reads your question once and puts a probability on every possible answer at the same time, so it can't return an answer you didn't allow.

    That opens up use cases that used to be too slow, too expensive, or too hard to personalize and fine-tune. And since Jev-like models can run on smaller weights, many of them could even be done on device.

    The field of decision models is moving fast. Try OpenAI’s Decisions API, Together’s model, Kev, d1, pplx-decider, or whatever comes next. We'll dig into how decision models work in a future post.

    Keep reading a similar post:

    Static Analysis for the Agentic Era

    We tested three specialized code-mapping methods with a coding agent. Performance improved but at the cost of more tokens and time. Specialized agent harnesses can help, though not for free.

    LLMs have excelled at software engineering tasks.

    On DeepSWE, a representational benchmark of software tasks, overall completion rates are trending toward 80%. With that said, the non-deterministic nature of LLMs means success rates can vary wildly across runs.

    Keep reading

    If this sparked an idea for your roadmap, let's talk.

    Rubric is a product studio helping companies design, build, and ship intelligent applications to production.