You can't run Jev locally, because TypeSafe has not released Jev's weights or any offline build, so if you want to run Jev locally the honest route is a Jev-compatible open model like Laya on your own machine.
I know that's not the answer you were hoping for.
But stick with me, because the workaround is genuinely good and I've run it on my own Mac.
Laya answers the same kind of question Jev answers, it speaks the same request format, and it costs nothing per question.
On my Mac it answered a single question in about 20 milliseconds.
In this walkthrough I'll show you exactly how to set it up, where it beats Jev, where it loses, and the free hosted doors you can use when you need the real Jev.
Can You Run Jev Locally? The Short Answer
No, you can't run the real Jev locally today.
Jev is a hosted model from a company called TypeSafe AI, and it launched on 15 September 2026.
TypeSafe serves it through its own API, and other platforms like OpenRouter, Vercel AI Gateway and OpenCode Zen pass your requests through to it.
TypeSafe's own website talks about early access, a console sign-in and a price per billion input tokens.
It says nothing about downloadable weights, an on-premise build or an offline version.
So there is no file you can pull down and load into Ollama or LM Studio.
Anyone promising you a "Jev local install" is selling you a different model with Jev's name on the box.
That's the bad news, and I'd rather you hear it in line one than after a 20-minute setup that goes nowhere.
The good news is that the job Jev does is small enough to run on a laptop.
What Jev Actually Does (And Why A Laptop Can Handle The Job)
Jev doesn't write anything.
You hand it some text, like an email or a support ticket, and a question with a fixed list of answers.
It picks one of those answers and tells you how sure it is.
For example, you might ask "which team should handle this?" with billing, technical, sales and other as the options.
Jev comes back with "billing" and a probability next to it.
Your agent then acts on its own when the number is high and asks you when the number is low.
That's the whole trick.
Because the job is choosing from a list rather than writing paragraphs, the model doing it can be tiny.
That's exactly why open models have appeared that copy Jev's job and run on normal hardware.
Your Three Real Options For Local Jev
Here is how I'd break down every route you've got as of October 2026.
| Option | Runs on your machine? | Cost per question | Speaks Jev's request format? | Have I tested it? |
|---|---|---|---|---|
| Jev (TypeSafe) | No, hosted only | $0.042 per million input tokens, output free | Yes, it is the original | Yes, through OpenCode Zen, OpenRouter and Vercel |
| Laya (Convai Innovations) | Yes, Apache 2.0 | Free | Yes, via laya-serve | Yes, on my own Mac |
| Clef (Cloudflare) | Yes, open weights on Hugging Face | Free if you host it yourself | Yes, Cloudflare says it is Jev-API compatible | Not yet |
Laya is the one I'd start with, because it's small, it's free, and I've seen it run with my own eyes.
Clef is the newer option and it's much bigger, so I'll cover it separately further down.
How To Run Jev Locally With Laya: Step By Step
This is the setup I'd hand to anyone on my team.
You need Python 3.10 or newer, because that's what the Laya package on PyPI asks for.
A GPU helps, but the Laya README says CPU-only inference works too, just slower.
Step 1: Install Laya
Open your terminal and run the install command from the Laya README.
python -m pip install laya
If you also want the local server that pretends to be Jev, install the serve extra instead.
pip install "laya[serve]"
I ran version 0.3.4 when I tested it on 21 September.
PyPI lists 0.3.27 as the latest version today, so you'll get a newer build than I did.
Step 2: Ask Your First Question In Python
Laya ships a Router that picks the right model for your text before it runs anything.
This is the quickstart from the Laya README, and it is the same shape of question I used on my Mac.
from laya import Router
router = Router()
result = router.predict(
{"body": "We were billed twice. Please refund."},
{"dept": {"type": "choice", "instructions": "Which team?",
"criteria": {"billing": "refunds", "tech": "bugs"}}}
)
print(result["answers"]["dept"]["choice"])
print(result["routing"]["model"])
The first line printed is the answer it picked.
The second line tells you which of the three Laya models the Router sent your text to.
Step 3: Learn The Three Question Types
Laya uses the same three question types as Jev, so anything you learn here carries straight over.
- A choice question picks one option from your list, like billing or technical.
- A score question rates something on your own scale, like how urgent an email is.
- A noul question is a yes or no answer with a probability between 0 and 1.
On my Mac I asked all four of my questions about one billing email in a single call.
It returned the department, the urgency, whether they'd leave, and whether they wanted a refund in 52.7 milliseconds.
Step 4: Turn On Preload Before You Do Anything Serious
This is the gotcha that nearly made me give up on it.
Out of the box, Laya keeps only some of its models in memory.
When my test messages flipped between languages, every switch reloaded a model, and that cost me about 20 seconds per answer.
The fix is one setting.
from laya import Router
router = Router(preload=True)
With every model loaded at the start, the same messages took about 19 milliseconds each on my Mac.
That's the difference between unusable and instant, so don't skip it.
Step 5: Run laya-serve So Your Jev Code Points At Your Own Machine
This is the step that makes Laya feel like local Jev.
The Laya README says its server implements Jev's request format at the same path, which is POST /v1/systemone.
You start it with this command from the README.
LAYA_DEVICE=cuda LAYA_PRELOAD=1 laya-serve
That example is for an NVIDIA card, and it binds to port 8000 with all three models preloaded.
Then you send it a request exactly like you'd send one to Jev.
curl -s localhost:8000/v1/systemone \
-H 'content-type: application/json' \
-d '{
"state": {"body": "billed twice, refund please"},
"questions": {"dept": {"type": "choice", "instructions": "which team?",
"criteria": {"billing": "refunds", "tech": "bugs"}}}
}'
If your agent already calls Jev, you change the URL and leave the rest alone.
The README is honest that there are three differences: the option limit is smaller, score levels need descriptions, and the confidence number is calculated differently.
🔥 Want me to walk you through wiring this into your own agents? Inside the AI Profit Boardroom, I share the decision-layer walkthroughs as I build them, including putting a fast local decision tool in front of Claude Code, Hermes and OpenClaw. Plus weekly coaching calls + 3,400+ members building real automations. → Get access here
How Fast Is Local Jev On A Normal Machine?
Speed is where the local route shines.
Here are the numbers, split between what the Laya team published and what I measured myself.
| Setup | Time per question | Who measured it |
|---|---|---|
| Laya, one question, with a GPU | 33 ms | Laya team |
| Laya, ten questions at once | 7 ms each | Laya team |
| Laya on a plain CPU, no GPU | 193 to 464 ms | Laya team |
| Laya on my Mac, one question | about 20 ms | Me |
| Laya on my Mac, preloaded, switching languages | about 19 ms | Me |
| Jev, one question | 236 to 276 ms | Published independent tests |
That's why my guide called Laya about seven times faster than Jev.
Even on a plain CPU with no graphics card, it lands at roughly Jev's speed.
The Catch: Untrained Laya Is Not As Accurate As Jev
I'm going to be straight with you here, because this is where most hype posts lie.
The headline score where Laya beat Jev, 0.766 against 0.727 on a 2,000-choice business test, came from the model that was trained on the practice version of that exact test.
When I ran untrained Laya on my own work, Jev was more accurate every time.
| My test on my Mac | Jev | Laya, untrained |
|---|---|---|
| 60 test emails sorted into folders | 54 out of 60 | 36 and 28 out of 60, across two models |
| 24 made-up contact-form messages: lead, vendor or noise | 24 out of 24 at about 392 ms each | 16 out of 24 at about 57 ms each |
| 10 article briefs matched to one of five sites | 10 out of 10 | 8 out of 10 |
There's one thing I really liked in those results, though.
Untrained Laya never once claimed to be 85% sure on the email test.
It knew it was guessing, which is exactly what you want a decision tool to admit.
Train It On Your Own Examples For The Real Win
The Laya team ships a fine-tuning notebook that runs on Kaggle's free GPUs.
It builds your training data, trains the model, tunes the sureness numbers and saves your finished model.
My guide estimated four to five hours for that run, with no card needed.
The README also mentions an Apple Silicon script if you'd rather train on a Mac.
If you skip training, my honest view is that you might be better off just using Jev through a free door.
When To Keep Using Hosted Jev Instead
There is one job where local Laya clearly loses.
On a banking test with 77 categories, Jev scored 0.870 and Laya scored 0.425.
Laya only has a fixed amount of room for your list of options, so 77 choices get squeezed down to three or four words each.
Jev handles up to 255 options without that problem.
My rule is simple.
If your question has fewer than 20 answers, test Laya.
If it has 50 or more, stay with Jev, or split it into two smaller questions.
What About Clef? The Bigger Open-Weight Option
Cloudflare released Clef on 1 October 2026.
Cloudflare's blog says it is open-sourcing the weights on Hugging Face under Apache 2.0, and that the models are fully Jev-API compatible.
There are two sizes.
Clef is post-trained from Qwen 3.8-27B and adds a vision encoder, and Clef-flash is a smaller 9B model.
Cloudflare lists a 64k context window, against Jev's 32k.
I have not run Clef myself yet, so I'm not going to tell you how it performs.
What I will say is that a 27B model is a very different beast to Laya's models, which are around 322 to 421 million parameters.
If you want something that runs comfortably next to everything else on a laptop, Laya is the lighter starting point.
The Free Hosted Doors When You Need Real Jev
Sometimes you need the real Jev rather than a local stand-in.
These are the doors I covered in my free Jev guide, updated for October.
- jevplayground.com is an independent test bench where you can run the real model with no signup, although usage limits apply.
- OpenCode Zen lists a model called
jev-1.13-freeathttps://opencode.ai/zen/v1/systemone, and its docs say it's free for a limited time. - Vercel AI Gateway had promotional free pricing that ended on 25 September 2026, and it now lists Jev at $0.042 per million input tokens.
- OpenRouter lists Jev at $0.042 per million input tokens with output free, which is how I ran my ten Jev builds.
The OpenCode Zen door is the one I'd build on while it lasts, because it uses the same request shape as Laya's local server.
That means you can switch between real Jev and local Laya by changing one URL.
How I'd Wire Local Jev Into An Agent Setup
Here is the pattern I'd use if I were setting this up today.
- Your main agent, whether that's Claude Code, Hermes or OpenClaw, stays in charge of the heavy thinking.
- Every small decision, like sorting an email or picking which agent gets a task, goes to the decision layer first.
- The decision layer is Laya running locally through laya-serve on port 8000.
- If the sureness number is high, the agent acts on its own.
- If the number is low, or the question has too many options, the request goes to real Jev or to you.
If you're running Hermes already, my one-click Hermes setup walkthrough gets the main agent side ready.
If you want your agents sharing tools across apps, my Hermes MCP server setup with Codex shows how I connect them.
And if you want your agent to remember what it decided last week, my Claude and Obsidian second brain setup is the memory layer I use.
Gotchas I Hit So You Don't Have To
- Forgetting preload cost me about 20 seconds per answer when languages switched, so always set
preload=TrueorLAYA_PRELOAD=1. - Trusting the number on unfamiliar text is dangerous, because the Laya team's own write-up showed an English model scoring zero right on Khmer text while claiming 95% sure.
- Using the Router fixes most of that, and the Laya team reported 45 of 51 languages working well with it against only 23 without it.
- Big option lists blur together in Laya, so count your options before you switch from Jev.
- Skipping the tuning step leaves the sureness number close to decoration, because the English model's calibration error fell from 0.466 before tuning to 0.081 after.
Related Reading
📺 Video notes + links to the tools 👉
🎥 Learn how I make these videos 👉
🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉
🔥 Ready to put a decision layer in front of your agents? The AI Profit Boardroom has the Agent OS walkthroughs, the decision-layer tutorials as they ship, and four coaching calls a week where you can bring your own setup. Over 3,400 members are building with it right now. → Join the AI Profit Boardroom
Also On Our Network
- 🌐 The business case for running Jev locally to cut agency costs
- 🌐 My verdict on local Jev: Laya vs Clef vs the real thing
- 🌐 Running a local Jev decision layer inside an Agent OS
- 🌐 My full Laya AI guide with every test I ran
FAQ: Running Jev Locally
Can I run Jev locally?
No, you can't run the real Jev locally, because TypeSafe has not published its weights or an offline build.
You can run Laya locally instead, which uses the same request format and runs free on your own machine.
What is the best local Jev alternative?
Laya is the one I've tested, and it's free under Apache 2.0, small enough for a laptop and fast at about 20 ms per question on my Mac.
Cloudflare's Clef is a newer open-weight option that Cloudflare says is Jev-API compatible, but I haven't tested it yet.
Does Laya work with code written for Jev?
Mostly, yes.
Laya's server answers at POST /v1/systemone like Jev, so you change the URL and keep your request.
The option limit is smaller, score levels need descriptions, and the confidence number is calculated differently.
Do I need a GPU to run Laya?
No, the Laya README says CPU-only inference is supported.
The Laya team published 193 to 464 ms per question on a plain CPU, which is roughly Jev's hosted speed.
Is local Laya as accurate as Jev?
Not without training.
On my 60-email test Jev got 54 right while untrained Laya got 36 and 28 across two models, so train it on your own examples before you trust it.
Is there a free way to use the real Jev?
Yes, jevplayground.com runs it with no signup, and OpenCode Zen lists jev-1.13-free as free for a limited time.
About Julian
I'm Julian Goldie, an AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom, which has 3,400+ members.
I help business owners scale with AI agents, automation, and SEO.
- I have 400,000+ YouTube subscribers who watch my AI tool tests every week.
- I built a 7-figure agency from the ground up.
- I run daily AI training inside the Boardroom.
- I wrote two Amazon best-sellers on SEO and agency growth.
→ Get my best AI training inside the AI Profit Boardroom
So, can you run Jev locally?
Not the real Jev, but you can run Jev locally in every way that matters for your agents by setting up Laya with laya-serve, training it on your own examples, and keeping a free Jev door open for the big lists.











