Kolibri AI is a free, open-weight AI model from the German company Aleph Alpha, and this is the operator's guide to setting it up and running it yourself.
You might have seen it spelled "Colibri" or even "Calibri" in comments and auto-captions, but the official name is Kolibri, which is the German word for a hummingbird.
It's a 78-billion-parameter model that only switches on about 3.46 billion parameters for each word it writes.
It speaks German and English natively, it ships under the Apache 2.0 licence, and anyone can download it from Hugging Face without asking permission.
The big question I keep getting is simple: can I actually run this thing on my own machine?
The honest answer is yes on the right hardware, and no on a normal laptop.
So in this guide I'll show you exactly what Kolibri is, what it needs, the three real ways to run it today, and how to plug it into Hermes Agent as a local brain.
What Is Kolibri AI? The Quick Answer
Kolibri AI is Aleph Alpha's first big open-weight reasoning model, released on 3 October 2026 under the name Kolibri 1.
Aleph Alpha is a German AI company based in Heidelberg, and this release is clearly built for European businesses that care about where their data goes.
The model card lists the developer as Aleph Alpha Research GmbH and the provider as Aleph Alpha GmbH.
Here's the spec sheet I pulled straight from the official Hugging Face model card.
| Spec | Kolibri 1 (official model card) |
|---|---|
| Hugging Face repo | Aleph-Alpha/Kolibri-1 (FP8), plus Aleph-Alpha/Kolibri-1-BF16 |
| Release date | 3 October 2026 |
| Licence | Apache 2.0, and the repo is not gated |
| Total parameters | 78.1 billion |
| Active parameters per token | About 3.46 billion |
| Architecture | Mixture-of-experts with 384 experts per layer, 6 routed plus 1 shared |
| Languages | German and English |
| Context window | 262,144 tokens natively, validated up to 1,048,576 tokens |
| Knowledge cutoff | 18 June 2026 |
| Modes | Reasoning mode with effort levels, plus tool calling |
| Input and output | Text only, so no images |
In my video I rounded the active parameters to "about 3 billion", and the exact official number is 3.46 billion.
I also said "up to 1 million tokens" of context, and that's right, but Aleph Alpha recommends staying at or under 262,144 tokens for speed and for complex tasks.
Who Made Kolibri AI And Why It Exists
Aleph Alpha built Kolibri as a "sovereign" model, which is their way of saying you own and control the whole thing.
Their launch blog says it was built with the EU AI Act, the General-Purpose AI Code of Practice and GDPR in mind from the ground up.
The model card also confirms Aleph Alpha is a signatory of the EU's General-Purpose AI Code of Practice.
The target customers they name are public administration, industrial companies and aerospace.
They picked depth in two languages over thin coverage of fifty languages.
They built a tokenizer that handles German word structure efficiently, which matters for long German compound words.
They trained it on roughly 20 trillion tokens in pre-training, with about 62.5% English, about 23.9% German and about 13.6% code.
On top of that came 3.44 trillion tokens of mid-training and 201 billion tokens of long-context training, which is where the "about 24 trillion tokens" figure from my video comes from.
The model also uses a hybrid attention design, where four out of every five layers look only at nearby text and the fifth looks across the whole context.
That's the trick that keeps very long documents affordable to process.
Step 1: Check Your Hardware Before You Download Anything
This is the step everyone skips, and it's the one that matters most for Kolibri AI.
A viewer under my video put it bluntly by saying a 78B model needs roughly 100 GB of memory to run properly, and most laptops simply can't do that.
They're broadly right, and here's why.
Kolibri only uses about 3.46 billion parameters per token, which makes it fast once it's loaded.
The catch is that all 78 billion parameters still have to sit in memory, because the model picks different experts for every token.
Aleph Alpha says this themselves on the model card, describing memory as the trade-off of the mixture-of-experts design.
Here's what the different versions actually need, using official numbers where they exist and the converters' own published numbers for the community builds.
| Version | Who made it | Size on disk | What it needs |
|---|---|---|---|
| Official FP8 | Aleph Alpha | About 78 GB | Minimum 2× A100 80 GB, 2× H100, 1× H200, 1× B200 or 1× B300 |
| Community MLX 4-bit | Community converter | About 41 GiB | A Mac with 64 GB or more of unified memory |
| Community MLX 3-bit | Community converter | About 33 GiB | A Mac with 48 GB or more |
| Community MLX 2-bit | Community converter | About 24 GiB | A Mac with 36 GB or more |
| Community GGUF Q4_K_M | Community converter | About 47.5 GB | Lots of system RAM plus a patched llama.cpp |
| Community GGUF Q2_K | Community converter | About 28.6 GB | Lots of RAM, with a bigger quality hit |
Another viewer asked whether it'll run on a 2 GB graphics card, and the answer is no.
If your machine has 16 GB of memory, Kolibri is not your local model, and that's fine.
The smaller models I'll mention later in this guide are a much better fit for that kind of machine.
Step 2: Download Kolibri AI From Hugging Face
The official home of Kolibri is the Hugging Face repo called Aleph-Alpha/Kolibri-1.
You don't need an account or sign-up to run it locally, because the repo is not gated.
The main repo holds the FP8 weights, and a second repo called Aleph-Alpha/Kolibri-1-BF16 holds the full-precision version for people with even more memory.
Within a day of launch the community also published MLX builds for Apple Silicon and GGUF builds for llama.cpp, in sizes from 2-bit up to 8-bit.
Treat those community builds as unofficial, because Aleph Alpha didn't make them and doesn't endorse them.
Step 3: Pick One Of The Three Ways To Run It
There are three working routes as of 7 October 2026, and which one you pick depends on your hardware.
Route A: The official way with vLLM
This is the route Aleph Alpha documents, and it's the one to use on a GPU server or a rented cloud GPU.
You install Aleph Alpha's inference package, which adds Kolibri support to vLLM.
pip install 'aleph-alpha-inference>=1'
Then you start the server with reasoning and tool calling switched on.
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
Aleph Alpha also publishes a ready-made container image if you'd rather use Docker.
Once it's running, you get an OpenAI-compatible API on your own machine at http://localhost:8000/v1.
The recommended sampling settings are a temperature of 1.0, a top-p of 0.97 and a top-k of 128.
Route B: A community MLX build on a big Mac
If you've got a Mac with 64 GB or more of memory, the 4-bit MLX build is the most practical route I've found.
The catch is that the standard mlx-lm library doesn't support Kolibri's architecture yet, and the request to add it is still open.
The community builds work around that by shipping their own model file and a small launcher script inside the download.
The converter of one popular 4-bit build reports around 52 to 56 tokens per second on an M1 Max, which is quick, but I haven't tested that myself.
Route C: A community GGUF with a patched llama.cpp
GGUF files exist, but stock llama.cpp can't load Kolibri yet.
The GGUF converter ships a patch you apply to llama.cpp before you build it.
The converter reports around 13 to 15 tokens per second on a CPU-only desktop with 128 GB of RAM using the 4-bit file.
What about LM Studio?
In my video I said you can run Kolibri through LM Studio, and that's where most people will want to run it.
I need to update that honestly.
As of today, LM Studio and Ollama won't load Kolibri out of the box, because both rely on llama.cpp or MLX support that hasn't landed yet.
There's an open request for llama.cpp support and an open pull request on Ollama, so this could change any week.
My advice is to search for Kolibri inside LM Studio after each update, and use Route A or Route B until it shows up.
🔥 Want my local model setups without the trial and error? Inside the AI Profit Boardroom, I've got step-by-step tutorials on running local models and plugging them into AI agents, plus four coaching calls a week and 3,400+ members comparing notes. → Get access here
Step 4: Use Reasoning Mode Properly
Kolibri AI has a reasoning mode, and how you set it changes both speed and quality.
You pass a reasoning effort of low, medium or high through the chat template.
You can also set it to none, or switch thinking off, and the model answers straight away.
Here's how I'd think about each setting.
- Use none for quick replies like rewriting a sentence or translating a paragraph.
- Use low for simple jobs where a short check helps, like pulling fields out of an email.
- Use medium for normal agent work, like planning a short task with a few tool calls.
- Use high for hard maths, tricky code or long document analysis.
Watch out for one trap with the community MLX builds.
Their converters note that if you don't set a reasoning effort, the chat template defaults to high, which thinks at length and feels slow.
So set the effort on every request until you know what you need.
Step 5: Turn On Tool Calling
Tool calling is what turns Kolibri from a chatbot into something an agent can use.
The model card says the vLLM serving command enables Hermes-style tool calling.
You pass your tools in the standard tools field of a chat request, and the model returns structured tool calls.
Your code runs the tool, sends the result back, and the model writes the final answer.
Aleph Alpha's own example is a weather lookup for Heidelberg, and it works the same way for searching files, calling a CRM or running a script.
Tool calling and reasoning mode can be used together, which is what you want for multi-step agent jobs.
Step 6: Plug Kolibri AI Into Hermes Agent
This is the setup I talked about at the end of my video, where Kolibri becomes a free local brain for Hermes Agent.
Because Kolibri runs behind an OpenAI-compatible server, Hermes can talk to it like any other model.
Here's the order I'd do it in.
- First, get Kolibri running with Route A or Route B and confirm it answers a simple chat request.
- Second, make a fresh Hermes profile so your main agent setup stays untouched.
- Third, run
hermes modelin that profile and point it at your local endpoint, which ishttp://localhost:8000/v1for vLLM. - Fourth, give it one small job with one tool, like reading a folder and summarising what's in it.
- Fifth, set the reasoning effort to medium and watch how long each step takes before you give it anything bigger.
If you've never set up Hermes before, my guide on setting up Hermes Agent in one click gets you to a working agent first.
I'll be honest about my own setup here.
I run a Mac Studio, and I don't actually run that much local AI day to day.
The best local model I've tested with Hermes so far is LFM 2.5 at 2.6 billion parameters, which is crazy fast and was trained with Hermes Agent in mind.
So Kolibri is not automatically the best local brain for Hermes.
It's a much bigger, smarter model that needs much more memory, so test it against your current local model on your own jobs before you switch.
How Kolibri AI Scores On Benchmarks
Every number in this section is reported by Aleph Alpha on their model card, so treat it as the vendor's own testing.
They ran Kolibri at high reasoning effort against other open models using their own evaluation framework.
| Benchmark group (vendor-reported) | Kolibri | Qwen3.6 35B-A3B | Gemma 4 26B-A4B | Nemotron 3 Super | Mistral Small 4 | GPT-OSS 120B |
|---|---|---|---|---|---|---|
| Overall, English | 75.5 | 71.4 | 71.9 | 73.0 | 63.1 | 72.3 |
| Overall, German | 70.8 | 67.3 | 66.3 | 67.9 | 61.4 | 70.2 |
| Agentic average, English | 63.4 | 62.1 | 54.6 | 54.9 | 40.7 | 54.0 |
| Code average, English | 89.3 | 87.7 | 89.0 | 88.3 | 82.0 | 90.8 |
That backs up the two claims I made in the video.
Kolibri has the highest German overall average among the mixture-of-experts models in Aleph Alpha's table.
It also beats Qwen 3.6, Nemotron 3 Super and Mistral Small 4 on the agentic tool-calling average, with Mistral Small 4 the weakest of that group.
Here's the part the launch hype leaves out.
In the same table, the dense Qwen 3.8 27B model scores higher overall, with 80.2 in English and 79.9 in German.
So the viewer who said Qwen 3.8 27B beats Kolibri on output quality has Aleph Alpha's own numbers on their side.
The difference is that Qwen 3.8 27B uses all 27 billion parameters on every token, while Kolibri uses about 3.46 billion.
If you want to compare more local options, my post on running Jev locally covers how I test local stand-ins against hosted models.
Kolibri AI Limits And Gotchas
These are the limits I'd want to know before building anything on it.
- Kolibri is text only, so it won't read images, screenshots or PDFs as pictures.
- Its built-in knowledge stops at 18 June 2026, so give it search or documents for anything newer.
- It's tuned for German and English, and Aleph Alpha says depth in two languages was a deliberate choice.
- Aleph Alpha says it's built for setups where a person reviews the output before it's acted on, not for fully unsupervised agents.
- Its own model card shows weaker scores on some fact-checking and closed-book tests, so don't use it as a trivia engine without sources.
- Speed numbers you see online come from community converters, so measure tokens per second on your own machine.
Who Should Set Up Kolibri AI?
Kolibri AI makes sense for some operators and not for others.
It's a strong fit if you work in German and English, have a 64 GB-plus Mac or a GPU server, and want an open model you fully control.
It's a poor fit if you're on a 16 GB laptop, because the model simply won't fit.
It's a poor fit if you want a one-click install today, because LM Studio and Ollama support haven't landed yet.
🔥 Want to see how I wire local models into real agent workflows? The AI Profit Boardroom has the Hermes and local model training, the Agent OS, and weekly coaching where you can ask me about your exact hardware. → Join the AI Profit Boardroom
Related Reading
- How to set up Hermes Agent in one click is the fastest way to get an agent ready for a local brain like Kolibri.
- Can you run Jev locally? walks through how I test local models against hosted ones.
- My Hermes MCP server setup with Codex shows you how to give a Hermes agent more tools to call.
Also On Our Network
- 🌐 How businesses and agencies can actually use Kolibri AI
- 🌐 My feature-by-feature Kolibri AI verdict with scores
- 🌐 A Kolibri AI quick-start for your Agent OS
- 🌐 Kolibri compared with the local models on GoldieBench
FAQ: Kolibri AI
What is Kolibri AI?
Kolibri AI is Kolibri 1, an open-weight mixture-of-experts model from Aleph Alpha in Germany.
It has 78.1 billion total parameters, about 3.46 billion active per token, and native German and English.
It was released on 3 October 2026 under the Apache 2.0 licence.
Is Kolibri AI free?
Yes, the weights are free to download from Hugging Face, and the repo isn't gated.
Your only cost is the hardware or rented GPU you run it on.
How much memory does Kolibri AI need?
The official FP8 version takes about 78 GB, and Aleph Alpha lists two 80 GB A100s or one H200 as the minimum.
Community 4-bit MLX builds need a 64 GB Mac or bigger, and even the smallest community builds need around 36 GB.
Can I run Kolibri AI in LM Studio or Ollama?
Not with the stock apps as of 7 October 2026, because support for Kolibri's new architecture hasn't landed yet.
Use vLLM with Aleph Alpha's plugin, a community MLX build or a patched llama.cpp until it does.
Can Kolibri AI power Hermes Agent?
Yes, because it supports tool calling and runs behind an OpenAI-compatible server.
Point a Hermes profile at that local endpoint and test it on small jobs first.
📺 Video notes + links to the tools 👉
🎥 Learn how I make these videos 👉
🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉
About Julian
I'm Julian Goldie, an SEO entrepreneur, author and founder of the AI Profit Boardroom, which has 3,400+ members.
I help business owners scale with AI agents, automation and SEO.
- I've built a 7-figure agency, Goldie Agency, with a team of around 50 people.
- I've grown a YouTube channel to 400,000+ subscribers.
- I wrote the Amazon best-sellers "SEO Link Building Mastery" and "Agency Marketing Mastery".
- My Udemy courses have taught over 50,000 students.
→ Get my best AI training inside the AI Profit Boardroom
Check your memory first, pick the route that fits your machine, and you'll have Kolibri AI running as a private local brain for your agents.











