Hello Mini Me: an agent born while preparing a talk for Devoxx
Meet Mini Me, a coding agent born from a Devoxx talk on the agent loop, built for small local models with guardrails against loops and hallucinations.

In two minutes
On Monday, I'm giving a tools in action session at Devoxx Belgium about the agent loop: One tool to rule them all: bash, a tiny LLM, and the birth of a coding agent
This agent loop is the foundation of AI agents, and especially of coding agents: the model proposes a tool call, the tool runs, the result goes back into the context, and we start again until the final answer. While digging into this loop with small local models (2 to 12 billion parameters, served by Docker Model Runner, llamacpp or llmman), I struggled with a list of failures that almost never show up with a big cloud model: hallucinations, loops, truncated tool calls, never-ending reasoning, tiny context window...
One thing leading to another, a side project was born: Mini Me (mm), a coding agent written specifically for small local models, with guardrails designed around their quirks.
I started my experiments with models like mellum2-12b-a2.5b-instruct-gguf or gemma-4-E4B-it-GGUF, but for about ten days now I've been testing it with bigger models like gemma-4-26B-A4B-it-GGUF, and the results are very encouraging.
So here is the first introductory post about Mini Me.
Where does mini-me come from
To prepare a talk about the agent loop, nothing beats implementing it and running it, again and again, on models that "fit" on my machine. And this is where things get interesting and instructive: models like those from Anthropic, OpenAI, Mistral, ... served through a cloud API rarely get the format of a tool call wrong, correctly re-read a block of code to edit, and know when to stop. A local model with 2 to 12B parameters, does not (actually, even with bigger ones, I run into this kind of problem).
It's while preparing the end (and the sequel?) of my talk: how to work around the weaknesses of small models (or how to improve my agent loop), that Mini Me was born, at first as a little brother next to the talk, and finally as a project in its own right.
Why an agent dedicated to small local models
The guiding idea of Mini Me: when the model gets it wrong, or "loops", give it back actionable information rather than a silent failure (or stop it, or have a way to stop it).
At the beginning: a single tool, by default: bash
Mini Me was born around a single tool, bash. The others (reading/writing/editing files, multi-step plan, reading skills, fetching web pages) were added little by little. They are optional and can be disabled one by one in the Mini Me configuration file.
When the model gets it wrong: a failure is text
An error (failed command, missing file, ambiguous edit, timeout) never surfaces as a system error that would cut the turn short. It goes back to the model, as text, written as a repair instruction rather than as a raw message. The goal being that the model can eventually correct itself (it can still be improved, but my history system lets me investigate to improve failure handling).
Against hallucinations, infinite loops and silent hangs
Several guardrails, from the gentlest to the most radical:
- A loop detector "warns" the model as soon as the same tool call repeats identically, then stops the turn if the warning is ignored;
- A watch on the arguments being generated, to immediately interrupt a tool call that goes off the rails by repeating a pattern while the model "spins out of control";
- And a watchdog that cancels the generation if no more tokens arrive (with a configurable delay).
ACP (Agent Client Protocol)
I had another wish: even though Mini Me comes with a TUI, I wanted a comfortable way to use it in an IDE. So Mini Me also speaks ACP (Agent Client Protocol), the protocol that lets you plug an agent directly into an editor like Zed.
I'll write a new post very soon to explain how to set it up.
Installing Mini Me
Mini Me comes as a binary, straight from the releases: https://rickub.com/bots-garden/mini-me/releases.
macOS (Apple Silicon)
# download the binary release in https://rickub.com/bots-garden/mini-me/releases
VERSION=0.2.6
curl -fsSLO "https://rickub.com/bots-garden/mini-me/releases/download/v${VERSION}/mm-${VERSION}-darwin-arm64"
chmod +x ./mm-${VERSION}-darwin-arm64
xattr -cr ./mm-${VERSION}-darwin-arm64
sudo rm -f /usr/local/bin/mm
sudo cp ./mm-${VERSION}-darwin-arm64 /usr/local/bin/mm
Linux
Same principle, without the xattr step (specific to macOS Gatekeeper quarantine):
# download the binary release in https://rickub.com/bots-garden/mini-me/releases
VERSION=0.2.6
curl -fsSLO "https://rickub.com/bots-garden/mini-me/releases/download/v${VERSION}/mm-${VERSION}-linux-amd64"
chmod +x ./mm-${VERSION}-linux-amd64
sudo rm -f /usr/local/bin/mm
sudo cp ./mm-${VERSION}-linux-amd64 /usr/local/bin/mm
Change
linux-amd64tolinux-arm64if you're on an ARM machine.
Windows
✋ I haven't tested Mini Me "for real" on Windows yet
No chmod, no sudo, no quarantine attribute to remove: download the executable and put it in a folder that is in your PATH.
# download the binary release in https://rickub.com/bots-garden/mini-me/releases
$VERSION = "0.2.6"
Invoke-WebRequest -Uri "https://rickub.com/bots-garden/mini-me/releases/download/v$VERSION/mm-$VERSION-windows-amd64.exe" -OutFile "mm.exe"
New-Item -ItemType Directory -Force -Path "$env:LOCALAPPDATA\mini-me" | Out-Null
Copy-Item ".\mm.exe" "$env:LOCALAPPDATA\mini-me\mm.exe" -Force
# Then, add $env:LOCALAPPDATA\mini-me to the user PATH.
You'll find all the available versions on the matching release page, for example for 0.2.6: https://rickub.com/bots-garden/mini-me/releases/v0.2.6.
First configuration
Mini Me plugs into a local OpenAI-compatible server. Here I'm using llmman, which I talked about in a previous post, but Docker Model Runner or llamacpp work just as well, the contract is the same.
Start the server and pull a model:
llmman serve
llmman pull huggingface.co/jetbrains/mellum2-12b-a2.5b-instruct-gguf-q4_k_m:Q4_K_M
llmman list
Then configure the agent in a YAML file, for example agent.llmman.yaml (you can name it whatever you like):
provider: llmman
model: huggingface.co/jetbrains/mellum2-12b-a2.5b-instruct-gguf-q4_k_m:Q4_K_M
baseUrl: http://127.0.0.1:17434/v1
fallback: http://host.docker.internal:17434/v1
bashTool: true
editTools: true
planTool: true
contextWindow: 0
maxOutput: 16000
maxTurns: 40
sessions:
enabled: true
dir: .mm/sessions
system: |
Your name is Bob.
You are a coding agent working in a terminal.
You have a "bash" tool to run shell commands.
Use it to explore files, run tests, inspect the repository, etc.
Chain several commands if needed, then answer clearly in English.
sampling:
temperature: 0.0
parallel_tool_calls: false
top_p: 0.9
max_tokens: 4096
watchdogTimeout: 80s
firstTokenTimeout: 3m
context:
enabled: true
threshold: 75
maxMessages: 80
keepLastTurns: 3
summaryMaxTokens: 2000
showStats: true
A few details:
baseUrlpoints to the OpenAI API exposed byllmman serve, andfallbacklets you fall back onhost.docker.internalif Mini Me itself runs in a container.samplingis deliberately conservative:temperature: 0.0, one tool call at a time (parallel_tool_calls: false). A small model produces a correct call more easily when it only has to handle one at a time.firstTokenTimeoutandwatchdogTimeoutare two separate delays: the first one covers reading the prompt (which can take several minutes on a large local context, before the first token), the second one covers silence once the answer has started.contextenables automatic compression of the history before the context window fills up (what the model is able to process), which is essential with a local model whose context window is fixed and more limited than huge models.
Then run mm in the project folder (mm --tui agent.llmman.yaml): Mini Me will read the configuration, connect to llmman, and you'll be able to start chatting with your agent locally.

✋ Warning: in TUI mode, when tools are enabled, they only work in YOLO mode (unlike ACP mode). (I'll change that in an upcoming release).
Mini Me is not here to replace the usual coding agents
Let me be clear: Mini Me does not aim to compete with existing coding agents (Docker Agent, Claude Code, Codex, OpenCode, etc. ...) running on huge models. That's not Mini Me's playground. The idea is to make the most of small local models, the ones running on your machine without a cloud API, with their own constraints (smaller context window, tendency to hallucinate, loops, fragile tool calls), by building an agent where every design choice starts from the intrinsic constraints of small models.
So, to sum up
- Mini Me was born while preparing my Devoxx talk about the agent loop.
- It's an agent dedicated to small local models (2 to 12B ... and even a bit more), with guardrails designed for their typical failures: hallucinations, loops, truncated tool calls, silent hangs.
- Mini Me also speaks ACP (detailed setup in an upcoming post, I promise).
- It installs as a single binary (Mini Me is written in Go, no dependencies), is configured in YAML, and plugs into any OpenAI-compatible server (llmman, Docker Model Runner, llamacpp...).
Feel free to ask questions 🙂
No comments yet. Be the first to comment!