tiny language models · local models · guardrails · go · 8 min read

Hello Mini Me: an agent born while preparing a talk for Devoxx

Meet Mini Me, a coding agent born from a Devoxx talk on the agent loop, built for small local models with guardrails against loops and hallucinations.

K
k33gOct 3, 2026 · 8 min
Hello Mini Me: an agent born while preparing a talk for Devoxx

In two minutes

On Monday, I'm giving a tools in action session at Devoxx Belgium about the agent loop: One tool to rule them all: bash, a tiny LLM, and the birth of a coding agent

This agent loop is the foundation of AI agents, and especially of coding agents: the model proposes a tool call, the tool runs, the result goes back into the context, and we start again until the final answer. While digging into this loop with small local models (2 to 12 billion parameters, served by Docker Model Runner, llamacpp or llmman), I struggled with a list of failures that almost never show up with a big cloud model: hallucinations, loops, truncated tool calls, never-ending reasoning, tiny context window...

One thing leading to another, a side project was born: Mini Me (mm), a coding agent written specifically for small local models, with guardrails designed around their quirks.

I started my experiments with models like mellum2-12b-a2.5b-instruct-gguf or gemma-4-E4B-it-GGUF, but for about ten days now I've been testing it with bigger models like gemma-4-26B-A4B-it-GGUF, and the results are very encouraging.

So here is the first introductory post about Mini Me.

Where does mini-me come from

To prepare a talk about the agent loop, nothing beats implementing it and running it, again and again, on models that "fit" on my machine. And this is where things get interesting and instructive: models like those from Anthropic, OpenAI, Mistral, ... served through a cloud API rarely get the format of a tool call wrong, correctly re-read a block of code to edit, and know when to stop. A local model with 2 to 12B parameters, does not (actually, even with bigger ones, I run into this kind of problem).

It's while preparing the end (and the sequel?) of my talk: how to work around the weaknesses of small models (or how to improve my agent loop), that Mini Me was born, at first as a little brother next to the talk, and finally as a project in its own right.

Why an agent dedicated to small local models

The guiding idea of Mini Me: when the model gets it wrong, or "loops", give it back actionable information rather than a silent failure (or stop it, or have a way to stop it).

At the beginning: a single tool, by default: bash

Mini Me was born around a single tool, bash. The others (reading/writing/editing files, multi-step plan, reading skills, fetching web pages) were added little by little. They are optional and can be disabled one by one in the Mini Me configuration file.

When the model gets it wrong: a failure is text

An error (failed command, missing file, ambiguous edit, timeout) never surfaces as a system error that would cut the turn short. It goes back to the model, as text, written as a repair instruction rather than as a raw message. The goal being that the model can eventually correct itself (it can still be improved, but my history system lets me investigate to improve failure handling).

Against hallucinations, infinite loops and silent hangs

Several guardrails, from the gentlest to the most radical:

  • A loop detector "warns" the model as soon as the same tool call repeats identically, then stops the turn if the warning is ignored;
  • A watch on the arguments being generated, to immediately interrupt a tool call that goes off the rails by repeating a pattern while the model "spins out of control";
  • And a watchdog that cancels the generation if no more tokens arrive (with a configurable delay).

ACP (Agent Client Protocol)

I had another wish: even though Mini Me comes with a TUI, I wanted a comfortable way to use it in an IDE. So Mini Me also speaks ACP (Agent Client Protocol), the protocol that lets you plug an agent directly into an editor like Zed.

I'll write a new post very soon to explain how to set it up.

Installing Mini Me

Mini Me comes as a binary, straight from the releases: https://rickub.com/bots-garden/mini-me/releases.

macOS (Apple Silicon)

bash
# download the binary release in https://rickub.com/bots-garden/mini-me/releases
VERSION=0.2.6
curl -fsSLO "https://rickub.com/bots-garden/mini-me/releases/download/v${VERSION}/mm-${VERSION}-darwin-arm64"
chmod +x ./mm-${VERSION}-darwin-arm64
xattr -cr ./mm-${VERSION}-darwin-arm64
sudo rm -f /usr/local/bin/mm
sudo cp ./mm-${VERSION}-darwin-arm64 /usr/local/bin/mm

Linux

Same principle, without the xattr step (specific to macOS Gatekeeper quarantine):

bash
# download the binary release in https://rickub.com/bots-garden/mini-me/releases
VERSION=0.2.6
curl -fsSLO "https://rickub.com/bots-garden/mini-me/releases/download/v${VERSION}/mm-${VERSION}-linux-amd64"
chmod +x ./mm-${VERSION}-linux-amd64
sudo rm -f /usr/local/bin/mm
sudo cp ./mm-${VERSION}-linux-amd64 /usr/local/bin/mm

Change linux-amd64 to linux-arm64 if you're on an ARM machine.

Windows

✋ I haven't tested Mini Me "for real" on Windows yet

No chmod, no sudo, no quarantine attribute to remove: download the executable and put it in a folder that is in your PATH.

powershell
# download the binary release in https://rickub.com/bots-garden/mini-me/releases
$VERSION = "0.2.6"
Invoke-WebRequest -Uri "https://rickub.com/bots-garden/mini-me/releases/download/v$VERSION/mm-$VERSION-windows-amd64.exe" -OutFile "mm.exe"
New-Item -ItemType Directory -Force -Path "$env:LOCALAPPDATA\mini-me" | Out-Null
Copy-Item ".\mm.exe" "$env:LOCALAPPDATA\mini-me\mm.exe" -Force
# Then, add $env:LOCALAPPDATA\mini-me to the user PATH.

You'll find all the available versions on the matching release page, for example for 0.2.6: https://rickub.com/bots-garden/mini-me/releases/v0.2.6.

First configuration

Mini Me plugs into a local OpenAI-compatible server. Here I'm using llmman, which I talked about in a previous post, but Docker Model Runner or llamacpp work just as well, the contract is the same.

Start the server and pull a model:

bash
llmman serve
llmman pull huggingface.co/jetbrains/mellum2-12b-a2.5b-instruct-gguf-q4_k_m:Q4_K_M
llmman list

Then configure the agent in a YAML file, for example agent.llmman.yaml (you can name it whatever you like):

yaml
provider: llmman
model: huggingface.co/jetbrains/mellum2-12b-a2.5b-instruct-gguf-q4_k_m:Q4_K_M

baseUrl: http://127.0.0.1:17434/v1
fallback: http://host.docker.internal:17434/v1

bashTool: true
editTools: true
planTool: true

contextWindow: 0
maxOutput: 16000
maxTurns: 40

sessions:
  enabled: true
  dir: .mm/sessions

system: |
  Your name is Bob.
  You are a coding agent working in a terminal.
  You have a "bash" tool to run shell commands.
  Use it to explore files, run tests, inspect the repository, etc.
  Chain several commands if needed, then answer clearly in English.

sampling:
  temperature: 0.0
  parallel_tool_calls: false
  top_p: 0.9
  max_tokens: 4096

watchdogTimeout: 80s
firstTokenTimeout: 3m

context:
  enabled: true
  threshold: 75
  maxMessages: 80
  keepLastTurns: 3
  summaryMaxTokens: 2000
  showStats: true

A few details:

  • baseUrl points to the OpenAI API exposed by llmman serve, and fallback lets you fall back on host.docker.internal if Mini Me itself runs in a container.
  • sampling is deliberately conservative: temperature: 0.0, one tool call at a time (parallel_tool_calls: false). A small model produces a correct call more easily when it only has to handle one at a time.
  • firstTokenTimeout and watchdogTimeout are two separate delays: the first one covers reading the prompt (which can take several minutes on a large local context, before the first token), the second one covers silence once the answer has started.
  • context enables automatic compression of the history before the context window fills up (what the model is able to process), which is essential with a local model whose context window is fixed and more limited than huge models.

Then run mm in the project folder (mm --tui agent.llmman.yaml): Mini Me will read the configuration, connect to llmman, and you'll be able to start chatting with your agent locally.

hello-mini-me
hello-mini-me

✋ Warning: in TUI mode, when tools are enabled, they only work in YOLO mode (unlike ACP mode). (I'll change that in an upcoming release).

Mini Me is not here to replace the usual coding agents

Let me be clear: Mini Me does not aim to compete with existing coding agents (Docker Agent, Claude Code, Codex, OpenCode, etc. ...) running on huge models. That's not Mini Me's playground. The idea is to make the most of small local models, the ones running on your machine without a cloud API, with their own constraints (smaller context window, tendency to hallucinate, loops, fragile tool calls), by building an agent where every design choice starts from the intrinsic constraints of small models.

So, to sum up

  • Mini Me was born while preparing my Devoxx talk about the agent loop.
  • It's an agent dedicated to small local models (2 to 12B ... and even a bit more), with guardrails designed for their typical failures: hallucinations, loops, truncated tool calls, silent hangs.
  • Mini Me also speaks ACP (detailed setup in an upcoming post, I promise).
  • It installs as a single binary (Mini Me is written in Go, no dependencies), is configured in YAML, and plugs into any OpenAI-compatible server (llmman, Docker Model Runner, llamacpp...).

Feel free to ask questions 🙂

K

Written by

k33g

Responses

0 Comments

No comments yet. Be the first to comment!

Related reading

From other blogs

How to cook a little coding agent with Docker Model Runner and Docker Agent (and `sbx`)

k33g_org's Blog