<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Mini Me</title>
        <link>https://mini-me.org</link>
        <description>Mini Me is an open-source coding agent built for small local LLMs, with guardrails against hallucinations, loops and silent failures. Runs fully offline.</description>
        <lastBuildDate>Sat, 03 Oct 2026 21:54:40 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>Writizzy</generator>
        <language>en</language>
        <copyright>All rights reserved 2026, Mini Me</copyright>
        <item>
            <title><![CDATA[Hello Mini Me: an agent born while preparing a talk for Devoxx]]></title>
            <link>https://mini-me.org/p/20261003-hello-mini-me</link>
            <guid>https://mini-me.org/p/20261003-hello-mini-me</guid>
            <pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Meet Mini Me, a coding agent born from a Devoxx talk on the agent loop, built for small local models with guardrails against loops and hallucinations.]]></description>
            <content:encoded><![CDATA[<h2>In two minutes</h2>
<p>On Monday, I&#39;m giving a <strong>tools in action</strong> session at <strong>Devoxx Belgium</strong> about the <strong>agent loop</strong>: <a href="https://m.devoxx.com/events/dvbe26/talks/22911/one-tool-to-rule-them-all-bash-a-tiny-llm-and-the-birth-of-a-coding-agent">One tool to rule them all: bash, a tiny LLM, and the birth of a coding agent</a></p>
<p>This agent loop is the foundation of AI agents, and especially of coding agents: the model proposes a tool call, the tool runs, the result goes back into the context, and we start again until the final answer. While digging into this loop with <strong>small local models</strong> (2 to 12 billion parameters, served by <strong><a href="https://docs.docker.com/ai/model-runner/">Docker Model Runner</a></strong>, <strong>llamacpp</strong> or <strong><a href="https://llmmanorg.github.io/">llmman</a></strong>), I struggled with a list of failures that almost never show up with a big cloud model: hallucinations, loops, truncated tool calls, never-ending reasoning, tiny context window...</p>
<p>One thing leading to another, a side project was born: <a href="https://rickub.com/bots-garden/mini-me">**Mini Me** (`mm`)</a>, a coding agent written <strong>specifically</strong> for small local models, with guardrails designed around their quirks.</p>
<blockquote>
<p>I started my experiments with models like <a href="https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M">mellum2-12b-a2.5b-instruct-gguf</a> or <a href="https://huggingface.co/unsloth/gemma-4-E4B-it-GGUF">gemma-4-E4B-it-GGUF</a>, but for about ten days now I&#39;ve been testing it with bigger models like <a href="https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF">gemma-4-26B-A4B-it-GGUF</a>, and the results are very encouraging.</p>
</blockquote>
<p>So here is the first introductory post about <strong>Mini Me</strong>.</p>
<h2>Where does mini-me come from</h2>
<p>To prepare a talk about the agent loop, nothing beats implementing it and running it, again and again, on models that &quot;fit&quot; on my machine. And this is where things get interesting and instructive: models like those from Anthropic, OpenAI, Mistral, ... served through a cloud API rarely get the format of a tool call wrong, correctly re-read a block of code to edit, and know when to stop. A local model with 2 to 12B parameters, does not (actually, even with bigger ones, I run into this kind of problem).</p>
<p>It&#39;s while preparing the end (and the sequel?) of my talk: how to work around the weaknesses of small models (or how to improve my agent loop), that <strong>Mini Me</strong> was born, at first as a little brother next to the talk, and finally as a project in its own right.</p>
<h2>Why an agent dedicated to small local models</h2>
<p>The guiding idea of <strong>Mini Me</strong>: when the model gets it wrong, or &quot;loops&quot;, give it back actionable information rather than a silent failure (or stop it, or have a way to stop it).</p>
<h3>At the beginning: a single tool, by default: bash</h3>
<p><strong>Mini Me</strong> was born around a single tool, <code>bash</code>. The others (reading/writing/editing files, multi-step plan, reading skills, fetching web pages) were added little by little. They are optional and can be disabled one by one in the <strong>Mini Me</strong> configuration file.</p>
<h3>When the model gets it wrong: a failure is text</h3>
<p>An error (failed command, missing file, ambiguous edit, timeout) never surfaces as a system error that would cut the turn short. It goes back <strong>to the model, as text</strong>, written as a repair instruction rather than as a raw message. The goal being that the model can eventually correct itself (it can still be improved, but my history system lets me investigate to improve failure handling).</p>
<h3>Against hallucinations, infinite loops and silent hangs</h3>
<p>Several guardrails, from the gentlest to the most radical:</p>
<ul>
<li>A loop detector &quot;warns&quot; the model as soon as the same tool call repeats identically, then stops the turn if the warning is ignored;</li>
<li>A watch on the arguments being generated, to immediately interrupt a tool call that goes off the rails by repeating a pattern while the model &quot;spins out of control&quot;;</li>
<li>And a watchdog that cancels the generation if no more tokens arrive (with a configurable delay).</li>
</ul>
<h3>ACP (Agent Client Protocol)</h3>
<p>I had another wish: even though <strong>Mini Me</strong> comes with a TUI, I wanted a comfortable way to use it in an IDE. So <strong>Mini Me</strong> also speaks <strong><a href="https://agentclientprotocol.com/get-started/introduction">ACP</a></strong> (Agent Client Protocol), the protocol that lets you plug an agent directly into an editor like <strong><a href="https://zed.dev/">Zed</a></strong>.</p>
<p>I&#39;ll write a new post very soon to explain how to set it up.</p>
<h2>Installing Mini Me</h2>
<p><strong>Mini Me</strong> comes as a binary, straight from the releases: <strong><a href="https://rickub.com/bots-garden/mini-me/releases">https://rickub.com/bots-garden/mini-me/releases</a></strong>.</p>
<h3>macOS (Apple Silicon)</h3>
<pre><code class="language-bash"># download the binary release in https://rickub.com/bots-garden/mini-me/releases
VERSION=0.2.6
curl -fsSLO &quot;https://rickub.com/bots-garden/mini-me/releases/download/v${VERSION}/mm-${VERSION}-darwin-arm64&quot;
chmod +x ./mm-${VERSION}-darwin-arm64
xattr -cr ./mm-${VERSION}-darwin-arm64
sudo rm -f /usr/local/bin/mm
sudo cp ./mm-${VERSION}-darwin-arm64 /usr/local/bin/mm
</code></pre>
<h3>Linux</h3>
<p>Same principle, without the <code>xattr</code> step (specific to macOS Gatekeeper quarantine):</p>
<pre><code class="language-bash"># download the binary release in https://rickub.com/bots-garden/mini-me/releases
VERSION=0.2.6
curl -fsSLO &quot;https://rickub.com/bots-garden/mini-me/releases/download/v${VERSION}/mm-${VERSION}-linux-amd64&quot;
chmod +x ./mm-${VERSION}-linux-amd64
sudo rm -f /usr/local/bin/mm
sudo cp ./mm-${VERSION}-linux-amd64 /usr/local/bin/mm
</code></pre>
<blockquote>
<p>Change <code>linux-amd64</code> to <code>linux-arm64</code> if you&#39;re on an ARM machine.</p>
</blockquote>
<h3>Windows</h3>
<blockquote>
<p>✋ I haven&#39;t tested <strong>Mini Me</strong> &quot;for real&quot; on Windows yet</p>
</blockquote>
<p>No <code>chmod</code>, no <code>sudo</code>, no quarantine attribute to remove: download the executable and put it in a folder that is in your <code>PATH</code>.</p>
<pre><code class="language-powershell"># download the binary release in https://rickub.com/bots-garden/mini-me/releases
$VERSION = &quot;0.2.6&quot;
Invoke-WebRequest -Uri &quot;https://rickub.com/bots-garden/mini-me/releases/download/v$VERSION/mm-$VERSION-windows-amd64.exe&quot; -OutFile &quot;mm.exe&quot;
New-Item -ItemType Directory -Force -Path &quot;$env:LOCALAPPDATA\mini-me&quot; | Out-Null
Copy-Item &quot;.\mm.exe&quot; &quot;$env:LOCALAPPDATA\mini-me\mm.exe&quot; -Force
# Then, add $env:LOCALAPPDATA\mini-me to the user PATH.
</code></pre>
<p>You&#39;ll find all the available versions on the matching release page, for example for 0.2.6: <strong><a href="https://rickub.com/bots-garden/mini-me/releases/v0.2.6">https://rickub.com/bots-garden/mini-me/releases/v0.2.6</a></strong>.</p>
<h2>First configuration</h2>
<p><strong>Mini Me</strong> plugs into a local OpenAI-compatible server. Here I&#39;m using <strong>llmman</strong>, which I talked about in a <a href="https://k33g.org/p/20261001-llmman">previous post</a>, but <a href="https://docs.docker.com/ai/model-runner/">Docker Model Runner</a> or <a href="https://llama.app/">llamacpp</a> work just as well, the contract is the same.</p>
<p>Start the server and pull a model:</p>
<pre><code class="language-bash">llmman serve
llmman pull huggingface.co/jetbrains/mellum2-12b-a2.5b-instruct-gguf-q4_k_m:Q4_K_M
llmman list
</code></pre>
<p>Then configure the agent in a YAML file, for example <code>agent.llmman.yaml</code> (you can name it whatever you like):</p>
<pre><code class="language-yaml">provider: llmman
model: huggingface.co/jetbrains/mellum2-12b-a2.5b-instruct-gguf-q4_k_m:Q4_K_M

baseUrl: http://127.0.0.1:17434/v1
fallback: http://host.docker.internal:17434/v1

bashTool: true
editTools: true
planTool: true

contextWindow: 0
maxOutput: 16000
maxTurns: 40

sessions:
  enabled: true
  dir: .mm/sessions

system: |
  Your name is Bob.
  You are a coding agent working in a terminal.
  You have a &quot;bash&quot; tool to run shell commands.
  Use it to explore files, run tests, inspect the repository, etc.
  Chain several commands if needed, then answer clearly in English.

sampling:
  temperature: 0.0
  parallel_tool_calls: false
  top_p: 0.9
  max_tokens: 4096

watchdogTimeout: 80s
firstTokenTimeout: 3m

context:
  enabled: true
  threshold: 75
  maxMessages: 80
  keepLastTurns: 3
  summaryMaxTokens: 2000
  showStats: true
</code></pre>
<p>A few details:</p>
<ul>
<li><code>baseUrl</code> points to the OpenAI API exposed by <code>llmman serve</code>, and <code>fallback</code> lets you fall back on <code>host.docker.internal</code> if <strong>Mini Me</strong> itself runs in a container.</li>
<li><code>sampling</code> is deliberately conservative: <code>temperature: 0.0</code>, one tool call at a time (<code>parallel_tool_calls: false</code>). A small model produces a correct call more easily when it only has to handle one at a time.</li>
<li><code>firstTokenTimeout</code> and <code>watchdogTimeout</code> are two separate delays: the <strong>first</strong> one covers reading the prompt (which can take several minutes on a large local context, before the first token), the <strong>second</strong> one covers silence once the answer has started.</li>
<li><code>context</code> enables automatic compression of the history before the context window fills up (what the model is able to process), which is essential with a local model whose context window is fixed and more limited than huge models.</li>
</ul>
<p>Then run <code>mm</code> in the project folder (<code>mm --tui agent.llmman.yaml</code>): <strong>Mini Me</strong> will read the configuration, connect to <code>llmman</code>, and you&#39;ll be able to start chatting with your agent locally.</p>
<p><img src="https://writizzy.b-cdn.net/blogs/b3ee339b-7400-4cf9-9f4a-ef02d709455f/media/1791042284872-20261003-hello-mini-me.png" alt="hello-mini-me" /></p>
<blockquote>
<p><strong>✋ Warning: in TUI mode, when tools are enabled, they only work in YOLO mode (unlike ACP mode).</strong> (I&#39;ll change that in an upcoming release).</p>
</blockquote>
<h2>Mini Me is not here to replace the usual coding agents</h2>
<p>Let me be clear: <strong>Mini Me</strong> does not aim to compete with existing coding agents (Docker Agent, Claude Code, Codex, OpenCode, etc. ...) running on huge models. That&#39;s not <strong>Mini Me</strong>&#39;s playground. The idea is to make the most of <strong>small local models</strong>, the ones running on your machine without a cloud API, with their own constraints (smaller context window, tendency to hallucinate, loops, fragile tool calls), by building an agent where every design choice starts from the intrinsic constraints of small models.</p>
<h2>So, to sum up</h2>
<ul>
<li><strong>Mini Me</strong> was born while preparing my <strong>Devoxx</strong> talk about the agent loop.</li>
<li>It&#39;s an agent dedicated to small local models (2 to 12B ... and even a bit more), with guardrails designed for their typical failures: hallucinations, loops, truncated tool calls, silent hangs.</li>
<li><strong>Mini Me</strong> also speaks <strong>ACP</strong> (detailed setup in an upcoming post, I promise).</li>
<li>It installs as a single binary (<strong>Mini Me</strong> is written in Go, no dependencies), is configured in YAML, and plugs into any OpenAI-compatible server (llmman, Docker Model Runner, llamacpp...).</li>
</ul>
<p>Feel free to ask questions 🙂</p>
]]></content:encoded>
            <category>tiny language models</category>
            <category>local models</category>
            <category>guardrails</category>
            <category>go</category>
            <enclosure url="https://writizzy.b-cdn.net/blogs/b3ee339b-7400-4cf9-9f4a-ef02d709455f/media/1791042286063-20261003-header.jpeg" length="0" type="image/jpeg"/>
        </item>
    </channel>
</rss>