Sunday, 4 October 2026

Running an LLM Is Not an AI Platform

Spend any time in self-hosting circles and you'll see the same post over and over. Someone has installed a model, pointed Continue.dev at it, and now they "run their own AI".

I know that setup well, because I built it. Back in April my first router did exactly that: a couple of models on two GPUs, an endpoint, and Continue.dev in the IDE. I wrote about it in Building a Self-Hosted LLM Router in Go.

It was a good first step. It was also short sighted, and I didn't realise how short sighted until I went further.

What You Actually Have

If you've loaded a model and connected an IDE plugin to it, here is what you have:

  • A model that knows nothing about your code beyond what the plugin pastes into the prompt
  • A model that knows nothing about your architecture, your conventions, or the decisions you made last year
  • A model that can't check its own work
  • A model that forgets everything the moment the session ends
  • A GPU that sits idle whenever you aren't typing

All the intelligence about how to use the model lives in someone else's plugin. You've built the engine and handed the steering wheel to whoever wrote the extension. When the plugin can't do something, neither can your "AI".

The Model Is the Commodity

Models are getting better and cheaper every month. The one you're running today will be replaced in a few months by something better that you can download in ten minutes. That is exactly why the model is the least valuable part of the setup.

The value is in what surrounds it:

Context. A model that can search your source code and documentation through RAG answers questions about your system, not about systems in general.

Knowledge. Your rules, procedures, runsheets and architecture decisions, stored somewhere the model can reach. This is the difference between an assistant that writes plausible code and one that writes your code.

Routing. One model doing everything is one model doing most things badly. Intent classification, tool calling, embeddings and code generation are different jobs, and smaller specialist models often do them better than one big generalist.

Tools. A model that can read logs, probe services, run builds and look at a rendered page can check its own work. A model that can only emit text can't.

Resilience. What happens when the GPU node goes offline? If the answer is "everything stops", you have a dependency, not a platform.

The Tool Calling Lesson

This is where I learned the most. My April router used Qwen2.5-Coder-7B for agent work. It could produce tool call JSON perfectly. It could not reason about what came back. Call ls, get a file listing, call ls again. Forever.

I added loop detection, next-file hints, JSON schema constraints and deduplication. Every one of them was scaffolding around a gap in the model. The real fix was architectural: give tool calling its own model, standardise tool output before any model sees it, and stop expecting one model to be good at everything.

You don't learn that by pointing a plugin at an endpoint. You learn it by owning the whole loop.

Local Is Not the Same as Secure

"It runs locally, so it's secure" is the other half of the short-sightedness.

Local inference fixes data egress. Your prompts stay on your network. That matters, and it is a big part of why I built Marvin.

But the moment you give a model the ability to act, a new problem appears. Many agent setups hand the model a shell, because a shell is the quickest way to make an agent useful. Now a language model, which can be wrong with total confidence, has arbitrary command execution on your machine.

Marvin has no shell. It has around 90 tools, each with defined inputs, outputs and scope. Deploying to Kubernetes, querying logs, reading a spreadsheet, pushing to git: each is a designed path. That is a trust boundary, and you can't get one from a plugin.

The Economics Don't Work Either

A GPU that only serves autocomplete is idle most of the day. You've bought hardware to save on subscriptions, then used a fraction of it.

Self-hosting only beats paying for tokens when the hardware is doing real work. That means summarising in the background, generating embeddings, routing intent, running agent tasks and indexing knowledge. A system keeps the hardware busy. A single endpoint mostly waits.

What I'm Not Saying

I'm not saying don't run a local model with Continue.dev. Do it. It is the right first step, and it teaches you a lot.

I'm saying don't stop there and call it done. A model behind an endpoint is a component. The platform is everything you build around it: the context, the knowledge, the routing, the tools, the boundaries, and the parts that keep it running when hardware goes away.

That is where the value is. It's also where the interesting engineering is.

Running something similar, or think I've got this wrong? The comments are open.

You Don't Know GPUs

In April I had a conversation about AI. Somewhere in it I was told I wouldn't be any good in the field because I "don't know GPUs".

I didn't know what that meant then. Six months later, I still don't.

What I do know is what happened next. I went home and started building. That project became Marvin.

Where It Started

The first version was not much. A couple of small models on a GTX 1650 and an RTX 3050, and a router I rewrote in Go over a weekend to send requests to the right one. I wrote about that in Building a Self-Hosted LLM Router in Go, all 18 phases of things breaking.

It worked. But it was a model behind an endpoint with some clever routing in front of it. It was not yet something I would trust with real work.

Where It Is Now

Marvin today is a self-hosted agentic AI platform. I use it to write production code. It replaces paid tokens, and it was designed security first from the start.

Nine services

  • Gateway router takes requests from the UI and sends them to the right backend service.
  • LLM router live-probes every model and cascades to the best one available.
  • Context chunks content for RAG ingest and handles retrieval.
  • Knowledge base holds long term rules, procedures and runsheets.
  • Data index is the only service that touches the database. It owns every schema and the full lifecycle of every NATS queue.
  • Workspace provides around 90 tools, all operating in confinement with no shell access.
  • Tool results processor turns raw tool output into something a model can actually use.
  • Agent service takes direct agent calls from design tools.
  • LLM proxies, one per model, in-cluster or on remote llama.cpp nodes.

Specialist models

Instead of one model doing everything badly, each job has its own model: coding, fallback inference, summarisation, embeddings, intent classification, tool calling, vision, and image generation. The right model for the right job, each isolated behind its own proxy.

Two outside dependencies

Apart from PostgreSQL with pgvector and NATS, every part of Marvin's platform is designed, engineered and written by me. When something goes wrong I know where to look, because I wrote the thing that broke.

Security First

Running a model locally solves one problem: your prompts don't leave your network. It doesn't solve the bigger one, which is what happens when that model starts acting on your behalf.

Marvin can read and write code, query cluster logs, deploy to Kubernetes, manage databases and push to git. That is a lot of power to hand to a language model. So Marvin has no shell. Every action goes through one of its tools, each with defined inputs, defined outputs and a defined scope. Nothing is uncontrolled. Nothing happens that I didn't design a path for.

That boundary was a decision made on day one, not a patch added later.

So What Does "Knowing GPUs" Mean?

I've thought about this a lot. Here is the closest I've come to an answer.

The GPU is the easy part. You buy one, you load a model, it runs. Anyone can do that in an afternoon.

The hard part is everything around it. Which model does which job. What happens when a node goes offline. Where the trust boundary sits. What the system is allowed to touch and what it isn't. How you keep the whole thing auditable when it is writing code on your behalf. How you make sure an idle GPU is still earning its keep.

None of that is GPU knowledge. It is engineering. It is the same engineering I have been doing for a long time, pointed at a new kind of problem.

What I Took From It

If someone tells you you're not cut out for something, it is worth asking what they actually mean. Sometimes there is a real gap and you learn something. Sometimes the answer is nothing at all.

I still couldn't tell you what "knowing GPUs" means. But I can tell you exactly what every part of Marvin does, because I wrote it.

Next post: why "I run an LLM" is not the same as having an AI platform.

Marvin Can Now See, Draw and Read

Back in June, Marvin became self aware. I asked it what it was, and it searched my own documentation and told me. It knew what it was because I had taught it what it was.

But it was blind. Everything Marvin understood arrived as text. Hand it a screenshot, a spreadsheet or a Word document and it had nothing to work with. That was one of the last real advantages the paid tools still had over it.

That gap is now closed. Marvin can see images, generate them, and read most of the document formats that show up in a working week.

Seeing

Vision arrived the same way every other capability in Marvin arrives: as a specialist. A new vision-language model joins the mesh with its own dedicated proxy, which makes it the seventh LLM in the system. The LLM router probes it like any other model, and requests that involve an image are sent to it.

Images reach Marvin from two places:

  • Uploads. Drop a screenshot, a diagram or a photo of a whiteboard into the chat and Marvin can describe it, reason about it, and act on it.
  • The test browser. Marvin already had a headless browser for probing services inside the cluster. It can now take a screen capture of what that browser rendered and look at it.

The second one is the one that matters most to me. Marvin builds Vue frontends. Until now it could check that a page returned a 200 and that the right elements were in the DOM, but it had no idea what the page actually looked like. A button can exist and still be hidden behind a modal. A layout can be valid HTML and still be broken. Now Marvin can deploy what it built, open it, capture it, and see it.

Drawing

Image generation runs on Z-Image through a Stable Diffusion backend.

It currently runs on CPU. That is a deliberate trade-off, not an accident. The GPUs are busy doing the work that matters most: code inference, tool calling, embeddings and intent routing. Pulling VRAM away from the coder so Marvin can draw faster would be the wrong priority. On CPU, generation is slower, but it works, and it costs the rest of the system nothing.

Because image generation sits behind the same proxy pattern as everything else, moving it onto a GPU later is a placement change, not a redesign.

Reading

Document conversion uses markitdown, a pure Go port of Microsoft's Python markitdown library. It converts a wide range of formats into clean Markdown:

  • Word (.docx), including headings, tables, lists, comments and equations
  • Excel (.xlsx and legacy .xls), with every sheet as a Markdown table
  • PowerPoint (.pptx), including slide notes
  • PDF, HTML, CSV, EPUB, Jupyter notebooks, RSS and plain text
  • ZIP archives, converting the supported files inside

The important part is not the conversion. It is where the result goes.

Converted documents are written back into the workspace as Markdown. From that point on they are just files, and every other tool in the workspace can use them. Search them, reference them while writing code, feed them into a build, compare versions of them. A document is converted once and becomes useful everywhere.

Why markitdown

Marvin is almost entirely my own code. The only outside software it depends on is PostgreSQL with pgvector, and NATS. Everything else is designed, engineered and written by me. So adding anything from outside is a decision I make consciously.

markitdown made the cut because of how it is built. It is pure Go with no CGO and no external runtime. PDF extraction runs PDFium as WebAssembly. It compiles straight into the workspace service as a library. There is no LibreOffice in a sidecar, no Python runtime, and no shell call to a converter.

That last point matters. Marvin's workspace has no shell access by design. Every action goes through a defined tool with known inputs and outputs. A document converter that needed to shell out to something would have broken that boundary. A library that lives inside the tool does not.

Writing my own parsers for a dozen file formats would have been the wrong trade-off. The line between build and buy is a judgement call. This one was easy.

Where Marvin Stands Now

Marvin now runs seven specialist language models plus an image generator, each with its own proxy:

RoleJob
CoderPrimary code inference
InferenceFallback coder and general chat
CompletionSummarisation
EmbeddingRAG embeddings
IntentClassifying what each prompt is asking for
ToolsAgent mode tool calling
VisionUnderstanding uploaded images and browser captures
Image generationZ-Image, currently on CPU

The pattern has not changed since June: the right model for the right job, each one isolated behind its own proxy, all of it running on my hardware.

What Changed

In June I wrote that Marvin knew what it was. Now it can look at what it made.

It can read the spec I wrote in Word, see the screenshot I took of the bug, and check that the page it just deployed actually looks right. None of that leaves my network. None of it is billed per token.

Marvin still costs me electricity. It just does a lot more with it.

Building something similar? Have questions about the stack? The comments are open.