Sunday, 4 October 2026

You Don't Know GPUs

In April I had a conversation about AI. Somewhere in it I was told I wouldn't be any good in the field because I "don't know GPUs".

I didn't know what that meant then. Six months later, I still don't.

What I do know is what happened next. I went home and started building. That project became Marvin.

Where It Started

The first version was not much. A couple of small models on a GTX 1650 and an RTX 3050, and a router I rewrote in Go over a weekend to send requests to the right one. I wrote about that in Building a Self-Hosted LLM Router in Go, all 18 phases of things breaking.

It worked. But it was a model behind an endpoint with some clever routing in front of it. It was not yet something I would trust with real work.

Where It Is Now

Marvin today is a self-hosted agentic AI platform. I use it to write production code. It replaces paid tokens, and it was designed security first from the start.

Nine services

  • Gateway router takes requests from the UI and sends them to the right backend service.
  • LLM router live-probes every model and cascades to the best one available.
  • Context chunks content for RAG ingest and handles retrieval.
  • Knowledge base holds long term rules, procedures and runsheets.
  • Data index is the only service that touches the database. It owns every schema and the full lifecycle of every NATS queue.
  • Workspace provides around 90 tools, all operating in confinement with no shell access.
  • Tool results processor turns raw tool output into something a model can actually use.
  • Agent service takes direct agent calls from design tools.
  • LLM proxies, one per model, in-cluster or on remote llama.cpp nodes.

Specialist models

Instead of one model doing everything badly, each job has its own model: coding, fallback inference, summarisation, embeddings, intent classification, tool calling, vision, and image generation. The right model for the right job, each isolated behind its own proxy.

Two outside dependencies

Apart from PostgreSQL with pgvector and NATS, every part of Marvin's platform is designed, engineered and written by me. When something goes wrong I know where to look, because I wrote the thing that broke.

Security First

Running a model locally solves one problem: your prompts don't leave your network. It doesn't solve the bigger one, which is what happens when that model starts acting on your behalf.

Marvin can read and write code, query cluster logs, deploy to Kubernetes, manage databases and push to git. That is a lot of power to hand to a language model. So Marvin has no shell. Every action goes through one of its tools, each with defined inputs, defined outputs and a defined scope. Nothing is uncontrolled. Nothing happens that I didn't design a path for.

That boundary was a decision made on day one, not a patch added later.

So What Does "Knowing GPUs" Mean?

I've thought about this a lot. Here is the closest I've come to an answer.

The GPU is the easy part. You buy one, you load a model, it runs. Anyone can do that in an afternoon.

The hard part is everything around it. Which model does which job. What happens when a node goes offline. Where the trust boundary sits. What the system is allowed to touch and what it isn't. How you keep the whole thing auditable when it is writing code on your behalf. How you make sure an idle GPU is still earning its keep.

None of that is GPU knowledge. It is engineering. It is the same engineering I have been doing for a long time, pointed at a new kind of problem.

What I Took From It

If someone tells you you're not cut out for something, it is worth asking what they actually mean. Sometimes there is a real gap and you learn something. Sometimes the answer is nothing at all.

I still couldn't tell you what "knowing GPUs" means. But I can tell you exactly what every part of Marvin does, because I wrote it.

Next post: why "I run an LLM" is not the same as having an AI platform.

No comments:

Post a Comment