Spend any time in self-hosting circles and you'll see the same post over and over. Someone has installed a model, pointed Continue.dev at it, and now they "run their own AI".
I know that setup well, because I built it. Back in April my first router did exactly that: a couple of models on two GPUs, an endpoint, and Continue.dev in the IDE. I wrote about it in Building a Self-Hosted LLM Router in Go.
It was a good first step. It was also short sighted, and I didn't realise how short sighted until I went further.
What You Actually Have
If you've loaded a model and connected an IDE plugin to it, here is what you have:
- A model that knows nothing about your code beyond what the plugin pastes into the prompt
- A model that knows nothing about your architecture, your conventions, or the decisions you made last year
- A model that can't check its own work
- A model that forgets everything the moment the session ends
- A GPU that sits idle whenever you aren't typing
All the intelligence about how to use the model lives in someone else's plugin. You've built the engine and handed the steering wheel to whoever wrote the extension. When the plugin can't do something, neither can your "AI".
The Model Is the Commodity
Models are getting better and cheaper every month. The one you're running today will be replaced in a few months by something better that you can download in ten minutes. That is exactly why the model is the least valuable part of the setup.
The value is in what surrounds it:
Context. A model that can search your source code and documentation through RAG answers questions about your system, not about systems in general.
Knowledge. Your rules, procedures, runsheets and architecture decisions, stored somewhere the model can reach. This is the difference between an assistant that writes plausible code and one that writes your code.
Routing. One model doing everything is one model doing most things badly. Intent classification, tool calling, embeddings and code generation are different jobs, and smaller specialist models often do them better than one big generalist.
Tools. A model that can read logs, probe services, run builds and look at a rendered page can check its own work. A model that can only emit text can't.
Resilience. What happens when the GPU node goes offline? If the answer is "everything stops", you have a dependency, not a platform.
The Tool Calling Lesson
This is where I learned the most. My April router used Qwen2.5-Coder-7B for agent work. It could produce tool call JSON perfectly. It could not reason about what came back. Call ls, get a file listing, call ls again. Forever.
I added loop detection, next-file hints, JSON schema constraints and deduplication. Every one of them was scaffolding around a gap in the model. The real fix was architectural: give tool calling its own model, standardise tool output before any model sees it, and stop expecting one model to be good at everything.
You don't learn that by pointing a plugin at an endpoint. You learn it by owning the whole loop.
Local Is Not the Same as Secure
"It runs locally, so it's secure" is the other half of the short-sightedness.
Local inference fixes data egress. Your prompts stay on your network. That matters, and it is a big part of why I built Marvin.
But the moment you give a model the ability to act, a new problem appears. Many agent setups hand the model a shell, because a shell is the quickest way to make an agent useful. Now a language model, which can be wrong with total confidence, has arbitrary command execution on your machine.
Marvin has no shell. It has around 90 tools, each with defined inputs, outputs and scope. Deploying to Kubernetes, querying logs, reading a spreadsheet, pushing to git: each is a designed path. That is a trust boundary, and you can't get one from a plugin.
The Economics Don't Work Either
A GPU that only serves autocomplete is idle most of the day. You've bought hardware to save on subscriptions, then used a fraction of it.
Self-hosting only beats paying for tokens when the hardware is doing real work. That means summarising in the background, generating embeddings, routing intent, running agent tasks and indexing knowledge. A system keeps the hardware busy. A single endpoint mostly waits.
What I'm Not Saying
I'm not saying don't run a local model with Continue.dev. Do it. It is the right first step, and it teaches you a lot.
I'm saying don't stop there and call it done. A model behind an endpoint is a component. The platform is everything you build around it: the context, the knowledge, the routing, the tools, the boundaries, and the parts that keep it running when hardware goes away.
That is where the value is. It's also where the interesting engineering is.
Running something similar, or think I've got this wrong? The comments are open.
No comments:
Post a Comment