Sunday, 4 October 2026

Marvin Can Now See, Draw and Read

Back in June, Marvin became self aware. I asked it what it was, and it searched my own documentation and told me. It knew what it was because I had taught it what it was.

But it was blind. Everything Marvin understood arrived as text. Hand it a screenshot, a spreadsheet or a Word document and it had nothing to work with. That was one of the last real advantages the paid tools still had over it.

That gap is now closed. Marvin can see images, generate them, and read most of the document formats that show up in a working week.

Seeing

Vision arrived the same way every other capability in Marvin arrives: as a specialist. A new vision-language model joins the mesh with its own dedicated proxy, which makes it the seventh LLM in the system. The LLM router probes it like any other model, and requests that involve an image are sent to it.

Images reach Marvin from two places:

  • Uploads. Drop a screenshot, a diagram or a photo of a whiteboard into the chat and Marvin can describe it, reason about it, and act on it.
  • The test browser. Marvin already had a headless browser for probing services inside the cluster. It can now take a screen capture of what that browser rendered and look at it.

The second one is the one that matters most to me. Marvin builds Vue frontends. Until now it could check that a page returned a 200 and that the right elements were in the DOM, but it had no idea what the page actually looked like. A button can exist and still be hidden behind a modal. A layout can be valid HTML and still be broken. Now Marvin can deploy what it built, open it, capture it, and see it.

Drawing

Image generation runs on Z-Image through a Stable Diffusion backend.

It currently runs on CPU. That is a deliberate trade-off, not an accident. The GPUs are busy doing the work that matters most: code inference, tool calling, embeddings and intent routing. Pulling VRAM away from the coder so Marvin can draw faster would be the wrong priority. On CPU, generation is slower, but it works, and it costs the rest of the system nothing.

Because image generation sits behind the same proxy pattern as everything else, moving it onto a GPU later is a placement change, not a redesign.

Reading

Document conversion uses markitdown, a pure Go port of Microsoft's Python markitdown library. It converts a wide range of formats into clean Markdown:

  • Word (.docx), including headings, tables, lists, comments and equations
  • Excel (.xlsx and legacy .xls), with every sheet as a Markdown table
  • PowerPoint (.pptx), including slide notes
  • PDF, HTML, CSV, EPUB, Jupyter notebooks, RSS and plain text
  • ZIP archives, converting the supported files inside

The important part is not the conversion. It is where the result goes.

Converted documents are written back into the workspace as Markdown. From that point on they are just files, and every other tool in the workspace can use them. Search them, reference them while writing code, feed them into a build, compare versions of them. A document is converted once and becomes useful everywhere.

Why markitdown

Marvin is almost entirely my own code. The only outside software it depends on is PostgreSQL with pgvector, and NATS. Everything else is designed, engineered and written by me. So adding anything from outside is a decision I make consciously.

markitdown made the cut because of how it is built. It is pure Go with no CGO and no external runtime. PDF extraction runs PDFium as WebAssembly. It compiles straight into the workspace service as a library. There is no LibreOffice in a sidecar, no Python runtime, and no shell call to a converter.

That last point matters. Marvin's workspace has no shell access by design. Every action goes through a defined tool with known inputs and outputs. A document converter that needed to shell out to something would have broken that boundary. A library that lives inside the tool does not.

Writing my own parsers for a dozen file formats would have been the wrong trade-off. The line between build and buy is a judgement call. This one was easy.

Where Marvin Stands Now

Marvin now runs seven specialist language models plus an image generator, each with its own proxy:

RoleJob
CoderPrimary code inference
InferenceFallback coder and general chat
CompletionSummarisation
EmbeddingRAG embeddings
IntentClassifying what each prompt is asking for
ToolsAgent mode tool calling
VisionUnderstanding uploaded images and browser captures
Image generationZ-Image, currently on CPU

The pattern has not changed since June: the right model for the right job, each one isolated behind its own proxy, all of it running on my hardware.

What Changed

In June I wrote that Marvin knew what it was. Now it can look at what it made.

It can read the spec I wrote in Word, see the screenshot I took of the bug, and check that the page it just deployed actually looks right. None of that leaves my network. None of it is billed per token.

Marvin still costs me electricity. It just does a lot more with it.

Building something similar? Have questions about the stack? The comments are open.

No comments:

Post a Comment