Because I've replaced every dependency I could, and I know what each one cost. Everyone has a stack. Almost nobody keeps a documented practice of catching their own tools lying to them.
B.S. Computer Science, Southern Oregon University, 2010. Independent technology founder ever since, and I've run my own S-corporation for sixteen years. Over the last few I rebuilt the entire operation around applied large language models.
I own the whole stack and ship end to end. Nothing I build is a demo. It runs my own business, and when it breaks it costs me. That is a harder test than any client review.
When I need to know whether something is actually true, I ask independent lineages once each. They fail differently, and that is the entire point. Agreement between models trained separately by rival labs is evidence. A model agreeing with itself is not. It's the same blind spot, twice, with more confidence.
The review panel is blind. Two lineages read the same work at the same time and neither one can see the other's answer. That ordering is the whole design. Agreement is only evidence if it was reached independently, and a shared thread manufactures the consensus it is supposed to be measuring. Whatever survives the panel then goes to a third lineage whose only job is to break it, and that seat had no hand in building the thing it is attacking.
Round one had it wrong. The correct answer did not exist until the fourth round. Four separate errors surfaced along the way, and not one of them was caught by the model that made it. A single pass would have shipped round one, with citations attached.
| Lineage | Model | Runs on | Seat |
|---|---|---|---|
| Anthropic | Claude Opus | cloud | Orchestrator |
| OpenAI | Codex | cloud | Build |
| MiniMax | MiniMax M3 | Mac Studio, 512 GB | Blind panel |
| Gemini 3.1 Pro | cloud | Blind panel | |
| xAI | Grok 4.6 | cloud | Teardown |
| Zhipu | GLM 4.7 Flash | Mac Studio, 512 GB | Local |
| Mistral | Mistral Nemo | i9 server, RTX 3060 | Local |
| Alibaba | Qwen 3 | Ryzen workstation, RTX 5080 | Local |
Two models from two labs is a second opinion. Eight lineages is a method. It is why disagreement counts as an instrument rather than noise.
Four of those seats run on three machines I own, and none of the silicon sits idle. The same inference server runs on all three boxes: the 512 GB Mac Studio holds the heavy model, the i9 server puts its RTX 3060 to work, and the Ryzen workstation carries a third. Every graphics card in the building has a job. That is not a cost decision, it is what makes a lineage mine instead of rented, and it is why a panel seat cannot be taken away from me.
The same gauntlet runs twice, at two settings. For verification work every seat runs at maximum reasoning. For coding it is the identical flow in the identical order with every seat at medium. One process, one place to change it, and no argument about which path a piece of work took.
The seats are named for the job, not the model. Break looks for an input that produces a wrong answer. Skeptic never looks for bugs at all, it asks whether the thing is worth building. That is the difference between asking several models and giving models a job to disagree with me.
A seat cannot lie about which lineage it is. Lineage is a property of the code that makes the call, not a label in a config file, so a Google call cannot be recorded as an Anthropic one. I learned that the hard way: in the system this replaced, a seat named for one lab was pointed at another, and it sat on the panel for weeks counting as an independent voice while being a second copy of a model already there.
No attacker may share the builder's lineage. A model reviewing its own output is not a review, and the seat count going up while independence stays flat is worse than having fewer seats, because the number looks better.
Coordinating frontier models with self-hosted open-weight LLMs over MCP. Agentic coding is the daily driver, not an experiment.
Real compute I own and run, not a rented endpoint. The same inference server runs on all three boxes, and every one of them holds a different model family, so three of the seats on my review panel cannot be taken away from me by a price change or a policy change. Nothing idles: the workstation and the server put their graphics cards to work instead of leaving them to collect dust. Speech is local too, an RTX 5080 runs Whisper large-v3 for dictation and video transcription, so none of it leaves the building either.
On the Mac the heavy model reasons and holds a seat on the review panel, and a light one sits beside it to answer instantly. The other two machines carry their own lineages rather than sitting idle between jobs. All four stay warm at once, because the memory is mine and nobody meters it, so I can leave a reasoning model loaded all day for work that would be economically irrational to run on rented inference.
Every one of these has shipped something that runs. A flat list invites a skeptic to pick its weakest member and discount the rest by it, so there isn't one.
Linux servers, plus a 2021 workstation I rebuilt into one. When it came out of daily service I wiped Windows off it, put Linux on, and gave an autonomous agent real work to do there. That is where the self-hosted side of this came from, and I learned it by operating it rather than reading about it. It runs my private Git and automation now. My own Git, my own servers, my own models.
Prompt and context engineering, evaluation, adversarial red-teaming, RAG, and technical and code data generation.
Self-hosted file sync instead of Drive. My own git instead of GitHub. My own calendar instead of Google's. My own relay instead of the vendor's. Local models instead of rented endpoints. A de-Microsofted editor build.
Each one was real work, and each one cost something. The editor migration alone dragged an entire status-bar patch lane with it, and the rebuild script shipped in a state where it couldn't parse, and the obvious fix would have armed a routine that killed thirty-five running processes.
That list isn't a set of tools. It's a position I keep paying for, which is the only reason it means anything.
And I won't pretend otherwise. I run on Claude and Codex. Two American AI companies, neither of which I can self-host, and both of which I depend on daily. Everything else on that list I moved off.
That trade is deliberate, and it's the one dependency I'd take again. A sovereignty claim that doesn't state its own ceiling is marketing.
What was believed, what was actually true, and how the gap was found. The failure is the product.
A self-healing watcher reported zero warnings and zero errors for weeks. It was healing a tree nobody was looking at, while the one on screen sat untouched. Green forever, and structurally incapable of telling me otherwise.
A comparison returned zero results for both candidates, which reads as clean on both. The search itself was broken. A control token caught it; re-reading the output never would have.
Security and quality updates deferred a month while feature updates shipped immediately. The inverse of what anyone would have chosen. Found because I went looking where nobody looks.
This page exists so you can judge the engineering. If you want local models running on hardware you own, on your own network, that is the work I would take, and the part I enjoy most.