Phil Springer · Independent AI Engineer
Independent AI Engineer / Phoenix / since 2010

Multi-agent AI systems that run real operations.

I orchestrate frontier and self-hosted models into agent fleets that ship code, run infrastructure, and execute operations. Nothing goes out on one model's say-so.

frontier
Claude · GPT / Codex
Claude GPT / Codex
local · Hermes
open-weight LLMs
Hermes Z.ai Moonshot
agent fleet · MCP
code · infrastructure · operations
before ship
adversarial thread · multiple lineages
1rounds until the disagreement resolves
arguing
who

Sixteen years, my own shop.

Computer scientist by training, B.S. Computer Science with a Business Administration minor from Southern Oregon University, 2010, and an independent technology founder ever since. I have run my own S-corporation for sixteen years, and over the last few I rebuilt the entire operation around applied large language models.

I own the whole stack and ship end to end. Nothing I build is a demo. It runs my own business, and when it breaks it costs me. That is a harder test than any client review.

method

One pass is a coin flip.

A model checking its own work agrees with itself, confidently, every time. So the reviewer is always a different lineage: different training, different failure modes. That much is table stakes.

What isn't: I don't stop at one exchange. Models from different lineages argue in a single shared thread, for as many rounds as the disagreement takes, until it resolves instead of averaging out.

It matters more than it sounds. In a recent run the first answer was wrong. The correct one did not exist until the fourth round. Four separate errors surfaced along the way, and not one of them was caught by the model that made it.

A single adversarial pass would have shipped round one, with citations attached.

stack

What I actually run.

Orchestration

Multi-agent systems

Coordinating frontier models with self-hosted open-weight LLMs over MCP. Agentic coding is the daily driver, not an experiment.

Claude
GPT / Codex
Hermes · local

Local inference

Hermes · 512 GB Apple silicon

Real compute I own and run, not a rented endpoint. 1.04 TB of open-weight models resident on the box, two of them loaded at once: one to reason, one to answer instantly. Speech transcription runs on the same node, so recorded calls never leave the building.

Mac Studio
LM Studio
Hugging Face
Speech · on device
loaded now · 222.94 GB held in memory
ZGLM 4.6198.58 GBlive
ZGLM 4.7 Flash · 30B24.36 GBlive

The large model reasons. The small one answers instantly. Both stay resident at the same time, because the memory is mine and nobody meters it.

on the node · ready to load
ZGLM 5.2 · 256×22B342.74 GBresident
Kimi K2.7 Code470.63 GBresident
NNomic Embed v1.584 MBresident

Languages

Polyglot and stack-agnostic.

I ship in whatever the problem needs, on a CS foundation that goes all the way down.

Python
PHP
Laravel
JavaScript
TypeScript
Next.js
React
Node
SQL
Shell
Java
C
C++
Go

Infrastructure

A self-hosted fleet.

Linux servers plus a rebuilt box running my private Git and automation. My own Git, my own servers, my own models.

Contabo
Linux
Forgejo
Tailscale
Caddy
nginx
Docker
Cloudflare
MariaDB

Model work

Where the leverage is.

Prompt and context engineering, evaluation, adversarial red-teaming, RAG, and technical and code data generation.

Red-teaming
RAG
Evaluation
Data generation
open to

AI engineering, model evaluation, applied LLM work.

Remote, or Phoenix hybrid. You get someone who already lives inside these systems at production stakes every day, not someone ramping into them.