Self-hosted inference, vector database, monitoring, and API layer on your VPS. No per-token bills. No vendor lock-in. Nobody can pull it out from under you. The smartest model you can economically own and operate.
The Matt Stack is not "cheaper than API." It is AI infrastructure you control. Your data stays on your metal. Your inference runs on your GPU. Your costs are flat and predictable. And nobody can change the pricing, throttle your access, or deprecate your model out from under you.
Your business data, your customer data, your prompts. Nothing routes through a third-party API. Full privacy by architecture, not by policy.
No per-token surprises. No usage spikes. Your GPU runs at a fixed rate whether you send 10 requests or 10,000. Budget once, run forever.
Models are interchangeable services behind an OpenAI-compatible endpoint. Swap Qwen for DeepSeek without changing application code. No lock-in at any layer.
Every component is open-source or Apache 2.0 licensed. Battle-tested. Replaceable. The stack deploys as a single Docker Compose on any GPU VPS provider.
| Layer | Component | Why |
|---|---|---|
| OS | Ubuntu 24.04 LTS | Massive driver/tooling ecosystem. Boring. Correct. |
| Container | Docker + Compose | The lingua franca. Every component ships a Dockerfile. |
| PaaS | Coolify v4 | 280+ templates, managed DBs, auto-SSL. Free, Apache 2.0. |
| Edge | Cloudflare | DNS, CDN, DDoS, WAF. Non-negotiable at this price point. |
| Proxy | Caddy | Auto-TLS, clean config. Coolify-native. |
| Database | PostgreSQL + pgvector | Relational + vector search in one engine. No extra dependency. |
| Cache | Redis | Session, queue, rate-limiting, KV cache coordination. |
| Inference | vLLM v0.6+ | OpenAI-compatible API. PagedAttention v2, speculative decoding. |
| Model | Qwen3-Coder-Next (default) | Best self-hosted coding model. Apache 2.0. Swappable. |
| API | FastAPI | Agent/API layer, business logic, auth, rate limiting. |
| Gateway | LiteLLM | Routes local + external APIs. Fallback, cost tracking, unified interface. |
| Monitor | Prometheus + Grafana | GPU utilization, inference latency, error rates, cost per request. |
Infrastructure cost depends on which GPU tier fits your workload. Client pays their VPS provider directly. Faction deploys and manages the stack on top of it.
24GB VRAM. Fits Qwen 27B dense at Q4. Good for general-purpose inference at lower volume.
$255 - $505/mo
48GB VRAM. Fits Qwen3-Coder-Next full precision. Larger MoE models at Q4.
$525 - $630/mo
80GB VRAM. Most MoE models, 70B dense at Q4. The workhorse for production agent workloads.
$440 - $780/mo
80GB HBM3. Any single-GPU-feasible model at full speed. For throughput-critical production.
$750 - $1,825/mo
Pricing from neo-cloud GPU providers (Vast.ai, RunPod, Spheron). Hyperscaler pricing (AWS, Azure) runs 2-5x higher. Stack certified on both. Add $20-50/month for the non-GPU application VPS.
Faction deploys and manages it for mid-market companies. The self-serve kit puts the same architecture in the hands of builders who can run it themselves.
Guided deployment: $1,500 one-time. Everything in the kit, plus a 2-hour guided deployment session, provider selection help, model benchmarking on your hardware, and a post-deploy health check.
The Matt Stack sells on sovereignty and control, not cost. But the economics do work for specific use cases.
Regulated industries, sensitive customer data, proprietary business logic. If the data can't leave your infrastructure, self-hosting is the only option.
Production agent workloads with high sustained utilization. Break-even against frontier APIs around 160-256M tokens/month at 60-70% utilization.
Model deprecation, pricing changes, rate limits, terms-of-service shifts. Your stack runs the same tomorrow regardless of what any provider announces today.
Production Discovery finds the problems. The Conveyor Belt fixes the code. The Matt Stack provides the infrastructure. FOE monitors everything.
Your team vibe-codes on the Matt Stack's inference layer. The Belt finishes, stages, and audits those builds into production. Self-hosted AI meets managed engineering.
Full AI operating platform builds can deploy on Matt Stack infrastructure instead of cloud APIs. Domain services, agents, and evaluation loops running on hardware you own.
Zero setup. Pay per token. Data leaves your network on every call. Provider controls pricing, model availability, rate limits, and terms. Fast to start, impossible to control at scale.
Your GPU, your model, your data. Flat monthly cost. Swap models without changing code. No rate limits, no per-token bills, no vendor lock-in. Takes a week to deploy. Runs as long as you want it to.
Tell us what you're running. We'll recommend the GPU tier, configure the stack, and hand you the keys to AI infrastructure nobody can take away.
Scope your stackThe Matt Stack runs the same methodology as every Faction build. The Conveyor Belt finishes your code. AIOS builds the platform. The Matt Stack provides the foundation.