Self-Hosting an LLM: What It Costs, What It Takes and When It Pays Off
Published on 8/13/2026 · André Hellmann
“We run our own LLM” sounds like independence. Technically, getting started is easy today: a machine with enough graphics memory, a tool like Ollama, thirty minutes. The hard question comes after. Anyone self-hosting an LLM buys more than compute. They buy an operating responsibility. This article does the math on what that costs, and when it pays off.
Discuss your next step in a free diagnosis call. Book a slot →
Contents
- What self-hosting an LLM really means
- Hardware: how much GPU a model needs
- European alternatives to buying hardware
- Tools: Ollama, Open WebUI and vLLM
- The math: self-hosted vs. API
- Conclusion: three questions before deciding
- Frequently asked questions
- Sources
What self-hosting an LLM really means
Self-hosted AI means the model runs on hardware the company controls. No request leaves its own infrastructure. That is one half of the story.
The other half is operations. Self-hosting a model means taking on permanent responsibility: for availability, updates, access rights, logging and security. Starting a model takes minutes. Keeping it in production is an ongoing job.
Three things get underestimated regularly. First, utilization: a GPU costs the same whether it computes or waits. Second, pace: open models ship weekly, and every update needs testing. Third, the interface: an API endpoint alone is not a working tool for employees.
Where a company stands on AI today can be mapped in a few minutes with the free Self-Check.
Hardware: how much GPU a model needs
The decisive figure is graphics memory, or VRAM. A reliable rule of thumb: parameters times bits per weight, divided by eight, gives the memory requirement in gigabytes. A model with 8 billion parameters at 4-bit quantization therefore needs roughly 4 GB, plus headroom for context and cache.
In practice, four tiers emerge:
- 8–12 GB VRAM: models with 7–8 billion parameters, quantized. Enough for text tasks, summaries and simple classification.
- 24 GB VRAM: models up to around 30 billion parameters, quantized. Noticeably better quality and more context.
- 48 GB VRAM: 70-billion-parameter models when quantized, or several smaller models in parallel.
- 80–96 GB VRAM: production use with large models and multiple concurrent users.
Quantization is the biggest lever here. It reduces the precision of model weights and with it the memory footprint, at a moderate cost in quality. A 70B model that occupies two specialist GPUs at full precision runs on one when quantized.
The difference between “it runs” and “it holds up” matters. A model on a laptop answers one person. Fifty concurrent requests are a different problem. That takes more than VRAM. It takes throughput, which is the real discipline of LLM inference.
The question is never whether a model runs. The question is whether the operation holds up.
European alternatives to buying hardware
Companies that do not want to buy hardware can rent it. European providers make that possible without routing through US data centers, a practical building block for digital sovereignty.
Hetzner offers the GEX131, a dedicated GPU server with 96 GB of graphics memory, from €889 per month, with no setup fee and in ISO 27001-certified data centers in Germany and Finland (Source: Hetzner, 2025). OVHcloud provides GPU instances by the hour; an L40S with 48 GB runs at around €1.40 per hour (Source: OVHcloud Public Cloud GPU, 2026). That is roughly €1,000 per month if run continuously, but cheap for time-boxed workloads.
The three realistic tiers:
| Scenario | Hardware / cloud | Setup effort | Monthly cost | Suitable for |
|---|---|---|---|---|
| Small | Existing workstation, 12–24 GB VRAM, Ollama | ½–1 day | Power and maintenance, no license cost | Tests, prototypes, single departments |
| Medium | Dedicated GPU server, 96 GB VRAM (e.g. Hetzner GEX131) | 2–5 days | from €889 | Steady internal load, sensitive data |
| Large | Multiple GPU servers or cloud GPUs with redundancy | Several weeks | from around €2,000 plus staffing | High load, failover, SLA |
The largest cost item is missing from that table: internal staff time. Setup, hardening, monitoring and updates consume capacity permanently, not once.
Tools: Ollama, Open WebUI and vLLM
Three tools cover almost every case. They solve different problems and complement each other.
| Tool | Job | Strength | Limit |
|---|---|---|---|
| Ollama | Load and run models locally | Fastest entry point, very simple to use | Not built for throughput across many users |
| Open WebUI | Chat interface with user management | Makes the model usable for employees | Does not replace governance |
| vLLM | High-performance inference server | High throughput under concurrent load | More setup and operating effort |
The typical path: Ollama for proof of concept, Open WebUI as the working surface, vLLM once several teams access it at the same time. All three expose an OpenAI-compatible interface. Swapping the server therefore does not force a rewrite of applications. That is what reduces dependency.
Which models are worth considering in the first place is covered in the open source LLM overview.
The math: self-hosted vs. API
Now the numbers. A GPU server costs a fixed amount per month regardless of usage. An API costs per token: little at low volume, a lot at high volume. Somewhere the two lines cross.
The calculation with documented prices: a dedicated GPU server costs €889 per month (Source: Hetzner, 2025). A frontier API such as Claude Sonnet costs $2 per million input and $10 per million output tokens (Source: Anthropic, 2026), around $4 per million tokens blended at a three-to-one ratio. A hosted open-model API sits far below that: Qwen3 235B costs $0.20 input and $0.60 output per million tokens at Together AI (Source: Together AI, 2026), roughly $0.30 blended.
The result is uncomfortable. Break-even against the frontier API sits at around 222 million tokens per month: roughly 7.4 million tokens per day, or about 3,700 substantial requests daily. Few mid-market companies reach that. Against a hosted open-model API, an in-house server practically never pays off, because providers spread GPU utilization across many customers.
Not included: operating effort, on-call coverage and failover. Price those honestly and break-even moves further right. How to reduce token cost in the first place is covered in the article on token-smart architecture.
The conclusion: self-hosting is rarely a cost decision. It is a control decision.
Conclusion: three questions before deciding
Three questions lead to a defensible answer.
- Must the data stay in-house? If yes, for legal or contractual reasons, self-hosting is often the only option, regardless of price.
- Is the load high and steady? Only then does the fixed monthly fee carry. Fluctuating usage clearly favors an API.
- Are there people to run it? Without named ownership and a time budget, an in-house server becomes a project nobody maintains.
Two yeses on questions 1 and 3 are usually enough. Question 2 alone does not carry the decision.
The pragmatic path starts small: an open model on existing hardware, one clearly defined use case, measured usage over three months. After that, the math rests on data instead of assumptions. This question belongs inside continuous AI Operations, not in a one-off decision.
Frequently asked questions
What does it cost to run your own LLM?
Initial tests on existing hardware cost little beyond electricity and time. A dedicated GPU server with 96 GB of graphics memory starts at €889 per month (Source: Hetzner, 2025). On top of that comes internal operating effort, which is missing from almost every calculation.
What is the minimum hardware an LLM needs?
Quantized models with 7 to 8 billion parameters run on 8 to 12 GB of VRAM. Quantized 70-billion-parameter models need around 48 GB. Memory requirements follow roughly from parameters times bits per weight, divided by eight.
Is Ollama suitable for production?
For individual users and prototypes, yes. Under many concurrent requests Ollama hits its limits. That is where vLLM fits. Both expose an OpenAI-compatible interface, so switching is manageable.
Is self-hosting cheaper than an API?
In most cases, no. Break-even against a frontier API sits at around 222 million tokens per month, and against low-cost open-model APIs it usually never arrives. The reason to self-host is data control, not price.
How do you prepare the decision properly?
With measurement instead of estimation: actual token volume, data protection requirements and available operating capacity. We map that out in a free diagnosis call.
Sources
- Hetzner: GEX131 GPU Server with NVIDIA RTX PRO 6000 Blackwell Max-Q, 2025
- Hetzner: Dedicated GPU-Line, 2026
- OVHcloud: L40S Cloud GPU Instance, 2026
- Anthropic: Claude API Pricing, 2026
- Together AI: Pricing for Open Models, 2026
- Ollama: Documentation and Model Library, 2026
- Open WebUI: Project Documentation, 2026
- vLLM: Project Documentation, 2026
Author & editorial responsibility
Founder & Managing Director
Founder and Managing Director of netzstrategen GmbH, on board since 2006. His focus: measurement, analytics and strategy definition. Today above all building AI Operations, from strategy to day-to-day operations. Industry experience in pharma, automotive and manufacturing.
How this article was produced
- Topic selection
- Source selection
- Fact-checking
- Approval
- Research
- Drafting
- Diagrams
- Publishing
This article was produced with AI support. Ideation, editorial planning, substantive review and approval rest with a human; copy-editing sits with the AI. Editorial responsibility is held by André Hellmann.
What's next
Assess AI potential
In 5 minutes: a concrete assessment of where the company stands with AI.
Start the Self-Check →Digital Impact straight to your inbox
One sign-up, three newsletters: the AI Insights Newsletter every week with the latest insights articles, the Digital Impact Longread Newsletter and the Digital Impact Update once a month each. Double opt-in, unsubscribe anytime.
AI Operations as a podcast
Experts including André Hellmann, Christina D'Ilio, Christian Sattel, Sarah Stock and regular guests from practice: all AI Operations topics as audio for on the go.
Assess the company's AI potential in 5 minutes
Start the Self-Check →Discuss the next step with an expert
Book a call →