netzstrategen
Technology

Self-Hosting an LLM: What It Costs, What It Takes and When It Pays Off

Published on 8/13/2026 · André Hellmann

“We run our own LLM” sounds like independence. Technically, getting started is easy today: a machine with enough graphics memory, a tool like Ollama, thirty minutes. The hard question comes after. Anyone self-hosting an LLM buys more than compute. They buy an operating responsibility. This article does the math on what that costs, and when it pays off.

Where do you stand?

Discuss your next step in a free diagnosis call. Book a slot →

Contents

What self-hosting an LLM really means

Self-hosted AI means the model runs on hardware the company controls. No request leaves its own infrastructure. That is one half of the story.

The other half is operations. Self-hosting a model means taking on permanent responsibility: for availability, updates, access rights, logging and security. Starting a model takes minutes. Keeping it in production is an ongoing job.

Three things get underestimated regularly. First, utilization: a GPU costs the same whether it computes or waits. Second, pace: open models ship weekly, and every update needs testing. Third, the interface: an API endpoint alone is not a working tool for employees.

Where a company stands on AI today can be mapped in a few minutes with the free Self-Check.

Hardware: how much GPU a model needs

The decisive figure is graphics memory, or VRAM. A reliable rule of thumb: parameters times bits per weight, divided by eight, gives the memory requirement in gigabytes. A model with 8 billion parameters at 4-bit quantization therefore needs roughly 4 GB, plus headroom for context and cache.

In practice, four tiers emerge:

  • 8–12 GB VRAM: models with 7–8 billion parameters, quantized. Enough for text tasks, summaries and simple classification.
  • 24 GB VRAM: models up to around 30 billion parameters, quantized. Noticeably better quality and more context.
  • 48 GB VRAM: 70-billion-parameter models when quantized, or several smaller models in parallel.
  • 80–96 GB VRAM: production use with large models and multiple concurrent users.

Quantization is the biggest lever here. It reduces the precision of model weights and with it the memory footprint, at a moderate cost in quality. A 70B model that occupies two specialist GPUs at full precision runs on one when quantized.

The difference between “it runs” and “it holds up” matters. A model on a laptop answers one person. Fifty concurrent requests are a different problem. That takes more than VRAM. It takes throughput, which is the real discipline of LLM inference.

The question is never whether a model runs. The question is whether the operation holds up.

European alternatives to buying hardware

Companies that do not want to buy hardware can rent it. European providers make that possible without routing through US data centers, a practical building block for digital sovereignty.

Hetzner offers the GEX131, a dedicated GPU server with 96 GB of graphics memory, from €889 per month, with no setup fee and in ISO 27001-certified data centers in Germany and Finland (Source: Hetzner, 2025). OVHcloud provides GPU instances by the hour; an L40S with 48 GB runs at around €1.40 per hour (Source: OVHcloud Public Cloud GPU, 2026). That is roughly €1,000 per month if run continuously, but cheap for time-boxed workloads.

The three realistic tiers:

ScenarioHardware / cloudSetup effortMonthly costSuitable for
SmallExisting workstation, 12–24 GB VRAM, Ollama½–1 dayPower and maintenance, no license costTests, prototypes, single departments
MediumDedicated GPU server, 96 GB VRAM (e.g. Hetzner GEX131)2–5 daysfrom €889Steady internal load, sensitive data
LargeMultiple GPU servers or cloud GPUs with redundancySeveral weeksfrom around €2,000 plus staffingHigh load, failover, SLA

The largest cost item is missing from that table: internal staff time. Setup, hardening, monitoring and updates consume capacity permanently, not once.

Tools: Ollama, Open WebUI and vLLM

Three tools cover almost every case. They solve different problems and complement each other.

ToolJobStrengthLimit
OllamaLoad and run models locallyFastest entry point, very simple to useNot built for throughput across many users
Open WebUIChat interface with user managementMakes the model usable for employeesDoes not replace governance
vLLMHigh-performance inference serverHigh throughput under concurrent loadMore setup and operating effort

The typical path: Ollama for proof of concept, Open WebUI as the working surface, vLLM once several teams access it at the same time. All three expose an OpenAI-compatible interface. Swapping the server therefore does not force a rewrite of applications. That is what reduces dependency.

Which models are worth considering in the first place is covered in the open source LLM overview.

The math: self-hosted vs. API

Now the numbers. A GPU server costs a fixed amount per month regardless of usage. An API costs per token: little at low volume, a lot at high volume. Somewhere the two lines cross.

The calculation with documented prices: a dedicated GPU server costs €889 per month (Source: Hetzner, 2025). A frontier API such as Claude Sonnet costs $2 per million input and $10 per million output tokens (Source: Anthropic, 2026), around $4 per million tokens blended at a three-to-one ratio. A hosted open-model API sits far below that: Qwen3 235B costs $0.20 input and $0.60 output per million tokens at Together AI (Source: Together AI, 2026), roughly $0.30 blended.

When your own server starts to pay off Monthly cost by token volume: model calculation €1,600 €889 €0 Break-even ≈ 222M tokens 0 200M 400M Tokens per month GPU server (€889/month) Frontier API (≈ $4/M) Open-model API (≈ $0.30/M) Illustrative · netzstrategen model calculation; prices: Hetzner 2025, Anthropic 2026, Together AI 2026 netzstrategen
Against a frontier API, an in-house GPU server pays off from around 222 million tokens per month. Against a hosted open-model API, it never does in this calculation. Dollar prices are treated one to one as euros for simplicity.
For presentations:

The result is uncomfortable. Break-even against the frontier API sits at around 222 million tokens per month: roughly 7.4 million tokens per day, or about 3,700 substantial requests daily. Few mid-market companies reach that. Against a hosted open-model API, an in-house server practically never pays off, because providers spread GPU utilization across many customers.

Not included: operating effort, on-call coverage and failover. Price those honestly and break-even moves further right. How to reduce token cost in the first place is covered in the article on token-smart architecture.

The conclusion: self-hosting is rarely a cost decision. It is a control decision.

Conclusion: three questions before deciding

Three questions lead to a defensible answer.

  1. Must the data stay in-house? If yes, for legal or contractual reasons, self-hosting is often the only option, regardless of price.
  2. Is the load high and steady? Only then does the fixed monthly fee carry. Fluctuating usage clearly favors an API.
  3. Are there people to run it? Without named ownership and a time budget, an in-house server becomes a project nobody maintains.

Two yeses on questions 1 and 3 are usually enough. Question 2 alone does not carry the decision.

The pragmatic path starts small: an open model on existing hardware, one clearly defined use case, measured usage over three months. After that, the math rests on data instead of assumptions. This question belongs inside continuous AI Operations, not in a one-off decision.

Frequently asked questions

What does it cost to run your own LLM?

Initial tests on existing hardware cost little beyond electricity and time. A dedicated GPU server with 96 GB of graphics memory starts at €889 per month (Source: Hetzner, 2025). On top of that comes internal operating effort, which is missing from almost every calculation.

What is the minimum hardware an LLM needs?

Quantized models with 7 to 8 billion parameters run on 8 to 12 GB of VRAM. Quantized 70-billion-parameter models need around 48 GB. Memory requirements follow roughly from parameters times bits per weight, divided by eight.

Is Ollama suitable for production?

For individual users and prototypes, yes. Under many concurrent requests Ollama hits its limits. That is where vLLM fits. Both expose an OpenAI-compatible interface, so switching is manageable.

Is self-hosting cheaper than an API?

In most cases, no. Break-even against a frontier API sits at around 222 million tokens per month, and against low-cost open-model APIs it usually never arrives. The reason to self-host is data control, not price.

How do you prepare the decision properly?

With measurement instead of estimation: actual token volume, data protection requirements and available operating capacity. We map that out in a free diagnosis call.

Sources

André Hellmann

Author & editorial responsibility

André Hellmann

Founder & Managing Director

Founder and Managing Director of netzstrategen GmbH, on board since 2006. His focus: measurement, analytics and strategy definition. Today above all building AI Operations, from strategy to day-to-day operations. Industry experience in pharma, automotive and manufacturing.

Profile & all posts Book a call LinkedIn

How this article was produced

Human
  • Topic selection
  • Source selection
  • Fact-checking
  • Approval
AI
  • Research
  • Drafting
  • Diagrams
  • Publishing

This article was produced with AI support. Ideation, editorial planning, substantive review and approval rest with a human; copy-editing sits with the AI. Editorial responsibility is held by André Hellmann.

How we produce our content →

What's next

Self-Check

Assess AI potential

In 5 minutes: a concrete assessment of where the company stands with AI.

Start the Self-Check →
Newsletter

Digital Impact straight to your inbox

One sign-up, three newsletters: the AI Insights Newsletter every week with the latest insights articles, the Digital Impact Longread Newsletter and the Digital Impact Update once a month each. Double opt-in, unsubscribe anytime.

Podcast

AI Operations as a podcast

Experts including André Hellmann, Christina D'Ilio, Christian Sattel, Sarah Stock and regular guests from practice: all AI Operations topics as audio for on the go.