llama.cpp hosting

The reference GGUF inference server, with nothing between it and the CPU.

LLM runtimesMITfrom $13.99/mo
In short

llama.cpp hosting at AIHostBox means a dedicated, single-tenant llama.cpp instance on our infrastructure: we install it, pin it to a reviewed release, put it behind TLS on your domain, back it up nightly and keep it patched. Plans start at $13.99 per month and the admin account, the data and any API keys are yours.

GGUFOpenAI-compatibleGrammarsLowest overhead

llama.cpp is the engine most other local runtimes are built on. Running its server directly gives you the lowest overhead of anything in this category, an OpenAI-compatible endpoint, grammar-constrained output and precise control over threads, context and quantization — at the cost of doing your own tuning.

llama.cpp hosting plans

Starter7-8B GGUF (Q4)
8 GB RAM · 4 vCPU · NVMe
$13.99/mo
Order
Plus13B GGUF (Q4)
12 GB RAM · 6 vCPU · NVMe
$24.99/mo
Order

Every plan is a dedicated single-tenant instance with NVMe storage, nightly off-box backups, TLS and managed patching. Quarterly, semi-annual and annual terms are discounted at checkout. Model tokens are not included — you bring your own key.

What people run llama.cpp for

Squeeze the most from CPU

Tune threads and batch size for your specific model.

Structured output

Grammar constraints force valid JSON without retry loops.

Licence and version policy

ItemPosition
Upstream licenceMIT
Sold as managed hostingYes
TenancySingle tenant — one customer per instance
PatchingCritical advisories within 24 hours, high severity within 72

Our full position on licences, including the applications we decline to sell, is on the licensing page.

Related in llm runtimes

How it works

Three steps, no Docker knowledge required

Pick the app and the size

Choose memory and storage from the plan table. Every plan is a single-tenant instance with its own volume, its own configuration and its own admin account.

We deploy and harden it

TLS on your domain or ours, firewall, a version pinned to a reviewed release, nightly off-box backups and isolated secrets.

You log in and build

The admin account is yours. Add your API keys, invite your team, export your data whenever you want. Patching stays with us.

Questions

llama.cpp hosting: your questions

Is this a shared account or my own instance?
Your own. Every plan is a single-tenant instance with its own storage, its own configuration and its own admin account. Nobody else’s workload runs inside it, and nothing you store is pooled with other customers.
Do I have to know Docker?
No. We install the application, put it behind TLS on your domain or a subdomain of ours, and hand you the login. If you do want shell-level access to your data we provide SFTP; you are never required to use it.
What is included in the monthly price?
The instance, the storage, the bandwidth, nightly off-box backups, monitoring, security patching and human support. Model tokens are not included — you bring your own provider key, so there is no markup on usage.
Can I move to a bigger plan later?
Yes, resizing is an in-place upgrade and your data stays where it is. If you outgrow the catalogue entirely we will help you export everything.

Run llama.cpp without running a server

A dedicated llama.cpp instance, deployed and maintained by us. Month to month, cancel whenever.