vLLM hosting
High-throughput model serving — GPU territory, listed for completeness.
vLLM is in our catalogue but not yet orderable: it needs GPU capacity, and we would rather list it honestly than sell a CPU plan that disappoints. Tell us what you need and we will tell you where it stands.
vLLM is the throughput leader for serving open-weight models, using paged attention and continuous batching to keep a GPU saturated. It is genuinely GPU-bound software: running it on CPU defeats the point. We list it because customers ask, and we will open it when we can price GPU capacity honestly rather than approximately.
Licence and version policy
| Item | Position |
|---|---|
| Upstream licence | Apache-2.0 |
| Sold as managed hosting | No |
| Tenancy | Single tenant — one customer per instance |
| Patching | Critical advisories within 24 hours, high severity within 72 |
Our full position on licences, including the applications we decline to sell, is on the licensing page.
Related in llm runtimes
vLLM hosting: your questions
Is this a shared account or my own instance?
Do I have to know Docker?
What is included in the monthly price?
Can I move to a bigger plan later?
Tell us what you need
If this is the app you need, say so — demand is how things move up the list.