@gpuwrld
Market/vLLM
Language models

vLLM

vLLM inference server. OpenAI-compatible /v1/chat/completions endpoint.

Details

Image
vllm/vllm-openai:latest
Ports
8000/http, 22/tcp
Setup
3-4 minutes
Difficulty
advanced
Features
OpenAI-compatible API, PagedAttention, Continuous batching
Compatible GPUs
RTX 4090
24GB GDDR6X
budget$0.44 / hr
RTX A6000
48GB GDDR6
standard$0.49 / hr
L40
48GB GDDR6
standard$0.79 / hr
A100 40GB
40GB HBM2e
premium$1.29 / hr
A100 80GB
80GB HBM2e
premium$1.89 / hr
H100 80GB
80GB HBM3
enterprise$2.99 / hr
Start
locked

Hold any amount of $GPU to start a treasury-funded session. Every session is capped at 1 hour.

Runtime 1 hour (fixed)
GPU NVIDIA GeForce RTX 4090
App vLLM
Image vllm/vllm-openai:latest
Ports 8000/http,22/tcp
Cost ~$0.44 covered if eligible

Connect a wallet on Robinhood Chain.