Serving Command Builder

vLLM, llama.cpp and Ollama flags, generated.

vllm serve meta-llama/Llama-2-7b-hf \
  --tensor-parallel-size 1 \
  --gpu-memory-utilization 0.90 \
  --max-model-len 4096

Flags explained

--tensor-parallel-size
— Number of GPUs to shard the model's weights across.
--gpu-memory-utilization
— Fraction of each GPU's memory vLLM is allowed to reserve for weights + KV cache.
--max-model-len
— Maximum context length (prompt + generation) the server will accept.

This command is generated in your browser from the options you pick — nothing is uploaded or sent anywhere.

What's next?