vLLM, llama.cpp and Ollama flags, generated.
vllm serve meta-llama/Llama-2-7b-hf \ --tensor-parallel-size 1 \ --gpu-memory-utilization 0.90 \ --max-model-len 4096
Flags explained
- --tensor-parallel-size
- — Number of GPUs to shard the model's weights across.
- --gpu-memory-utilization
- — Fraction of each GPU's memory vLLM is allowed to reserve for weights + KV cache.
- --max-model-len
- — Maximum context length (prompt + generation) the server will accept.
This command is generated in your browser from the options you pick — nothing is uploaded or sent anywhere.