Compare Q4, Q5, Q8 and FP16 trade-offs.
python -c "
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model = AutoAWQForCausalLM.from_pretrained('meta-llama/Llama-2-7b-hf', device_map='cuda:0')
tokenizer = AutoTokenizer.from_pretrained('meta-llama/Llama-2-7b-hf')
model.quantize(tokenizer, quant_config={'zero_point': True, 'q_group_size': 128, 'w_bit': 4, 'version': 'GEMM'})
model.save_quantized('Llama-2-7b-hf-awq')
"Quality / speed / VRAM trade-off
AWQ (Activation-aware Weight Quantization) at 4-bit is commonly reported in public benchmarks as staying close to FP16 quality — often a small perplexity-point delta — while cutting VRAM roughly 3-4x, and it's generally competitive with GPTQ on GPU throughput. These are general ranges from published community benchmarks, not guaranteed numbers for every model.
This command is generated in your browser from the options you pick — nothing is uploaded or sent anywhere.