LLM Weights & Inference VRAM Estimator
Estimate inference VRAM from weights, KV cache, and runtime overhead and compare it with aggregate GPU VRAM.
Parameters
Results
Assumptions
- 顯存由權重、KV Cache 和運行時開銷共同估算;KV Cache 需按上下文長度、批量和量化方式輸入,并假設可跨 GPU 分片。
Important notes
- 未單獨建模激活值、通信緩沖、框架碎片和具體推理引擎實現差異;不能只用權重大小選 GPU。