CVE-2026-53923
vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor processing. The output tensor is allocated at full size via torch::empty (uninitialized memory), but the dequantize CUDA kernel processes only a truncated number of elements. The unfilled portion of the output tensor retains whatever was previously in GPU memory. In multi-tenant inference deployments, this residual GPU memory may contain tensor data from other users' inference requests, constituting information disclosure. This vulnerability is fixed in 0.23.1rc0.
- Affected products
- Vllm
- Vllm
- < 0.23.1
- Fix
- Available
- CVSS 3.1
- 7.5 HIGH
- EPSS
- 0.3% (20th percentile)
- Weakness
- CWE-681, CWE-200
- NVD status
- Analyzed
- Published
- 2026-06-22
No indexed exploits for CVE-2026-53923 yet
Our index is partial: it proves presence, never absence
No exploit for CVE-2026-53923 has been indexed yet. Our index is built from live traffic and upstream syncs, so this page can only say what it knows — not that no exploit exists.