CVE-2026-24271
Official description Straight from the sourceThe vendor's or NVD's own wording, published unedited. Authoritative, but often terse — it says what broke, rarely what to do.
NVD · uneditedNVIDIA TensorRT-LLM contains a vulnerability in the OpenAI-compatible inference API, where an attacker could cause allocation of GPU resources without limits or throttling. A successful exploit of this vulnerability might lead to denial of service.
Technical summary Written by usOur analysis, written from the advisory, the CVSS vector and the affected-version data. It adds context the advisory leaves out, and never invents facts that are not in the source.
dbcve analysis · moderate confidenceNVIDIA TensorRT-LLM's OpenAI-compatible inference API lacks proper resource limits or throttling on GPU memory allocation, allowing an attacker to exhaust GPU resources by requesting excessive allocations. This creates a denial-of-service condition where legitimate requests cannot be serviced due to resource starvation.
Verify against the referenced sources before acting — the references below are authoritative for this CVE, this summary is not.
CVSS breakdown How the score is builtThe industry scoring standard. It rates how the flaw is reached, what it takes to exploit, and what an attacker gains — the score is derived from those, not the other way round.
From the vector- Attack vector
- Local
- Complexity
- Low
- Privileges
- None
- User interaction
- None
- Scope
- Unchanged
- Confidentiality
- None
- Integrity
- None
- Availability
- High
CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
Am I affected? How to checkSteps we derive from the advisory and the affected-version data, so you can decide whether this CVE reaches your setup. They are a guide, not a scan — your own configuration is the authority.
dbcve checksWork through these to decide whether this CVE applies to you.
-
Identify TensorRT-LLM installationRun 'pip list | grep tensorrt' or check for tensorrt-llm package installation directory. If using container, check 'nvidia-smi' for TensorRT-LLM processes.Affected if TensorRT-LLM is installed and running
-
Confirm OpenAI-compatible API is enabledCheck if the TensorRT-LLM server is running with the OpenAI API endpoint (typically port 8000, /v1/completions or /v1/chat/completions paths). Review startup scripts or docker-compose for '--api_model' or 'trt_llm_serve' with OpenAI flags.Affected if The OpenAI-compatible inference API endpoint is exposed and accessible
-
Verify GPU memory limit configurationReview TensorRT-LLM configuration files or runtime flags for memory-related settings such as '--max_num_tokens', '--max_batch_size', or 'gpu_memory_limit'. Check if any memory bounds are set on the inference server.Affected if No GPU memory allocation limits are configured or enforced per request
-
Check for request throttling settingsInspect if rate limiting, request queuing, or throttling is enabled on the API endpoint. Look for configurations like 'max_requests_per_minute', 'request_queue_size', or external rate limiters (e.g., nginx rate limits, token bucket settings).Affected if No throttling or rate-limiting mechanism is configured on the inference API
-
Monitor GPU memory behavior under loadDuring active inference usage, run 'nvidia-smi' or 'nvtop' to observe GPU memory allocation patterns. Check if a single request or burst of requests can consume all available GPU memory.Affected if Single requests can consume all available GPU memory causing legitimate requests to fail
A user is affected if they run TensorRT-LLM with the OpenAI-compatible API enabled and have not configured GPU memory limits or request throttling, allowing resource exhaustion.
Generated from the published advisory. Verify against your own configuration.
Remediation Closing itWhat it takes to close this. Where a vendor fix exists we point at it; where none exists we say so plainly, and can build one. Effort estimates are scoped from the advisory, not from your codebase.
From vendor dataImplement GPU resource limits and request throttling within the TensorRT-LLM inference API to bound memory allocation per request and enforce queueing or rate-limiting on incoming requests.
- Consultation4.0 h
- Implementation8.0 h
- Testing6.0 h
- Review / QA2.0 h
An estimate, not a bill — we confirm scope with you before any work starts. Need it this week? Rush from $5,600.
Scan for this in your stack
Free · runs locallyCheck whether your project pulls in CVE-2026-24271 — or any other known-vulnerable package — straight from your lock files. Free and open source; it runs locally and uploads nothing.
References Go to the primary sourcePrimary sources — vendor advisories, patches and trackers. Where our summary and a reference disagree, the reference wins.
Primary sourcesPractitioner notes
ContributedPeer-ranked notes from engineers who’ve handled CVE-2026-24271 in production — separate from our analysis above.
The advisory tells you what broke. It rarely tells you what actually worked. If you’ve dealt with this one, that detail is what the next engineer is searching for.
- The version that genuinely resolved it — not the one the vendor claimed
- A config change or rule that shut the vector down
- A gotcha in the upgrade path that cost you an afternoon
No notes yet
Be the first to add a field note for this CVE — a mitigation you’ve verified, a version caveat, or a link to a working fix. Sign in above to contribute.
A place for practitioners to share what actually worked: a mitigation you’ve tested, a configuration change, a version- or environment-specific caveat, or a link to a verified patch. The most useful notes rise to the top as peers upvote them, so the signal stays high.
- Verified mitigations, workarounds, and config changes
- Version or environment caveats, and links to real fixes
- No weaponised exploit code, or anything meant to cause harm
- No spam, self-promotion, credentials, or personal data