GHSA-2823-QMQ8-RWVJ
Vulnerability from github – Published: 2026-10-05 23:42 – Updated: 2026-10-05 23:42Affected
- Ecosystem / package: pip /
vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit
752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the loosecache_saltvalidator and unguarded scheduling-path lookup reach.
Summary
vLLM's OpenAI-compatible request models (Completions, Chat Completions, Responses) accept a client-supplied cache_salt field and validate it only as "must be a non-empty string" — no character or length restrictions. On a deployment with the built-in LMCache-MP KV connector enabled, that value is stored verbatim on the request tracker and forwarded unguarded as a keyword argument into the scheduler's per-step cache lookup. The downstream LMCache library applies a stricter check in IPCCacheServerKey.__post_init__, which raises ValueError for any cache_salt containing @, /, \, or NUL (or longer than 128 characters).
Neither the LMCache-MP connector lookup call site nor Scheduler.schedule() wraps that call in a request-scoped try/except, so the ValueError propagates uncaught into EngineCore's top-level handler, which treats any uncaught exception as fatal and kills the whole engine process. A single publicly reachable request with, for example, cache_salt="/" therefore takes down the engine for all concurrent users. vLLM's boundary validator is looser than the downstream consumer's, and the gap is never converted into a request-scoped failure on the scheduling path.
Affected code
Links pinned to the confirmed commit 752a3a504485 (v0.25.1):
- The three
check_cache_salt_supportvalidators require only a non-empty string — no character or length bound:vllm/entrypoints/openai/completion/protocol.py#L502-L508(field at#L172),vllm/entrypoints/openai/chat_completion/protocol.py#L913-L919(field at#L425), andvllm/entrypoints/openai/responses/protocol.py#L459-L465(field at#L235). The loose test itself is atcompletion #L503-L505,chat_completion #L914-L916,responses #L460-L462. - Two further request models accept
cache_saltwith the same or weaker checking, and should be hardened at the same time:vllm/entrypoints/pooling/base/protocol.py#L74-L85carries the identical non-empty-string-only validator (field at#L59), and the token-in-token-out scale-outGenerateRequestexposes the field with nocache_saltvalidator at all:vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L110. - The LMCache-MP request tracker stores the salt verbatim:
vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L189(LMCacheMPRequestTracker), assignment at#L223. - The lookup call forwards
cache_saltwith no surroundingtry:vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L733-L774(get_num_new_matched_tokens→maybe_submit_lookup_request(...)at#L770). - The scheduler invokes the connector unguarded:
vllm/v1/core/sched/scheduler.py#L739(self.connector.get_num_new_matched_tokens(...), insideschedule()at #L396). - The generic top-level handler treats the exception as fatal:
vllm/v1/engine/core.py#L1229-L1235— theexcept Exceptioninrun_engine_core(#L1154) logsEngineCore encountered a fatal error., calls_send_engine_dead(), and re-raises. - Downstream strict validator (external LMCache library, not vLLM):
IPCCacheServerKey.__post_init__inlmcache/v1/multiprocess/custom_types.pyrejects@ / \NUL and >128-char salts (_SALT_FORBIDDEN_CHARS = frozenset("@/\\\x00")) by raisingValueError.
The three OpenAI check_cache_salt_support validators are identical; the completion one is representative — a non-empty-string test with no character or length bound:
# vllm/entrypoints/openai/completion/protocol.py Lines 503-509
if data.get("cache_salt") is not None and (
not isinstance(data["cache_salt"], str) or not data["cache_salt"]
):
raise VLLMValidationError(
"Parameter 'cache_salt' must be a non-empty string if provided.",
parameter="cache_salt",
)
On the scheduling path the connector is invoked with no surrounding try — a ValueError from the downstream salt check propagates straight out of schedule():
# vllm/v1/core/sched/scheduler.py Lines 736-742
# Get externally-cached tokens if using a KVConnector.
if self.connector is not None:
ext_tokens, load_kv_async = (
self.connector.get_num_new_matched_tokens(
request, num_new_local_computed_tokens
)
)
run_engine_core's generic handler — the next except up the stack — treats that as fatal, marks the engine dead, and re-raises:
# vllm/v1/engine/core.py Lines 1229-1235
except Exception as e:
if engine_core is None:
logger.exception("EngineCore failed to start.")
else:
logger.exception("EngineCore encountered a fatal error.")
engine_core._send_engine_dead()
raise e
Impact
Availability only. cache_salt is an attacker-controlled, publicly reachable request field that vLLM validates too loosely. A value such as "/" passes vLLM's check, reaches the stricter downstream validator, and its ValueError is never converted into a request-scoped failure — instead it kills the EngineCore process, a denial of service for every concurrent request on that server (HTTP failures, then /health failing).
Applicability: the built-in LMCache-MP KV connector must be enabled (lmcache >= 0.4.4), which is itself an opt-in KV-connector boundary. On such deployments no other special configuration is required, and the crash is a resource-availability failure rather than expected behavior of the opt-in feature.
Suggested Fix
Two independent fixes:
- Tighten admission — add one shared
validate_cache_salt()helper (for example invllm/entrypoints/openai/engine/protocol.py) that matches or exceeds the downstreamIPCCacheServerKeyrules — reject@,/,\, NUL, and >128-character salts at the HTTP boundary with a 4xx — and route every request model that exposescache_saltthrough it: the threecheck_cache_salt_supportvalidators above, the pooling base request, and the token-in-token-outGenerateRequest(which has no validator today). - Defense in depth — wrap the LMCache-MP lookup call reached from
Scheduler.schedule()so a downstream validatorValueErrorbecomes a request-scoped failure instead of an EngineCore-fatal exception. This is the fix that also covers future divergence between vLLM's and LMCache's salt rules; the right failure semantics (fail the request vs. fall back to a cold lookup) is a maintainer design call.
The core of fix 1 is a single shared helper; each check_cache_salt_support body then becomes validate_cache_salt(data.get("cache_salt")), and GenerateRequest gains an equivalent mode="before" validator:
# vllm/entrypoints/openai/engine/protocol.py — new shared helper
_CACHE_SALT_FORBIDDEN_CHARS = frozenset("@/\\\x00")
_MAX_CACHE_SALT_LENGTH = 128
def validate_cache_salt(cache_salt: object) -> None:
"""Validate cache salts before they reach downstream cache backends."""
if cache_salt is None:
return
if not isinstance(cache_salt, str) or not cache_salt:
raise VLLMValidationError(
"Parameter 'cache_salt' must be a non-empty string if provided.",
parameter="cache_salt",
)
if len(cache_salt) > _MAX_CACHE_SALT_LENGTH or any(
char in _CACHE_SALT_FORBIDDEN_CHARS for char in cache_salt
):
raise VLLMValidationError(
"Parameter 'cache_salt' must be at most 128 characters and must "
"not contain '@', '/', '\\\\', or NUL.",
parameter="cache_salt",
)
This distinguishes the finding from GHSA-6qc9-v4r8-22xg: that advisory fixed only the guided_json/xgrammar trigger of the same schedule()-into-run_engine_core fatal-handler family, so the cache_salt admission gap and the schedule()-level defense-in-depth (fix 2) survive its published fix. A patch implementing fix 1 across all five request models, with a regression test covering the rejected ("/", 129-char) and accepted salt shapes, applies to v0.25.1 with line offsets and no fuzz.
Credit
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51444
{
"affected": [
{
"package": {
"ecosystem": "PyPI",
"name": "vllm"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "0.30.0"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-105756"
],
"database_specific": {
"cwe_ids": [
"CWE-20",
"CWE-248"
],
"github_reviewed": true,
"github_reviewed_at": "2026-10-05T23:42:34Z",
"nvd_published_at": null,
"severity": "MODERATE"
},
"details": "## Affected\n\n- **Ecosystem / package:** pip / `vllm`\n- **Affected versions:** vLLM \u2264 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the loose `cache_salt` validator and unguarded scheduling-path lookup reach.\n\n## Summary\n\nvLLM\u0027s OpenAI-compatible request models (Completions, Chat Completions, Responses) accept a client-supplied `cache_salt` field and validate it only as \"must be a non-empty string\" \u2014 no character or length restrictions. On a deployment with the built-in LMCache-MP KV connector enabled, that value is stored verbatim on the request tracker and forwarded unguarded as a keyword argument into the scheduler\u0027s per-step cache lookup. The downstream LMCache library applies a *stricter* check in `IPCCacheServerKey.__post_init__`, which raises `ValueError` for any `cache_salt` containing `@`, `/`, `\\`, or NUL (or longer than 128 characters).\n\nNeither the LMCache-MP connector lookup call site nor `Scheduler.schedule()` wraps that call in a request-scoped `try`/`except`, so the `ValueError` propagates uncaught into `EngineCore`\u0027s top-level handler, which treats any uncaught exception as fatal and kills the whole engine process. A single publicly reachable request with, for example, `cache_salt=\"/\"` therefore takes down the engine for all concurrent users. vLLM\u0027s boundary validator is looser than the downstream consumer\u0027s, and the gap is never converted into a request-scoped failure on the scheduling path.\n\n## Affected code\n\nLinks pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1):\n\n- The three `check_cache_salt_support` validators require only a non-empty string \u2014 no character or length bound: [`vllm/entrypoints/openai/completion/protocol.py#L502-L508`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/completion/protocol.py#L502-L508) (field at [`#L172`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/completion/protocol.py#L172)), [`vllm/entrypoints/openai/chat_completion/protocol.py#L913-L919`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/chat_completion/protocol.py#L913-L919) (field at [`#L425`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/chat_completion/protocol.py#L425)), and [`vllm/entrypoints/openai/responses/protocol.py#L459-L465`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/protocol.py#L459-L465) (field at [`#L235`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/protocol.py#L235)). The loose test itself is at [`completion #L503-L505`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/completion/protocol.py#L503-L505), [`chat_completion #L914-L916`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/chat_completion/protocol.py#L914-L916), [`responses #L460-L462`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/protocol.py#L460-L462).\n- Two further request models accept `cache_salt` with the same or weaker checking, and should be hardened at the same time: [`vllm/entrypoints/pooling/base/protocol.py#L74-L85`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/pooling/base/protocol.py#L74-L85) carries the identical non-empty-string-only validator (field at [`#L59`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/pooling/base/protocol.py#L59)), and the token-in-token-out scale-out `GenerateRequest` exposes the field with **no** `cache_salt` validator at all: [`vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L110`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L110).\n- The LMCache-MP request tracker stores the salt verbatim: [`vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L189`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L189) (`LMCacheMPRequestTracker`), assignment at [`#L223`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L223).\n- The lookup call forwards `cache_salt` with no surrounding `try`: [`vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L733-L774`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L733-L774) (`get_num_new_matched_tokens` \u2192 `maybe_submit_lookup_request(...)` at [`#L770`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L770)).\n- The scheduler invokes the connector unguarded: [`vllm/v1/core/sched/scheduler.py#L739`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/core/sched/scheduler.py#L739) (`self.connector.get_num_new_matched_tokens(...)`, inside [`schedule()` at #L396](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/core/sched/scheduler.py#L396)).\n- The generic top-level handler treats the exception as fatal: [`vllm/v1/engine/core.py#L1229-L1235`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/engine/core.py#L1229-L1235) \u2014 the `except Exception` in `run_engine_core` ([`#L1154`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/engine/core.py#L1154)) logs `EngineCore encountered a fatal error.`, calls `_send_engine_dead()`, and re-raises.\n- Downstream strict validator (external LMCache library, not vLLM): `IPCCacheServerKey.__post_init__` in `lmcache/v1/multiprocess/custom_types.py` rejects `@ / \\` NUL and \u003e128-char salts (`_SALT_FORBIDDEN_CHARS = frozenset(\"@/\\\\\\x00\")`) by raising `ValueError`.\n\nThe three OpenAI `check_cache_salt_support` validators are identical; the completion one is representative \u2014 a non-empty-string test with no character or length bound:\n\n```python\n# vllm/entrypoints/openai/completion/protocol.py Lines 503-509\n if data.get(\"cache_salt\") is not None and (\n not isinstance(data[\"cache_salt\"], str) or not data[\"cache_salt\"]\n ):\n raise VLLMValidationError(\n \"Parameter \u0027cache_salt\u0027 must be a non-empty string if provided.\",\n parameter=\"cache_salt\",\n )\n```\n\nOn the scheduling path the connector is invoked with no surrounding `try` \u2014 a `ValueError` from the downstream salt check propagates straight out of `schedule()`:\n\n```python\n# vllm/v1/core/sched/scheduler.py Lines 736-742\n # Get externally-cached tokens if using a KVConnector.\n if self.connector is not None:\n ext_tokens, load_kv_async = (\n self.connector.get_num_new_matched_tokens(\n request, num_new_local_computed_tokens\n )\n )\n```\n\n`run_engine_core`\u0027s generic handler \u2014 the next `except` up the stack \u2014 treats that as fatal, marks the engine dead, and re-raises:\n\n```python\n# vllm/v1/engine/core.py Lines 1229-1235\n except Exception as e:\n if engine_core is None:\n logger.exception(\"EngineCore failed to start.\")\n else:\n logger.exception(\"EngineCore encountered a fatal error.\")\n engine_core._send_engine_dead()\n raise e\n```\n\n## Impact\n\nAvailability only. `cache_salt` is an attacker-controlled, publicly reachable request field that vLLM validates too loosely. A value such as `\"/\"` passes vLLM\u0027s check, reaches the stricter downstream validator, and its `ValueError` is never converted into a request-scoped failure \u2014 instead it kills the EngineCore process, a denial of service for every concurrent request on that server (HTTP failures, then `/health` failing).\n\nApplicability: the built-in LMCache-MP KV connector must be enabled (`lmcache \u003e= 0.4.4`), which is itself an opt-in KV-connector boundary. On such deployments no other special configuration is required, and the crash is a resource-availability failure rather than expected behavior of the opt-in feature.\n\n\n## Suggested Fix\n\nTwo independent fixes:\n\n1. **Tighten admission** \u2014 add one shared `validate_cache_salt()` helper (for example in `vllm/entrypoints/openai/engine/protocol.py`) that matches or exceeds the downstream `IPCCacheServerKey` rules \u2014 reject `@`, `/`, `\\`, NUL, and \u003e128-character salts at the HTTP boundary with a 4xx \u2014 and route every request model that exposes `cache_salt` through it: the three `check_cache_salt_support` validators above, the pooling base request, and the token-in-token-out `GenerateRequest` (which has no validator today).\n2. **Defense in depth** \u2014 wrap the LMCache-MP lookup call reached from `Scheduler.schedule()` so a downstream validator `ValueError` becomes a request-scoped failure instead of an EngineCore-fatal exception. This is the fix that also covers future divergence between vLLM\u0027s and LMCache\u0027s salt rules; the right failure semantics (fail the request vs. fall back to a cold lookup) is a maintainer design call.\n\nThe core of fix 1 is a single shared helper; each `check_cache_salt_support` body then becomes `validate_cache_salt(data.get(\"cache_salt\"))`, and `GenerateRequest` gains an equivalent `mode=\"before\"` validator:\n\n```python\n# vllm/entrypoints/openai/engine/protocol.py \u2014 new shared helper\n_CACHE_SALT_FORBIDDEN_CHARS = frozenset(\"@/\\\\\\x00\")\n_MAX_CACHE_SALT_LENGTH = 128\n\n\ndef validate_cache_salt(cache_salt: object) -\u003e None:\n \"\"\"Validate cache salts before they reach downstream cache backends.\"\"\"\n if cache_salt is None:\n return\n if not isinstance(cache_salt, str) or not cache_salt:\n raise VLLMValidationError(\n \"Parameter \u0027cache_salt\u0027 must be a non-empty string if provided.\",\n parameter=\"cache_salt\",\n )\n if len(cache_salt) \u003e _MAX_CACHE_SALT_LENGTH or any(\n char in _CACHE_SALT_FORBIDDEN_CHARS for char in cache_salt\n ):\n raise VLLMValidationError(\n \"Parameter \u0027cache_salt\u0027 must be at most 128 characters and must \"\n \"not contain \u0027@\u0027, \u0027/\u0027, \u0027\\\\\\\\\u0027, or NUL.\",\n parameter=\"cache_salt\",\n )\n```\n\nThis distinguishes the finding from **[GHSA-6qc9-v4r8-22xg](https://github.com/vllm-project/vllm/security/advisories/GHSA-6qc9-v4r8-22xg)**: that advisory fixed only the `guided_json`/`xgrammar` trigger of the same `schedule()`-into-`run_engine_core` fatal-handler family, so the `cache_salt` admission gap and the `schedule()`-level defense-in-depth (fix 2) survive its published fix. A patch implementing fix 1 across all five request models, with a regression test covering the rejected (`\"/\"`, 129-char) and accepted salt shapes, applies to v0.25.1 with line offsets and no fuzz.\n\n## Credit\n\n**Reported by:** Patch the Planet (Trail of Bits + OpenAI collaboration)\n\nThis vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.\n\n---\n\n**Proposed fix:** a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51444",
"id": "GHSA-2823-qmq8-rwvj",
"modified": "2026-10-05T23:42:34Z",
"published": "2026-10-05T23:42:34Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/vllm-project/vllm/security/advisories/GHSA-2823-qmq8-rwvj"
},
{
"type": "WEB",
"url": "https://github.com/vllm-project/vllm/pull/51444"
},
{
"type": "WEB",
"url": "https://github.com/vllm-project/vllm/commit/e962733e08d10f7ca65dac4df99e116460b8b174"
},
{
"type": "PACKAGE",
"url": "https://github.com/vllm-project/vllm"
},
{
"type": "WEB",
"url": "https://github.com/vllm-project/vllm/releases/tag/v0.30.0"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
],
"summary": "vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP deployments \u2014 uncaught downstream `ValueError` denial of service"
}
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.
The approach is described in our paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.
Browse all ATT&CK techniques and the vulnerabilities related to each.
Related by attack behaviour
Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.