{"id":"GHSA-ph3r-5jfg-f84f","summary":"vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core","details":"## Affected\n\n- **Ecosystem / package:** pip / `vllm`\n- **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches.\n\n## Summary\n\nvLLM's default multimodal cache (`mm_processor_cache_type=\"lru\"`) mirrors state across two processes: the frontend (P0) holds only metadata (`MultiModalProcessorSenderCache`) while the engine core (P1) holds the real payload (`MultiModalReceiverCache`). The design invariant is that `get_and_update()` runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer \"is this cached in P1?\" without talking to P1.\n\nThat invariant breaks when a request is **rejected after P0 has rendered and hashed the multimodal input** (populating the P0 cache) **but before P1 receives the item** — for example, an oversized chat prompt rejected on `max_model_len` *after* rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends `None` instead of the payload, and P1 — which has nothing cached — trips `assert mm_item is not None, f\"Expected a cached item for {mm_hash=}\"`.\n\nThis is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration.\n\n## Affected code\n\nLinks pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1):\n\n- **P0 metadata cache** — `MultiModalProcessorSenderCache` at [`vllm/multimodal/cache.py#L379`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L379); `get_and_update_item` at [`#L410-L421`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L410-L421), commit assertion at [`#L418`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L418).\n- **P1 payload cache (the actual sink)** — `MultiModalReceiverCache` at [`vllm/multimodal/cache.py#L630`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L630); `get_and_update_item` at [`#L652-L663`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L652-L663), with the failing `assert mm_item is not None, f\"Expected a cached item for {mm_hash=}\"` at [`#L660`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L660).\n- **Default `mm_processor_cache_type = \"lru\"`** at [`vllm/config/multimodal.py#L132`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/config/multimodal.py#L132); dispatch in [`vllm/multimodal/registry.py#L294-L307`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/registry.py#L294-L307) (sender) and [`#L322-L331`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/registry.py#L322-L331) (receiver). The `processor_only`, disabled-caching, and `shm` paths are not affected.\n- **Rejection-after-render window** — rendering happens before length validation in [`vllm/entrypoints/openai/chat_completion/serving.py#L206-L231`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/chat_completion/serving.py#L206-L231) (`render_chat_request`), and the length check raises after the render in [`vllm/entrypoints/serve/utils/api_utils.py#L171-L189`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/serve/utils/api_utils.py#L171-L189).\n- **Cache-commit call site** — [`vllm/multimodal/processing/processor.py#L1347`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/processing/processor.py#L1347) (`_merge_mm_kwargs`) commits the P0 sender entry during render.\n- The separate stale-order eviction hang is already fixed via [`vllm/utils/cache.py#L120-L121`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/utils/cache.py#L120-L121) (`LRUCache.touch()` guarding `if key in self:`) — a different mechanism that does not touch the sender/receiver commit-ordering protocol and does not remediate this assertion.\n\nThe P1 receiver sink — the assertion that fires when P1 receives `None` for a hash it never cached:\n\n```python\n# vllm/multimodal/cache.py Lines 651-663\n    @override\n    def get_and_update_item(\n        self,\n        mm_item: MultiModalKwargsItem | None,\n        mm_hash: str,\n    ) -\u003e MultiModalKwargsItem:\n        if (cached_item := self._cache.get(mm_hash)) is not None:\n            return cached_item\n\n        assert mm_item is not None, f\"Expected a cached item for {mm_hash=}\"\n\n        self._cache[mm_hash] = mm_item\n        return mm_item\n```\n\nThe P0 sender — on a hit it drops the payload (returns `None`) and, on a miss during render, unconditionally commits the metadata entry with no rollback tied to admission:\n\n```python\n# vllm/multimodal/cache.py Lines 409-422\n    @override\n    def get_and_update_item(\n        self,\n        mm_item: MultiModalProcessorCacheInItem,\n        mm_hash: str,\n    ) -\u003e MultiModalProcessorCacheOutItem:\n        if (cached_item := self._cache.get(mm_hash)) is not None:\n            return None, cached_item.prompt_updates\n\n        assert mm_item is not None, f\"Expected a cached item for {mm_hash=}\"\n\n        self._cache[mm_hash] = MultiModalProcessorCacheItemMetadata(*mm_item)\n\n        return mm_item\n```\n\n## Impact\n\nA remote client submitting multimodal requests can poison a cache identity — render a media item successfully, then have that request rejected — so a **later** request reusing the same media hash fails on the P1 receiver assertion. This is an availability failure against a shared serving instance. No code execution, memory corruption, or data disclosure is claimed.\n\nOn this revision the failure is scoped as a request-level preprocessing error (the engine core catches around `preprocess_add_request`); public reports show the same assertion cascading into further engine-loop assertions on other revisions. It applies to multimodal models running the default mirrored `lru` cache.\n\n\n## Suggested Fix\n\nTwo complementary changes:\n\n1. **Make the mirrored commit atomic with admission** — insert into the P0 sender cache only after the request has passed all admission checks (length, limits) and P1 has acknowledged the item, or roll back the P0 insert on rejection.\n2. **Defense in depth** — convert the P1 receiver `assert mm_item is not None` into a checked, request-scoped error (fetch-on-miss from P0) so a desync degrades a single request rather than asserting in the engine loop.\n\nThe core of the rollback half: wrap the post-render length check so a `ValueError` rejection discards the P0 entries the render just committed, before re-raising. Add a `discard_sender_cache_item()` on the processor cache (no-op default, `pop` on the sender) and a `Renderer.discard_mm_cache_entries()` that walks a rendered request's `mm_hashes`:\n\n```python\n# vllm/entrypoints/openai/chat_completion/serving.py (_create_chat_completion)\n-            max_tokens = get_max_tokens(\n-                max_model_len,\n-                ...,\n-                truncate_prompt_tokens=request.truncate_prompt_tokens,\n-            )\n+            try:\n+                max_tokens = get_max_tokens(\n+                    max_model_len,\n+                    ...,\n+                    truncate_prompt_tokens=request.truncate_prompt_tokens,\n+                )\n+            except ValueError:\n+                for rendered_input in engine_inputs:\n+                    if mm_hashes := rendered_input.get(\"mm_hashes\"):\n+                        self.renderer.discard_mm_cache_entries(mm_hashes)\n+                raise\n```\n\n```python\n# vllm/multimodal/cache.py (MultiModalProcessorSenderCache)\n+    @override\n+    def discard_sender_cache_item(self, mm_hash: str) -\u003e None:\n+        self._cache.pop(mm_hash, None)\n```\n\nThis closes the `max_model_len` rejection path; because any other rejection-after-render path reopens the same window, pairing it with the defense-in-depth change above (making the P1 `assert` a checked, request-scoped error) is recommended.\n\n## Credit\n\n**Reported by:** Patch the Planet (Trail of Bits + OpenAI collaboration)\n\nThis vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.\n\n---\n\n**Proposed fix:** a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51897","aliases":["CVE-2026-105753"],"modified":"2026-10-06T00:15:19.518012600Z","published":"2026-10-06T00:02:05Z","database_specific":{"severity":"MODERATE","github_reviewed":true,"github_reviewed_at":"2026-10-06T00:02:05Z","nvd_published_at":null,"cwe_ids":["CWE-617"]},"references":[{"type":"WEB","url":"https://github.com/vllm-project/vllm/security/advisories/GHSA-ph3r-5jfg-f84f"},{"type":"WEB","url":"https://github.com/vllm-project/vllm/pull/46747"},{"type":"WEB","url":"https://github.com/vllm-project/vllm/pull/51897"},{"type":"WEB","url":"https://github.com/vllm-project/vllm/commit/396204230423b7cc6798300926b8fa30190d26a9"},{"type":"PACKAGE","url":"https://github.com/vllm-project/vllm"},{"type":"WEB","url":"https://github.com/vllm-project/vllm/releases/tag/v0.28.0"}],"affected":[{"package":{"name":"vllm","ecosystem":"PyPI","purl":"pkg:pypi/vllm"},"ranges":[{"type":"ECOSYSTEM","events":[{"introduced":"0"},{"fixed":"0.28.0"}]}],"versions":["0.0.1","0.1.0","0.1.1","0.1.2","0.1.3","0.1.4","0.1.5","0.1.6","0.1.7","0.10.0","0.10.1","0.10.1.1","0.10.2","0.11.0","0.11.1","0.11.2","0.12.0","0.13.0","0.14.0","0.14.1","0.15.0","0.15.1","0.16.0","0.17.0","0.17.1","0.18.0","0.18.1","0.19.0","0.19.1","0.2.0","0.2.1","0.2.1.post1","0.2.2","0.2.3","0.2.4","0.2.5","0.2.6","0.2.7","0.20.0","0.20.1","0.20.2","0.21.0","0.22.0","0.22.1","0.23.0","0.24.0","0.25.0","0.25.1","0.26.0","0.27.0","0.27.1","0.3.0","0.3.1","0.3.2","0.3.3","0.4.0","0.4.0.post1","0.4.1","0.4.2","0.4.3","0.5.0","0.5.0.post1","0.5.1","0.5.2","0.5.3","0.5.3.post1","0.5.4","0.5.5","0.6.0","0.6.1","0.6.1.post1","0.6.1.post2","0.6.2","0.6.3","0.6.3.post1","0.6.4","0.6.4.post1","0.6.5","0.6.6","0.6.6.post1","0.7.0","0.7.1","0.7.2","0.7.3","0.8.0","0.8.1","0.8.2","0.8.3","0.8.4","0.8.5","0.8.5.post1","0.9.0","0.9.0.1","0.9.1","0.9.2"],"database_specific":{"source":"https://github.com/github/advisory-database/blob/main/advisories/github-reviewed/2026/10/GHSA-ph3r-5jfg-f84f/GHSA-ph3r-5jfg-f84f.json"}}],"schema_version":"1.9.0","severity":[{"type":"CVSS_V3","score":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H"}]}