vllm752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches.vLLM's default multimodal cache (mm_processor_cache_type="lru") mirrors state across two processes: the frontend (P0) holds only metadata (MultiModalProcessorSenderCache) while the engine core (P1) holds the real payload (MultiModalReceiverCache). The design invariant is that get_and_update() runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1.
That invariant breaks when a request is rejected after P0 has rendered and hashed the multimodal input (populating the P0 cache) but before P1 receives the item — for example, an oversized chat prompt rejected on max_model_len after rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends None instead of the payload, and P1 — which has nothing cached — trips assert mm_item is not None, f"Expected a cached item for {mm_hash=}".
This is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration.
Links pinned to the confirmed commit 752a3a504485 (v0.25.1):
MultiModalProcessorSenderCache at vllm/multimodal/cache.py#L379; get_and_update_item at #L410-L421, commit assertion at #L418.MultiModalReceiverCache at vllm/multimodal/cache.py#L630; get_and_update_item at #L652-L663, with the failing assert mm_item is not None, f"Expected a cached item for {mm_hash=}" at #L660.mm_processor_cache_type = "lru" at vllm/config/multimodal.py#L132; dispatch in vllm/multimodal/registry.py#L294-L307 (sender) and #L322-L331 (receiver). The processor_only, disabled-caching, and shm paths are not affected.vllm/entrypoints/openai/chat_completion/serving.py#L206-L231 (render_chat_request), and the length check raises after the render in vllm/entrypoints/serve/utils/api_utils.py#L171-L189.vllm/multimodal/processing/processor.py#L1347 (_merge_mm_kwargs) commits the P0 sender entry during render.vllm/utils/cache.py#L120-L121 (LRUCache.touch() guarding if key in self:) — a different mechanism that does not touch the sender/receiver commit-ordering protocol and does not remediate this assertion.The P1 receiver sink — the assertion that fires when P1 receives None for a hash it never cached:
# vllm/multimodal/cache.py Lines 651-663
@override
def get_and_update_item(
self,
mm_item: MultiModalKwargsItem | None,
mm_hash: str,
) -> MultiModalKwargsItem:
if (cached_item := self._cache.get(mm_hash)) is not None:
return cached_item
assert mm_item is not None, f"Expected a cached item for {mm_hash=}"
self._cache[mm_hash] = mm_item
return mm_item
The P0 sender — on a hit it drops the payload (returns None) and, on a miss during render, unconditionally commits the metadata entry with no rollback tied to admission:
# vllm/multimodal/cache.py Lines 409-422
@override
def get_and_update_item(
self,
mm_item: MultiModalProcessorCacheInItem,
mm_hash: str,
) -> MultiModalProcessorCacheOutItem:
if (cached_item := self._cache.get(mm_hash)) is not None:
return None, cached_item.prompt_updates
assert mm_item is not None, f"Expected a cached item for {mm_hash=}"
self._cache[mm_hash] = MultiModalProcessorCacheItemMetadata(*mm_item)
return mm_item
A remote client submitting multimodal requests can poison a cache identity — render a media item successfully, then have that request rejected — so a later request reusing the same media hash fails on the P1 receiver assertion. This is an availability failure against a shared serving instance. No code execution, memory corruption, or data disclosure is claimed.
On this revision the failure is scoped as a request-level preprocessing error (the engine core catches around preprocess_add_request); public reports show the same assertion cascading into further engine-loop assertions on other revisions. It applies to multimodal models running the default mirrored lru cache.
Two complementary changes:
assert mm_item is not None into a checked, request-scoped error (fetch-on-miss from P0) so a desync degrades a single request rather than asserting in the engine loop.The core of the rollback half: wrap the post-render length check so a ValueError rejection discards the P0 entries the render just committed, before re-raising. Add a discard_sender_cache_item() on the processor cache (no-op default, pop on the sender) and a Renderer.discard_mm_cache_entries() that walks a rendered request's mm_hashes:
# vllm/entrypoints/openai/chat_completion/serving.py (_create_chat_completion)
- max_tokens = get_max_tokens(
- max_model_len,
- ...,
- truncate_prompt_tokens=request.truncate_prompt_tokens,
- )
+ try:
+ max_tokens = get_max_tokens(
+ max_model_len,
+ ...,
+ truncate_prompt_tokens=request.truncate_prompt_tokens,
+ )
+ except ValueError:
+ for rendered_input in engine_inputs:
+ if mm_hashes := rendered_input.get("mm_hashes"):
+ self.renderer.discard_mm_cache_entries(mm_hashes)
+ raise
# vllm/multimodal/cache.py (MultiModalProcessorSenderCache)
+ @override
+ def discard_sender_cache_item(self, mm_hash: str) -> None:
+ self._cache.pop(mm_hash, None)
This closes the max_model_len rejection path; because any other rejection-after-render path reopens the same window, pairing it with the defense-in-depth change above (making the P1 assert a checked, request-scoped error) is recommended.
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51897
{
"cwe_ids": [
"CWE-617"
],
"github_reviewed": true,
"github_reviewed_at": "2026-10-06T00:02:05Z",
"nvd_published_at": null,
"severity": "MODERATE"
}