{"id":"GHSA-x6mc-67gf-chw4","summary":"vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the num_frames ceiling does not reach","details":"### Summary\n\nAn unauthenticated remote attacker can exhaust the memory of the vLLM API-server process by raising the request-level `media_io_kwargs.video.max_frames` and `fps` knobs on any deployment serving a Qwen2-VL or Qwen3-VL model. 74 extra bytes of JSON took the server's peak RSS from 2 271 MiB to 13 629 MiB over unauthenticated `POST /tokenize`.\n\nThe `num_frames` ceiling reported in GHSA-vxqj-p4gw-9h4c and fixed by open PR #51969 does not reach this path: the Qwen samplers do not read `num_frames` at all. The same knobs were already capped upstream for `GLMGAVideoBackend` as an accepted security fix in `8b6de0eb9` (PR #54935, merged 2026-09-04); that cap never reached Qwen.\n\n### Details\n\n#### Relationship to GHSA-vxqj-p4gw-9h4c and PR #51969 (read this first)\n\nGHSA-vxqj-p4gw-9h4c reported that request-level `media_io_kwargs.video.num_frames` overrides the engine frame-count ceiling, and open PR #51969 fixes it by clamping `num_frames` inside `VideoMediaIO.merge_kwargs`.\n\nThat clamp does not reach the Qwen samplers. `Qwen2VLVideoBackend` and `Qwen3VLVideoBackend` do not read `num_frames` at all — `Qwen2VLVideoBackend`'s own docstring says so (\"``num_frames`` is ignored (fps-driven, like the Qwen3-VL loader)\"). They bound on `max_frames`, read from the same merged dict with no ceiling:\n\n```python\n# vllm/multimodal/video.py — Qwen3VLVideoBackend.compute_frames_index_to_sample\nmin_frames = kwargs.get(\"min_frames\", 4)\nmax_frames = kwargs.get(\"max_frames\", 768)\nnum_frames = int(total_frames_num / original_fps * fps)\nnum_frames = min(max(num_frames, min_frames), max_frames, total_frames_num)\n```\n\nWith `max_frames` raised from the request, the only remaining bound is `total_frames_num` — every frame in the container.\n\nI applied PR #51969's patch locally and re-ran both paths through the real merge layer (`merge_media_io_kwargs` → `VideoMediaIO.merge_kwargs` → `MediaConnector`), against `main` @ `b23433088`:\n\n| request `media_io_kwargs.video` | merged kwargs after #51969 | frames decoded | peak RSS |\n|---|---|---|---|\n| *(absent)* | `None` | 32 | 706 MiB |\n| `{\"video_backend\":\"opencv\",\"num_frames\":-1}` | `{…,\"num_frames\":32}` | **32 — fixed** | 706 MiB |\n| `{\"video_backend\":\"qwen3_vl\"}` | `{…,\"num_frames\":32}` | 60 | 779 MiB |\n| `{\"video_backend\":\"qwen3_vl\",\"max_frames\":1e9,\"fps\":1e6}` | `{…,\"max_frames\":1000000000,\"fps\":1000000,\"num_frames\":32}` | **900 — survives** | **2 994 MiB** |\n\n#51969 does exactly what it claims for `num_frames`; the clamp writes `num_frames: 32` into the merged dict and the Qwen sampler ignores it, while `max_frames` and `fps` pass through untouched.\n\n#### The codebase already has the fix pattern, on other backends\n\nThis is not a new control being proposed. Commit `8b6de0eb9` — \"[Security] Cap GLMGA video sampling to prevent request-driven resource exhaustion\" (PR #54935, merged 2026-09-04, same author as #51969) — caps precisely these two knobs for `GLMGAVideoBackend`:\n\n```python\n_MAX_FRAMES: ClassVar[int] = 640\n_MAX_FPS: ClassVar[int] = 30\n...\ntarget_fps = min(target.fps, cls._MAX_FPS)\nmax_frames = min(kwargs.get(\"max_frames\", cls._MAX_FRAMES), cls._MAX_FRAMES)\n```\n\nIts description states the root cause as \"both `target_fps` and `max_frames` are controllable via request-level `media_io_kwargs`\", and says it follows \"the pattern established by `GLM46VVideoBackend`\" (which caps via `_MAX_FRAME_COUNT_DYNAMIC = 640` and `_MAX_DURATION = 2400`).\n\nSo two backends cap request-controlled `fps`/`max_frames` as an accepted security measure. `Qwen2VLVideoBackend` and `Qwen3VLVideoBackend` — the most widely deployed video models on vLLM — cap neither. In GLMGA the uncapped knobs sized an intermediate *index list*; in the Qwen samplers they size the *decoded frame buffer*, which is larger by the per-frame pixel count.\n\n#### Reachability — default configuration, no authentication\n\n* `media_io_kwargs` is a request body field on `ChatCompletionRequest` and is carried by `/v1/chat/completions`, `/v1/embeddings`, `/v1/responses`, `/tokenize` and `/invocations`. No flag gates it.\n* `api_key` defaults to `None` (`vllm/entrypoints/launchers/cli_args.py`), and `AuthenticationMiddleware` is installed only when a key is configured — a default `vllm serve` is entirely unauthenticated.\n* Even with `--api-key` set, `GUARDED_PREFIX = (\"/v1\", \"/v2\", \"/inference\", \"/cohere\")` (`vllm/entrypoints/serve/middleware/authenticate.py:11`), so `/invocations` — which validates the same `ChatCompletionRequest` body — and `/tokenize` — which performs full media ingestion — remain unauthenticated.\n* `MediaConnector.fetch_video` applies the model's registered sampler only when the request did not name one (`if \"video_backend\" not in video_io_kwargs`), so `video_backend: \"qwen3_vl\"` is selectable on any deployment; request-level selection of stock sampler subclasses is already established as reachable by GHSA-j682-9xp5-rrf3. On a Qwen deployment no `video_backend` key is needed at all.\n* Default-on for any video-capable Qwen model.\n\n**The Rust frontend is not affected.** `rust/src/server/src/routes/openai/chat_completions/validate.rs:90` rejects `media_io_kwargs` with \"media_io_kwargs is not supported.\" No second front is needed.\n\n### Impact\n\nUnauthenticated remote denial of service by memory exhaustion of the API-server process. The decode runs in the frontend during chat parsing, before scheduling or admission control, so every tenant on the instance is affected. `--limit-mm-per-prompt` does not apply — it bounds media *items*, not frames within an item. CWE-770 / CWE-400.\n\nMeasured over unauthenticated HTTP: 1.43 MiB request body, 74 extra bytes of JSON, server peak RSS 2 271 → 13 629 MiB. In-process, the same request takes the decode from 60 to 900 frames. The frame count is bounded only by `total_frames_num`, which is the attacker's choice of video, and decoded bytes are `frames × H × W × 3`.\n\n**Context, measured on the `num_frames` path (GHSA-vxqj-p4gw-9h4c's path), not this one** — these figures show what an unbounded frame count costs once the source video is chosen for it, and they transfer to this path because both converge on the same `_read_frames_no_recovery` allocation:\n\n* 5.74 MiB request body → 9.27 GiB decoded, 9.73 GiB of new resident memory (1 736×).\n* 117.6 MiB `video/jpeg` payload → process OOM-killed: `Out of memory: Killed process 182667 (python) total-vm:38665704kB, anon-rss:21121616kB`\n\nThe two were not re-run at the larger sizes through the Qwen sampler; the 900-frame figure above is what I measured on this path.\n\n### Suggested fix\n\nExtend the ceiling to the sampler-side knobs rather than clamping the single `num_frames` key. Two options, either acceptable:\n\n1. **Per-backend caps, matching `8b6de0eb9`.** Give `Qwen2VLVideoBackend` and `Qwen3VLVideoBackend` the `_MAX_FRAMES` / `_MAX_FPS` treatment already applied to `GLMGAVideoBackend`, so `kwargs.get(\"max_frames\", …)` and `target.fps` are clamped to class constants. Smallest change; consistent with the accepted precedent. It leaves `Glm5NextVideoBackend`, `Molmo2VideoBackend`, `NemotronVLVideoBackend`, `DynamicVideoBackend` and `OpenCVDynamicOpenPanguVideoBackend` to be audited one by one, which is the current trajectory (#55727, #56207, #56390).\n2. **Strip the frame-count knobs at the merge boundary.** In `VideoMediaIO.merge_kwargs`, drop `max_frames` / `min_frames` / `fps` from `runtime_kwargs` the way `hw_decoders`, `pool_size`, `device` and unconfigured GPU backends are already dropped there. That treats the whole frame-count family as startup-only in one place and is robust to future sampler subclasses, at the cost of removing a request-level knob some users may rely on. A clamp-don't-strip variant (request may lower, never raise) preserves the feature.\n\nOption 2 composes with #51969 and needs no per-backend audit; I would favour it, but option 1 is the more conservative change and matches what has already been merged.\n\n### Affected versions\n\n`\u003e= 0.24.0`, through v0.29.1rc0 and `main` @ `b23433088`.\n\nLower bound established by probing release tags through the GitHub contents API; no version below is inferred, each was read out of the file at that tag.\n\n**Confirmed present — Qwen samplers reading unclamped `max_frames`** (`max_frames = kwargs.get(\"max_frames\", 768)` inside `Qwen2VLVideoBackend` / `Qwen3VLVideoBackend` in `vllm/multimodal/video.py`):\n\n| ref | `Qwen3VLVideoBackend` present | unclamped `max_frames` |\n|---|---|---|\n| v0.23.0 | no (class does not exist) | n/a |\n| v0.24.0 | yes | yes |\n| v0.25.0, v0.26.0, v0.27.0, v0.28.0, v0.29.0, v0.29.1rc0 | yes | yes |\n| `main` @ `b23433088` | yes | yes |\n\n**Confirmed present — request-level `media_io_kwargs`** (field on the chat request model, and `VideoMediaIO.merge_kwargs` present): every ref probed, v0.19.0 through v0.29.1rc0 and `main` @ `b23433088`.\n\n**Confirmed absent — any `num_frames` ceiling clamp** (PR #51969 unmerged): every ref probed, v0.19.0 through v0.29.1rc0 and `main` @ `b23433088`.\n\n**Not resolved, and why:**\n\n1. The introducing commit/PR for `Qwen2VLVideoBackend` / `Qwen3VLVideoBackend`. My clone is shallow (`--depth=300`), so `git log -S'max_frames'` cannot reach it; the v0.23.0 → v0.24.0 boundary above is the tightest bound I established by probing release tags.\n2. The rc tags between v0.23.0 and v0.24.0, to tighten the lower bound to a specific release candidate.\n3. Whether a `max_frames` knob on a *differently named* pre-v0.24.0 backend is separately affected — at v0.23.0 `GLMGAVideoBackend` already carried `max_frames = kwargs.get(\"max_frames\", 640)` with no clamp, and that clamp was only added on 2026-09-04 by `8b6de0eb9`. Versions between are plausibly affected through GLMGA rather than Qwen; I did not test that path.\n4. Whether GHSA-vxqj-p4gw-9h4c's own affected range differs, which I cannot see — the advisory returns 404 to me.\n\nNo dependency versions are asserted anywhere in this report.\n\n### Weaknesses of this report, stated upfront\n\n* **No GPU was used.** This host has no CUDA device, so the HTTP results come from the in-tree GPU-less render server rather than a full `vllm serve`. That server runs the real frontend — the same request model, the same `merge_media_io_kwargs` → `VideoMediaIO.merge_kwargs` → `MediaConnector.fetch_video` chain, the same `/tokenize` route — and media decoding happens entirely in the frontend, so I do not believe the engine's presence changes the result. I have not confirmed that on a GPU deployment, and a reviewer may reasonably want that repeated under `vllm serve`.\n* The in-process measurements (the #51969 comparison table) were taken by importing the tree directly at `b23433088`, verified by `__file__`, with no `vllm` wheel installed.\n* **Amplification figures are a floor.** The OpenCV build available here offers only `mp4v`/`XVID`, giving ~2 200× compression on static content. An attacker using x264/x265 would do materially better for the same frame count.\n* The two large figures in the Impact section (9.73 GiB RSS; the OOM kill) were measured on the `num_frames` path, not this one, and are labelled as such.\n* Per-frame pixels remain bounded by `VLLM_MAX_IMAGE_PIXELS`; this concerns the unbounded frame *count*, which multiplies it.\n* `VLLM_MAX_MEDIA_DOWNLOAD_SIZE_MB` (default 256) caps the compressed size of an HTTP-fetched video but not a `data:` URI, which `MediaConnector.load_from_url` dispatches before any size logic is reached.\n\n\n---\n\n*Prepared with AI assistance (Claude), per the repository's contributing guidance on disclosing AI-assisted contributions. All findings were verified by executing vLLM's own code at the commits cited; the analysis and the claims are my own.*\n\n*Reported by Eva Crystal / 0xiviel (XSource Security).*","aliases":["CVE-2026-105758"],"modified":"2026-10-06T00:00:11.054910444Z","published":"2026-10-05T23:42:59Z","database_specific":{"severity":"MODERATE","github_reviewed":true,"github_reviewed_at":"2026-10-05T23:42:59Z","nvd_published_at":null,"cwe_ids":["CWE-770"]},"references":[{"type":"WEB","url":"https://github.com/vllm-project/vllm/security/advisories/GHSA-x6mc-67gf-chw4"},{"type":"WEB","url":"https://github.com/vllm-project/vllm/pull/56729"},{"type":"WEB","url":"https://github.com/vllm-project/vllm/commit/ea723c81c3ea26425cb69503a5d5e90822a04a45"},{"type":"PACKAGE","url":"https://github.com/vllm-project/vllm"},{"type":"WEB","url":"https://github.com/vllm-project/vllm/releases/tag/v0.30.0"}],"affected":[{"package":{"name":"vllm","ecosystem":"PyPI","purl":"pkg:pypi/vllm"},"ranges":[{"type":"ECOSYSTEM","events":[{"introduced":"0.24.0"},{"fixed":"0.30.0"}]}],"versions":["0.24.0","0.25.0","0.25.1","0.26.0","0.27.0","0.27.1","0.28.0","0.29.0"],"database_specific":{"source":"https://github.com/github/advisory-database/blob/main/advisories/github-reviewed/2026/10/GHSA-x6mc-67gf-chw4/GHSA-x6mc-67gf-chw4.json"}}],"schema_version":"1.9.0","severity":[{"type":"CVSS_V3","score":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L"}]}