PT-2026-64142 · Pypi · Vllm
Published
2026-07-23
·
Updated
2026-07-23
CVSS v3.1
7.5
High
| Vector | AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H |
Summary
A frontend-legal multi-request speculative workload can make vLLM produce an out-of-vocabulary recovered token equal to
vocab size, convert that value to -1 when choosing the next live token for a request, and then feed that -1 back into the next drafter input ids. On Qwen3 GPTQ this reaches the worker-side drafting / attention path and crashes the engine with a GPU device-side assert.The same issue is reachable through the public gRPC request surface by sending a specific overlapping
Generate / Abort sequence.Impact
- A remote client that can send public gRPC generation requests can crash the shared vLLM engine worker
- The triggering request sequence aborts concurrent requests and prevents later requests from completing until the worker is restarted
- In shared deployments, this is a service-wide denial of service for other clients, not just a failure isolated to the attacking requests
- The failure is reproducible, so repeated request sequences can sustain the outage
Affected version
- Confirmed on vLLM
0.17.1 - Earlier and later versions have not been checked yet in this report
Repro model
- Official Hugging Face repo:
Qwen/Qwen3-0.6B-GPTQ-Int8- Anyone wants to reproduce the bug with my PoC scripts should download
Qwen3-0.6B-GPTQ-Int8first
Trigger chain
- A legal multi-request speculative workload keeps structured-output state, speculative decoding, overlap, and request cancellation active in the same live engine.
- During rejection sampling, vLLM produces a recovered token equal to the
model
vocab sizeboundary value. - That recovered token appears in position 0 of the sampled speculative row
for a live request. The same row also contains trailing padding entries
equal to
-1, but those padding entries are not the key fault by themselves. - The next-token preparation step treats the position-0 recovered token as the
real next token for that request and converts that out-of-vocabulary value
to
-1. - The drafter writes that converted
-1back into the live next-step input-id row for the request. - The drafting / embedding / attention path later consumes that live invalid token and the worker crashes on GPU.
Details
Simple example
The important distinction is:
- trailing
-1values in a speculative row can be ordinary padding - the bug appears when the first live token for a request becomes
151936 == vocab size, and that live token is then converted into-1
In simplified form, the bad transition looks like this:
text
sampled speculative row:
[151936, -1, -1, -1, ...]At this point, the trailing
-1 values are only padding. The critical problem
is that the first position holds 151936, which is out of vocabulary and is
being treated as the request's real next token.Then vLLM prepares the next-token buffer:
text
next token ids:
[-1, ...]Finally, that converted
-1 is written back into the live model input ids:text
input ids after:
[-1, 0, 0, 0, ...]The crash happens because the live next token became
-1 and was later consumed by the drafting / embedding / attention path, not merely because the speculative row contained padded -1 entries.Trigger path in code
- The workload is frontend-legal. The requests use normal
SamplingParamsfeatures such as structured outputs,stop,bad words,min tokens, and streaming overlap. No malformed token-id list is required at the request boundary. - In speculative decoding, the rejection sampler can generate recovered tokens when drafted tokens are rejected.
python
# vllm/v1/sample/rejection sampler.py
def sample recovered tokens(...):
recovered token ids = torch.empty like(draft token ids)
sample recovered tokens kernel[(batch size, max spec len)](...)
return recovered token idsOn the verified Qwen3 run, the recovered-token trace shows
recovered token ids[0] = 151936, which is exactly vocab size for this
checkpoint.
3. The speculative proposer then prepares the next-token row from the sampled
speculative row.python
# vllm/v1/spec decode/eagle.py
def prepare next token ids padded(...):
...
eagle prepare next token padded kernel[grid](
sampled token ids,
discard request mask,
backup tokens gpu,
next token ids,
valid sampled tokens count,
gpu input batch.vocab size,
...
)
return next token ids, valid sampled tokens countIn the verified trace, this step receives a sampled row beginning with
151936, followed by -1 padding. The important point is that 151936
occupies the first live token position for the request. This step then
produces next token ids[0] = -1, meaning the live next token for the
request has been converted to -1.
4. The drafter then rotates the draft input ids and inserts those
next token ids back into the live input-id buffer.python
# vllm/v1/spec decode/eagle.py
def set inputs first pass(...):
...
self.input ids[token indices to sample] = next token idsIn the verified trace, this produces
input ids after[0] = -1.
5. The model-side embed path later consumes those input ids.python
# vllm/model executor/models/qwen2.py
def embed input ids(self, input ids: torch.Tensor) -> torch.Tensor:
return self.embed tokens(input ids)In the verified trace, this is the first point where the converted
-1
becomes visible as a real model input. The bug is not merely that the
sampled speculative row contained padding -1; the bug is that the live
next token for the request became -1 and was written back into input ids.
6. After that point, the visible sink depends on timing and backend state. On
the attached Qwen3 reproducer, the engine commonly dies later in the
drafting / attention path with CUDA error: device-side assert triggered,
for example under flash attn varlen func(...).Local script breakdown
repro g4 recovered minus1 local.py is a standalone local reproducer.- It reads the Qwen3 checkpoint path from
VLLM POC G4 MODELor the built-in/path/to/qwen3placeholder - It creates
EngineCoredirectly without any external helper dependency - It submits one fixed multi-request workload that preserves the same overlap and speculative-decoding state needed for the bug
- It writes:
request payloads.jsonrepro config.jsontimeline.jsonresponses.jsonerror.txtrecovered chain trace.jsonlrecovered chain trace.jsonlis the key attribution artifact. It records the recovered-token chain directly from the standalone reproducer
gRPC script breakdown
repro g4 recovered minus1 grpc.py is a standalone public gRPC reproducer.- It reads the Qwen3 checkpoint path from
VLLM POC G4 MODELor the built-in/path/to/qwen3placeholder - It starts a temporary
vllm.entrypoints.grpc serverprocess - It sends only public
GenerateandAbortRPCs - It submits one fixed overlapping request sequence that preserves the same speculative-decoding state needed for the bug
- After the crash window, it sends one more public
Generateprobe request to confirm that later gRPC requests also fail after the worker dies - It writes:
request payloads.jsontimeline.jsonserver command.jsonresponses.jsonpost crash probe.jsonserver.stdout.logserver.stderr.log
Observed result
Local repro typically ends with:
- a recovered-token trace showing:
sample recovered tokens return -> recovered token ids[0] = 151936prepare next token ids padded -> next token ids[0] = -1set inputs first pass -> input ids after[0] = -1embed input ids out of range -> input ids[0] = -1CUDA error: device-side assert triggered- a fatal engine-side failure
gRPC repro typically ends with:
- the triggering gRPC requests failing with
INTERNAL: EngineCore encountered an issue. See stack trace (above) for the root cause. - server logs showing the worker dies with
CUDA error: device-side assert triggered - a later public probe request also failing after the worker is dead
This demonstrates that the issue is reachable through the public gRPC request surface, not only through a local reproducer.
Log snippets
Local recovered-chain trace
text
sample recovered tokens return:
recovered token ids = [151936, ...]
vocab size = 151936
prepare next token ids padded:
sampled token ids head = [[151936, -1, -1, ...], ...]
next token ids = [-1, ...]
set inputs first pass:
input ids after = [-1, 0, 0, 0, ...]
embed input ids out of range:
input ids = [-1, 0, 0, 0, ...]gRPC server log
text
torch.AcceleratorError: CUDA error: device-side assert triggered
...
vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
...
Error in Generate for request post crash probe
vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.Root cause
This is a speculative-decoding state-handling bug, not an invalid frontend token-id input bug.
The root cause is that a recovered speculative token can become equal to
vocab size, then be selected as the live next token for a request, then be converted to -1, and that converted -1 is still written back into live drafter input ids and later consumed by the drafting / embedding / attention path.For the Qwen3 checkpoint used here:
151936 == vocab size
This value should be described as the model
vocab size boundary value, not as a legal token id.Attachments
The attached bundle for this report should contain:
repro g4 recovered minus1 local.pyrepro g4 recovered minus1 grpc.py
These two standalone scripts are sufficient to reproduce the issue and its public gRPC reachability.
Fix
A fix for this vulnerability has been merged in: https://github.com/vllm-project/vllm/pull/44744
Fix
Found an issue in the description? Have something to add? Feel free to write us 👾
Related Identifiers
Affected Products
Vllm