Speculative Decoding Under Real Serving Load
Speculative decoding promises faster text generation. This preprint measures what happens once real traffic, batching and memory limits enter the picture.
In one paragraph
A preprint measures speculative decoding on a production-style serving stack with mixed request lengths and continuous batching. Speedups of more than 2 times in isolated tests shrank to 1.1 to 1.4 times under load. The authors trace the gap to draft acceptance rates and verification cost at larger batch sizes. This is a lab experiment on one hardware configuration and has not been peer reviewed.
Briefing reviewed by an editor before publication. The paper belongs to its authors; this page summarises it and links to the original. Reviewed by Dr. V. Santoro. Paper licence: CC BY 4.0. Corrections: support@papersays.com.