InfrastructureFaster Than Thought: How Speculative Decoding Speeds Up AILarge language models (LLMs) like GPT are powerful, but they’re notoriously slow at inference time.July 25, 20263 min read