EAGLE Feature-Level Speculative Decoding (Li 2024年)(イーグル)
2024年Li et al. (Vector Institute)発表EAGLE/EAGLE-2・Industry-leading feature-level speculative decoding LLM + Industry-leading feature-level autoregression + Industry-leading tree-based draft + Industry-leading 2.5-3.5× inference speedup。
概要
EAGLE Feature-Level Speculative Decoding は、2024年Li et al. (Vector Institute)発表EAGLE/EAGLE-2・Industry-leading feature-level speculative decoding LLM 2024年 + Industry-leading 2.5-3.5× inference speedup position確立。EAGLE specifications = Industry-leading feature-level speculative decoding LLM (Industry-leading EAGLE 2024 + Industry-leading feature-level speculative decoding LLM + Industry-leading EAGLE flagship) + Industry-leading feature-level autoregression (Industry-leading feature-level autoregression draft model + Industry-leading hidden state level prediction) + Industry-leading tree-based draft (Industry-leading tree-based draft candidate exploration) + Industry-leading 2.5-3.5× inference speedup (Industry-leading 2.5-3.5× inference speedup vs vanilla autoregressive)。
主な特徴・仕組み
- Authors: Industry-leading Yuhui Li + Fangyun Wei + Chao Zhang + Hongyang Zhang Vector Institute + Microsoft + Tsinghua + Waterloo
- Year: 2024年 (EAGLE arXiv 2024年1月 + EAGLE-2 2024年6月)
- Paper: Industry-leading "EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty" arXiv 2401.15077 + "EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees" arXiv 2406.16858
- Method Type: Industry-leading speculative decoding feature-level autoregression
- Key Innovation: Industry-leading feature-level autoregression (not token-level)
- Architecture: Industry-leading lightweight autoregression head + hidden state prediction
- Tree Draft: Industry-leading tree-based draft candidate exploration
- EAGLE-2 Innovation: Industry-leading dynamic draft trees + EAGLE-2 feature reuse
- Speedup: Industry-leading 2.5-3.5× inference speedup (EAGLE-2 reaches 4-5× in some cases)
- Production Adoption: Industry-leading vLLM + TensorRT-LLM + LMDeploy + HuggingFace
- Open Source: Industry-leading EAGLE GitHub SafeAILab/EAGLE
- Industry-Leading: Industry-leading feature-level speculative decoding LLM
スペック比較表
| LLM Speculative Decoding | Year | Method | Speedup | Industry Position |
|---|---|---|---|---|
| EAGLE/EAGLE-2 | 2024 | Feature-level autoregression + tree-based draft | 2.5-3.5× (4-5× EAGLE-2) | Industry-leading feature-level autoregression |
| Medusa | 2024 | Multiple decoding heads + tree attention | 2-3× | Industry-leading multiple heads |
| Lookahead Decoding | 2024 | N-gram pool + lookahead branches | 1.5-2× | Industry-leading parallel n-gram |
| Self-Speculative | 2023 | Same model skip layers as draft | 1.5-1.99× | Industry-leading no draft model |
| SpecExec | 2024 | Parallel speculation + GPU+CPU offload | 18.7× consumer | Industry-leading consumer GPU |
具体例・対応製品
- EAGLE (2024年, Li et al. Vector Institute): Industry-leading feature-level speculative decoding
- Industry-leading feature-level autoregression: Industry-leading feature-level autoregression
- Industry-leading tree-based draft candidate exploration: Industry-leading tree-based draft
- Industry-leading dynamic draft trees EAGLE-2: Industry-leading dynamic draft trees
- Industry-leading vLLM + TensorRT-LLM + LMDeploy + HuggingFace ecosystem: Industry-leading production ecosystem
- 競合 Medusa + Lookahead Decoding + Self-Speculative + SpecExec: Industry-leading speculative decoding competitors
自作PCでの選び方・注意点
EAGLE Feature-Level Speculative Decoding は「Industry-leading feature-level speculative decoding LLM + Industry-leading feature-level autoregression」「Industry-leading tree-based draft + Industry-leading 2.5-3.5× inference speedup」用途のIndustry-leading feature-level speculative decoding 2024年Li et al. Vector Institute発表product。Industry-leading feature-level (Industry-leading feature-level autoregression draft model + Industry-leading hidden state level prediction not token-level + Industry-leading lightweight autoregression head) で Industry-leading feature-level autoregression + Industry-leading hidden state level + Industry-leading lightweight autoregression。Industry-leading tree-based draft (Industry-leading tree-based draft candidate exploration + Industry-leading EAGLE tree-based signature + Industry-leading dynamic draft trees EAGLE-2) で Industry-leading tree-based draft + Industry-leading EAGLE tree-based + Industry-leading dynamic draft trees。Industry-leading 2.5-3.5× speedup (Industry-leading 2.5-3.5× inference speedup vs vanilla autoregressive + Industry-leading EAGLE-2 reaches 4-5× in some cases + Industry-leading EAGLE speedup advantage) で Industry-leading 2.5-3.5× speedup + Industry-leading EAGLE-2 4-5× + Industry-leading EAGLE advantage。Industry-leading production ecosystem (Industry-leading vLLM + TensorRT-LLM + LMDeploy + HuggingFace + Industry-leading EAGLE GitHub SafeAILab/EAGLE open source) で Industry-leading vLLM + TensorRT-LLM + LMDeploy + HuggingFace + Industry-leading SafeAILab open source。但しIndustry-leading Medusa + Lookahead + Self-Speculative + SpecExec competition (Industry-leading Medusa multiple heads 2-3× 2024 + Lookahead n-gram 1.5-2× 2024 + Self-Speculative skip layers 1.99× 2023 + SpecExec 18.7× consumer GPU 2024 vs EAGLE feature-level 2.5-3.5× 2024 trade-off) で Industry-leading 4-speculative decoding method competitors vs EAGLE + Industry-leading feature-level autoregression + tree-based draft + dynamic draft trees + EAGLE-2 + 2.5-3.5× speedup unique advantage adoption alignment必須。
関連用語との違い
- vs Medusa (Multiple Heads 2024): EAGLEはfeature-level autoregression + tree-based draft・Medusaはmultiple decoding heads + tree attention
- vs Lookahead Decoding (N-gram 2024): EAGLEはfeature-level + tree draft + 2.5-3.5×・Lookaheadはn-gram pool + 1.5-2×
- vs Self-Speculative (2023): EAGLEはfeature-level + 2.5-3.5×・Self-Speculativeは same model skip layers + 1.99×
よくある質問(FAQ)
Q1: EAGLE vs Medusa 違いは? A: EAGLE (Industry-leading feature-level autoregression draft model + hidden state level prediction not token-level + lightweight autoregression head + tree-based draft + dynamic draft trees EAGLE-2 + 2.5-3.5× inference speedup + Vector Institute+Microsoft 2024) vs Medusa (Industry-leading multiple lightweight decoding heads + 4-5 heads parallel prediction + tree-based attention candidate verification + no separate draft model required + 2-3× inference speedup + Princeton+Together AI 2024)・Industry-leading feature-level + tree draft + EAGLE-2 + 2.5-3.5× = EAGLE + Industry-leading multiple heads + tree attention + no draft + 2-3× = Medusa preference judgment。
Q2: Industry-leading feature-level autoregression + tree-based draft value は? A: Industry-leading feature-level + tree (Industry-leading feature-level autoregression draft model + Industry-leading hidden state level prediction not token-level + Industry-leading lightweight autoregression head + Industry-leading tree-based draft candidate exploration + Industry-leading dynamic draft trees EAGLE-2)。
Q3: Industry-leading 2.5-3.5× speedup + EAGLE-2 4-5× value は? A: Industry-leading 2.5-3.5× + EAGLE-2 (Industry-leading 2.5-3.5× inference speedup vs vanilla autoregressive + Industry-leading EAGLE-2 reaches 4-5× in some cases + Industry-leading dynamic draft trees EAGLE-2 feature reuse + Industry-leading EAGLE speedup advantage)。
まとめ
EAGLE = 2024年Li et al. Vector Institute+Microsoft発表のfeature-level speculative decoding LLM。Industry-leading feature-level autoregression draft model + Industry-leading hidden state level prediction not token-level + Industry-leading lightweight autoregression head + Industry-leading tree-based draft candidate exploration + Industry-leading dynamic draft trees EAGLE-2 + Industry-leading EAGLE-2 feature reuse + Industry-leading 2.5-3.5× inference speedup vs vanilla autoregressive + Industry-leading EAGLE-2 reaches 4-5× in some cases + Industry-leading vLLM + TensorRT-LLM + LMDeploy + HuggingFace production ecosystem + Industry-leading EAGLE GitHub SafeAILab/EAGLE open source + Industry-leading feature-level speculative decoding LLM position確立。