<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>稀疏注意力 on Yiwen Cai</title><link>https://yiwen-cai.github.io/tags/%E7%A8%80%E7%96%8F%E6%B3%A8%E6%84%8F%E5%8A%9B/</link><description>Recent content in 稀疏注意力 on Yiwen Cai</description><generator>Hugo -- gohugo.io</generator><language>zh-cn</language><managingEditor>caiyiwen.cs@foxmail.com (Yiwen Cai)</managingEditor><webMaster>caiyiwen.cs@foxmail.com (Yiwen Cai)</webMaster><copyright>© 2026 Yiwen Cai</copyright><lastBuildDate>Fri, 10 Jul 2026 21:12:36 +0800</lastBuildDate><atom:link href="https://yiwen-cai.github.io/tags/%E7%A8%80%E7%96%8F%E6%B3%A8%E6%84%8F%E5%8A%9B/index.xml" rel="self" type="application/rss+xml"/><item><title>SparseSpec：加速推理模型的稀疏自推测解码</title><link>https://yiwen-cai.github.io/notes/papers/sparsec-speculative-decoding/</link><pubDate>Sat, 20 Jun 2026 00:00:00 +0000</pubDate><author>caiyiwen.cs@foxmail.com (Yiwen Cai)</author><guid>https://yiwen-cai.github.io/notes/papers/sparsec-speculative-decoding/</guid><description>针对推理语言模型长输出的 memory-bound 瓶颈，用同一模型做 self-speculative decoding——verification 阶段顺手 dump 出 attention scores 做 Top-K，作为后续 draft 的动态稀疏模式，零训练、无损、最高 2.13× 加速。</description></item></channel></rss>