Breaking the Exactness Barrier: Interleaved DeepSeek Sparse Attention for Efficient Long Context Reasoning
Yifan Guo, Wei Cui
Blog
NeurIPS 2026
TL;DR: Top-k selection in sparse attention is difficult to optimize in distributed systems because it introduces global dependencies. We interleave the context across GPUs and perform top-k selection locally on each GPU, eliminating the global dependency.
SEDG: Stitch-Compatible End-to-End Layout Decomposition Based on Graph Neural Network
Yifan Guo, Jiawei Chen, Yexin Li, Yunxiang Zhang, Qing Zhang, Yuhang Zhang, Yongfu Li
Paper
DATE 2025
TL;DR: Layout decomposition is an important step in chip manufacturing. We explore a more efficient decomposition method using heterogeneous graph neural networks.