Fast Weight Attention for Continual Learning

Authors: Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao
Venue: arXiv preprint arXiv:2608.27763 (2026)

Falcon-1, Falcon-2, and Falcon-3 fast-weight update methods Overview

Fast Weight Attention studies recurrent memory updates as online learning rules under read-after-write autoregressive semantics. It introduces the Falcon family of normalized updates, spanning scalar, per-channel, and sliding-window variants, and derives recurrent, masked-parallel, and chunk-parallel implementations. The proposed methods remain competitive in language modeling while improving length extrapolation on variable-digit addition.

Citation
@misc{zhang2026fastweightattention,
  title        = {Fast Weight Attention for Continual Learning},
  author       = {Zhang, Yifan and Ta, Steve and Zhang, Jasper and Feng, Jichen and Li, Shuzhen and Zhang, Yongxin and Liu, Yifeng and Yuan, Huizhuo and Wang, Mengdi and Gu, Quanquan and Yao, Andrew Chi-Chih},
  year         = {2026},
  eprint       = {2608.27763},
  archivePrefix= {arXiv},
  primaryClass = {cs.LG},
  url          = {https://arxiv.org/abs/2608.27763}
}