|
Authors: Yifan Zhang, Jichen Feng, Shihan Qin PDF | arXiv | Project | GitHub | DOI
Architecture figure from the official project. Copyright 2026 Yifan Zhang; Apache License 2.0. RLT combines parallel causal encoding with token-by-token recurrent decoding. A gated connection carries the previous decoder output into the next token, extending the model’s computation path as sequences grow while keeping the number of blocks evaluated for each token fixed. The paper studies six algorithmic tasks across encoder/decoder depth splits and feedback schedules. On selected parity and permutation-tracking settings, recurrent feedback substantially improves length extrapolation over an eight-layer Transformer. The ablations also show a trade-off: updating feedback less frequently enables parallel chunk processing, but hurts tasks that require fine-grained state updates. @misc{zhang2026recurrentlooped,
title = {Recurrent Looped Transformer},
author = {Zhang, Yifan and Feng, Jichen and Qin, Shihan},
year = {2026},
eprint = {2610.07591},
archivePrefix= {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2610.07591}
}
|