Tag: attention
All the articles with the tag "attention".
-
从 GPT-2 到 KimiK3:七年 22580 倍的架构演进
Published: at 11:56 PM从 GPT-2 的 softmax attention 出发,历经 Linear Attention、DeltaNet、Gated DeltaNet、Kimi Linear,最终到达 KimiK3 的混合架构。七年内模型参数增长 22580 倍,但真正的变化发生在注意力机制如何存储、更新和检索信息的方式上。