Sparse is Enough in Scaling Transformers (aka Terraformer) | ML Research Paper Explained
#scalingtransformers #terraformer #sparsity
Transformers keep pushing the state of the art in language and other domains, mainly due to their ability to scale to ever more parameters. However, this scaling has made it prohibitively expensive to run a lot of inference requests against a Transformer, both in terms of compute and memory requirements. Scaling Transformers are a new kind of architecture that leverage sparsity in the Transformer blocks to massively speed up inference, and by including additional ideas from other architectures, they create the Terraformer, which is both fast, accurate, and consumes very little memory.
OUTLINE:
0:00 - Intro & Overview
4:10 - Recap: Transformer stack
6:55 - Sparse Feedforward layer
19:20 - Sparse QKV Layer
43:55 - Terraformer architecture
55:05 - Experimental Results & Conclusion
Paper: https://arxiv.org/abs/2111.12763
Code: https://github.com/google/trax/blob/m...
Abstract:
Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging sparsity. We study sparse variants for all layers in the Transformer and propose Scaling Transformers, a family of next generation Transformer models that use sparse layers to scale efficiently and perform unbatched decoding much faster than the standard Transformer as we scale up the model size. Surprisingly, the sparse layers are enough to obtain the same perplexity as the standard Transformer with the same number of parameters. We also integrate with prior sparsity approaches to attention and enable fast inference on long sequences even with limited memory. This results in performance competitive to the state-of-the-art on long text summarization.
Authors: Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Łukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, Jonni Kanerva
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
-
28:21
Lloyd And Mandy
9 hours agoThe INCREDIBLE Hack Every Online Business Owner MUST KNOW In 2024..
23.7K3 -
0:51
scoutthedoggie
1 day agoWhat's in your Northeast UZI Rob?
17.5K1 -
12:49
Misha Petrov
18 hours agoI Triggered The Furries…
47K73 -
6:22:05
Akademiks
19 hours agoNicki Minaj EXPOSES Jay Z over Finessing HER! Diddy got all his Peers Hiring Lawyers & on the RUN!
141K80 -
1:04:17
WarRoom Films
10 days agoGovernment Gangsters - WarRoom Film
328K66 -
8:02
Colion Noir
20 hours agoWow, $67 Million Spent On Mandatory Gun Buy Backs, Not One Gun Collected
121K146 -
1:27:27
Man in America
1 day ago🔴 LIVE: Diddy & the Hip Hop Cabal—Sodomy, Satan & Selling Souls EXPOSED
144K375 -
14:53
Winston Marshall
5 days agoTrump Just Said THIS On X...It Will Surprise You!
147K137 -
38:37
Patriots With Grit
20 hours agoRestoring Our Republic: Taking Back Our Border | Kim Yeater, Eddie Cornell & Scotty Saks
107K24 -
51:54
TheTapeLibrary
1 day ago $5.56 earnedDisturbing Haunting of a Witches' Prison | The True Story of The Cage
88.7K18