Sitemap

Beyond BERT

9 min readNov 29, 2020

--

Press enter or click to view image in full size
Alammar, Jay (2018). The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning) [Blog post]. Retrieved from http://jalammar.github.io/illustrated-bert/
Press enter or click to view image in full size
From ‘RoBERTa: A Robustly Optimized BERT Pretraining Approach’ by Liu et al.
Press enter or click to view image in full size
From ‘BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension’ by Lewis et al.
Press enter or click to view image in full size
From ‘BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension’ by Lewis et al.
Press enter or click to view image in full size
From ‘DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter’ by Sanh et al.
Press enter or click to view image in full size
Distilling the knowledge from a teacher to a student. Source: Neural Network Distiller 2019
The softmax function with temperature
From ‘DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter’ by Sanh et al.

--

--