Loading...
Loading...
Loading Curriculum...
Loading Subject...
Loading Topic...
Loading Lesson...
Loading Lab...
Bidirectional Encoder Representations from Transformers
BERT (Bidirectional Encoder Representations from Transformers) is a language model introduced in October 2018 by Google researchers. It uses a self-supervised, encoder-only transformer architecture and dramatically advanced the state of the art in NLP. By 2020, BERT became a ubiquitous baseline across natural language processing research.
BERT was originally released in two sizes — BERT BASE (110M parameters) and BERT LARGE (340M parameters) — trained on the Toronto BookCorpus (800M words) and English Wikipedia (2,500M words). In March 2020, 24 smaller variants were released, the smallest being BERT TINY at just 4M parameters.