As the book states, until transformer architecture and attention were developed natural language processing (NLP) had hit a roadblock. The state-of-the-art recurrent neural networks (RNNs) and the RNN variant, Long Short-Term Memory (LSTM) networks, could only achieve a certain level of accuracy in text processing tasks like sentence completion and summarization. The two main issues with RNNs and its variants are that input is processed sequentially and that their context window or “memory” is limited. Enter transformer architecture and attention, which ultimately led to the huge success of the NLP models we have today and their associated proprietary user products like ChatGPT. Transformer architecture and attention allowed for inputs to be processed in parallel which provided an increase in speed and efficiency when training and using these models, but also enabled models to consider the positions and features (e.g., nouns, verbs, numbers, etc) of words or subwords (i.e. “tokens”) in relation to other tokens in the input text, something we as humans know is extremely important to our understanding of language. The results are obvious. Talking to chatbots like ChatGPT can feel like talking to a real person. While transformer architecture and attention have revolutionized NLP and enabled transformative artificial intelligence (AI) tools like ChatGPT, I’ve left out a fundamental component of their success – scale. The scale of the training data used to determine the weights of the model parameters and the scale of the model parameters themselves. For example, GPT4 which ChatGPT runs on, has ~1.8 trillion parameters and was trained on approximately 13 trillion tokens. Models like these are referred to as large language models (LLMs) and are essentially trained on a data set consisting of a large chunk of the internet plus books and textbooks. Some are even trained on images and videos. I have my own reservations about the public domain of human knowledge being exploited for personal financial gain, but without vacuuming up the majority of the internet and feeding it into these foundational models, we would not be where we are today. The reason for that is not entirely clear, but it has been empirically observed that above a certain threshold of training data size and number of parameters, models become significantly more accurate and display emergent abilities – skills that the models were not specifically trained for such as following instructions and summarizing text. These abilities essentially come for free if the size of the model and its training data are large enough. Yes, models can be trained for specific tasks, which requires fewer parameters and a smaller training set than GPT4 for example, but if the goal is a generalized AI tool that can “do it all”, then you need the benefits of scale.

As I make my way through this book, I’m running all the code examples. Here’s the latest, which simply demonstrates another English to French translation using 3 different LLMs. Rather than utilizing a proprietary model via their API such as OpenAI, these models are open (although not necessarily open-source) and are available to be downloaded and run on your local machine. Note that these models only have several billion parameters compared to the 1.8 trillion of GPT4. That’s important because if the model were too large I wouldn’t be able to run it on my setup which only has 1 GPU.