Posted on

Large language models

A large language model (LLM) is a type of artificial intelligence algorithm that uses deep learning techniques and massively large data sets to understand, summarise, generate and predict new content. Often described as generative AI, LLMs are specifically engineered to produce text-based content by mimicking patterns and structures derived from vast amounts of existing data. Generative AI refers to AI systems capable of creating diverse types of content, including text, images, music, and videos. These systems analyse extensive datasets to recognize underlying patterns, enabling them to generate new content that shares similar characteristics with the input data.

LLMs mark a significant advancement in AI language models by using massive datasets for both training and inference, substantially enhancing their capabilities. While there is no standardised measure for the necessary size of these datasets, LLMs typically incorporate at least one billion parameters. In the context of machine learning, parameters are the variables within a model that are fine-tuned during training to facilitate the generation of new insights or content.

The emergence of modern LLMs began around 2017 with the introduction of transformer architectures—advanced neural networks known as transformers. These models use a large number of parameters and the transformer framework to quickly understand and generate accurate responses, making them highly adaptable across various industries (Hagos, Battle, & Rawat, 2024).

Developing an LLM involves several intricate steps. Watch our short slide presentation outlining the steps:

Use the dots to progress through the steps.

Once trained, the LLM serves as a robust foundation for practical applications. Users can interact with the model by providing prompts, allowing the AI to generate responses such as answers to questions, new text, summaries, or sentiment analyses.

LLMs gained widespread public attention in 2022 with the launch of ChatGPT by OpenAI. However, the development of ChatGPT began earlier with the establishment of OpenAI in December 2015 by leaders including Sam Altman and Elon Musk. The first iteration, GPT-1, introduced in June 2018, featured 117 million parameters and established the foundational architecture for subsequent models. GPT-1 showcased the effectiveness of unsupervised learning by predicting the next word in sentences derived from books, demonstrating significant potential in language understanding tasks.

GPT-2, released in February 2019, marked a substantial upgrade with 1.5 billion parameters. It exhibited a dramatic improvement in text generation capabilities, producing coherent multi-paragraph text. Due to concerns about potential misuse, GPT-2 was initially withheld from the public but was eventually released in November 2019 after a staged rollout to study and mitigate associated risks (Hagos et al., 2024).

The launch of GPT-3 in June 2020 represented a monumental leap forward with 175 billion parameters. GPT-3’s advanced text-generation abilities enabled a wide range of applications, from drafting emails and writing articles to creating poetry and generating programming code. It also demonstrated proficiency in answering factual questions and translating languages, making it a pivotal moment in the recognition and adoption of LLM technology (Hagos et al., 2024).

With GPT-3, the public began interacting directly with AI models like ChatGPT, highlighting the profound impact and potential of LLMs. This interaction demonstrated the versatility and transformative capabilities of generative AI, paving the way for future advancements and broader applications across various sectors.

Following the success of Chat GPT, many other LLMs are now becoming available.

References and Further Reading

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is All You Need. NeurIPS.

Hagos, D. H., Battle, R., & Rawat, D. B. (2024). Recent Advances in Generative AI and Large Language Models: Current Status, Challenges, and Perspectives. Journal of IEEE Transactions on Artificial Intelligence (TAI).

Where have you witnessed large language models in action?
What were you impressed by?
Were there things that didn’t impress you?

If you don’t think that you have seen large language models in action, then go to Google Gemini or Chat GPT and ask it to create a report on a topic of your choice.

Record your thoughts in your learning diary before moving on…