Bite-Sized AI
Understand the foundations
An overview of how large language models work during inference
How text is converted into tokens and IDs before reaching the model
How token IDs are converted into vectors that are refined inside the model
How transformer layers progressively refine token representations during inference
How tokens retrieve information from other tokens in the context
How the MLP transforms each token’s hidden state after attention
How the KV cache helps an LLM avoid recomputing information during decode
Explore the ecosystelm
Where open-source models live: how to find a model, read a model card, and download what you need
Run open-source LLMs locally using Ollama and add a web UI with Open WebUI
Deploy Ollama with persistent model storage and a web UI, using an initContainer to pull models automatically
PagedAttention, continuous batching, and why a GPU is essential for production inference
Install the NVIDIA GPU Operator and serve your first model via the OpenAI-compatible API
How to make LLMs answer questions about your own data: the indexing pipeline, embeddings, vector stores, and retrieval
How LLMs can call functions and take actions: defining tools, handling responses, and feeding results back
Introduction to the protocol that allows an LLM to access external services