Posts about tutorial — from the AI Labs team.
QLoRA fits a 13B LLM on one L4 GPU by pairing 4-bit quantization with LoRA adapters. Here's exactly how it works and how to run it yourself.
Most ML models die in notebooks. Here's exactly what it takes to get one running on Google Vertex AI Pipelines — from first component to live endpoint.
Week 8 of our NLP & LLM Engineering course ends with students shipping a working RAG chatbot that answers questions over their own PDFs. Here's the full build.
We teach retrieval augmented generation with pgvector in week 7 of NLP & LLM Engineering. Here's exactly what the build looks like.
Every post on this blog comes out of a real course we teach. Browse the catalog, pick a track, and ship something that holds up in interviews.