Demystifying Mixed Precision Training: Speeding Up Deep Learning with FP16 and BF16

As deep learning models continue to scale into billions of parameters, training them has become an incredibly resource-intensive endeavor. If you have ever trained a large convolutional neural network or a Transformer model, you have likely run into the dreaded out-of-memory (OOM) error on your GPU. Traditionally, deep learning computations are performed in Single Precision (also known as FP32), where every weight, activation, and gradient is represented by a 32-bit floating-point number. While FP32 provides high numerical precision, it consumes significant memory and bandwidth. ...

July 27, 2026 · 9 min · Pranav Buradkar

A Simple Guide to LLM Serving: Quantization, KV Caching, and Continuous Batching

When you type a prompt into ChatGPT or Claude, the model generates a response word-by-word (or token-by-token) in real-time. Behind this smooth user interface lies a massive engineering challenge: LLM Inference is incredibly expensive and slow. If you run a base LLM without optimizations, it will devour your GPU memory, process requests one-by-one, and make your users wait. To solve this, developers and researchers use three core serving techniques that work in tandem to speed up generation by up to 10x while drastically cutting hardware costs. ...

May 10, 2026 · 5 min · Pranav Buradkar

How to Fine-Tune Llama 3 on Your Own Data: A Practical Guide

In our Simple Guide to LoRA, we explored the theory behind Low-Rank Adaptation and how it slashes the GPU memory requirements for model training. Now, let’s put theory into practice. In this guide, we’ll walk through a step-by-step, hands-on tutorial to fine-tune Meta’s Llama 3 (8B) on a custom dataset using Hugging Face, PEFT, and QLoRA (Quantized LoRA) on a single GPU. Step 1: Format Your Dataset To train Llama 3, you need prompt-response pairs. Llama 3 uses a specific chat template format. For custom datasets, the easiest approach is to structure your data as a JSON file containing lists of messages: ...

April 12, 2026 · 4 min · Pranav Buradkar

A Simple Guide to LoRA (Low-Rank Adaptation)

Have you ever tried to download or fine-tune a modern Large Language Model (LLM) like Llama 3 or Mistral? If so, you probably ran into a massive wall: GPU memory. Fine-tuning models with billions of parameters requires specialized, high-end hardware, costing thousands of dollars. But what if you could achieve the exact same performance by training less than 1% of the model’s parameters? That is the magic of LoRA (Low-Rank Adaptation). It is currently the most popular technique for making fine-tuning fast, cheap, and accessible to everyone. In this guide, we’ll break down how it works, why it is so effective, and how you can use it. ...

March 20, 2026 · 4 min · Pranav Buradkar

Deep Learning Fundamentals: Attention Mechanisms and Vision Transformers

Breaking down the mechanics of Self-Attention and understanding how the Transformer architecture leaped from Natural Language Processing to dominating Computer Vision with ViTs.

December 15, 2025 · 4 min · Pranav Buradkar

A Simple Guide to Variational Autoencoders (VAEs)

Have you ever wondered how AI can create new images, music, or even text that feels original? One powerful tool in the generative AI toolbox is the Variational Autoencoder (VAE). While GANs (Generative Adversarial Networks) pit two networks against each other, VAEs take a different approach, focusing on learning the underlying structure of data. In this guide, we’ll break down VAEs in simple terms, explore how they work, and see their real-world applications. ...

November 30, 2025 · 3 min · Pranav Buradkar

A Beginner's Guide to LLM Quantization: GGUF, GPTQ, and AWQ

Expanding on LLM optimization by explaining how massive models are compressed using techniques like GGUF, GPTQ, and AWQ to run on consumer hardware.

November 20, 2025 · 5 min · Pranav Buradkar

A Simple Guide to GANs

Have you ever seen a photo of a person who doesn’t exist? Or a painting so unique you can’t believe a human didn’t make it? Chances are, you were looking at the work of a GAN, one of the most creative and fascinating ideas in modern artificial intelligence. But what is a GAN? The name, Generative Adversarial Network, sounds incredibly complex. In reality, the core idea is a brilliant and surprisingly simple story of competition. ...

August 21, 2025 · 5 min · Pranav Buradkar