Demystifying Mixed Precision Training: Speeding Up Deep Learning with FP16 and BF16

As deep learning models continue to scale into billions of parameters, training them has become an incredibly resource-intensive endeavor. If you have ever trained a large convolutional neural network or a Transformer model, you have likely run into the dreaded out-of-memory (OOM) error on your GPU. Traditionally, deep learning computations are performed in Single Precision (also known as FP32), where every weight, activation, and gradient is represented by a 32-bit floating-point number. While FP32 provides high numerical precision, it consumes significant memory and bandwidth. ...

July 27, 2026 · 9 min · Pranav Buradkar

A Simple Guide to LLM Serving: Quantization, KV Caching, and Continuous Batching

When you type a prompt into ChatGPT or Claude, the model generates a response word-by-word (or token-by-token) in real-time. Behind this smooth user interface lies a massive engineering challenge: LLM Inference is incredibly expensive and slow. If you run a base LLM without optimizations, it will devour your GPU memory, process requests one-by-one, and make your users wait. To solve this, developers and researchers use three core serving techniques that work in tandem to speed up generation by up to 10x while drastically cutting hardware costs. ...

May 10, 2026 · 5 min · Pranav Buradkar

A Beginner's Guide to LLM Quantization: GGUF, GPTQ, and AWQ

Expanding on LLM optimization by explaining how massive models are compressed using techniques like GGUF, GPTQ, and AWQ to run on consumer hardware.

November 20, 2025 · 5 min · Pranav Buradkar