Demystifying Mixed Precision Training: Speeding Up Deep Learning with FP16 and BF16
As deep learning models continue to scale into billions of parameters, training them has become an incredibly resource-intensive endeavor. If you have ever trained a large convolutional neural network or a Transformer model, you have likely run into the dreaded out-of-memory (OOM) error on your GPU. Traditionally, deep learning computations are performed in Single Precision (also known as FP32), where every weight, activation, and gradient is represented by a 32-bit floating-point number. While FP32 provides high numerical precision, it consumes significant memory and bandwidth. ...