Low Precision Flash Attention 4: End-to-End Block-Scaled Attention for Blackwell
TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes. On our internal shapes…
What moved
TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes. On our internal shapes, FA4 MX8 reaches 2.54 PF/s... PyTorch issued this as an official post on 2026-09-16. The desk files the title and summary from the allow-listed official source, not a rewrite of claims the source did not make. No second outlet is added, and no launch is invented.
Why it matters
TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes. That is a public note from PyTorch, dated 2026-09-16. Tagged LLM / Hardware.
On the record
- Filed from the PyTorch official RSS on 2026-09-16.
- Primary source host: pytorch.org.
- TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes.
pytorch.org
Logged as brief 003