Lab LedgerDesk

PyTorch · Note ·

Low Precision Flash Attention 4: End-to-End Block-Scaled Attention for Blackwell

TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes. On our internal shapes…

What moved

TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes. On our internal shapes, FA4 MX8 reaches 2.54 PF/s... PyTorch issued this as an official post on 2026-09-16. The desk files the title and summary from the allow-listed official source, not a rewrite of claims the source did not make. No second outlet is added, and no launch is invented.

Why it matters

TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes. That is a public note from PyTorch, dated 2026-09-16. Tagged LLM / Hardware.

On the record

  • Filed from the PyTorch official RSS on 2026-09-16.
  • Primary source host: pytorch.org.
  • TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes.
Primary source
pytorch.org
Desk
Logged as brief 003