CURRENT SNAPSHOT
Benchmarks tied to the implementation that produced them.
These are selected internal/research ESA measurements already surfaced on the MLBricks homepage. They are not universal performance guarantees.
AT A GLANCE
ESA vs FA4 throughput + quality
| Context | ESA throughput | FA4 throughput | Speedup | ESA PPL | FA4 PPL |
|---|---|---|---|---|---|
| 1K | 16.22M | 4.38M | 3.70× | 9.50 | 12.09 |
| 16K | 28.62M | 7.06M | 4.05× | 9.08 | 11.88 |
| 64K | 29.57M | 2.67M | 11.06× | 9.18 | 11.95 |
ESA vs FA4 peak memory
| Context | ESA peak | FA4 peak | Memory reduction |
|---|---|---|---|
| 1K | 182.4 MB | 978.1 MB | 81.4% lower |
| 16K | 454.5 MB | 12,213.4 MB | 96.3% lower |
| 64K | 1,325.3 MB | 48,166.6 MB | 97.2% lower |
A focused comparison of ESA and FA4 using the measured rows from the current H100 benchmark report. Internal Thunder C16 Compile Reduce Overhead results are presented here as ESA, while FA4 uses the native FA4 baseline.
CONTEXT 1,024
ESA vs FA4
NVIDIA H100 · training · batch 64 · FP16 · three-seed mean · parameters 555,520 · tokens trained 13.11M.
| Method / runtime | Validation loss | PPL | Training throughput | Reported peak memory |
|---|---|---|---|---|
| ESA · C16 Compile Reduce OH | 2.2502 | 9.50 | 16.22M | 182.4 MB |
| FA4 · Native | 2.4912 | 12.09 | 4.38M | 978.1 MB |
1K takeaway:ESA reaches 3.70× the FA4 training throughput and reports 81.4% lower peak memory, with PPL 9.50 versus 12.09 for FA4.
CONTEXT 16,384
ESA vs FA4
NVIDIA H100 · training · batch 64 · FP16 · three-seed mean · parameters 2,521,600 · tokens trained 209.72M.
| Method / runtime | Validation loss | PPL | Training throughput | Reported peak memory |
|---|---|---|---|---|
| ESA · C16 Compile Reduce OH | 2.2062 | 9.08 | 28.62M | 454.5 MB |
| FA4 · Native | 2.4747 | 11.88 | 7.06M | 12,213.4 MB |
16K takeaway:ESA reaches 4.05× the FA4 training throughput and reports 96.3% lower peak memory, with PPL 9.08 versus 11.88 for FA4.
CONTEXT 65,536
ESA vs FA4
NVIDIA H100 · training · batch 64 · FP16 · three-seed mean · parameters 8,813,056 · tokens trained 838.86M.
| Method / runtime | Validation loss | PPL | Training throughput | Reported peak memory |
|---|---|---|---|---|
| ESA · C16 Compile Reduce OH | 2.2166 | 9.18 | 29.57M | 1,325.3 MB |
| FA4 · Native | 2.4807 | 11.95 | 2.67M | 48,166.6 MB |
64K takeaway:ESA reaches 11.06× the FA4 training throughput and reports 97.2% lower peak memory, with PPL 9.18 versus 11.95 for FA4.
DECODER BENCHMARK
Decode throughput and recurrent-state memory
| Prompt tokens | ESA Reduce OH tok/s | Attention KV tok/s | ESA speedup | ESA state | Attention KV memory |
|---|---|---|---|---|---|
| 512 | 6,658.09 | 1,103.19 | 6.04× | 0.000488 MB | 0.750000 MB |
| 1,024 | 6,574.91 | 1,118.56 | 5.88× | 0.000488 MB | 1.250000 MB |
| 4,096 | 6,518.19 | 1,109.38 | 5.88× | 0.000488 MB | 4.250000 MB |
| 8,192 | 6,569.44 | 1,062.37 | 6.18× | 0.000488 MB | 8.250000 MB |
Decode takeaway:Across the four displayed prompt lengths (512, 1K, 4K, and 8K), ESA Compile Reduce Overhead averages6,580.16 tok/sversus1,098.38 tok/sfor native Attention KV — a5.99×ratio of average throughput. ESA recurrent state remains0.000488 MBacross the displayed prompt lengths.
INTERPRETATION
