Projects/ESA/Documentation

[ ESA / MLBRICKS KIT 1.0.0B1 ]

ESA Benchmarks

Focused H100 benchmark comparison of ESA and FlashAttention-4.

DISTRIBUTIONmlbricks-kitDOCS VERSION1.0.0b1

CURRENT SNAPSHOT

Benchmarks tied to the implementation that produced them.

These are selected internal/research ESA measurements already surfaced on the MLBricks homepage. They are not universal performance guarantees.

THROUGHPUT29.57Mtok/s · ESA training · 64K context
MEMORY1,325.3 MBESA reported peak · 64K context
PPL9.08best ESA perplexity · 16K context
DECODE6,580.16tok/s · ESA average over 512, 1K, 4K, and 8K

AT A GLANCE

ESA vs FA4 throughput + quality

ContextESA throughputFA4 throughputSpeedupESA PPLFA4 PPL
1K16.22M4.38M3.70×9.5012.09
16K28.62M7.06M4.05×9.0811.88
64K29.57M2.67M11.06×9.1811.95

ESA vs FA4 peak memory

ContextESA peakFA4 peakMemory reduction
1K182.4 MB978.1 MB81.4% lower
16K454.5 MB12,213.4 MB96.3% lower
64K1,325.3 MB48,166.6 MB97.2% lower

A focused comparison of ESA and FA4 using the measured rows from the current H100 benchmark report. Internal Thunder C16 Compile Reduce Overhead results are presented here as ESA, while FA4 uses the native FA4 baseline.

Method mapping:ESA= Thunder C16 Compile Reduce OH (compass 16);FA4= FA4 Native. Training results use NVIDIA H100, batch 64, FP16, and three-seed means. Throughput is million tokens per second; reported peak memory is MB.

CONTEXT 1,024

ESA vs FA4

NVIDIA H100 · training · batch 64 · FP16 · three-seed mean · parameters 555,520 · tokens trained 13.11M.

Method / runtimeValidation lossPPLTraining throughputReported peak memory
ESA · C16 Compile Reduce OH2.25029.5016.22M182.4 MB
FA4 · Native2.491212.094.38M978.1 MB

1K takeaway:ESA reaches 3.70× the FA4 training throughput and reports 81.4% lower peak memory, with PPL 9.50 versus 12.09 for FA4.

CONTEXT 16,384

ESA vs FA4

NVIDIA H100 · training · batch 64 · FP16 · three-seed mean · parameters 2,521,600 · tokens trained 209.72M.

Method / runtimeValidation lossPPLTraining throughputReported peak memory
ESA · C16 Compile Reduce OH2.20629.0828.62M454.5 MB
FA4 · Native2.474711.887.06M12,213.4 MB

16K takeaway:ESA reaches 4.05× the FA4 training throughput and reports 96.3% lower peak memory, with PPL 9.08 versus 11.88 for FA4.

CONTEXT 65,536

ESA vs FA4

NVIDIA H100 · training · batch 64 · FP16 · three-seed mean · parameters 8,813,056 · tokens trained 838.86M.

Method / runtimeValidation lossPPLTraining throughputReported peak memory
ESA · C16 Compile Reduce OH2.21669.1829.57M1,325.3 MB
FA4 · Native2.480711.952.67M48,166.6 MB

64K takeaway:ESA reaches 11.06× the FA4 training throughput and reports 97.2% lower peak memory, with PPL 9.18 versus 11.95 for FA4.

DECODER BENCHMARK

Decode throughput and recurrent-state memory

Prompt tokensESA Reduce OH tok/sAttention KV tok/sESA speedupESA stateAttention KV memory
5126,658.091,103.196.04×0.000488 MB0.750000 MB
1,0246,574.911,118.565.88×0.000488 MB1.250000 MB
4,0966,518.191,109.385.88×0.000488 MB4.250000 MB
8,1926,569.441,062.376.18×0.000488 MB8.250000 MB

Decode takeaway:Across the four displayed prompt lengths (512, 1K, 4K, and 8K), ESA Compile Reduce Overhead averages6,580.16 tok/sversus1,098.38 tok/sfor native Attention KV — a5.99×ratio of average throughput. ESA recurrent state remains0.000488 MBacross the displayed prompt lengths.

INTERPRETATION

These values describe the displayed H100 training configurations only. Performance can change with model size, dataset, context, batch, precision, backend, compile mode, software version, and hardware. “Reported peak memory” should not be interpreted as total GPU memory unless the measurement scope is defined by the benchmark procedure.