GATED MULTI-HEAD LATENT ATTENTION

Per-token KV-cache footprint

Adjust a representative attention configuration. Standard multi-head attention caches keys and values for every head; MLA caches one compact KV latent.

Standard MHA: 2 nₕ dₕ
MLA: d_c