Categories
1 page
Transformer Architecture
Tensor Parallelism II: GQA Attention, One Head Group at a Time