cudatraining · lab notes

2026 · APR · 18 → 20 · 3 DAYS · T4 → L4

From a single line of vector_add
to 80% of cuDNN FA-2.

9CUDA / Triton kernels
4.79×Flash Attention speedup
65×HBM traffic reduction
54 TFLOPSTriton fp16 matmul
0.86×vs cuDNN FA-2
PART I · Memory
01

vector_add — what the first kernel taught

230 GB/s on T4. But the real bottleneck wasn't the kernel — it was the copy across PCIe.

LESSON 01
02

The true cost of copies — pageable vs pinned

One reason D2H is 12.7× slower. And the 5.6× end-to-end gap.

LESSON 02
PART II · Parallelism
03

Reduction — why atomic is 100× slower

Shared memory, warp shuffle, and the 29% a single __syncthreads bought.

LESSON 03
04

Matmul — from memory-bound to Tensor Cores

0.4 → 7.9 TFLOPS. A 20× roofline journey, in four layers.

LESSON 04
PART III · Fusion & Attention
05

Softmax & Fusion — the mathematical half of Flash Attention

Fuse three kernels into one and get exactly 2× speedup. The birth of online softmax.

LESSON 05
06

Flash Attention — five lessons converging into 80 lines

4.79× faster at N=4096, 65× less HBM traffic. And the line we hit.

LESSON 06 · CAPSTONE
PART IV · Integration
07

PyTorch Custom Op — one layer broken through in 50 lines

torch.ops.mylib.flash_attention. From this moment on, vLLM source becomes readable.

LESSON 07
08

Triton port — what does abstraction cost?

Jump to L4. Four kernels into Triton. Found a zone where we narrowly beat cuBLAS.

LESSON 08
09

MHA + Causal Flash Attention

4-D shape + causal + torch.compile(fullgraph=True). 80–90% of cuDNN FA-2.

LESSON 09
PART V · Profiling & Production
10

Nine kernels under the knife — nsys · ncu

Beyond "fast/slow," reading lessons 1–9 through lg_throttle, wait, math_pipe_throttle.

LESSON 10
11

Paged Attention — rewriting vLLM in Triton

contiguous → paged. A GQA grid refactor beats cuDNN FA-2 by 14 %. Axis-for-axis match with vLLM source.

LESSON 11
12

Split-K — closing the MQA SM-occupancy hole

ctx-axis partition + reduce kernel. 1.68× speedup. Remaining gap is L2 residency.

LESSON 12
PART VI · Essays
§

Triton vs CUDA — where does the cost of abstraction show up?

Can 50 lines replace 5000? The answer from four kernels.

ESSAY
§

How a 300-line Triton FA closes in on 80–90% of cuDNN

IS_CAUSAL: tl.constexpr, loop-skip, custom_op. Three tricks combined.

ESSAY
BONUS · 《The Stack》 Podcast
42

Why did your GPT end up using my GPU this way?

Sam × Jensen (fictional). Replay the 7-week path in reverse and Flash Attention looks inevitable.

EP.42 · 38 MIN
43

What Triton hides, and what it exposes

Seven files side by side, line by line. When CUDA folds into Triton, what vanishes and what remains.

EP.43 · 42 MIN