Posts
GPUMODE Matmul v2: Custom kernels only where they mattered
I left PyTorch in charge of the shapes it already handled well, and replaced only the shapes where a Triton kernel could actually win.
CUDA Training — 3 Days, 11 Parts, from vector_add to Flash Attention
A reconstructed learning log from 72 hours on Apple Silicon and GCP T4/L4 spot VMs, moving from a one-line vector_add to 80% of cuDNN Flash Attention-2 across 11 parts.
The Economy Runs on Four Currencies
From collateral chains to time-lag credit and the twelve exchange houses — a single question followed all the way through.
The AI Era, the Pissing Contests, and the Next New World
Five years from now, everyone uses AI. Machines run on APIs. Spreadsheets run on Claude. Kiosks are fully automated. Productivity? Who knows. People scroll Reels for eight hours a day. Everything feels easy, and everyone is arguing about nothing. This essay started with ...