Aaron Brooks
I design end-to-end AI systems — training infrastructure, LLM platforms, and CUDA-level optimization. Notes from production.
Recent
- Retro Engineering
- ONNX Runtime vs. TensorRT in Triton: Where the Crossover Actually Is
- Why Our ONNX Model Was 3× Slower in Triton Than in a Local Script
- A Training Run Ought to Leave Tracks
- Your CUDA Kernel Is Probably Paying a Memory Tax
- LLM Platforms Need Operational Memory
- I Was Wrong to Sleep on JAX
- Checkpointing Is a Distributed Systems Problem