GPU Tooling · Inference
An Easier Way to Diagnose flash-attn Installation Problems
A practical guide to separating FlashAttention wheel, PyTorch, CUDA, platform, and GPU compatibility problems.
Archive
Personal and practical notes from my day-to-day work.
GPU Tooling · Inference
A practical guide to separating FlashAttention wheel, PyTorch, CUDA, platform, and GPU compatibility problems.
Transformers · Inference
Build a reusable prefix cache for independent Transformers requests, verify token boundaries, and check cached output against an uncached baseline.
African NLP · Tokenization
How Unicode encoding, normalization, and byte-level tokenization interact across Swahili, Yorùbá, and Amharic text.