GPU Tooling · Inference
An Easier Way to Diagnose flash-attn Installation Problems
A practical guide to separating FlashAttention wheel, PyTorch, CUDA, platform, and GPU compatibility problems.
LLM Researcher Applied Scientist
I’m an Applied Scientist at Amazon working on Alexa Personalization. Part of my work focuses on helping Alexa learn useful information from a user’s interactions and apply the right user memories at the right moment.
Before Amazon, I was a Senior Data Scientist at LexisNexis, where I worked on retrieval-augmented generation, conversational search, entity linking, and named entity recognition.
I have a Master’s degree in Computer Science from the University of Waterloo.
I also collaborate on research and side projects focused on improving NLP systems for African languages. This site is where I write about what I learn while building these systems.

Launch collection
GPU Tooling · Inference
A practical guide to separating FlashAttention wheel, PyTorch, CUDA, platform, and GPU compatibility problems.
Transformers · Inference
Build a reusable prefix cache for independent Transformers requests, verify token boundaries, and check cached output against an uncached baseline.
African NLP · Tokenization
How Unicode encoding, normalization, and byte-level tokenization interact across Swahili, Yorùbá, and Amharic text.
Research
Selected publications and public research projects spanning multilingual benchmarks, retrieval, and low-resource languages.
View selected research →About
Applied Scientist working on Alexa Personalization, with a background in retrieval, conversational search, and multilingual NLP research.
More about me →