Kannav Sethi
I'm an AI applications developer. I build tools and workflows around large language models, with a focus on agents, reliable software, and inference. I write about how these systems work, from attention and floating point numbers to running models on limited compute.
Currently at Locate Alpha. Previously at Seneca Polytechnic and Project Human City.
research
Do Query-Agnostic Memory Structures Earn Their Budget? A 10M-Token Audit with Exact Read LimitsAccepted at the REALM Workshop at EMNLP 2026, forthcoming
Research with Amit Maraj on whether organizing an agent's memories improves its answers. We compared memory organization approaches with direct passage search under the same reading budget.
writing
- Controlling Probabilistic Outcomes by Ensuring a Deterministic SystemJul 2026
- Local Models, Unified Memory, and Why Quantization MattersMay 2026
- A (not so big) Primer on GPU InternalsMar 2026
- Floats and Precision in LLM TrainingFeb 2026
- Normalization TechniquesDec 2025
- Rotary Positional Embeddings and its MathDec 2025
- KV Cache, MQA, GQA and Attention!!!!Nov 2025
work
- Software Engineer, Locate Alphacurrent
- AI Engineer, Acclaim Abilitycurrent
- Product Builder, Seneca Polytechnicpreviously
- Software Dev Intern, Project Human Citypreviously
projects
- Inference Engineering Companion
a reading companion to Philip Kiely's Inference Engineering, with original labs on local inference, KV caching, memory budgets, and benchmarking
- X Bookmark Engine
a local bookmark CLI combining keyword and semantic search, with local models for embeddings and tagging
- tiny-doc-cli
a TypeScript CLI that extracts readable content from PDF, DOCX, XLSX, and CSV files for agents and scripts
- DialectMorph
an AI-assisted CLI for translating code between Python, JavaScript, Java, and C++