Skip to content

Archives

All the articles I've archived.

2026 4
September 2
  • Optimization of the Rust LLM engine

    Published: at 10:00 AM

    rvllm generated 0.9 tokens per second from a 360M model. Profiling it showed almost all the time went to memcpy, not math.

  • rvllm - A Small LLM Inference Engine in Rust

    Published: at 10:00 AM

    A basic local LLM inference engine written from scratch in Rust, running SmolLM2 on CPU with continuous batching and paged KV cache, and where CUDA and Metal support fit in next.

July 1
June 1
  • Building a KV Cache Block Scheduler in Rust

    Published: at 10:00 AM

    A from-scratch PagedAttention-style KV cache block manager in Rust - reference counting, prefix caching via radix trie, LRU eviction, and copy-on-write for beam search.

2025 6
October 2
September 1
June 1
  • Data capture for ML endpoints

    Published: at 04:01 PM

    One of the approaches of how how you can add data capture to your ML endpoints

January 2
2024 1
November 1