Tag: rust
All the articles with the tag "rust".
Optimization of the Rust LLM engine
Published: at 10:00 AMrvllm generated 0.9 tokens per second from a 360M model. Profiling it showed almost all the time went to memcpy, not math.
rvllm - A Small LLM Inference Engine in Rust
Published: at 10:00 AMA basic local LLM inference engine written from scratch in Rust, running SmolLM2 on CPU with continuous batching and paged KV cache, and where CUDA and Metal support fit in next.
Building a KV Cache Block Scheduler in Rust
Published: at 10:00 AMA from-scratch PagedAttention-style KV cache block manager in Rust - reference counting, prefix caching via radix trie, LRU eviction, and copy-on-write for beam search.