llamini.cpp is a from-scratch implementation of the llama.cpp inference core written in pure C99 (POSIX + libc) with zero third-party dependencies. One command (make) builds the binary, and it supports TinyLlama 1.1B, Qwen2.5 0.5B, GPT-2 124M, and Gemma 2B with Q4_K, Q6_K, and Q5 quantization formats.
The project started as an exercise in understanding how modern LLM inference actually works under the hood — memory mapping, KV cache management, attention computation, and quantized weight loading — without leaning on existing abstractions.
Building inference from scratch forces you to confront every design decision: how tensors are laid out in memory, how RoPE is applied, how the sampling loop interacts with the token buffer. The result is a minimal but functional inference engine that runs real models on modest hardware.