/
← Accept All   Archive
llama.cpp releasesTools

v0.4.0

September 4
v0.4.0

Overview llama.cpp 0.4.0 adds initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support, on-demand tensor reading, per-slot server context limits, video input options, and a ggml update to 0.23.0 with major sparse flash a

Models & releases
Read at llama.cpp releases ↗

Related

More from llama.cpp releases on Accept All.