/

llama.cpp releases

10 stories

b10820

llama.cpp releasesToolsb10820 Github: limit blank issues to maintainers ( #28435 )

September 5
b10819

llama.cpp releasesToolsb10819 metal : fix memory leak in early return ( #28399 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/45438612 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, K

September 5
b10818

llama.cpp releasesToolsb10818 sycl : fix test-backend-ops CI break && restore Kronecker product FWHT support ( #28016 ) ( #28254 ) Reapply "sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 12…" ( #28184 ) This reverts commit c845263

September 5
b10817

llama.cpp releasesToolsb10817 sycl: attribute device allocations by site (GGML_SYCL_MEMTRACE) ( #27631 ) define two new environment variables to better understand how much memory is being allocated, and when. This has been invaluable in inproving the

September 5
b10816

llama.cpp releasesToolsb10816 metal : add remaining fa-vec tunings for M3 ( #28396 ) addition of m3 in fa_vec_tuned_table adding q4_0,q4_1,q5_0,q5_1 in ggml-metal-tuning Fix formatting in ggml-metal-tuning.cpp Website: https://llama.app Attestations:

September 4
v0.4.0

llama.cpp releasesToolsv0.4.0 Overview llama.cpp 0.4.0 adds initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support, on-demand tensor reading, per-slot server context limits, video input options, and a ggml update to 0.23.0 with major sparse flash a

September 4
b10814

llama.cpp releasesToolsb10814 opencl: extend the elementwise and data‐movement op coverage ( #27633 ) opencl: add extended elementwise unary ops (sgn, step, elu, hardswish, hardsigmoid, floor, ceil, round, trunc) Adds nine GGML_UNARY_OP_* elementwise

September 4
b10813

llama.cpp releasesToolsb10813 opencl: add Adreno xmem SDPA path ( #26331 ) opencl: add Adreno xmem SDPA path Assisted-by: Codex Removed the Adreno-specific queue profiling override Clean up formatting 修复数值误差优化gqa/mask attn Assisted-by: Codex add env

September 4
b10809

llama.cpp releasesToolsb10809 llama.cpp : bump version to 0.4.0 ( #28386 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/45314398 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiA

September 4
b10798

llama.cpp releasesToolsb10798 common : make build info output stream configurable ( #28322 ) Let llama_print_build_info write to a caller-provided FILE* instead of hardcoding stderr. The parameter defaults to stderr so existing callers keep their cur

September 4