Ollama Local AI Runtimes v0.40.2-rc0: server: hide duplicate and downgrade guards from list (#18874)
Executive Summary
When a legacy GGUF which requires our llama.cpp patch set is loaded, we do a lazy conversion on disk to make it llama.cpp compatible. To support downgrades, we temporarily hold both the new and old GGUFs and create a shadow v1 manifest tag to prevent the downgraded server from deleting the new blobs. Showing these in the list output is confusing to users and API consumers. The legacy copies will be removed in a future release to recover the disk space.
Artificial Intelligence Architecture & Model Evaluation
From an artificial intelligence architecture, model weights governance, and inference efficiency perspective: - **Weights Accessibility & Sovereignty:** Evaluates whether weights are open for private self-hosting or locked behind centralized cloud APIs. - **Quantization & Edge Performance:** Kernel optimizations (4-bit/8-bit GGUF, AWQ, EXL2) allow high tokens-per-second on consumer GPUs and Apple Silicon. - **Reasoning & Architectural Scaling:** Scrutinizes mixture-of-experts (MoE), attention mechanisms, and fine-tuning datasets against open community benchmarks.
Impact on the Open Ecosystem
Protects developers and enterprises from proprietary black-box entrapment, fostering auditable, sovereign AI infrastructure.
Support Independent, Ad-Free Open Source Journalism
FOSSWire runs autonomous analysis pipelines without selling your attention to commercial advertisers.