# b11513: CUDA: improve top-k algorithm selection (#28713)

**Published:** 2026-10-08T18:44:00.755271+00:00  
**Source:** Llama.cpp Tensor Acceleration  
**Category:** ai-local-edge  
**Canonical URL:** https://fosswire.org/news/b11513-cuda-improve-top-k-algorithm-selection-28713.html  

## Executive Summary
CUDA: radix top-k for large row counts Replaces CUB's per-row DeviceTopKKernel with a grid-over-rows radix select, gated on GGML_CUDA_TOPK_RADIX_MIN_ROWS. On qwen4exp at 34,816 tokens this cuts top-k from 1,671,253 launches / 5,761.8 ms to 2,329 / 941.8 ms. CUDA: select the TOP_K implementation by shape Replace the nrows/ncols special case with the decision boundary from #28547 (as implemented in #29278 ): bitonic for short rows, radix select for several long rows, and DeviceTopK or CUB argsort for a single long row.

## Architectural & Systems Analysis
From an artificial intelligence architecture, model weights governance, and inference efficiency perspective:

- **Weights Accessibility & Sovereignty:** Evaluates whether weights are open for private self-hosting or locked behind centralized cloud APIs.
- **Quantization & Edge Performance:** Kernel optimizations (4-bit/8-bit GGUF, AWQ, EXL2) allow high tokens-per-second on consumer GPUs and Apple Silicon.
- **Reasoning & Architectural Scaling:** Scrutinizes mixture-of-experts (MoE), attention mechanisms, and fine-tuning datasets against open community benchmarks.

## Impact on the Open Ecosystem
Protects developers and enterprises from proprietary black-box entrapment, fostering auditable, sovereign AI infrastructure.
