trunk/1b51ed6169475ed863f4ee43b7eaa52ecda0f862: Register fused AdamW Meta kernel for tensor learning rates (#200362)
Executive Summary
Linked issue or supporting maintainer @sanketpurandare @tianyu-l @weifengpy Summary (human written only) To enable optimizer CUDA graph capture, we need to change learning rate to be a tensor. However, with fsdp2, the optimizer states are DTensor, which cause DTensor sharding propagation materializing the GPU tensors. The root cause is that there is no meta registration for optimizer with learning rate being a tensor.
Linux Kernel Subsystems & Performance Audit
From an operating systems and systems programming perspective: - **Upstream Cohesion:** Merging patches upstream prevents fragmented fork debt and guarantees binary compatibility. - **Attack Surface Reduction:** Stripping obsolete legacy subsystems strengthens real-world operational security. - **Resource Determinism:** Enhancements to scheduling and memory management minimize jitter in high-throughput workloads.
Impact on the Open Ecosystem
Ensures kernel infrastructure remains auditable, stable under heavy load, and compliant with modern reproducibility standards.