一种面向大语言模型稀疏推理的计算与存储方法及系统
By partitioning the sparse activation matrix data and optimizing the dynamic scheduling impedance, the problems of computational load imbalance and memory congestion in sparse inference of large language models are solved, thereby improving the performance of parallel inference.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies lack a dynamic matching mechanism between sparse data topology and hardware physical state in large language model sparse inference, leading to unbalanced computational load and memory access congestion, which affects parallel inference performance.
By acquiring sparse activation matrix data, dividing it into micro data blocks according to the block step size parameter, extracting the spatial span of adjacent non-zero elements, and combining the idle resource quantity of thread bundles and the computational resource demand, dynamically scheduling impedance, generating the optimal thread allocation quantity, and optimizing task allocation.
It achieves dynamic balanced allocation of computing power, eliminates thread divergence and memory access congestion, and improves the parallel inference performance of large language models.
Smart Images

Figure CN122195684B_ABST