计算装置、计算设备以及用于线程组累加的方法
By decoupling the accumulation operation within the thread group to dedicated hardware for parallel processing, the problem of high instruction overhead in the accumulation operation within the thread group is solved, achieving more efficient accumulation performance and computational efficiency.
CN112817735BActive Publication Date: 2026-07-17SHANGHAI BIREN TECH CO LTD
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI BIREN TECH CO LTD
- Filing Date
- 2021-03-08
- Publication Date
- 2026-07-17
AI Technical Summary
Technical Problem
In existing technologies, the accumulation operation within a thread group requires a large amount of instruction overhead, resulting in low efficiency.
Method used
The accumulation operation within the thread group is decoupled to dedicated hardware for processing. It is then processed in parallel by the accumulation calculation unit and storage unit in the computing device to generate and store the current accumulation result.
Benefits of technology
It significantly improves the accumulation performance within the thread group, reduces hardware costs and the number of operations on storage units, and improves computational efficiency.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN112817735B_ABST
Abstract
本公开的实施例涉及计算装置、计算设备以及用于线程组累加的方法,涉及计算机领域。计算装置包括:存储单元;以及累加计算单元,与存储单元相耦接,被配置为:从与计算装置相耦接的向量处理单元接收第一线程组累加指令、与线程组通道数相对应的多个第一值和第一存储地址;响应于第一线程组累加指令,基于多个第一值生成当前累加结果;以及在存储单元中的第一存储地址中存储当前累加结果,以用于向量处理单元读取。由此,能够将线程组内的累加解耦到专用硬件进行处理,从而显著提升整体累加性能。
Need to check novelty before this filing date? Find Prior Art