规约请求执行方法、缓存、计算设备及计算系统
By implementing a method that reads data from memory and storage buffers into cache lines in parallel, the problem of insufficient resources for last-level cache reduction operations is solved, improving the processing efficiency of reduction requests and the utilization rate of cache resources, thus meeting the needs of high-concurrency, low-latency AI computing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI BIREN TECH CO LTD
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-17
AI Technical Summary
In existing technologies, the reduction operation of the last-level cache is constrained by physical area, resulting in insufficient dedicated temporary storage resources. This makes it impossible to adapt to the multi-core parallel, high-throughput, and low-latency computing requirements of artificial intelligence chips. Consequently, the reduction operator has low execution efficiency, reduced chip computing power utilization, increased power consumption, and increased latency, making it difficult to meet the performance requirements of large-scale, high-density AI computing scenarios.
By executing the steps of reading the first source operand of the target reduction request from memory and the second source operand of the target reduction request from the storage buffer in parallel and writing them into the target cache line, the two core data preparation steps are synchronized, shortening the overall data preparation time.
It improves the processing efficiency of protocol requests, reduces data preparation time, increases cache resource utilization, reduces latency, and enhances the application advantages of computing devices in high-concurrency, low-latency computing scenarios.
Smart Images

Figure CN122019419B_ABST