规约请求执行方法、缓存、计算设备及计算系统

By implementing a method that reads data from memory and storage buffers into cache lines in parallel, the problem of insufficient resources for last-level cache reduction operations is solved, improving the processing efficiency of reduction requests and the utilization rate of cache resources, thus meeting the needs of high-concurrency, low-latency AI computing.

CN122019419BActive Publication Date: 2026-07-17SHANGHAI BIREN TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI BIREN TECH CO LTD
Filing Date
2026-04-08
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, the reduction operation of the last-level cache is constrained by physical area, resulting in insufficient dedicated temporary storage resources. This makes it impossible to adapt to the multi-core parallel, high-throughput, and low-latency computing requirements of artificial intelligence chips. Consequently, the reduction operator has low execution efficiency, reduced chip computing power utilization, increased power consumption, and increased latency, making it difficult to meet the performance requirements of large-scale, high-density AI computing scenarios.

Method used

By executing the steps of reading the first source operand of the target reduction request from memory and the second source operand of the target reduction request from the storage buffer in parallel and writing them into the target cache line, the two core data preparation steps are synchronized, shortening the overall data preparation time.

Benefits of technology

It improves the processing efficiency of protocol requests, reduces data preparation time, increases cache resource utilization, reduces latency, and enhances the application advantages of computing devices in high-concurrency, low-latency computing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019419B_ABST
    Figure CN122019419B_ABST
Patent Text Reader

Abstract

本公开涉及一种规约请求执行方法、缓存、计算设备及计算系统,该方法应用于缓存,包括:接收目标计算单元发送的目标规约请求;在目标规约请求的第一源操作数未命中缓存的情况下,为目标规约请求分配对应的目标缓存行;并行执行从内存中读取第一源操作数的第一步骤,及从存储缓冲区中与目标规约请求对应的目标缓存区域读取目标规约请求的第二源操作数、并将第二源操作数写入目标缓存行的第二步骤;在接收到第一源操作数的情况下,对第一源操作数和第二源操作数执行与目标规约请求相匹配的规约运算,得到目标规约请求的规约运算结果;将规约运算结果返回目标计算单元。这样,可以实现规约请求的两个源操作数的准备时间,提升规约请求的处理效率。
Need to check novelty before this filing date? Find Prior Art