Accumulating memory, processor, device and data processing method

CN122822007APending Publication Date: 2026-09-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510353658.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,相关技术存在过多访问存储器的问题,存储器的过多访问容易带来较大的功耗开销,这不利于累加存储器的性能提升

Benefits of technology

[0015]所述数据处理单元在所述第一存储单元中存储有所述第二操作数的情况下,从所述第一存储单元中读取所述第二操作数;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822007A_ABST
    Figure CN122822007A_ABST
Patent Text Reader

Abstract

An accumulation memory, a processor, a device and a data processing method, relate to the technical field of computers. The accumulation memory comprises a data processing unit, a first storage unit and a second storage unit, the first storage unit is used for assisting the second storage unit to store an accumulation number; the data processing unit is used for receiving a data accumulation request, the data accumulation request is used for requesting to store an accumulation number of a first operand and a second operand in the accumulation memory; in the case that the second operand is stored in the first storage unit, reading the second operand from the first storage unit; obtaining the accumulation number according to the first operand and the second operand; and storing the accumulation number in the first storage unit. By storing the accumulation number in the first storage unit in priority and reading the second operand for calculating the accumulation number from the first storage unit, the number of access to the second storage unit can be reduced, thereby reducing the power consumption overhead caused by accessing the second storage unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an accumulation memory, processor, device, and data processing method. Background Technology

[0002] Accumulator memory is widely used in processors to store operands and results of logical operations. For example, when a processor performs operations such as convolution or matrix multiplication, it needs to continuously accumulate intermediate results until the final accumulated result is obtained. This accumulation process is usually performed by accumulator memory.

[0003] In related technologies, accumulators use memory to store the accumulated result, and each addition operation involves two memory accesses: a read access to the operand and a write access to the accumulated result. However, these technologies suffer from excessive memory access, which can lead to significant power consumption and hinder performance improvements. Summary of the Invention

[0004] This application provides an accumulation memory, a processor, a device, and a data processing method. The technical solutions provided by this application include the following:

[0005] According to one aspect of the embodiments of this application, an accumulation memory is provided, the accumulation memory comprising: a data processing unit, a first storage unit, and a second storage unit, wherein the first storage unit is used to assist the second storage unit in storing the accumulated number; the data processing unit is used for:

[0006] A data accumulation request is received, the data accumulation request being used to request that the sum of the first operand and the second operand in the accumulation memory be stored in the accumulation memory;

[0007] If the second operand is stored in the first storage unit, the second operand is read from the first storage unit;

[0008] The accumulated number is obtained based on the first operand and the second operand;

[0009] The accumulated number is stored in the first storage unit.

[0010] According to one aspect of the embodiments of this application, a processor is provided, the processor including the accumulation memory as described above.

[0011] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor, the processor including the accumulated memory as described above.

[0012] According to one aspect of the embodiments of this application, a server cluster is provided, the server cluster including at least two servers, at least one of the at least two servers including the accumulation memory as described above.

[0013] According to one aspect of the embodiments of this application, a data processing method in an accumulation memory is provided, the accumulation memory comprising: a data processing unit, a first storage unit, and a second storage unit, wherein the first storage unit is used to assist the second storage unit in storing accumulated numbers; the method includes:

[0014] The data processing unit receives a data accumulation request, which requests that the sum of the first operand and the second operand in the accumulation memory be stored in the accumulation memory.

[0015] When the data processing unit has stored the second operand in the first storage unit, it reads the second operand from the first storage unit.

[0016] The data processing unit obtains the accumulated number based on the first operand and the second operand;

[0017] The data processing unit stores the accumulated number into the first storage unit.

[0018] The technical solutions provided in this application embodiment may include the following beneficial effects.

[0019] By configuring a first storage unit and a second storage unit in the accumulator memory, and prioritizing the storage of the accumulated number in the first storage unit, the first storage unit assists the second storage unit in storing the accumulated number, which helps reduce the storage pressure on the second storage unit. Furthermore, by prioritizing the reading of the second operand used to calculate the accumulated number from the first storage unit, the number of accesses to the second storage unit can be reduced, thereby reducing the power consumption overhead associated with accessing the second storage unit and ultimately improving the performance of the accumulator memory. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the accumulator memory provided by related technologies;

[0022] Figure 2This is a schematic diagram of an accumulation memory provided in one embodiment of this application;

[0023] Figure 3 This is a schematic diagram of a first storage unit provided in one embodiment of this application;

[0024] Figure 4 This is a schematic diagram of an accumulation memory provided in another embodiment of this application;

[0025] Figure 5 This is a schematic diagram of a data processing method in an accumulator memory provided in one embodiment of this application;

[0026] Figure 6 This is a schematic diagram of a data processing method in an accumulator memory provided in another embodiment of this application;

[0027] Figure 7 This is a flowchart of a data processing method in an accumulator memory provided in one embodiment of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0029] Accumulator memory is a special type of memory used to provide a working area for the Arithmetic and Logic Unit (ALU) to perform arithmetic or logical operations and to temporarily store the results of the operations. This avoids frequent read and write operations to the main memory, thereby improving operational efficiency.

[0030] For example, an accumulation memory can be used to store an operand and the result of an arithmetic or logical operation. For instance, in the accumulation process involved in an arithmetic or logical operation, the accumulation memory can provide an operand for performing the accumulation operation, and can be used to store the accumulation result. The aforementioned operation can include at least one of the following: accumulation operation, convolution operation, and matrix multiplication operation.

[0031] In this context, the accumulation operation refers to the process of adding a series of operands one by one to obtain a sum (i.e., the accumulated result). The accumulated number is the number obtained by successively adding the operands. The sum corresponding to each accumulation operation can be called the accumulated number or the accumulated result. The accumulation operation includes at least one addition operation, that is, the series of operands includes at least two operands. The addition operation is the process of performing addition operations.

[0032] In the embodiments of this application, the result of the addition operation involved in the accumulation process can also be referred to as the accumulated number. For example, for the numerical sequence: 1, 2, 3, and 4, the accumulation process includes two addition operations: 1+2=3, 3+4=7, and both 3 and 7 can be referred to as the accumulated number. Further, 3 can be referred to as the historical accumulated number (i.e., the intermediate result), and 7 can be referred to as the final accumulated number.

[0033] Operands are data used in computer programs to perform operations. For example, an operand is the entity on which an operator acts; it is a component of an expression that specifies the numerical values ​​to be operated on. An expression is a combination of operands and operators.

[0034] Accumulator memory has a wide range of applications in processors, such as being applicable to at least one of the following processors: Central Processing Unit (CPU), Graphics Processing Unit (GPU), General-Purpose Computing on Graphics Processing Units (GPGPU), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Tensor Processing Unit (TPU), and Field Programmable Gate Array (FPGA).

[0035] Please refer to Figure 1 The diagram illustrates an accumulation memory provided by related technologies. The accumulation memory 100 includes a data reading unit 101, a data processing unit 102, a data writing unit 103, and a data storage unit 104. The data storage unit 104 is implemented as a memory used to store historical accumulation numbers (i.e., historical accumulation results or intermediate results), which can be used as operands in subsequent accumulation processes. Upon completion of the operation, the historical accumulation numbers in the data storage unit 104 can be used as the final accumulation number.

[0036] For each addition operation in the accumulation process, the operation of the accumulation register 100 can be described as follows:

[0037] 1. The data reading unit 101 obtains a data accumulation request from the arithmetic logic unit. The data accumulation request includes a first operand and a first storage address. The first storage address is the storage address of the second operand in the accumulation register 100, such as the storage address of the second operand in the data storage unit 104.

[0038] The first operand is an operand to be added. For example, the first operand can be the result of a calculation by an arithmetic logic unit, such as the result of a multiply-adder.

[0039] 2. The data reading unit 101 reads the corresponding historical accumulated number from the data storage unit 104 according to the first storage address, and uses it as the second operand.

[0040] 3. The data reading unit 101 sends the first operand and the second operand to the data processing unit 102.

[0041] 4. The data processing unit 102 adds the first operand and the second operand to obtain the accumulated number, and sends the accumulated number to the data writing unit 103.

[0042] 5. The data writing unit 103 stores the accumulated number into the data storage unit 104 according to the first storage address to replace the second operand.

[0043] It is evident that to complete an addition operation, the relevant technology requires two accesses to the data storage unit 104 (i.e., memory), namely, reading the second operand and storing the accumulated number, which can easily lead to significant power consumption.

[0044] Furthermore, the memory is typically single-port RAM (Random Access Memory), meaning only one read / write request is allowed at any given time. Therefore, each addition operation requires two clock cycles of memory access. If a memory access conflict occurs during this process, such as two requests accessing the same memory block simultaneously, the two requests must be controlled to access the memory sequentially, with one request being delayed. Frequent memory accesses increase the probability of memory access conflicts, which is detrimental to improving the performance of the accumulation memory.

[0045] To address the problems existing in related technologies, this application embodiment adds a storage unit to the accumulator memory to assist the memory in storing accumulated numbers. This storage unit is used to cache recently accessed accumulated numbers. When a data accumulation request is received, the accumulator memory first reads the second operand from this storage unit. If the second operand is not found in the storage unit, it then reads the second operand from the memory. Furthermore, the accumulator memory prioritizes storing accumulated numbers in this storage unit and transfers the least recently accessed accumulated number from this storage unit to the memory. This significantly reduces the number of memory accesses, thereby reducing the power consumption overhead associated with memory access and improving the performance of the accumulator memory. Additionally, the reduced number of memory accesses helps decrease the probability of memory access conflicts, further improving the access performance of the accumulator memory.

[0046] Please refer to Figure 2 The diagram illustrates an accumulation memory provided in one embodiment of this application. The accumulation memory 200 includes a data processing unit 201, a first storage unit 202, and a second storage unit 203.

[0047] The first storage unit 202 is used to assist the second storage unit 203 in storing the accumulated number. Optionally, the first storage unit 202 is used to cache the recently accessed accumulated number, such as the recently calculated accumulated number.

[0048] For example, the accumulator memory 200 preferentially reads the second operand from the first storage cell 202. If the second operand is not present in the first storage cell 202, it then reads the second operand from the second storage cell 203. Furthermore, the accumulator memory 200 preferentially stores the accumulated number in the first storage cell 202 and transfers the least recently accessed accumulated number from the first storage cell 202 to the second storage cell 203. This significantly reduces the number of accesses to the second storage cell 203, thereby reducing the power consumption overhead associated with accessing the second storage cell 203 and improving the performance of the accumulator memory 200. Additionally, the reduced number of accesses to the second storage cell 203 helps decrease the probability of access conflicts, further improving the access performance of the accumulator memory 200.

[0049] In one example, the power consumption of accessing the first memory cell 202 is less than that of accessing the second memory cell 203, which helps to further reduce the overall power consumption of the accumulator memory 200.

[0050] For example, the first storage unit 202 is constructed based on registers, and the second storage unit 203 is constructed based on memory. That is, the second storage unit 203 is an existing storage unit in the accumulator memory 200, and the first storage unit 202 is a newly added storage unit in the accumulator memory.

[0051] Registers and memory are two different types of storage units in a computer system. For example, memory can be storage devices such as RAM and hard disks. Registers can be a high-speed storage unit within the processor. Compared to memory, registers have faster access speeds and lower power consumption. In this embodiment, by setting the first storage unit 202 as a register, it is beneficial to reduce the overall power consumption of the accumulation memory 200 and to improve the access efficiency of the accumulation memory 200.

[0052] In one example, the first storage unit 202 includes at least one entry, which includes a first field segment, a second field segment, and a third field segment. The first field segment is used to store data, the second field segment stores the storage address of the data stored in the entry, and the third field segment stores valid identification information for indicating whether data is stored in the entry.

[0053] An entry is a data item or record stored in computer memory. For example, each entry includes information from three fields. A field is a memory segment, such as a memory segment in a register, used to store information. Each entry corresponds to an accumulator.

[0054] Optionally, at least one entry is distributed in a table format, with each row corresponding to one entry. For each entry, the information stored in the first field, second field, and third field are arranged sequentially. In this embodiment, the first field can be used to store the accumulated number, and the second field can be used to store the storage address of the accumulated number in the accumulation memory 200, such as the storage address of the accumulated number in the second storage unit 203.

[0055] Optionally, for any entry, if the entry stores an accumulated number, the valid identification information stored in the entry can be the first parameter; if the entry does not store an accumulated number, the valid identification information stored in the entry can be the second parameter. The first and second parameters can be set and adjusted according to actual usage requirements, and this application embodiment does not limit this.

[0056] For example, refer to Figure 3 The first storage unit 202 includes multiple entries, which are distributed in a table style. Each entry corresponds to three domain segments: the first domain segment 2021, the second domain segment 2022, and the third domain segment 2023.

[0057] In a feasible example, each entry may consist of only the first and second field segments. The accumulator 200 can retrieve data by traversing the information stored in the second field segment.

[0058] This application embodiment stores the accumulated numbers in an orderly manner in the first storage unit 202 in the form of entries, which facilitates the reading or writing of the accumulated numbers.

[0059] In one example, the second storage unit 203 includes multiple memory banks, each corresponding to a block of memory within the second unit 203. Since multiple requests can access different banks in parallel, this allows multiple requests to access the second storage unit 203 as simultaneously as possible, thus improving the access efficiency of the second storage unit 203.

[0060] The aforementioned data processing unit 201 is responsible for data processing, such as request processing, data reading, data calculation, and data storage. This data processing unit 201 can be implemented as a logic circuit for processing data.

[0061] In one example, for each addition operation in the accumulation process, the data processing unit 201 can be used to implement at least one of the following functions:

[0062] 1. Receive a data accumulation request. The data accumulation request is used to request that the sum of the first operand and the second operand in the accumulation memory 200 be stored in the accumulation memory 200.

[0063] Optionally, the aforementioned data accumulation request can be generated by the arithmetic logic unit in the processor and sent to the accumulation memory 200. The first operand is an operand to be added. For example, the first operand can be the calculation result of the arithmetic logic unit (such as an intermediate calculation result). The second operand is another operand to be added, which can be specified by the arithmetic logic unit.

[0064] Optionally, the data accumulation request includes a first operand and a first storage address, where the first storage address is the storage address of the second operand in the accumulation register 200, such as the storage address of the second operand in the second storage unit 203.

[0065] 2. If the second operand is stored in the first storage unit 202, the second operand is read from the first storage unit 202.

[0066] Optionally, if the first storage unit 202 stores the second operand and the first storage address of the second operand is stored in the first storage unit 202, then the data processing unit 201 can determine that the first storage unit 202 stores the second operand if it determines that the first storage address is stored in the first storage unit 202.

[0067] For example, the data processing unit 201 is configured to determine a first entry in the first storage unit 202 that stores a first storage address, and to read the accumulated number of history from the first field segment of the first entry as a second operand.

[0068] 3. If the second operand is not stored in the first storage unit 202, the second operand is read from the second storage unit 203.

[0069] Optionally, if the first storage unit 202 does not store the second operand, and therefore does not store the first storage address, the data processing unit 201 can determine that the second storage unit 203 stores the second operand if the first storage address is not stored in the first storage unit 202. Optionally, the arithmetic logic unit will only send the aforementioned data accumulation request to the accumulation memory 200 after determining that the second operand exists in the accumulation memory 200.

[0070] For example, the data processing unit 201 is further configured to read the corresponding historical accumulated number from the second storage unit 203 according to the first storage address, as the second operand.

[0071] This application embodiment reads data from the second storage unit 203 only when there is no second operand in the first storage unit 202, which helps to reduce the number of accesses to the second storage unit 203.

[0072] 4. Obtain the sum based on the first operand and the second operand.

[0073] Optionally, the data processing unit 201 is also used to add the first operand and the second operand to obtain the accumulated number.

[0074] 5. Store the accumulated number into the first storage unit 202.

[0075] Optionally, the data processing unit 201 is further configured to store the accumulated number into the first storage unit 202 to replace the second operand. For example, the data processing unit 201 may store the accumulated number into the first field of the aforementioned first entry to overwrite the storage of the second operand.

[0076] Optionally, if the operation is completed, the accumulated number is updated to the second storage unit 203 as the final accumulated number; if the operation is not completed, the accumulated number continues to participate in subsequent accumulation operations as a historical accumulated number.

[0077] 6. Transfer the data in the eliminated entries to the second storage unit 203.

[0078] The eliminated item is the item that has been accessed the least among at least one items within a preset time period. The preset time period can be set and adjusted according to actual usage needs, and this application embodiment does not limit it. For example, the preset time period is a time period ending at the current time, and the duration of the preset time period is a set value.

[0079] In this embodiment, if the data in an entry participates in an addition operation, the entry can be determined to have been accessed once, regardless of whether the data is used as an operand or an accumulator. Alternatively, if the data in an entry is read once or updated once, the entry can be determined to have been accessed once; this embodiment does not limit this.

[0080] Optionally, the data processing unit 201 is configured to determine an item to be discarded when at least one item has been fully used and there is a need to store a new accumulator, and to transfer the data in the discarded item to the second storage unit 203. The transferred data is the accumulator stored in the first field of the discarded item.

[0081] For example, if there is a need to store a new accumulated number in the first storage unit 202, and if the first entry in at least one entry is accessed the least number of times within a preset time period, the data processing unit 201 will transfer the data in the first entry to the second storage unit 203.

[0082] Optionally, the data processing unit 201 is further configured to transfer the transferred data to the corresponding storage area in the second storage unit 203 according to the storage address stored in the second field segment of the eliminated entry.

[0083] In one example, in the initial state, both the first storage unit 202 and the second storage unit 203 are empty, meaning neither stores any data. The data processing unit 201 is used to, upon acquiring the first data, store the first data in the first storage unit 202 and allocate a storage address for the first data in the second storage unit 203.

[0084] In summary, the technical solution provided in this application, by setting a first storage unit and a second storage unit in the accumulator memory, and prioritizing the storage of the accumulated number in the first storage unit, achieves the goal of using the first storage unit to assist the second storage unit in storing the accumulated number. This helps reduce the storage pressure on the second storage unit. Furthermore, by prioritizing the reading of the second operand used to calculate the accumulated number from the first storage unit, the number of accesses to the second storage unit can be reduced, thereby reducing the power consumption overhead caused by accessing the second storage unit and ultimately improving the performance of the accumulator memory.

[0085] In addition, as the number of memory accesses decreases, the probability of memory access conflicts is reduced, allowing requests to access memory more quickly, which in turn improves the access performance of the accumulator memory.

[0086] Please refer to Figure 4 The diagram illustrates an accumulation memory provided in another embodiment of this application. The data processing unit 201 described above may include a request processing subunit 2011 and a data update subunit 2012.

[0087] The request processing subunit 2011 is coupled to the first storage unit 202 and has access to the first storage unit 202. The request processing subunit 2011 is coupled to the data update subunit 2012 and can transmit data to the data update subunit 2012. The data update subunit 2012 is coupled to the first storage unit 202 and has access to the first storage unit 202.

[0088] For example, the request processing subunit 2011 includes three types of interfaces. The first type of interface of the request processing subunit 2011 is used to receive data accumulation requests, such as being coupled with the interface of the arithmetic logic unit. The second type of interface of the request processing subunit 2011 is used to acquire data, and the third type of interface of the request processing subunit 2011 is used to send data. The data update subunit 2012 also includes three types of interfaces. The first type of interface of the data update subunit 2012 is used to receive data, the second type of interface of the data update subunit 2012 is used to acquire data, and the third type of interface of the data update subunit 2012 is used to send data. The first storage unit 202 includes two types of interfaces. The first type of interface of the first storage unit 202 is used to write data, and the second type of interface of the first storage unit 202 is used to read data.

[0089] For example, the second type of interface of the request processing subunit 2011 is coupled to the first type of interface of the first storage unit 202 for reading data from the first storage unit 202. The third type of interface of the request processing subunit 2011 is coupled to the first type of interface of the data update subunit 2012 for passing data to the data update subunit 2012.

[0090] The second type of interface of the data update subunit 2012 is coupled to the first type of interface of the first storage unit 202 for reading data from the first storage unit 202. The third type of interface of the data update subunit 2012 is coupled to the second type of interface of the first storage unit 202 for storing data into the first storage unit 202.

[0091] The interface in the embodiments of this application is used for coupling between various units, and it can be implemented as at least one of the following: pin, terminal, connector.

[0092] The aforementioned request processing subunit 2011 is a unit for processing requests, and can be implemented as a logic circuit for processing requests. Optionally, the request processing subunit 2011 can be used to extract data from the request.

[0093] For example, the request processing subunit 2011 is configured to receive a data accumulation request, and after receiving the data accumulation request, extract a first operation and a first storage address from the data accumulation request, wherein the first storage address is the storage address of the second operand in the accumulation memory 200.

[0094] In one example, the request processing subunit 2011 can also be used to determine the storage location of the second operand in the accumulator memory 200. For example, the request processing subunit 2011 can also be used to determine whether the second operand is stored in the first storage unit 202 or in the second storage unit 203.

[0095] For example, the request processing subunit 2011 is configured to determine at least one target entry from at least one entry based on valid identification information, wherein the target entry stores data; read the storage addresses stored in the at least one target entry from the first storage unit 202 to obtain a set of storage addresses; and determine that the first storage unit 202 stores a second operand if the first storage address exists in the set of storage addresses.

[0096] Optionally, the request processing subunit 2011 traverses the third field of at least one entry in the first storage unit 202, and for any entry, if the valid identification information is the first parameter, then the entry is determined as the target entry.

[0097] This application embodiment, by adding a setting request processing subunit 2011, can realize the determination of the storage location of the second operand, which is beneficial to improving the accuracy of obtaining the second operand.

[0098] Optionally, the request processing subunit 2011 is further configured to send the first operand in the data accumulation request to the data update subunit 2012 if the second operand is stored in the first storage unit 202.

[0099] Optionally, the request processing subunit 2011 can also be used to send the first operand to the data update subunit 2012 if a first storage address exists in the storage address set.

[0100] This application embodiment can quickly determine whether the second operand is stored in the first storage unit 202 by matching the storage address cached in the first storage unit 202, which helps to improve the efficiency of determining the storage location of the second operand.

[0101] In one example, the data update subunit 2012 can be used to perform read operations, addition operations, and storage operations of the second operand when the second operand is stored in the first storage unit 202.

[0102] For example, the data update subunit 2012 is used to read the second operand from the first storage unit 202; add the first operand and the second operand to obtain an accumulated number; and store the accumulated number into the first storage unit 202.

[0103] Optionally, the data update subunit 2012 may read the second operand from a first entry storing a first storage address. The data update subunit 2012 may perform an addition operation on the first operand and the second operand to obtain an accumulated number. The data update subunit 2012 may store the accumulated number into the first entry to replace the second operand, such as overwriting the second operand in the first field of the first entry with the accumulated number.

[0104] The above describes the implementation process of the addition operation when the second operand is stored in the first storage unit 202. The following will explain the implementation process of the addition operation when the second operand is not stored in the first storage unit 202.

[0105] In one example, the data processing unit 201 further includes a data access subunit 2013. The data access subunit 2013 is coupled to the request processing subunit 2011 and the second storage unit 203, respectively, and can receive data from the request processing subunit 2011 and access the second storage unit 203.

[0106] For example, the data access subunit 2013 includes three types of interfaces: a first type of interface for receiving data, a second type of interface for acquiring data, and a third type of interface for sending data. The second storage unit 203 includes two types of interfaces: a first type of interface for writing data and a second type of interface for reading data.

[0107] For example, the third type interface of the request processing subunit 2011 is coupled to the first type interface of the data access subunit 2013 for passing data to the data access subunit 2013.

[0108] The second type interface of the data access subunit 2013 is coupled to the first type interface of the second storage unit 203 for reading data from the second storage unit 203.

[0109] The third type interface (such as a portion) of the data access subunit 2013 is coupled to the second type interface of the second storage unit 203 for storing data into the second storage unit 203.

[0110] Optionally, the request processing subunit 2011 is used to send a first storage address to the data access subunit 2013 when the second operand is not stored in the first storage unit 202. The first storage address is the storage address of the second operand in the accumulator memory 200.

[0111] Optionally, the request processing subunit 2011 can be used to generate a data read request based on the first storage address and send the data read request to the data access subunit 2013. The data read request includes the first storage address, which is used to request the retrieval of the second operand.

[0112] After obtaining a data read request, the data access subunit 2013 can extract a first storage address from the data read request. After obtaining the first storage address, the data access subunit 2013 can be used to read a second operand from the second storage unit 203 according to the first storage address.

[0113] Optionally, the data access subunit 2013 may determine the data (such as historical accumulated data) stored at the first storage address in the second storage unit 203 as the second operation data.

[0114] In one feasible example, after reading the second operation data, the data access subunit 2013 can delete the second operand from the second storage unit 203 to reduce the storage pressure on the second storage unit 203. Alternatively, in the subsequent accumulation process, the second operand can be overwritten with a new accumulated number; this embodiment of the application does not limit this approach.

[0115] In this embodiment, data is read from the second storage unit 203 by the data access subunit 2013, which can effectively reduce the workload of the data update subunit 2012, thereby improving the performance of the accumulator memory 200.

[0116] In one example, the data processing unit 201 further includes a data selection subunit 2014 and an adder 2015. The data selection subunit 2014 is coupled to the data memory access subunit 2013 and can be used to receive data from the data memory access subunit 2013. The data selection subunit 2014 is coupled to the adder 2015 and can be used to send data to the adder 2015.

[0117] For example, the data selection subunit 2014 includes two types of interfaces: a first type of interface for receiving data and a second type of interface for sending data. The adder 2015 includes two types of interfaces: a first type of interface for receiving data and a second type of interface for sending data.

[0118] For example, a third type interface (such as a portion) of the data access subunit 2013 is coupled to a first type interface of the data selection subunit 2014 for sending data to the data selection subunit 2014.

[0119] The second type interface of the data selection subunit 2014 is coupled to the first type interface of the adder 2015 for transmitting data to the adder 2015. If the adder 2015 includes two first type interfaces, the second type interface of the data selection subunit 2014 is coupled to one of the two first type interfaces of the adder 2015.

[0120] Optionally, the data access subunit 2013 is further configured to send the second operand to the data selection subunit 2014. Optionally, the data access subunit 2013 may send multiple acquired data to the data selection subunit 2014; the data selection subunit 2014 is configured to select the second operand from the received multiple data.

[0121] Optionally, the data selection subunit 2014 is used to send the second operand to the adder 2015 if the second operand is selected.

[0122] Adder 2015 is used to perform addition operations. Optionally, adder 2015 can be used to perform an addition algorithm on two operands. For example, after obtaining the first operand and the second operand, adder 2015 is used to add the first operand and the second operand to obtain an accumulated number.

[0123] In this embodiment of the application, when the second operand is not stored in the first storage unit 202, the addition operation is performed by the adder, which helps to further reduce the working pressure of the data update subunit 2012, thereby improving the performance of the accumulation memory 200.

[0124] In one example, the accumulator memory 200 further includes a third storage unit 204. The third storage unit 204 is used to temporarily store the first operand. The third storage unit 204 is coupled to the request processing subunit 2011 and the adder 2015, respectively.

[0125] For example, the third storage unit 204 includes two types of interfaces: a first type of interface for writing data and a second type of interface for popping data.

[0126] For example, the third type interface of the request processing subunit 2011 is coupled to the first type interface of the third storage unit 204 for transmitting data to the third storage unit 204. The second type interface of the third storage unit 204 is coupled to the first type interface of the adder 2015 for transmitting data to the adder 2015. If the adder 2015 includes two first type interfaces, the second type interface of the third storage unit 204 is coupled to one of the two first type interfaces of the adder 2015 (which is different from the interface coupled to the data selection subunit 2014).

[0127] When the second operand needs to be read from the second storage unit 203, the reading process of the second operand requires several clock cycles. To address this problem, this embodiment stores the first operand in the third storage unit 204, waits for the second operand to be read from the second storage unit 203, and then pops the first operand from the third storage unit 204 and sends them together to the adder 2015 for accumulation. This ensures that the timing of the first operand matches the timing of the second operand.

[0128] Optionally, the third storage unit 204 can be implemented as a FIFO (First In First Out) queue. FIFO is a type of buffer with first-in-first-out characteristics.

[0129] For example, the request processing subunit 2011 is further configured to send the first operand to the third storage unit 204 if the second operand is not stored in the first storage unit. The third storage unit 204 is configured to temporarily store the first operand. The data selection subunit 2014 is further configured to send a data pop notification to the third storage unit 204 when the second operand is selected. The data pop notification is used to trigger the popping of the first operand. The third storage unit 204 is configured to pop the first operand to the adder 2015 upon receiving the data pop notification.

[0130] In one example, the data processing unit 201 further includes a data writing subunit 2016. The data writing subunit 2016 is coupled to the adder 2015 and the first storage unit 202, respectively, and can be used to obtain data sent by the adder 2015 and access the first storage unit 202.

[0131] For example, the data writing unit 2016 includes two types of interfaces: a first type of interface for receiving data and a second type of interface for sending data.

[0132] For example, the second type interface of adder 2015 is coupled to (as partly) the first type interface of data writing unit 2016 for transferring data to data writing unit 2016. The second type interface of data writing unit 2016 is coupled to the first type interface of first storage unit 202 for transferring data to first storage unit 202.

[0133] For example, adder 2015 is also used to send the accumulated number to data writing subunit 2016; data writing subunit 2016 is used to store the accumulated number into first storage unit 202.

[0134] Optionally, the data writing sub-unit 2016 can be used to store the accumulated number into the first field of the idle entry when there is an idle entry in the first storage unit 202, wherein the idle entry is an entry in which at least one entry does not store data.

[0135] The data writing sub-unit 2016 can also be used to store the accumulated number into the first field of the eliminated entry when there are no idle entries in the first storage unit 202, wherein the eliminated entry is the entry that has been accessed the least among at least one entry within a preset time period.

[0136] When storing the accumulated number in an idle entry or a eviction entry, the data writing subunit 2016 can also be used to store the first storage address of the second operand into the second field of the idle entry or the second field of the eviction entry. For idle entries, the data writing subunit 2016 can also be used to adjust the valid identification information from the second parameter to the first parameter.

[0137] In this embodiment of the application, by storing the accumulated number obtained by the adder in the first storage unit 202, the number of accesses to the second storage unit 203 is reduced, thereby reducing the power consumption overhead caused by accessing the second storage unit 203.

[0138] In addition, by storing the accumulated number in the form of entries, the accumulated number can be stored in an ordered manner, which is conducive to fast memory access of the accumulated number.

[0139] In one example, the data processing unit 201 further includes a data transfer subunit 2017. The data transfer subunit 2017 is coupled to the first storage unit 202 and the data access subunit 2013, respectively, and can access the first storage unit 202 and send data to the data access subunit 2013.

[0140] For example, the data transfer subunit 2017 includes three types of interfaces: the first type of interface of the data transfer subunit 2017 is used to receive data, the second type of interface of the data transfer subunit 2017 is used to acquire data, and the third type of interface of the data transfer subunit 2017 is used to send data.

[0141] For example, the second type of interface of the data transfer subunit 2017 is coupled to the first type of interface of the first storage unit 202 for reading data from the first storage unit 202. The third type of interface of the data transfer subunit 2017 is coupled to the first type of interface of the data access subunit 2013 for transferring data to the data access subunit 2013.

[0142] For example, the data transfer subunit 2017 is used to read the transfer data stored in the elimination entry from the first storage unit 202 and send the transfer data to the data access subunit 2013; the data access subunit 2013 is used to store the transfer data into the second storage unit 203.

[0143] Optionally, the data transfer subunit 2017 is used to read the transferred data stored in the elimination entry from the first storage unit 202 when the data writing subunit 2016 stores the accumulated number in the first field of the elimination entry. For example, while the data transfer subunit 2017 reads data from the elimination entry, the data writing subunit 2016 can write data to the elimination entry.

[0144] For example, the data transfer subunit 2017, after acquiring the transfer data, generates a data write request based on the transfer data. The data write request includes the transfer data and its storage address, and requests the transfer data to be stored in the second storage unit 203. The data write request is then sent to the data access subunit 2013. The data access subunit 2013 stores the transfer data in the corresponding storage area of ​​the second storage unit 203 according to the storage address in the data write request.

[0145] In this embodiment, by transferring the data in the elimination entry to the second storage unit 203, the probability of accessing the data in the elimination entry is low, which helps to reduce the number of accesses to the second storage unit 203, thereby reducing the power consumption overhead caused by accessing the second storage unit 203.

[0146] In one example, the data processing unit 201 further includes an entry determination subunit 2018. The entry determination subunit 2018 is coupled to the data writing subunit 2016 and the data transfer subunit 2017, respectively, and can send data to the data writing subunit 2016 and the data transfer subunit 2017, respectively.

[0147] For example, the entry determination subunit 2018 includes two types of interfaces. The first type of interface of the entry determination subunit 2018 is used to acquire data, and the second type of interface of the entry determination subunit 2018 is used to send data.

[0148] For example, the entry identifies that a first type of interface of subunit 2018 is coupled to a first type of interface of first storage unit 202 for retrieving data from first storage unit 202. The entry identifies that a portion of a second type of interface of subunit 2018 is coupled to data writing subunit 2016 for transferring data to writing subunit 2016. The entry identifies that a portion of a second type of interface of subunit 2018 is coupled to a first type of interface of data transfer subunit 2017 for transferring data to data transfer subunit 2017.

[0149] For example, the entry determination subunit 2018 is used to determine the elimination entry based on the number of times each of the at least one entry is accessed within a preset time period when there are no idle entries in the first storage unit 202; send the identification information of the elimination entry to the data writing subunit 2016; and send the identification information of the elimination entry to the data transfer subunit 2017.

[0150] Optionally, the entry determination subunit 2018 can also be used to determine the elimination entries based on the number of times each of at least one entry has been accessed within a preset time period upon receiving an entry determination request. The entry determination request is used to request the determination of elimination entries, and this request can be generated by the data writing subunit 2016 if there are no idle entries in the first storage unit 202.

[0151] The identification information of the eliminated entry is used to uniquely identify the eliminated entry. For example, the address information stored in the eliminated entry can be used as the identification information of the eliminated entry.

[0152] After obtaining the identification information of the eliminated item, the data writing subunit 2016 can store the accumulated number into the eliminated item according to the identification information. After obtaining the identification information of the eliminated item, the data transfer subunit 2017 can read the transferred data from the first storage unit 202 according to the identification information of the eliminated item.

[0153] Optionally, the item determination subunit 2018 can be used to determine the item that has been accessed the least during a preset time period as the eliminated item.

[0154] For example, the item determination subunit 2018 can be a logic circuit implemented based on LRU (Least Recently Used) replacement technology. The item determination subunit 2018 maintains several status registers, which record the access status of items, such as the number of times an item has been accessed within a preset time period. When an item in the first storage unit 202 is accessed, the item determination subunit 2018 updates these status registers to record the number of times at least one item has been accessed within the preset time period. If an item has not been accessed for a long time, it is highly likely to be identified as an obsolete item.

[0155] This application embodiment identifies entries that have not been accessed for a long time as eviction entries and transfers the data in the eviction entries to the second storage unit 203. This helps to reduce the number of accesses to the second storage unit 203 and alleviate the storage pressure on the first storage unit 202.

[0156] The accumulation of the accumulator exhibits temporal locality, meaning that within a certain time period, multiple accumulation operations may occur targeting the same memory address in the accumulator memory. Since the accumulator is cached in the first memory cell 202 during the accumulation process, most read and write operations only need to be performed on the first memory cell 202, resulting in minimal power consumption for the first memory cell 202. A read operation on the second memory cell 203 is only required when the first memory cell 202 is missed (i.e., no second operand is stored there), and a write operation on the second memory cell 203 is only required when an entry in the first memory cell 202 needs to be replaced. This significantly reduces the number of accesses to the second memory cell 203, thereby reducing the power consumption associated with accessing the second memory cell 203 and ultimately improving the performance of the accumulator memory 200.

[0157] In summary, the technical solution provided in this application, by setting a first storage unit and a second storage unit in the accumulator memory, and prioritizing the storage of the accumulated number in the first storage unit, achieves the goal of using the first storage unit to assist the second storage unit in storing the accumulated number. This helps reduce the storage pressure on the second storage unit. Furthermore, by prioritizing the reading of the second operand used to calculate the accumulated number from the first storage unit, the number of accesses to the second storage unit can be reduced, thereby reducing the power consumption overhead caused by accessing the second storage unit and ultimately improving the performance of the accumulator memory.

[0158] In some embodiments, reference Figure 5 (shown) Figure 4 (Part of the content), when the first storage cell 202 is hit (i.e., the second operand is stored), the operation of the accumulator memory 200 may include the following:

[0159] 1. Request processing subunit 2011 is used to receive data accumulation request; compare the first storage address in the data accumulation request with the storage address stored in each entry in the first storage unit 202; if the first storage address is found to be the same as the storage address stored in a certain entry in the first storage unit 202, determine that the first storage unit 202 is hit (i.e., it stores the second operand).

[0160] The first storage unit 202 is implemented as a register.

[0161] 2a. Request processing subunit 2011 is used to send the first operand and the first storage address in the data accumulation request to the data update subunit 2012 in preparation for the addition operation.

[0162] 2b. Data update subunit 2012 is used to read the second operand from the first storage unit 202 according to the first storage address.

[0163] 3. The data update subunit 2012 is used to perform an addition operation on the first operand and the second operand to obtain an accumulated number, and to write the accumulated number into the first storage unit 202 to replace the second operand. The storage address of the accumulated number is reserved as the first storage address of the second operand.

[0164] In some embodiments, reference Figure 6 (shown) Figure 4 In the event that the first storage cell 202 is not hit (i.e., the second operand is stored there), the operation of the accumulator memory 200 may also include the following:

[0165] 1a. Request processing subunit 2011 is used to receive a data accumulation request; compare the first storage address in the data accumulation request with the storage addresses stored in each entry in the first storage unit 202; if it is detected that the first storage unit 202 does not contain a storage address that is the same as the first storage address, determine that the first storage unit 202 is a miss (i.e., the second operand is not stored); and store the first operand in the data accumulation request into the third storage unit 204.

[0166] The third storage unit 204 is implemented as a FIFO.

[0167] 1b. Request processing subunit 2011 is used to initiate a data read request to the second storage unit 203 based on the first storage address. For example, request processing subunit 2011 is used to send a data read request to data access subunit 2013, and the data read request includes the first storage address.

[0168] 2a. Data access subunit 2013 is used to read the second operand from the second storage unit 203 according to the first storage address, and to send the second operand to the data selection subunit 2014; the data selection subunit 2014 is used to send the second operand to the adder 2015, and to send a data pop notification to the third storage unit 204, the data pop notification being used to trigger the pop of the first operand.

[0169] The second storage unit 203 is implemented as a memory.

[0170] 2b. The third storage unit 204 is used to pop the first operand to the adder 2015 according to the data pop notification. The adder 2015 is used to add the first operand and the second operand to obtain the accumulated number, and to send the accumulated number to the data writing subunit 2016.

[0171] 3a. Data writing sub-unit 2016 is used to store the accumulated number into the first field of the eliminated entry when there is no idle entry in the first storage unit 202. The eliminated entry is the entry that has been accessed the least among at least one entry within a preset time period.

[0172] 3b. Data transfer subunit 2017 is used to read the transferred data in the elimination entry into data access subunit 2013; data access subunit 2013 is used to store the transferred data into the second storage unit 203.

[0173] Steps 3a and 3b can be performed simultaneously. The elimination entry can be determined by the entry determination subunit 2018, and the elimination entry can be notified to the data writing subunit 2016 and the data transfer subunit 2017 by the entry determination subunit 2018.

[0174] In summary, the technical solution provided in this application, by setting a first storage unit and a second storage unit in the accumulator memory, and prioritizing the storage of the accumulated number in the first storage unit, achieves the goal of using the first storage unit to assist the second storage unit in storing the accumulated number. This helps reduce the storage pressure on the second storage unit. Furthermore, by prioritizing the reading of the second operand used to calculate the accumulated number from the first storage unit, the number of accesses to the second storage unit can be reduced, thereby reducing the power consumption overhead caused by accessing the second storage unit and ultimately improving the performance of the accumulator memory.

[0175] The following are embodiments of the method of this application. For details not disclosed in the embodiments of the method of this application, please refer to the module embodiments of this application.

[0176] Please refer to Figure 7The diagram illustrates a flowchart of a data processing method in an accumulator memory provided in an embodiment of this application. This method can be applied to the accumulator memory 200 described above, which includes a data processing unit, a first storage unit, and a second storage unit. The method may include the following steps (701-704).

[0177] Step 701: The data processing unit receives a data accumulation request. The data accumulation request is used to request that the sum of the first operand and the second operand in the accumulation memory be stored in the accumulation memory.

[0178] Step 702: If the data processing unit has stored the second operand in the first storage unit, it reads the second operand from the first storage unit.

[0179] Step 703: The data processing unit obtains the accumulated number based on the first operand and the second operand.

[0180] Step 704: The data processing unit stores the accumulated number into the first storage unit.

[0181] In some embodiments, the data processing unit includes: a request processing subunit and a data update subunit;

[0182] The request processing subunit receives a data accumulation request, and if the second operand is stored in the first storage unit, it sends the first operand in the data accumulation request to the data update subunit.

[0183] The data update subunit reads the second operand from the first storage unit; adds the first operand and the second operand to obtain the accumulated number; and stores the accumulated number into the first storage unit.

[0184] In one example, the first storage unit includes at least one entry, which includes a first domain segment, a second domain segment, and a third domain segment. The first domain segment is used to store data, the second domain segment stores the storage address of the data stored in the entry, and the third domain segment stores valid identification information for indicating whether the entry stores data.

[0185] In one example, the data accumulation request also includes a first storage address, which is the storage address of the second operand in the accumulation memory;

[0186] The request processing subunit determines at least one target entry from at least one entry based on valid identification information, and the target entry stores data.

[0187] The request processing subunit reads the storage addresses stored in at least one target entry from the first storage unit to obtain a set of storage addresses;

[0188] If the request processing subunit has a first storage address in the storage address set, it will send the first operand to the data update subunit.

[0189] In some embodiments, if the data processing unit does not store the second operand in the first storage unit, it reads the second operand from the second storage unit.

[0190] In one example, the data processing unit includes a request processing subunit and a data access subunit;

[0191] If the request processing subunit does not store the second operand in the first storage unit, it sends the first storage address to the data access subunit. The first storage address is the storage address of the second operand in the accumulator memory.

[0192] The data access subunit reads the second operand from the second memory unit based on the first memory address.

[0193] In one example, the data processing unit also includes: a data selection subunit and an adder;

[0194] The data memory access subunit sends the second operand to the data selection subunit;

[0195] If the data selection subunit selects the second operand, it sends the second operand to the adder.

[0196] The adder adds the first operand and the second operand to obtain the sum.

[0197] In one example, the accumulator memory also includes: a third memory unit;

[0198] If the request processing subunit does not store the second operand in the first storage unit, it will send the first operand to the third storage unit.

[0199] When the data selection subunit selects the second operand, it sends a data pop notification to the third storage unit. The data pop notification is used to trigger the pop of the first operand.

[0200] Upon receiving a data pop notification, the third storage unit pops the first operand into the adder.

[0201] In one example, the data processing unit further includes a data writing subunit;

[0202] The adder sends the accumulated number to the data writing sub-unit;

[0203] The data writing sub-unit stores the accumulated number into the first storage unit.

[0204] In one example, the first storage unit includes at least one entry, and the first field segment of the entry is used to store data;

[0205] When there are idle entries in the first storage unit, the accumulated number is stored in the first field of the idle entry, wherein the idle entry is at least one entry that does not store data.

[0206] Alternatively, if there are no idle entries in the first storage unit, the accumulated number is stored in the first field of the eviction entry, where the eviction entry is the entry that has been accessed the least among at least one entries within a preset time period.

[0207] In one example, the data processing unit further includes: a data transfer subunit and a data memory access subunit;

[0208] The data transfer subunit reads the transfer data stored in the elimination entry from the first storage unit and sends the transfer data to the data access subunit;

[0209] The data access subunit stores the transferred data into the second storage unit.

[0210] In one example, the data processing unit further includes: an entry determination subunit;

[0211] If there are no idle entries in the first storage unit, the entry determination sub-unit determines the eviction entries based on the number of times each entry is accessed within a preset time period.

[0212] The item determination subunit sends the identification information of the eliminated item to the data writing subunit;

[0213] The entry determination subunit sends the identification information of the eliminated entry to the data transfer subunit.

[0214] In one example, the first storage unit is built based on registers, and the second storage unit is built based on memory.

[0215] In summary, the technical solution provided in this application, by setting a first storage unit and a second storage unit in the accumulator memory, and prioritizing the storage of the accumulated number in the first storage unit, achieves the goal of using the first storage unit to assist the second storage unit in storing the accumulated number. This helps reduce the storage pressure on the second storage unit. Furthermore, by prioritizing the reading of the second operand used to calculate the accumulated number from the first storage unit, the number of accesses to the second storage unit can be reduced, thereby reducing the power consumption overhead caused by accessing the second storage unit and ultimately improving the performance of the accumulator memory.

[0216] In some embodiments, this application also provides a processor, which includes the accumulator memory described in the above embodiments.

[0217] In some embodiments, this application also provides a computer device, the computer device including a processor, the processor including the accumulator memory described in the above embodiments.

[0218] In some embodiments, the processor described above is an AI (Artificial Intelligence) processor, or other processor that requires the use of accumulator memory; this application does not limit this.

[0219] In some embodiments, the computer device described above may be a server, or a terminal device such as a mobile phone, tablet computer, vehicle terminal, wearable device, smart home device, or any device that uses a processor, such as a robot or base station. This application embodiment does not limit this.

[0220] In some embodiments, this application provides a server cluster, which includes at least two servers, and at least one of the at least two servers includes the accumulator memory described in the above embodiments.

[0221] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0222] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An accumulator memory, characterized in that, The accumulator memory includes: a data processing unit, a first storage unit, and a second storage unit. The first storage unit is used to assist the second storage unit in storing the accumulated number. The data processing unit is used for: A data accumulation request is received, the data accumulation request being used to request that the sum of the first operand and the second operand in the accumulation memory be stored in the accumulation memory; If the second operand is stored in the first storage unit, the second operand is read from the first storage unit; The accumulated number is obtained based on the first operand and the second operand; The accumulated number is stored in the first storage unit.

2. The accumulator memory according to claim 1, characterized in that, The data processing unit includes: a request processing subunit and a data update subunit; The request processing subunit is configured to receive the data accumulation request, and, if the second operand is stored in the first storage unit, send the first operand in the data accumulation request to the data update subunit. The data update subunit is used to read the second operand from the first storage unit; add the first operand and the second operand to obtain the accumulated number; and store the accumulated number into the first storage unit.

3. The accumulator memory according to claim 2, characterized in that, The first storage unit includes at least one entry, the entry including a first domain segment, a second domain segment and a third domain segment, the first domain segment being used to store data, the second domain segment storing the storage address of the data stored in the entry, and the third domain segment storing valid identification information for indicating whether the entry stores data.

4. The accumulator memory according to claim 3, characterized in that, The data accumulation request also includes a first storage address, which is the storage address of the second operand in the accumulation memory; the request processing subunit is further configured to: Based on the valid identification information, at least one target entry is determined from the at least one entry, wherein the target entry stores data; Read the storage addresses stored in the at least one target entry from the first storage unit to obtain a set of storage addresses; If the first storage address exists in the set of storage addresses, the first operand is sent to the data update subunit.

5. The accumulator memory according to any one of claims 1 to 4, characterized in that, The data processing unit is further configured to: If the second operand is not stored in the first storage unit, the second operand is read from the second storage unit.

6. The accumulator memory according to claim 5, characterized in that, The data processing unit includes: a request processing subunit and a data access subunit; The request processing subunit is configured to send a first storage address to the data access subunit when the second operand is not stored in the first storage unit. The first storage address is the storage address of the second operand in the accumulator memory. The data access subunit is used to read the second operand from the second storage unit according to the first storage address.

7. The accumulator memory according to claim 6, characterized in that, The data processing unit further includes: a data selection subunit and an adder; The data access subunit is further configured to send the second operand to the data selection subunit; The data selection subunit is used to send the second operand to the adder when the second operand is selected. The adder is used to add the first operand and the second operand to obtain the accumulated number.

8. The accumulator memory according to claim 7, characterized in that, The accumulator memory further includes: a third storage unit; The request processing subunit is further configured to send the first operand to the third storage unit if the second operand is not stored in the first storage unit; The data selection subunit is further configured to send a data pop-up notification to the third storage unit when the second operand is selected, and the data pop-up notification is used to trigger the pop-up of the first operand; The third storage unit is used to pop the first operand to the adder upon receiving the data pop notification.

9. The accumulator memory according to claim 7 or 8, characterized in that, The data processing unit further includes: a data writing subunit; The adder is also used to send the accumulated number to the data writing subunit; The data writing sub-unit is used to store the accumulated number into the first storage unit.

10. The accumulator memory according to claim 9, characterized in that, The first storage unit includes at least one entry, the first field of which is used to store data; the data writing subunit is further used for: If there is an idle entry in the first storage unit, the accumulated number is stored in the first field of the idle entry, wherein the idle entry is an entry in which no data is stored in the at least one entry; or, If there are no idle entries in the first storage unit, the accumulated number is stored in the first field of the eviction entry, wherein the eviction entry is the entry that has been accessed the least among the at least one entries within a preset time period.

11. The accumulator memory according to claim 10, characterized in that, The data processing unit further includes: a data transfer subunit and a data memory access subunit; The data transfer subunit is used to read the transfer data stored in the elimination entry from the first storage unit, and to send the transfer data to the data access subunit. The data access subunit is used to store the transferred data into the second storage unit.

12. The accumulator memory according to claim 10 or 11, characterized in that, The data processing unit further includes: an entry determination subunit; the entry determination subunit is used for: If there are no idle entries in the first storage unit, the discarded entries are determined based on the number of times each of the at least one entry is accessed within the preset time period; The identification information of the eliminated item is sent to the data writing subunit; The identification information of the eliminated entry is sent to the data transfer subunit.

13. The accumulator memory according to any one of claims 1 to 12, characterized in that, The first storage unit is built based on registers, and the second storage unit is built based on memory.

14. A processor, characterized in that, The processor includes the accumulator memory as described in any one of claims 1 to 13.

15. A computer device, characterized in that, The computer device includes a processor, the processor including an accumulation memory as described in any one of claims 1 to 13.

16. A server cluster, characterized in that, The server cluster includes at least two servers, and at least one of the at least two servers includes an accumulation memory as described in any one of claims 1 to 13.

17. A data processing method in an accumulator memory, characterized in that, The accumulator memory includes: a data processing unit, a first storage unit, and a second storage unit, wherein the first storage unit is used to assist the second storage unit in storing the accumulated number; the method includes: The data processing unit receives a data accumulation request, which requests that the sum of the first operand and the second operand in the accumulation memory be stored in the accumulation memory. When the data processing unit has stored the second operand in the first storage unit, it reads the second operand from the first storage unit. The data processing unit obtains the accumulated number based on the first operand and the second operand; The data processing unit stores the accumulated number into the first storage unit.