Atomic operation processing system, method and computer device
Patent Information
- Application Number
- CN202611349800.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-02
- Publication Date
- 2026-09-29
AI Technical Summary
这使得原子操作的效率较低
[0007]本申请提供的技术方案带来的有益效果至少包括:
Smart Images

Figure CN122837758A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of chips, and in particular to an atomic manipulation processing system, method and computer device. Background Technology
[0002] Atomic operations are operations that cannot be interrupted during execution and are one of the important operations in a GPU (Graphics Processing Unit). The atomic operation processing units (such as the ALU (Arithmetic Logic Unit)) used to perform atomic operations are also important units in a GPU. During the execution of an atomic operation, data needs to be read from or written to memory.
[0003] In related technologies, when reading data, the memory returns the read data to the atomic operation processing unit according to the read request. The data transmission process is similar when writing data. During both reading and writing, to ensure the exclusivity of atomic operations, the next atomic operation can only be executed after the data reading and writing are complete. This makes atomic operations relatively inefficient. Summary of the Invention
[0004] This application provides an atomic manipulation processing system, method, and computer device, the technical solution of which is as follows: According to one aspect of this application, an atomic operation processing system is provided, the atomic operation processing system including an atomic operation processing unit, an auxiliary storage unit, and a data storage unit; the access latency of the auxiliary storage unit is lower than the access latency of the data storage unit; The atomic operation processing unit is configured to read initial data from the auxiliary storage unit when the auxiliary storage unit is hit, wherein the initial data is read from the data storage unit to the auxiliary storage unit based on historical atomic operation requests; The atomic operation processing unit is used to execute the atomic operation based on the atomic operation request and the initial data, and obtain the atomic operation result; The atomic operation processing unit is used to write the atomic operation result back to the auxiliary storage unit.
[0005] According to one aspect of this application, an atomic operation processing method is provided, the method being executed by an atomic operation processing system, the atomic operation processing system including an atomic operation processing unit, an auxiliary storage unit, and a data storage unit; the access latency of the auxiliary storage unit is lower than the access latency of the data storage unit; When the atomic operation processing unit is hit in the auxiliary storage unit, it reads initial data from the auxiliary storage unit. The initial data is read from the data storage unit to the auxiliary storage unit based on historical atomic operation requests. The atomic operation processing unit executes the atomic operation based on the atomic operation request and the initial data to obtain the atomic operation result; The atomic operation processing unit writes the atomic operation result back to the auxiliary storage unit.
[0006] According to one aspect of this application, a computer device is provided, the computer device including an atomic operation processing system.
[0007] The beneficial effects of the technical solution provided in this application include at least the following: The atomic operation processing system employs two storage units, with the access latency of the auxiliary storage unit being lower than that of the data storage unit. Compared to traditional atomic operation processing systems that only have one data storage unit, this system utilizes an additional auxiliary storage unit with lower access latency. This eliminates the need for excessive time spent waiting for data reads and writes during atomic operations, improving execution efficiency, especially for multiple consecutive atomic operations. Furthermore, this system is a modification of a traditional atomic operation processing system, achieving efficiency improvements with minor alterations while maintaining a small footprint and low production costs. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 A schematic diagram of the architecture of an atomic operation processing system in the related technology is shown; Figure 2 This invention illustrates a schematic diagram of the architecture of an atomic operation processing system provided in an exemplary embodiment of this application. Figure 3 A schematic diagram of the architecture of an atomic operation processing system provided in another exemplary embodiment of this application is shown; Figure 4 A schematic diagram of an atomic operation process provided in an exemplary embodiment of this application is shown; Figure 5 A schematic diagram of an atomic operation process provided by another exemplary embodiment of this application is shown; Figure 6 A flowchart of an atomic operation processing method provided in an exemplary embodiment of this application is shown; Figure 7 A structural block diagram of a computer device provided in an exemplary embodiment of this application is shown. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0011] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0012] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0013] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the settings and operation information involved in this application were obtained with full authorization.
[0014] It should be understood that although the terms first, second, etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0015] First, let me introduce the relevant terms used in this application: An instruction is a command that directs a computer to perform a specific operation; it is the smallest functional unit of computer operation. An instruction is a statement in machine language, or a set of meaningful binary code. The collection of all the instructions of a computer constitutes its instruction set, also known as its instruction system.
[0016] Instruction Format: A readable representation of an instruction. An instruction typically includes an opcode and operands. The opcode describes the type of operation the instruction will perform, while the operands provide the data or addresses of data required to execute the instruction. The opcode is indispensable in an instruction, but operands are optional, and there can be one or two operands.
[0017] An atomic operation is a single operation or group of operations that exists in only two states: either not yet executed or fully executed. There are no intermediate states. It can also be understood as an independent and indivisible operation composed of multiple steps. In a single-core environment, threads cannot be switched during atomic operations. Thread switching can only occur before or after the atomic operation. In a multi-core environment, this means that when one core performs an atomic operation on a specific memory location, it must ensure that other cores do not simultaneously modify the data in the same memory location. In other words, an atomic operation refers to one or more operations that must be completed as a whole. If an atomic operation is interrupted before any step is completed, all completed operations must be rolled back to ensure that either all operations are incomplete or all operations are completed.
[0018] In a GPU, the hardware unit that performs atomic operations can be called an Arithmetic Logic Unit (ALU), atomic operation processing unit, or data processing unit, etc., and usually requires several atomic operands as input. Atomic operations include xchg, add, sub, and, or, xor, inc, dec, cmp, cmpxchg (Compare and Exchange, also known as CAS, or Compare-And-Swap), etc.
[0019] Here, xchg represents a swap operation, which requires one atomic operand as input. This operation means replacing the data at a certain address in a storage unit (such as main memory or video memory) with the input atomic operand.
[0020] The `add` operation represents the addition operation, which requires two atomic operands. This operation means adding the two input atomic operands together.
[0021] The sub operation represents a subtraction operation, which requires two atomic operands. This operation means subtracting the two input atomic operands.
[0022] The AND operation requires two atomic operands and represents the logical sum of the two input atomic operands.
[0023] The OR operation requires two atomic operands and performs a logical OR on the two input atomic operands.
[0024] XOR stands for Exclusive OR operation, which requires two atomic operands. This operation represents the logical XOR of the two input atomic operands.
[0025] `inc` represents the increment operation, which requires two atomic operands. This operation first compares the atomic operand passed from the bus with the atomic operand read from the memory location. If the atomic operand passed from the bus is greater than the read atomic operand, then no increment occurs; otherwise, the increment occurs. It should be noted that the above example uses the case where the atomic operand passed from the bus equals the read atomic operand. In actual use, if the atomic operand passed from the bus is less than or equal to the read atomic operand, no increment is necessary. The decrement operation follows the same logic and will not be elaborated further here.
[0026] `dec` represents a decrement operation, which requires two atomic operands. This operation first compares the atomic operand passed from the bus with the atomic operand read from the memory cell. If the atomic operand passed from the bus is less than the atomic operand read, then the decrement does not occur; if the atomic operand passed from the bus is greater than or equal to the atomic operand read, then the decrement occurs.
[0027] cmp represents the comparison operation, which requires two atomic operands and is used to compare the size of the two atomic operands.
[0028] float represents floating-point sum and requires two atomic operands. This operation is used to add two atomic operands based on the precision of floating-point numbers.
[0029] CAS stands for Compare and Swap, which requires three atomic operands. Atomic operands 1 and 2 are passed from the bus, and atomic operand 3 is read from the memory location. This operation compares atomic operands 1 and 3. If atomic operands 1 and 3 are equal, atomic operand 2 replaces atomic operand 3 in the memory location; if atomic operands 1 and 3 are not equal, no replacement is performed.
[0030] Each type of atomic operation can be divided into 16-bit, 32-bit, and 64-bit based on the bit width of the atomic operand.
[0031] It should be noted that the atomic operations and their functions described above are for illustrative purposes only, and can be adjusted according to the needs of different processors in actual use.
[0032] Atomic operation instructions: Instructions used to instruct the execution of atomic operations.
[0033] Atomic operands in atomic operations can be further divided into raw data and read data, depending on their source.
[0034] Original data: Atomic operands carried in atomic operation instructions, or atomic operands read from memory locations based on addresses carried in atomic operation instructions but not modified in the atomic operation.
[0035] Reading data: This refers to reading the atomic operand from the memory location based on the address specified in the atomic operation instruction. Alternatively, it means reading the atomic operand from the memory location based on the address specified in the atomic operation instruction, where this address is also the address where the result of the atomic operation will be written after completion. Or, it can be further defined as reading the atomic operand from the memory location based on the address specified in the atomic operation instruction, where this address serves as both the read and write address.
[0036] In some embodiments, to reduce the latency of atomic operations, a common approach is to add a set of computational units in front of memory to perform atomic operations. Each atomic operation consists of three steps: reading memory, performing the operation, and writing the result.
[0037] In a cached system, this operation can also be performed in the last-level global cache. That is, a set of computational units for performing atomic operations is added before the last-level global cache. The atomic operations performed by these computational units include three steps: reading the cached SRAM (Static Random-Access Memory), performing the operation, and writing to the cached SRAM.
[0038] For example, atomic manipulation processing systems in related technologies, such as Figure 1 As shown. The atomic manipulation processing system includes: a computing unit 10 and a data storage unit 20.
[0039] Among them, data storage 20 can be the memory mentioned above, the last-level global cache, or other storage units.
[0040] For computation unit 10, it needs to read the data 44 before the atomic operation from data storage 20. Then, it performs computation based on the data 44 before the atomic operation to obtain the atomic operation result 45. Finally, the atomic operation result 45 is stored in data storage 20.
[0041] In some embodiments, the atomic operation processing system further includes a cache tag storage 30. The cache tag storage 30 is used to record relevant information about the data stored in the data storage 20. This relevant information includes at least one of the following: the data address of the data, the address tag of the data, and a status bit corresponding to the data address. The address tag is used to determine whether the data corresponding to the data address is in the data storage 20. The status bit is used to indicate the status of the data corresponding to the data address, such as whether it is valid or dirty data (i.e., modified but not synchronized to memory).
[0042] For example, for an atomic operation request 41, the cache tag store 30 is first checked for a hit.
[0043] If cache tag 30 is hit, it means that data storage 20 contains the data corresponding to atomic operation request 41. Atomic operator 42 is sent to computation unit 10, and data address 43 is sent to data storage 20 to read data 44 before the atomic operation. Data 44 before the atomic operation is the memory data before the atomic operation. After reading data 44 before the atomic operation, data storage 20 also sends data 44 before the atomic operation to computation unit 10.
[0044] Optionally, the calculation unit 10 calculates the atomic operation result 45 based on the atomic operator 42 and the data 44 before the atomic operation. Alternatively, the calculation unit 10 calculates the atomic operation result 45 based on the atomic operator 42, the data 44 before the atomic operation, and the atomic operation data 46.
[0045] Here, atomic operation data 46 refers to data transferred from an external source into the atomic operating system. For example, atomic operation data 46 can be data carried in an atomic operation request, or data read from a register based on the atomic operation request. Atomic operation data 46 can also be called raw data. Data before atomic operation 44 is data read from data storage 20. Generally, the address corresponding to atomic operation data 44 is the write address of the atomic operation result 45. That is, atomic operation data 44 can also be called read data.
[0046] If cache tag 30 is not found, it means that data corresponding to atomic operation request 41 is not stored in data storage 20. Data needs to be read from other storage units (e.g., retrieved from memory from the cache line) before performing the data reading and calculation steps described above.
[0047] However, the above scheme has the following problems: For consecutive atomic operations at the same address, such as atomic accumulation of multiple data, many cavitations are generated between two atomic operations because of the waiting time for write-back and read-back; for atomic operations at the same address, memory or cached SRAM needs to be read and written once, resulting in additional bandwidth consumption and power consumption. It should be understood that when performing atomic operations, in order to ensure the exclusivity and correctness of atomic operations, the computing unit must pause the execution of subsequent instructions, which will generate a blank cycle without valid instruction execution, i.e., "cavitation".
[0048] With the development of modern applications, the demand for reduction operations is increasing daily, making optimization for this scenario particularly important. A reduction operation combines a set of data into a single value through a binary operation (such as addition, bitwise AND, or maximum value). An atomic reduction operation is one where the binary operations within the reduction operation remain atomic. That is, an atomic reduction operation may involve multiple atomic operation requests performing multiple reads and writes to the same address, leading to significantly increased bandwidth consumption and power consumption, and reducing the execution efficiency of atomic operation requests.
[0049] Figure 2 A schematic diagram of the architecture of an atomic operation processing system provided in an exemplary embodiment of this application is shown. The atomic operation processing system includes an atomic operation processing unit 110, an auxiliary storage unit 120, and a data storage unit 130.
[0050] Optionally, the read / write speed of the auxiliary storage unit 120 is faster than that of the data storage unit 130. Alternatively, the storage capacity of the auxiliary storage unit 120 is less than that of the data storage unit 130. Or, the access speed of the auxiliary storage unit 120 is faster than that of the data storage unit 130. Or, the access latency of the auxiliary storage unit 120 is lower than that of the data storage unit 130.
[0051] Optionally, the atomic operation processing unit 110 is connected to the auxiliary storage unit 120, and the auxiliary storage unit 120 is connected to the data storage unit 130.
[0052] Atomic operation processing unit 110 is configured to read initial data from auxiliary storage unit 120 when an auxiliary storage unit 120 is hit. This initial data is read from data storage unit 130 to auxiliary storage unit 120 based on historical atomic operation requests. Optionally, historical atomic operation requests include: previous atomic operation requests.
[0053] In some embodiments, the auxiliary storage unit 120 is used to read initial data from the data storage unit 130 and store the initial data in the auxiliary storage unit 120 when the auxiliary storage unit 120 is not hit.
[0054] The initial data is the data value stored in the data address corresponding to the atomic operation request before the atomic operation is performed.
[0055] Optionally, an atomic operation request can also be called an atomic operation instruction. The data address corresponding to the atomic operation request is the opcode in the atomic operation request. That is, the atomic operation request carries a data address; the data address is used to locate the initial data corresponding to the atomic operation request.
[0056] It should be understood that the initial data read from the auxiliary storage unit 120 or the data storage unit 130 can be one initial data or multiple initial data. That is, if the atomic operation processing unit 110 hits the auxiliary storage unit 120, it reads at least one initial data from the auxiliary storage unit 120. Optionally, if the atomic operation processing unit 110 misses the auxiliary storage unit 120, it reads at least one initial data from the data storage unit 130.
[0057] The initial data is the data read from the storage unit, which can also be called read data.
[0058] Optionally, in addition to the corresponding initial data (i.e., read data), the atomic operation request may also contain corresponding raw data, such as data stored in the registers of the atomic operation processing unit 110. This raw data participates in the calculation process of the atomic operation. The source of the raw data is usually carried by the atomic operation request itself and does not need to be read from the storage unit.
[0059] For example, the atomic operation request carries a data address. The atomic operation processing unit 110 determines whether the auxiliary storage unit 120 stores the data value corresponding to the data address based on the data address. If the auxiliary storage unit 120 stores the data value corresponding to the data address, it indicates that the auxiliary storage unit 120 has a hit. The atomic operation processing unit 110 reads the initial data from the auxiliary storage unit 120 according to the data address. This initial data is read from the data storage unit 130 to the auxiliary storage unit 120 based on historical atomic operation requests. If the auxiliary storage unit 120 does not store the data value corresponding to the data address, it indicates that the auxiliary storage unit 120 has a miss. The auxiliary storage unit 120 reads the initial data from the data storage unit 130 according to the data address, stores the initial data in the auxiliary storage unit 120, and sends the initial data to the atomic operation processing unit 110.
[0060] The atomic operation processing unit 110 is used to perform atomic operations based on atomic operation requests and initial data to obtain atomic operation results.
[0061] The initial data is used to perform atomic operations. For example, if an atomic operation request corresponds to an addition operation, the initial data is used as the addend to perform the atomic operation; if an atomic operation request corresponds to a subtraction operation, the initial data is used as the subtrahend and / or minuend to perform the atomic operation; if an atomic operation request corresponds to a swap operation, the initial data is the data value waiting to be stored in the new data address, and so on.
[0062] Optionally, the atomic operation request may also include atomic operands (or raw data). Raw data and / or initial data are used to perform the atomic operation.
[0063] The atomic operation can be any one or more of the exchange operation, addition operation, subtraction operation, sum operation, OR operation, XOR operation, increment operation, decrement operation, comparison operation, floating-point sum operation, and comparison sum operation mentioned above. It can also be an atomic operation not mentioned above, and the embodiments of this application do not limit it.
[0064] The result of an atomic operation is the result obtained after performing the atomic operation, such as the sum obtained after an addition operation, the difference obtained after a subtraction operation, and so on.
[0065] The atomic operation processing unit 110 is used to write the atomic operation results back to the auxiliary storage unit 120.
[0066] Optionally, the write address of the atomic operation result is the read address of the initial data. Optionally, the data address written to the atomic operation result is the data address carried in the atomic operation request, which is the address where the initial data was read. Optionally, the write address of the atomic operation result is the read address of one of the multiple initial data sets.
[0067] Optionally, when the initial data is read from data storage unit 130, the cache line storing the atomic operation result is determined from the auxiliary storage unit 120. This determination can be made using a pseudo-random algorithm from free cache lines, or based on the data address of the initial data. Free cache lines are either cache lines that do not store data, or cache lines whose stored data has expired. Alternatively, when the initial data is read from data storage unit 130, the cache line storing the atomic operation result is determined based on the write address of the atomic operation result indicated in the atomic operation request.
[0068] Optionally, when the initial data is read from auxiliary storage unit 120, the cache line storing the atomic operation result is a cache line in auxiliary storage unit 120 used to store the initial data. Alternatively, when the initial data is read from auxiliary storage unit 120, the cache line storing the atomic operation result is a cache line in auxiliary storage unit 120 determined from free cache lines using a pseudo-random algorithm. Alternatively, when the initial data is read from auxiliary storage unit 120, the cache line storing the atomic operation result is a cache line determined based on the write address of the atomic operation result indicated in the atomic operation request.
[0069] In summary, the atomic operation processing system provided in this application embodiment has two storage units, with the access latency of the auxiliary storage unit being lower than that of the data storage unit. Compared to traditional atomic operation processing systems that only have one data storage unit, this system additionally includes an auxiliary storage unit with lower access latency. This eliminates the need to spend excessive time waiting for data reading and writing during the execution of atomic operations, improving the execution efficiency of atomic operations, especially for multiple consecutive atomic operations. Furthermore, this system is a modification of a traditional atomic operation processing system, achieving improved execution efficiency with minor changes, and helps maintain a small area and low production cost.
[0070] Next, we will further introduce the processing procedure of the atomic manipulation system.
[0071] In some embodiments, the auxiliary storage unit 120 and the data storage unit 130 are respectively connected to the atomic operation processing unit 110. Optionally, the auxiliary storage unit 120 and the data storage unit 130 are connected.
[0072] In some embodiments, such as Figure 2 As shown, the auxiliary storage unit 120 is located between the atomic operation processing unit 110 and the data storage unit 130. That is, the atomic operation processing unit 110 is connected to the auxiliary storage unit 120, and the auxiliary storage unit 120 is connected to the data storage unit 130. Therefore, when reading initial data from the data storage unit 130, the auxiliary storage unit 120 supports receiving and storing the initial data.
[0073] Optionally, the auxiliary storage unit 120 is used to read initial data from the data storage unit 130 when the auxiliary storage unit 120 is not hit; and to store the initial data in the auxiliary storage unit 120.
[0074] For example, if the auxiliary storage unit 120 misses a data cache, it reads the initial data from the data storage unit 130 and stores the initial data in the auxiliary storage unit 120. The auxiliary storage unit 120 then transmits the initial data to the atomic operation processing unit 110.
[0075] In the case of reading initial data from data storage unit 130, auxiliary storage unit 120 also saves the read initial data so that subsequent operations on the initial data or on the data address corresponding to the initial data can be performed based on auxiliary storage unit 120, thereby accelerating the overall atomic operation process and improving the execution efficiency of atomic operations.
[0076] A reduction operation is a process that combines a set of data into a single value through a binary operation (such as addition, bitwise AND, or maximum value). An atomic reduction operation is one in which the binary operations within the reduction operation remain atomic.
[0077] In some embodiments, the atomic operation request corresponding to the atomic operation processing unit 110 includes multiple atomic operation requests. The multiple atomic operation requests are multiple sub-atomic operation requests corresponding to the atomic reduction operation and are all directed to a first address. The first address is the data address corresponding to each of the multiple atomic operation requests.
[0078] The atomic operation processing unit 110 is configured to perform an atomic operation based on the i-th atomic operation request among multiple atomic operation requests, obtain the i-th atomic operation result, and write the i-th atomic operation result back to the storage location corresponding to the first address in the auxiliary storage unit 120, where i is an integer greater than or equal to 1.
[0079] The atomic operation processing unit 110 is configured to: read the result of the i-th atomic operation from the auxiliary storage unit according to the first address based on the (i+1)-th atomic operation request among multiple atomic operation requests, wherein the result of the i-th atomic operation is the initial data corresponding to the (i+1)-th atomic operation request; perform atomic operations based on the (i+1)-th atomic operation request and the result of the i-th atomic operation to obtain the result of the (i+1)-th atomic operation; and write the result of the (i+1)-th atomic operation back to the storage location in the auxiliary storage unit corresponding to the first address.
[0080] For example, multiple atomic operation requests are represented as three atomic operation requests. For the first atomic operation request, if the auxiliary storage unit 120 misses, the auxiliary storage unit 120 reads initial data from the data storage unit 130, stores the initial data in the auxiliary storage unit 120, and sends the initial data to the atomic operation processing unit 110. The atomic operation processing unit 110 performs an atomic operation based on the initial data and the first atomic operation request to obtain the first atomic operation result, and writes the first atomic operation result back to the storage location in the auxiliary storage unit 120 corresponding to the first address. For the second atomic operation request, the atomic operation processing unit 110 reads the first atomic operation result from the auxiliary storage unit 120, where the first atomic operation result is the initial data corresponding to the second atomic operation request; performs an atomic operation based on the second atomic operation request and the first atomic operation result to obtain the second atomic operation result, and writes the second atomic operation result back to the storage location in the auxiliary storage unit 120 corresponding to the first address. In response to the third atomic operation request, the atomic operation processing unit 110 reads the result of the second atomic operation from the auxiliary storage unit 120; performs the atomic operation based on the third atomic operation request and the result of the second atomic operation to obtain the result of the third atomic operation, and writes the result of the third atomic operation back to the storage location in the auxiliary storage unit 120 corresponding to the first address.
[0081] Optionally, in addition to the data related to the first address mentioned above, atomic operations may also contain data related to other addresses. For example, besides reading the data corresponding to the first address, it is also necessary to read the data corresponding to the second address, third address, etc., to participate in the calculation of the atomic operation. Alternatively, the atomic operation processing unit 110 may also include multiple registers for storing operands corresponding to the atomic operation. These operands may be read into the registers during the compilation stage or other stages. These operands are temporary data that does not need to be stored in memory or read from memory.
[0082] For example, the calculation formula for multiple atomic operations is x = x + a + b + c, where the first atomic operation is x = x + a, the second atomic operation is x = x + b, and the third atomic operation is x = x + c. The x on the right-hand side of the equation for the first atomic operation is the initial data read from the first address. This initial data can be read from the auxiliary storage unit 120 or the data storage unit 130. The x on the left-hand side of the equation for the first atomic operation indicates that the result of the first atomic operation is stored at the first address or the storage location corresponding to the first address. The x on the right-hand side of the equation for the second atomic operation is the initial data read from the first address. Due to the execution of the first atomic operation, the second atomic operation is highly likely to obtain the result of the first atomic operation from the auxiliary storage unit 120, which is faster than reading the corresponding data from the data storage unit 130. The x on the left-hand side of the equation for the second atomic operation is similar to the x on the left-hand side of the equation for the first atomic operation; it is the result of the second atomic operation. After the second atomic operation is completed, it will also be stored at the first address or the storage location corresponding to the first address. The same applies to the third atomic operation.
[0083] It should be understood that the aforementioned multiple atomic operation requests can be multiple consecutive and sequentially executed atomic operation requests, or multiple sequentially executed atomic operation requests. That is, atomic operation requests targeting other addresses can also be interspersed among multiple atomic operation requests. For example, the execution order of the atomic operation processing unit 110 is: atomic operation request 1 corresponding to the first address, atomic operation request 2 corresponding to the first address, atomic operation request 3 corresponding to the second address, atomic operation request 4 corresponding to the second address, atomic operation request 5 corresponding to the first address, and so on.
[0084] Optionally, the aforementioned multiple atomic operation requests can be referred to as multiple atomic operation requests corresponding to atomic reduction operations.
[0085] The atomic operating system described above, when performing atomic reduction operations, utilizes auxiliary storage units to read and write data at the same address for multiple atomic operation requests. Since the access latency of the auxiliary storage unit is lower than that of the data storage unit, the execution of atomic operations does not require excessive time spent waiting for data reading and writing, thus improving the execution efficiency of atomic operations, especially for multiple consecutive atomic operations. Furthermore, this system is a modification of a traditional atomic operation processing system, achieving improved execution efficiency with minor changes while maintaining a small footprint and low production cost.
[0086] In some embodiments, the atomic operation processing unit 110 is configured to write the atomic operation result back to the auxiliary storage unit 120 and the data storage unit 130 when the initial data is read from the data storage unit 130.
[0087] In some embodiments, the atomic operation processing unit 110 is configured to write the atomic operation result back to the auxiliary storage unit 120 when the initial data is read from the auxiliary storage unit 120.
[0088] Optionally, the write address of the atomic operation result is the read address of the initial data. Optionally, the data address written to the atomic operation result is the data address carried in the atomic operation request, which is the address where the initial data was read. Optionally, the write address of the atomic operation result is the read address of one of the multiple initial data sets.
[0089] For example, if the initial data is read from data address 1 of data storage unit 130, the result of the atomic operation is written back to data address 1 of data storage unit 130. As another example, if the initial data is read from data address 1 of data storage unit 130, and the auxiliary storage unit 120 writes the initial data to data address 2, the result of the atomic operation is written back to data address 2 of the auxiliary storage unit 120, and the result of the atomic operation is also written back to data address 1 of data storage unit 130. Yet another example, if the initial data is read from data address 1 of data storage unit 130, and the auxiliary storage unit 120 writes the initial data to cache line 1, the result of the atomic operation is written back to cache line 1 of the auxiliary storage unit 120, and the result of the atomic operation is also written back to data address 1 of data storage unit 130. Alternatively, if the initial data is read from data address 3 of auxiliary storage unit 120, the result of the atomic operation is written back to data address 3 of auxiliary storage unit 120; or, if the initial data is read from cache line 2 of auxiliary storage unit 120, the result of the atomic operation is written back to cache line 2 of auxiliary storage unit 120. For example, if the initial data is read from cache line 3 of data storage unit 130, the result of the atomic operation is written back to cache line 3 of data storage unit 130.
[0090] It should be understood that the smallest operating unit in the auxiliary storage unit 120 and the data storage unit 130 can be a row or a block, that is, a row or a block mapped to in a many-to-one manner based on an address. Alternatively, the smallest operating unit can also be an address unit, that is, a storage location that can be addressed one-to-one based on an address. This application embodiment does not limit this. Optionally, the storage hierarchy of the auxiliary storage unit 120 and the data storage unit 130 may be the same or different. For example, both the auxiliary storage unit 120 and the data storage unit 130 may be caches; or, the auxiliary storage unit 120 may be a cache and the data storage unit 130 may be memory; or, both the auxiliary storage unit 120 and the data storage unit 130 may be memory.
[0091] In some embodiments, the auxiliary storage unit 120 includes an access index, which is used to locate data in the auxiliary storage unit 120 according to the data address corresponding to the atomic operation request. The access index is determined based on all or part of the addresses of the data storage unit 130. Data in the data storage unit 130 is also located using the access index. For the data storage unit 130, its access index is an address or data address. The structure and composition of the access index, or address, of the data storage unit 130 are related to the mapping technology of the data storage unit 130. The cache capacity is smaller than memory, and the content it stores is only a subset of the memory content. To put data into these storage units, a mapping function must be applied to locate the memory address in the cache; this process is called address mapping. After the data is loaded into the cache according to this mapping relationship, when the processor executes the program, it transforms the memory address in the program into the cache address; this transformation process is called address translation. Cache address mapping methods include direct mapping, fully associative mapping, and set-associative mapping. Direct mapping maps each data block in one storage unit to a unique location in another storage unit. For example, if the first storage unit has 8 cache lines, then data blocks 0, 8, 16, 24... in the second storage unit will be mapped to cache line 0, and similarly, data blocks 1, 9, 17... will be mapped to cache line 1. When the read order is data block 0-data block 8-data block 0-data block 8, since cache line 0 can only cache one data block at a time, a cache miss will occur when reading data block 8. That is, the required data block cannot be found in cache line 0 of the first storage unit, and the search must be conducted in the second storage unit. Fully associative mapping means that any data block in one storage unit can be mapped to any free location in another storage unit. For example, if the first storage unit has 8 cache lines, then any data block in the second storage unit can be mapped to any free cache line among these 8 cache lines. Set-associative mapping means that a storage unit is divided into several groups, each group containing multiple locations (called "paths"). Data blocks in another storage unit can only be mapped to a specific set, but can be stored arbitrarily within that set (direct mapping between sets, fully associative within a set). For example, the first storage unit includes N ways, and each way includes M sets. Each set contains N cache lines. For instance, there are two ways, way0 and way1, each with 8 lines, corresponding to 8 sets. Each set contains 2 cache lines, meaning line 0 of way0 and line 0 of way1 form a set. Thus, any two data blocks from the second storage unit (0, 8, 16, 24...) can be simultaneously stored in the two lines 0 of the first storage unit.
[0092] In direct mapping and set-associative mapping, the address is typically divided into three segments: Tag, Index, and Line offset. The Line offset indicates the offset of the address within the cache line; the Index indicates which set (in set-associative mapping) or line (in direct mapping); and the Tag distinguishes different blocks within the same set, or can be understood as determining whether a data block has been hit.
[0093] For example, data storage unit 130 is a 32-group, 16-way set-associative mapping memory unit. Auxiliary storage unit 120 is a fully associative mapping memory unit with 16 cache lines. For data storage unit 130, its address representation is typically the aforementioned Tag, Index, and Line offset. For auxiliary storage unit 120, a fully associative mapping memory unit, any data block in one memory unit can be mapped to any free location in another memory unit; theoretically, an "access index" is not needed (because full associativity allows storage in any location). However, to speed up hardware lookups or implement a cache inclusion strategy, a pseudo-index can be artificially introduced. This access index can be set based on the address of data storage unit 130. For example, all addresses of data storage unit 130 can be directly set as the access index of auxiliary storage unit 120. For example, the index and line offset of data storage unit 130 can be set as the access index of auxiliary storage unit 120, that is, the group / way and cache line offset of data storage unit 130 can be used as the access index. Another example is setting the index of data storage unit 130 as the access index of auxiliary storage unit 120.
[0094] Specifically, for the initial data read from the data storage unit 130, after the atomic operation is performed and the result is obtained, the atomic operation result is written back to the data storage unit 130 and the auxiliary storage unit 120. On one hand, writing the atomic operation result back to the auxiliary storage unit 120 ensures that subsequent atomic operations on the same address can be performed directly on the auxiliary storage unit 120, i.e., read from the auxiliary storage unit 120, thereby reducing the processing latency of subsequent atomic operations. On the other hand, writing the atomic operation result back to the data storage unit 130 ensures that the data in the data storage unit 130 is relatively new for a period of time, completing the persistent storage of the atomic operation result and preventing data loss due to anomalies in the auxiliary storage unit 120.
[0095] The auxiliary storage unit 120 has lower access latency and smaller storage capacity. Therefore, the data stored in the auxiliary storage unit 120 needs to be written back to the data storage unit 130 in a timely manner to ensure that at least one of the auxiliary storage unit 120 and the data storage unit 130 stores the latest data.
[0096] In some embodiments, the auxiliary storage unit 120 includes at least one cache line. The auxiliary storage unit 120 is configured to write data from the first cache line to be replaced in the at least one cache line back to the data storage unit 130 when at least one cache line stores data; and to write the initial data read into the first cache line to be replaced.
[0097] That is, when a cache line replacement occurs in the auxiliary storage unit 120, the data in the cache line to be replaced (i.e., the data in the first cache line) is written back to the data storage unit 130, and then the new data to be stored (i.e., the initial data read) is written to the first cache line.
[0098] Optionally, cache line replacement may occur in situations such as: at least one cache line contains data before or during the reading of initial data, or at least one cache line contains valid cache lines before or during the reading of initial data. A valid cache line is a cache line marked as valid.
[0099] Optionally, after the first cache line writes data back to the data storage unit 130, the first cache line is updated from a valid cache line to an invalid cache line. An invalid cache line is a cache line marked as invalid, used to indicate that new data can be stored.
[0100] Specifically, cache line replacement in the secondary storage unit ensures that at least one storage unit in both the secondary storage unit and the data storage unit contains the latest data. Cache line replacement occurs when all cache lines in at least one cache line contain data, or in other words, all contain valid data. This means that data is stored in the secondary storage unit as preferentially as possible, thereby leveraging the low access latency of the secondary storage unit to improve the execution efficiency of atomic operation requests, especially for multiple atomic operation requests targeting the same address.
[0101] In some embodiments, in order to further improve the execution efficiency of atomic operations, the execution process in the above-mentioned atomic operation processing system can be further improved, as shown below.
[0102] In some embodiments, the atomic operation processing unit 110 is configured to send query requests for the data addresses corresponding to the atomic operation requests to the auxiliary storage unit 120 and the data storage unit 130 respectively, so as to query the initial data in parallel; if the auxiliary storage unit 120 is hit, the initial data is obtained from the auxiliary storage unit 120; if the auxiliary storage unit 120 is not hit, the initial data is obtained from the data storage unit 130.
[0103] That is, the atomic operation processing unit 110 sends query requests for the same data address to both the auxiliary storage unit 120 and the data storage unit 130, enabling them to execute the queries in parallel. If the auxiliary storage unit 120 hits the query, the atomic operation processing unit 110 retrieves the initial data from it; the data returned by the data storage unit 130 can be discarded or used for consistency checks. If the auxiliary storage unit 120 misses the query, the atomic operation processing unit 110 waits and retrieves the initial data from the data storage unit 130; optionally, the auxiliary storage unit 120 also stores this initial data for use by subsequent atomic operation requests.
[0104] Optionally, if both auxiliary storage unit 120 and data storage unit 130 are hit, the atomic operation processing unit 110 obtains the initial data from auxiliary storage unit 120 and the initial data from data storage unit 130, respectively. If the initial data in auxiliary storage unit 120 and the initial data in data storage unit 130 are the same, an atomic operation is performed based on the initial data to obtain the atomic operation result. If the initial data in auxiliary storage unit 120 and the initial data in data storage unit 130 are different, an anomaly is reported.
[0105] Parallel querying minimizes the query time for initial data in the worst-case scenario, improving the execution efficiency of atomic operations.
[0106] In some embodiments, the atomic operation processing unit 110 first sends a query request to the auxiliary storage unit 120; if the auxiliary storage unit 120 is hit, it reads the initial data corresponding to the atomic operation request from the auxiliary storage unit 120; if the auxiliary storage unit 120 is not hit, it then sends a query request to the data storage unit 130 and retrieves the initial data from the data storage unit 130. This serial query method can save bandwidth resources as much as possible during the execution of atomic operations.
[0107] In some embodiments, the atomic operation processing system executes multiple atomic operation requests. For example, the atomic operation requests to be processed by the atomic operating system include a first atomic operation request and a second atomic operation request. The first atomic operation request is the atomic operation request executed before the second atomic operation request. The connection between the first atomic operation request and the second atomic operation request can be handled in the following manner.
[0108] Optionally, the auxiliary storage unit 120 is used to query the initial data corresponding to the second atomic operation request based on the second atomic operation request after the atomic operation result corresponding to the first atomic operation request has been written back.
[0109] Optionally, the auxiliary storage unit 120 is configured to, after the atomic operation result corresponding to the first atomic operation request is written back, query the initial data corresponding to the second atomic operation request based on the second atomic operation request. If the auxiliary storage unit 120 is hit, it returns the initial data of the second atomic operation request to the atomic operation processing unit 110. If the auxiliary storage unit 120 is not hit, the data storage unit 130 queries the initial data corresponding to the second atomic operation request based on the second atomic operation request.
[0110] Optionally, after the atomic operation result corresponding to the first atomic operation request is written back, the auxiliary storage unit 120 and the data storage unit 130 query the initial data corresponding to the second atomic operation request in parallel based on the second atomic operation request.
[0111] In other words, multiple atomic operation requests are executed serially. Only after all operations corresponding to the first atomic operation request have been completed does the operation for the second atomic operation request begin. This ensures that only one instruction occupies the atomic operation processing system at any given time, reducing the execution complexity of the atomic operation processing system, improving overall stability, and minimizing overall energy consumption per unit time by eliminating the need for additional control operations.
[0112] It should be understood that the operation addresses corresponding to the first and second atomic operation requests are not limited. That is, the following situations exist: the write address of the first atomic operation request is a first address, and the write address of the second atomic operation request is a second address; or, the write addresses of both the first and second atomic operation requests are the first address. In other words, the first and second atomic operation requests can be two sub-operation requests within a single atomic reduction operation, or they may not belong to the same atomic reduction operation. In other embodiments, the third atomic operation request is an atomic operation request executed before the fourth atomic operation request, and the initial data of the fourth atomic operation request is the atomic operation result corresponding to the third atomic operation request. The atomic operation processing unit 110 is used to execute the fourth atomic operation request when the atomic operation result for the third atomic operation request is triggered and written back to the auxiliary storage unit 120.
[0113] For example, during the cycle that triggers the write-back of the atomic operation result for the third atomic operation request to the auxiliary storage unit 120, the auxiliary storage unit 120 is executed to read the initial data corresponding to the fourth atomic operation request. That is, the execution of the fourth atomic operation can begin after the atomic operation result for the third atomic operation request has been completed and the atomic operation result has been obtained, rather than after the atomic operation for the third atomic operation request has been completed.
[0114] Optionally, the third atomic operation request and the fourth atomic operation request are two sub-operation requests in an atomic reduction operation, and the execution order of the third atomic operation request and the fourth atomic operation request is adjacent and executed sequentially. When the atomic operation result of the third atomic operation request is triggered to be written back to the auxiliary storage unit 120, there is no need to trigger a query operation for the fourth atomic operation request, or no need to trigger a query operation for the first address corresponding to the fourth atomic operation request. The atomic operation result corresponding to the third atomic operation request is directly used to execute the fourth atomic operation request, further accelerating the execution efficiency of the fourth atomic operation request.
[0115] Optionally, the auxiliary storage unit 120 is used to query the initial data corresponding to the second atomic operation request based on the second atomic operation request after the initial data corresponding to the first atomic operation request is sent to the atomic operation processing unit.
[0116] Optionally, the auxiliary storage unit 120 is configured to, after the initial data corresponding to the first atomic operation request is sent to the atomic operation processing unit, query the initial data corresponding to the second atomic operation request based on the second atomic operation request. If the auxiliary storage unit 120 is hit, it returns the initial data of the second atomic operation request to the atomic operation processing unit 110. If the auxiliary storage unit 120 is not hit, the data storage unit 130 queries the initial data corresponding to the second atomic operation request based on the second atomic operation request.
[0117] Optionally, after the initial data corresponding to the first atomic operation request is sent to the atomic operation processing unit, the auxiliary storage unit 120 and the data storage unit 130 query the initial data corresponding to the second atomic operation request in parallel based on the second atomic operation request.
[0118] In other words, multiple atomic operation requests are executed in parallel. While the first atomic operation request is being computed in the atomic operation processing unit, the data query for the second atomic operation request can be initiated simultaneously, meaning that the auxiliary storage unit and / or data storage unit also remain operational. This allows the auxiliary storage unit and / or data storage unit, which were originally idle and waiting, to be effectively utilized. Especially when memory access latency (such as cache misses) is long, the atomic operation processing unit can continuously process subsequent tasks with prepared data, thus "hiding" this waiting time. For example, there are three atomic operation requests. During the execution of the first atomic operation request, the initial data query for the second atomic operation request has been completed, and the initial data query for the third atomic operation request is in progress. The initial data query for the third atomic operation request needs to be performed through the data storage unit, which requires a significant access latency. At this point, after the first atomic operation request is completed, the atomic operation processing unit can first execute the second atomic operation request based on the initial data of the second atomic operation request that has been queried. During the execution of the first atomic operation request and the second atomic operation request, the query for the initial data of the third atomic operation request is completed. From an overall perspective, the entire atomic operation processing system does not experience an increase in overall processing time due to the longer access latency corresponding to the third atomic operation request, thus achieving the hiding of the waiting time for querying the initial data of the atomic operation request.
[0119] In some embodiments, the atomic operation processing unit 110 is a unit that is independent of the processor.
[0120] Optionally, the processor sends an atomic operation request to the atomic operation processing unit 110; if the auxiliary storage unit 120 is hit, the atomic operation processing unit 110 reads initial data from the auxiliary storage unit 120, which is read from the data storage unit 130 to the auxiliary storage unit 120 based on historical atomic operation requests; if the auxiliary storage unit 120 is not hit, the auxiliary storage unit 120 reads the initial data from the data storage unit 130 and stores the initial data in the auxiliary storage unit 120; the atomic operation processing unit 110 performs atomic operations based on the atomic operation request and the initial data, obtains the atomic operation result, and writes the atomic operation result to the auxiliary storage unit 120.
[0121] Optionally, the processor is a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning. This application uses a GPU as an example, but does not limit the specific type of processor.
[0122] In this system, the atomic operation processing unit that supports atomic operations is set up independently of the processor. When performing atomic operations, the atomic operation processing unit only needs to interact with the auxiliary storage unit or the data storage unit. Compared with the atomic operation processing unit located in the processor in related technologies, which needs to communicate with the auxiliary storage unit or the data storage unit through an additional transmission path, the transmission latency of this system is lower than that of related technologies. The lower transmission latency can indirectly improve the operating efficiency of the atomic operation processing system.
[0123] In some embodiments, the atomic operation processing system provided in this application is as follows: Figure 3 As shown. The atomic operation processing system includes an atomic operation processing unit 110, an auxiliary storage unit 120, a data storage unit 130, and a cache tag storage 140.
[0124] The details of the atomic operation processing unit 110, auxiliary storage unit 120, and data storage unit 130 can be found in the above description. Figure 2 The relevant content is shown below. The cache tag storage 140 is used to generate and send query instructions to the auxiliary storage unit 120 and / or the data storage unit 130 in response to query requests from the atomic operation processing unit 110.
[0125] For example, for an atomic operation request 51, the atomic operation processing unit 110 can send query requests for data address 53 to the auxiliary storage unit 120 and the data storage unit 130 respectively through the cache tag storage 140 to query the data 54 before the atomic operation in parallel. If the auxiliary storage unit 120 is hit, the atomic operation processing unit 110 obtains the data 54 before the atomic operation from the auxiliary storage unit 120; this data 54 before the atomic operation is read from the data storage unit 130 to the auxiliary storage unit 120 based on historical atomic operation requests. If the auxiliary storage unit 120 is not hit, the atomic operation processing unit 110 obtains the data 54 before the atomic operation from the data storage unit 130 and can store the data 54 before the atomic operation in the auxiliary storage unit 120. The cache tag storage 140 is also used to send the atomic operator 52 to the atomic operation processing unit 110. The atomic operation processing unit 110 calculates the atomic operation result 55 based on the atomic operator 52 and the data 54 before the atomic operation; or, it calculates the atomic operation result 55 based on the atomic operator 52, the data 54 before the atomic operation, and the data 56 before the atomic operation.
[0126] Here, atomic operation data 56 refers to data transferred from an external source into the atomic operating system. For example, atomic operation data 56 can be data carried in an atomic operation request, or data read from a register based on the atomic operation request. Atomic operation data 56 can also be called raw data. Data before atomic operation 54 is data read from auxiliary storage unit 120 and / or data storage unit 130. Typically, the address corresponding to atomic operation data 54 is the write address of the atomic operation result 55. That is, atomic operation data 54 can also be called read data.
[0127] In the parallel query embodiment, the cache tag storage 140 responds to the query request from the atomic operation processing unit 110 by sending a query instruction carrying the data address 53 to the auxiliary storage unit 120 and the data storage unit 130 in parallel. When the auxiliary storage unit 120 is hit, the atomic operation processing unit 110 retrieves the data 54 before the atomic operation from the auxiliary storage unit 120; when the auxiliary storage unit 120 is not hit, the atomic operation processing unit 110 retrieves the data 54 before the atomic operation from the data storage unit 130, and the auxiliary storage unit 120 can store the data 54 before the atomic operation.
[0128] For example, data storage unit 130 is a 32-set, 16-way set-associative cache with a cache line size of 128 bytes. The bit width of data storage unit 130 is 32 bytes. Auxiliary storage unit 120 is fully associative with 16 cache lines, each with a cache line size of 32 bytes.
[0129] The pipeline length for storing 140 cached tags is 1 cycle.
[0130] The auxiliary storage unit 120 consists of a register array, with a query cycle of 1 cycle and a data access cycle of 1 cycle.
[0131] The data storage unit 130 is organized by static random access memory, with an access latency of 3 cycles.
[0132] The computation unit has a 1-cycle delay for integer operations and a 3-cycle delay for floating-point operations.
[0133] The access request sequence is as follows: Request 1: Floating-point atomic operation at address 0; Request 2: Integer atomic operations at address 1; Request 3: Integer atomic operations at address 1; Request 4: Floating-point atomic operation at address 0; Request 5: Atomic integer operations at address 1; Request 6: Floating-point atomic operations at address 0.
[0134] Then compare with auxiliary storage units (i.e. Figure 2 and Figure 3 (corresponding structure) and no auxiliary storage unit (i.e.) Figure 1 For the corresponding structure scenarios, the processing flow for the two scenarios is as follows: Figure 4 and Figure 5 As shown in the comparison.
[0135] Figure 4 For scenarios without auxiliary storage units: the total time is 30 cycles, resulting in 12 data storage unit reads and writes. Figure 5 For scenarios with auxiliary storage units: the total time is 25 cycles, resulting in 4 data storage unit read / write operations.
[0136] In other words, the configuration of the auxiliary storage unit 120 greatly reduces the frequency of reading and writing to the data storage unit, thereby reducing the overall atomic operation processing time and improving the execution efficiency of atomic operations.
[0137] Figure 6 A flowchart illustrating an exemplary embodiment of the atomic operation processing method provided in this application is shown. The method comprises, as... Figure 2 The atomic operation processing system shown is executed. The atomic operation processing system includes an atomic operation processing unit, an auxiliary storage unit, and a data storage unit.
[0138] Optionally, the read / write speed of the secondary storage unit is faster than that of the data storage unit. Alternatively, the storage capacity of the secondary storage unit is smaller than that of the data storage unit. Or, the access speed of the secondary storage unit is faster than that of the data storage unit. Or, the access latency of the secondary storage unit is lower than that of the data storage unit.
[0139] Optionally, the atomic operation processing unit is connected to the auxiliary storage unit, and the auxiliary storage unit is connected to the data storage unit.
[0140] Step 210: When the atomic operation processing unit hits the auxiliary storage unit, it reads the initial data from the auxiliary storage unit. This initial data is read from the data storage unit to the auxiliary storage unit based on the historical atomic operation requests.
[0141] In this application, the initial data is read from the data storage unit to the auxiliary storage unit based on historical atomic operation requests. It can also be understood that the initial data is read from the data storage unit to the auxiliary storage unit based on the previous atomic operation request. This application does not limit this.
[0142] It should be understood that the initial data read from the auxiliary storage unit or the data storage unit can be one initial data or multiple initial data. That is, if the atomic operation processing unit finds a match in the auxiliary storage unit, it reads at least one initial data from the auxiliary storage unit. Optionally, if the atomic operation processing unit does not find a match in the auxiliary storage unit, it reads at least one initial data from the data storage unit.
[0143] Optionally, in addition to the corresponding initial data (i.e., read data), an atomic operation request may also contain corresponding raw data, such as data stored in the registers of the atomic operation processing unit. This raw data participates in the calculation process of the atomic operation. The source of the raw data is usually carried by the atomic operation request itself and does not need to be read from the storage unit.
[0144] For example, the atomic operation request carries a data address. The atomic operation processing unit determines whether the auxiliary storage unit stores the data value corresponding to the data address based on the data address. If the auxiliary storage unit stores the data value corresponding to the data address, it indicates that the auxiliary storage unit has a hit. The atomic operation processing unit reads the initial data from the auxiliary storage unit according to the data address. This initial data is read from the data storage unit to the auxiliary storage unit based on historical atomic operation requests. If the auxiliary storage unit does not store the data value corresponding to the data address, it indicates that the auxiliary storage unit has a miss. In some embodiments, the auxiliary storage unit reads the initial data from the data storage unit, stores the initial data in the auxiliary storage unit, and sends the initial data to the atomic operation processing unit.
[0145] Step 220: The atomic operation processing unit performs atomic operations based on the atomic operation request and initial data to obtain the atomic operation results.
[0146] The initial data is the data value stored at the data address corresponding to the atomic operation request before the atomic operation is performed. The initial data is used to perform the atomic operation. For example, if the atomic operation request corresponds to an addition operation, the initial data is used as the addend; if the atomic operation request corresponds to a subtraction operation, the initial data is used as the subtrahend and / or minuend; if the atomic operation request corresponds to a swap operation, the initial data is the data value waiting to be stored at the new data address, and so on.
[0147] Optionally, an atomic operation request can also be called an atomic operation instruction. The data address corresponding to the atomic operation request is the opcode in the atomic operation request. Optionally, the atomic operation request may also carry atomic operands (or raw data). Raw data and / or initial data are used to perform the atomic operation.
[0148] The atomic operation can be any one or more of the exchange operation, addition operation, subtraction operation, sum operation, OR operation, XOR operation, increment operation, decrement operation, comparison operation, floating-point sum operation, and comparison sum operation mentioned above. It can also be an atomic operation not mentioned above, and the embodiments of this application do not limit it.
[0149] The result of an atomic operation is the result obtained after performing the atomic operation, such as the sum obtained after an addition operation, the difference obtained after a subtraction operation, and so on. Optionally, an auxiliary storage unit is used to read initial data from the data storage unit in the event of a miss; and to store the initial data in the auxiliary storage unit.
[0150] Step 230: The atomic operation processing unit writes the atomic operation result back to the auxiliary storage unit.
[0151] Optionally, the write address of the atomic operation result is the read address of the initial data. Optionally, the data address written to the atomic operation result is the data address carried in the atomic operation request, which is the address where the initial data was read. Optionally, the write address of the atomic operation result is the read address of one of the multiple initial data sets.
[0152] Optionally, when the initial data is read from the data storage unit, the cache line storing the atomic operation result is determined from the auxiliary storage unit. This determination can be made using a pseudo-random algorithm from free cache lines, or based on the data address of the initial data. Free cache lines are either cache lines that do not store data, or cache lines whose stored data has expired. Alternatively, when the initial data is read from the data storage unit, the cache line storing the atomic operation result is determined based on the write address of the atomic operation result indicated in the atomic operation request.
[0153] Optionally, if the initial data is read from the auxiliary storage unit, the cache line storing the atomic operation result is a cache line in the auxiliary storage unit used to store the initial data. Alternatively, if the initial data is read from the auxiliary storage unit, the cache line storing the atomic operation result is a cache line in the auxiliary storage unit determined from free cache lines using a pseudo-random algorithm. Alternatively, if the initial data is read from the auxiliary storage unit, the cache line storing the atomic operation result is a cache line determined based on the write address of the atomic operation result indicated in the atomic operation request.
[0154] In summary, the method provided in this application provides two storage units, with the access latency of the auxiliary storage unit being lower than that of the data storage unit. Compared to traditional atomic operation processing systems that only have one data storage unit, this system additionally uses an auxiliary storage unit with lower access latency. This eliminates the need to spend excessive time waiting for data reading and writing during the execution of atomic operations, improving the execution efficiency of atomic operations, especially for multiple consecutive atomic operations. Furthermore, this system is a modification of a traditional atomic operation processing system, achieving improved execution efficiency with minor changes, while ensuring that the modified atomic operating system maintains a small area and low production cost.
[0155] In some embodiments, the auxiliary storage unit and the data storage unit are respectively connected to the atomic operation processing unit. Optionally, the auxiliary storage unit and the data storage unit are connected together.
[0156] In some embodiments, such as Figure 2 As shown, the auxiliary storage unit is located between the atomic operation processing unit and the data storage unit; that is, the atomic operation processing unit is connected to the auxiliary storage unit, and the auxiliary storage unit is connected to the data storage unit. Therefore, when reading initial data from the data storage unit, the auxiliary storage unit supports receiving and storing the initial data.
[0157] Optionally, the auxiliary storage unit is used to read initial data from the data storage unit in the event of a miss in the auxiliary storage unit; and to store the initial data in the auxiliary storage unit.
[0158] For example, if the secondary storage unit misses a data cache miss, it reads initial data from the data storage unit and stores the initial data in the secondary storage unit. Then, the secondary storage unit transmits the initial data to the atomic operation processing unit.
[0159] In the case of reading initial data from the data storage unit, the auxiliary storage unit also saves the read initial data so that subsequent operations on the initial data or on the data address corresponding to the initial data can be performed based on the auxiliary storage unit, thereby accelerating the overall atomic operation process and improving the execution efficiency of atomic operations.
[0160] A reduction operation is a process that combines a set of data into a single value through a binary operation (such as addition, bitwise AND, or maximum value). An atomic reduction operation is one in which the binary operations within the reduction operation remain atomic.
[0161] In some embodiments, the atomic operation request corresponding to the atomic operation processing unit includes multiple atomic operation requests. The multiple atomic operation requests are multiple sub-atomic operation requests corresponding to the atomic reduction operation and are all directed to a first address. The first address is the data address corresponding to each of the multiple atomic operation requests.
[0162] The atomic operation processing unit is used to perform an atomic operation based on the i-th atomic operation request among multiple atomic operation requests, to obtain the i-th atomic operation result; and to write the i-th atomic operation result back to the storage location corresponding to the first address in the auxiliary storage unit, where i is an integer greater than or equal to 1.
[0163] An atomic operation processing unit is configured to: read the result of the i-th atomic operation from an auxiliary storage unit based on the (i+1)-th atomic operation request among multiple atomic operation requests, wherein the result of the i-th atomic operation is the initial data corresponding to the (i+1)-th atomic operation request; perform an atomic operation based on the (i+1)-th atomic operation request and the result of the i-th atomic operation to obtain the result of the (i+1)-th atomic operation; and write the result of the (i+1)-th atomic operation back to the storage location in the auxiliary storage unit corresponding to the first address.
[0164] For example, multiple atomic operation requests are represented as three atomic operation requests. For the first atomic operation request, if the secondary storage unit misses, the secondary storage unit reads initial data from the data storage unit, stores the initial data in the secondary storage unit, and sends the initial data to the atomic operation processing unit. The atomic operation processing unit performs the atomic operation based on the initial data and the first atomic operation request, obtains the result of the first atomic operation, and writes the result back to the storage location in the secondary storage unit corresponding to the first address. For the second atomic operation request, the atomic operation processing unit reads the result of the first atomic operation from the secondary storage unit, where the result of the first atomic operation is the initial data corresponding to the second atomic operation request; it performs the atomic operation based on the second atomic operation request and the result of the first atomic operation, obtains the result of the second atomic operation, and writes the result back to the storage location in the secondary storage unit corresponding to the first address. In response to the request for the third atomic operation, the atomic operation processing unit reads the result of the second atomic operation from the auxiliary storage unit; performs the atomic operation based on the request for the third atomic operation and the result of the second atomic operation to obtain the result of the third atomic operation, and writes the result of the third atomic operation back to the storage location in the auxiliary storage unit corresponding to the first address.
[0165] Optionally, in addition to the data related to the first address mentioned above, atomic operations may also contain data related to other addresses. For example, besides reading the data corresponding to the first address, it may also be necessary to read data corresponding to the second address, third address, etc., to participate in the calculation of the atomic operation. Alternatively, the atomic operation processing unit may also include multiple registers for storing operands corresponding to the atomic operation. These operands may be read into the registers during the compilation stage or other stages, and these operands are temporary data that does not need to be stored in memory or read from memory.
[0166] For example, the calculation formula for multiple atomic operations is x = x + a + b + c, where the first atomic operation is x = x + a, the second atomic operation is x = x + b, and the third atomic operation is x = x + c. The x on the right-hand side of the equation for the first atomic operation represents the initial data read from the first address. This initial data can be read from an auxiliary storage unit or a data storage unit. The x on the left-hand side of the equation for the first atomic operation indicates that the result of the first atomic operation is stored at the first address or the storage location corresponding to the first address. The x on the right-hand side of the equation for the second atomic operation represents the initial data read from the first address. Due to the execution of the first atomic operation, the second atomic operation is highly likely to obtain the result of the first atomic operation from the auxiliary storage unit, which is faster than reading the corresponding data from the data storage unit. The x on the left-hand side of the equation for the second atomic operation is similar to the x on the left-hand side of the equation for the first atomic operation; it represents the result of the second atomic operation. After the second atomic operation is completed, it will also be stored at the first address or the storage location corresponding to the first address. The same applies to the third atomic operation.
[0167] It should be understood that the aforementioned multiple atomic operation requests can be multiple consecutive and sequentially executed atomic operation requests, or multiple sequentially executed atomic operation requests. That is, atomic operation requests targeting other addresses can also be interspersed among multiple atomic operation requests. For example, the execution order of the atomic operation processing unit is: atomic operation request 1 corresponding to the first address, atomic operation request 2 corresponding to the first address, atomic operation request 3 corresponding to the second address, atomic operation request 4 corresponding to the second address, atomic operation request 5 corresponding to the first address, and so on.
[0168] Optionally, the aforementioned multiple atomic operation requests can be referred to as multiple atomic operation requests corresponding to atomic reduction operations.
[0169] The atomic operating system described above, when performing atomic reduction operations, utilizes auxiliary storage units to read and write data at the same address for multiple atomic operation requests. Since the access latency of the auxiliary storage unit is lower than that of the data storage unit, the execution of atomic operations does not require excessive time spent waiting for data reading and writing, thus improving the execution efficiency of atomic operations, especially for multiple consecutive atomic operations. Furthermore, this system is a modification of a traditional atomic operation processing system, achieving improved execution efficiency with minor changes while maintaining a small footprint and low production cost.
[0170] In some embodiments, the atomic operation processing unit is configured to write the atomic operation result back to the auxiliary storage unit and the data storage unit when the initial data is read from the data storage unit.
[0171] In some embodiments, the atomic operation processing unit is configured to write the atomic operation result back to the auxiliary storage unit when the initial data is read from the auxiliary storage unit.
[0172] Optionally, the write address of the atomic operation result is the read address of the initial data. Optionally, the data address written to the atomic operation result is the data address carried in the atomic operation request, which is the address where the initial data was read. Optionally, the write address of the atomic operation result is the read address of one of the multiple initial data sets.
[0173] For example, if the initial data is read from data address 1 of the data storage unit, the result of the atomic operation is written back to data address 1 of the data storage unit. As another example, if the initial data is read from data address 1 of the data storage unit, and the secondary storage unit writes the initial data to data address 2, the result of the atomic operation is written back to data address 2 of the secondary storage unit, and also to data address 1 of the data storage unit. Yet another example, if the initial data is read from data address 1 of the data storage unit, and the secondary storage unit writes the initial data to cache line 1, the result of the atomic operation is written back to cache line 1 of the secondary storage unit, and also to data address 1 of the data storage unit. Alternatively, if the initial data is read from data address 3 of the secondary storage unit, the result of the atomic operation is written back to data address 3 of the secondary storage unit; or, if the initial data is read from cache line 2 of the secondary storage unit, the result of the atomic operation is written back to cache line 2 of the secondary storage unit. For example, if the initial data is read from cache line 3 of the data storage unit, the result of the atomic operation is written back to cache line 3 of the data storage unit.
[0174] It should be understood that the smallest unit of operation in the auxiliary storage unit and the data storage unit can be a row or a block, that is, a row or a block mapped to in a many-to-one manner based on an address. Alternatively, the smallest unit of operation can also be an address unit, that is, a storage location that can be addressed one-to-one based on an address. This application does not limit this. Optionally, the storage hierarchy of the auxiliary storage unit and the data storage unit can be the same or different. For example, both the auxiliary storage unit and the data storage unit are caches; or, the auxiliary storage unit is a cache and the data storage unit is memory; or, both the auxiliary storage unit and the data storage unit are memory.
[0175] In some embodiments, the auxiliary storage unit includes an access index, which is used to locate data in the auxiliary storage unit according to the data address corresponding to the atomic operation request. The access index is determined based on all or part of the addresses of the data storage unit. Data in the data storage unit is also located using the access index; for the data storage unit, its access index is an address or data address. The structure and composition of the access index, or address, of the data storage unit are related to the mapping technology of the data storage unit. The cache capacity is smaller than memory, storing only a subset of the memory content. To place data into these storage units, a mapping function must be applied to locate the memory address in the cache; this process is called address mapping. After data is loaded into the cache according to this mapping relationship, when the processor executes the program, it transforms the memory address in the program into the cache address; this transformation process is called address translation. Cache address mapping methods include direct mapping, fully associative mapping, and set-associative mapping. Direct mapping maps each data block in one storage unit to a unique location in another storage unit. For example, if the first storage unit has 8 cache lines, then data blocks 0, 8, 16, 24... in the second storage unit will be mapped to cache line 0, and similarly, data blocks 1, 9, 17... will be mapped to cache line 1. When the read order is data block 0-data block 8-data block 0-data block 8, since cache line 0 can only cache one data block at a time, a cache miss will occur when reading data block 8. That is, the required data block cannot be found in cache line 0 of the first storage unit, and the search must be conducted in the second storage unit. Fully associative mapping means that any data block in one storage unit can be mapped to any free location in another storage unit. For example, if the first storage unit has 8 cache lines, then any data block in the second storage unit can be mapped to any free cache line among these 8 cache lines. Set-associative mapping means that a storage unit is divided into several groups, each group containing multiple locations (called "paths"). Data blocks in another storage unit can only be mapped to a specific set, but can be stored arbitrarily within that set (direct mapping between sets, fully associative within a set). For example, the first storage unit includes N ways, and each way includes M sets. Each set contains N cache lines. For instance, there are two ways, way0 and way1, each with 8 lines, corresponding to 8 sets. Each set contains 2 cache lines, meaning line 0 of way0 and line 0 of way1 form a set. Thus, any two data blocks from the second storage unit (0, 8, 16, 24...) can be simultaneously stored in the two lines 0 of the first storage unit.
[0176] In direct mapping and set-associative mapping, the address is typically divided into three segments: Tag, Index, and Line offset. The Line offset indicates the offset of the address within the cache line; the Index indicates which set (in set-associative mapping) or line (in direct mapping); and the Tag is used to distinguish different blocks within the same set, or to determine whether a data block has been hit.
[0177] For example, consider a 32-set, 16-way set-associative memory cell. The secondary memory cell is a fully associative memory cell with 16 cache lines. The address representation for the data storage cell typically uses the aforementioned Tag, Index, and Line offset. For the fully associative secondary memory cell, any data block in one memory cell can be mapped to any free location in another memory cell; theoretically, an "access index" is unnecessary (because full associativity allows storage at any location). However, to speed up hardware lookups or implement cache inclusion strategies, a pseudo-index can be artificially introduced. This access index can be set based on the address of the data storage cell. For example, the entire address of the data storage cell can be directly set as the access index of the secondary memory cell. Another example is setting the Index and Line offset of the data storage cell as the access index of the secondary memory cell; that is, using the set / way and cache line offset of the data storage cell as the access index. Yet another example is setting the Index of the data storage cell as the access index of the secondary memory cell.
[0178] Specifically, for the initial data read from the data storage unit, after performing the atomic operation and obtaining the result, the result is written back to both the data storage unit and the auxiliary storage unit. On one hand, writing the atomic operation result back to the auxiliary storage unit ensures that subsequent atomic operations on the same address can be performed directly from the auxiliary storage unit, reducing processing latency. On the other hand, writing the atomic operation result back to the data storage unit guarantees that the data in the data storage unit is relatively new for a period of time, thus achieving persistent storage of the atomic operation result and preventing data loss due to anomalies in the auxiliary storage unit.
[0179] Secondary storage units have lower access latency and smaller storage capacity. Therefore, data stored in secondary storage units needs to be written back to the data storage unit in a timely manner to ensure that at least one of the secondary storage units and the data storage unit stores the latest data.
[0180] In some embodiments, the auxiliary storage unit includes at least one cache line. The auxiliary storage unit is configured to, when at least one cache line stores data, write data from a first cache line to be replaced back to the data storage unit; and write initial data read into the first cache line to be replaced.
[0181] In other words, when a cache line replacement occurs in the auxiliary storage unit, the data in the cache line to be replaced (i.e., the data in the first cache line) is written back to the data storage unit, and then the new data to be stored (i.e., the initial data read) is written to the first cache line.
[0182] Optionally, cache line replacement may occur in situations such as: at least one cache line contains data before or during the reading of initial data, or at least one cache line contains valid cache lines before or during the reading of initial data. A valid cache line is a cache line marked as valid.
[0183] Optionally, after the first cache line writes the data back to the data storage unit, the first cache line is updated from a valid cache line to an invalid cache line. An invalid cache line is a cache line marked as invalid, used to indicate that new data can be stored.
[0184] Specifically, cache line replacement in the secondary storage unit ensures that at least one storage unit in both the secondary storage unit and the data storage unit contains the latest data. Cache line replacement occurs when all cache lines in at least one cache line contain data, or in other words, all contain valid data. This means that data is stored in the secondary storage unit as preferentially as possible, thereby leveraging the low access latency of the secondary storage unit to improve the execution efficiency of atomic operation requests, especially for multiple atomic operation requests targeting the same address.
[0185] In some embodiments, in order to further improve the execution efficiency of atomic operations, the execution process in the above-mentioned atomic operation processing system can be further improved, as shown below.
[0186] In some embodiments, the atomic operation processing unit is configured to send query requests for the data address corresponding to the atomic operation request to the auxiliary storage unit and the data storage unit respectively, so as to query the initial data in parallel; if the auxiliary storage unit is hit, the initial data is obtained from the auxiliary storage unit; if the auxiliary storage unit is not hit, the initial data is obtained from the data storage unit.
[0187] In other words, the atomic operation processing unit sends query requests for the same data address to both the auxiliary storage unit and the data storage unit, enabling them to execute the queries in parallel. If the auxiliary storage unit is hit, the atomic operation processing unit retrieves the initial data from it; the data returned by the data storage unit can be discarded or used for consistency checks. If the auxiliary storage unit is not hit, the atomic operation processing unit waits and retrieves the initial data from the data storage unit; optionally, the auxiliary storage unit also stores this initial data in its own storage unit for use by subsequent atomic operation requests.
[0188] Optionally, if both the auxiliary storage unit and the data storage unit are hit, the atomic operation processing unit obtains the initial data from the auxiliary storage unit and the initial data from the data storage unit, respectively. If the initial data in the auxiliary storage unit and the initial data in the data storage unit are the same, the atomic operation is performed based on the initial data to obtain the atomic operation result. If the initial data in the auxiliary storage unit and the initial data in the data storage unit are different, an exception is reported.
[0189] Parallel querying minimizes the query time for initial data in the worst-case scenario, improving the execution efficiency of atomic operations.
[0190] In some embodiments, the atomic operation processing unit first sends a query request to the auxiliary storage unit; if the auxiliary storage unit is matched, it reads the initial data corresponding to the atomic operation request from the auxiliary storage unit; if the auxiliary storage unit is not matched, it sends a query request to the data storage unit and retrieves the initial data from the data storage unit. This serial query method can save bandwidth resources as much as possible during the execution of atomic operations.
[0191] In some embodiments, the atomic operation processing system executes multiple atomic operation requests. For example, the atomic operation requests to be processed by the atomic operating system include a first atomic operation request and a second atomic operation request. The first atomic operation request is the atomic operation request executed before the second atomic operation request. The connection between the first atomic operation request and the second atomic operation request can be handled in the following manner.
[0192] Optionally, the auxiliary storage unit is used to query the initial data corresponding to the second atomic operation request based on the second atomic operation request after the atomic operation result corresponding to the first atomic operation request has been written back.
[0193] Optionally, the auxiliary storage unit is used to query the initial data corresponding to the second atomic operation request based on the second atomic operation request after the atomic operation result corresponding to the first atomic operation request has been written back. If the auxiliary storage unit is hit, the auxiliary storage unit returns the initial data of the second atomic operation request to the atomic operation processing unit. If the auxiliary storage unit is not hit, the data storage unit queries the initial data corresponding to the second atomic operation request based on the second atomic operation request.
[0194] Optionally, after the atomic operation result corresponding to the first atomic operation request is written back, the auxiliary storage unit and the data storage unit query the initial data corresponding to the second atomic operation request in parallel based on the second atomic operation request.
[0195] In other words, multiple atomic operation requests are executed serially. Only after all operations corresponding to the first atomic operation request have been completed does the operation for the second atomic operation request begin. This ensures that only one instruction occupies the atomic operation processing system at any given time, reducing the execution complexity of the atomic operation processing system, improving overall stability, and minimizing overall energy consumption per unit time by eliminating the need for additional control operations.
[0196] It should be understood that the operation addresses corresponding to the first and second atomic operation requests are not limited. That is, the following situations exist: the write address of the first atomic operation request is the first address, and the write address of the second atomic operation request is the second address; or, the write addresses of both the first and second atomic operation requests are the first address. In other words, the first and second atomic operation requests can be two sub-operation requests within a single atomic reduction operation, or they may not belong to the same atomic reduction operation. In other embodiments, the third atomic operation request is an atomic operation request executed before the fourth atomic operation request, and the initial data of the fourth atomic operation request is the atomic operation result corresponding to the third atomic operation request. The atomic operation processing unit is used to execute the fourth atomic operation request when the atomic operation result for the third atomic operation request is triggered and written back to the auxiliary storage unit.
[0197] For example, during the cycle that triggers the writing back of the atomic operation result for the third atomic operation request to the auxiliary storage unit, the auxiliary storage unit reads the initial data corresponding to the fourth atomic operation request. That is, the execution of the fourth atomic operation can begin, such as starting the query of the initial data for the fourth atomic operation, not after the writing back of the atomic operation result for the third atomic operation request is completed and the atomic operation result is obtained, rather than after the atomic operation for the third atomic operation request is completed.
[0198] Optionally, the third atomic operation request and the fourth atomic operation request are two sub-operation requests within an atomic reduction operation, and their execution order is adjacent and sequential. When the atomic operation result for the third atomic operation request is triggered and written back to the auxiliary storage unit, there is no need to trigger a query operation for the fourth atomic operation request, or a query operation for the first address corresponding to the fourth atomic operation request. The atomic operation result corresponding to the third atomic operation request is directly used to execute the fourth atomic operation request, further accelerating the execution efficiency of the fourth atomic operation request.
[0199] Optionally, the auxiliary storage unit is used to query the initial data corresponding to the second atomic operation request based on the second atomic operation request after the initial data corresponding to the first atomic operation request is sent to the atomic operation processing unit.
[0200] Optionally, the auxiliary storage unit is configured to, after the initial data corresponding to the first atomic operation request is sent to the atomic operation processing unit, query the initial data corresponding to the second atomic operation request based on the second atomic operation request. If the auxiliary storage unit is hit, the auxiliary storage unit returns the initial data of the second atomic operation request to the atomic operation processing unit. If the auxiliary storage unit is not hit, the data storage unit queries the initial data corresponding to the second atomic operation request based on the second atomic operation request.
[0201] Optionally, after the initial data corresponding to the first atomic operation request is sent to the atomic operation processing unit, the auxiliary storage unit and the data storage unit query the initial data corresponding to the second atomic operation request in parallel based on the second atomic operation request.
[0202] In other words, multiple atomic operation requests are executed in parallel. While the first atomic operation request is being computed in the atomic operation processing unit, the data query for the second atomic operation request can be initiated simultaneously, meaning that the auxiliary storage unit and / or data storage unit also remain operational. This allows the auxiliary storage unit and / or data storage unit, which were originally idle and waiting, to be effectively utilized. Especially when memory access latency (such as cache misses) is long, the atomic operation processing unit can continuously process subsequent tasks with prepared data, thus "hiding" this waiting time. For example, there are three atomic operation requests. During the execution of the first atomic operation request, the initial data query for the second atomic operation request has been completed, and the initial data query for the third atomic operation request is in progress. The initial data query for the third atomic operation request needs to be performed through the data storage unit, which requires a significant access latency. At this point, after the first atomic operation request is completed, the atomic operation processing unit can first execute the second atomic operation request based on the initial data of the second atomic operation request that has been queried. During the execution of the first atomic operation request and the second atomic operation request, the query for the initial data of the third atomic operation request is completed. From an overall perspective, the entire atomic operation processing system does not experience an increase in overall processing time due to the longer access latency corresponding to the third atomic operation request, thus achieving the hiding of the waiting time for querying the initial data of the atomic operation request.
[0203] In some embodiments, the atomic operation processing unit is a unit that is independent of the processor.
[0204] Optionally, the processor sends an atomic operation request to the atomic operation processing unit; if the auxiliary storage unit is hit, the atomic operation processing unit reads initial data from the auxiliary storage unit, which is read from the data storage unit to the auxiliary storage unit based on historical atomic operation requests; if the auxiliary storage unit is not hit, the auxiliary storage unit reads the initial data from the data storage unit and stores the initial data in the auxiliary storage unit; the atomic operation processing unit performs the atomic operation based on the atomic operation request and the initial data, obtains the atomic operation result, and writes the atomic operation result to the auxiliary storage unit.
[0205] Optionally, the processor is a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning. This application uses a GPU as an example, but does not limit the specific type of processor.
[0206] In this system, the atomic operation processing unit that supports atomic operations is set up independently of the processor. When performing atomic operations, the atomic operation processing unit only needs to interact with the auxiliary storage unit or the data storage unit. Compared with the atomic operation processing unit located in the processor in related technologies, which needs to communicate with the auxiliary storage unit or the data storage unit through an additional transmission path, the transmission latency of this system is lower than that of related technologies. The lower transmission latency can indirectly improve the operating efficiency of the atomic operation processing system.
[0207] On the other hand, embodiments of this application provide a graphics card that includes the atomic operation processing system described in the above embodiments. The atomic operation processing unit in the atomic operation processing system is independent of the GPU core, and the GPU core is used to send atomic operation requests to the atomic operation processing unit. The data storage unit in the atomic operation processing system is connected to the video memory controller and is used for data interaction with the video memory through the video memory controller. An auxiliary storage unit is located between the atomic operation processing unit and the data storage unit and is used to cache the initial data and atomic operation results corresponding to the atomic operation requests.
[0208] When the GPU core executes parallel computing tasks containing numerous atomic reduction operations, the atomic operation processing unit receives atomic operation requests from the GPU core. It prioritizes reading initial data from the auxiliary storage unit, and if the auxiliary storage unit is not found, it reads initial data from the data storage unit. After executing the atomic operation, the result is written back to the auxiliary storage unit. Because the access latency of the auxiliary storage unit is lower than that of the data storage unit, consecutive atomic reduction operations can be completed quickly, reducing the number of reads and writes to the data storage unit and improving the overall computing performance of the graphics card.
[0209] On the other hand, embodiments of this application provide a computer device, which includes the atomic operation processing system described above. This computer device can be at least one of a portable computer, desktop computer, server, server cluster, artificial intelligence (AI) computing cluster, and cloud computing cluster. The AI computing cluster can also be simply referred to as an intelligent computing cluster or smart computing cluster.
[0210] For example, Figure 7 A structural block diagram of a computer device provided in an exemplary embodiment of this application is shown.
[0211] The computer device 800 includes a central processing unit (CPU) 801, a system memory 804 including random access memory (RAM) 802 and read-only memory (ROM) 803, and a system bus 805 connecting the system memory 804 and the processor 801. The computer device 800 also includes a basic input / output system (I / O system) 806 to facilitate information transfer between various devices within the computer device, and a mass storage device 807 for storing the operating system 813, application programs 814, and other program modules 815.
[0212] The processor can be a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning. This application uses a GPU as an example, but does not limit the specific type of processor.
[0213] The basic input / output system 806 includes a display 808 for displaying information and an input device 809 for user input, such as a mouse or keyboard. Both the display 808 and the input device 809 are connected to the processor 801 via an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include the input / output controller 810 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 810 also provides output to a display screen, printer, or other types of output devices.
[0214] The mass storage device 807 is connected to the processor 801 via a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable storage media provide non-volatile storage for the computer device 800. That is, the mass storage device 807 may include computer-readable storage media (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0215] Without loss of generality, the computer-readable storage medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable storage instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage medium is not limited to the above-mentioned types. The system memory 804 and mass storage device 807 described above can be collectively referred to as memory.
[0216] The memory stores one or more programs, which are configured to be executed by one or more processors 801. The one or more programs contain instructions for implementing the above method embodiments, and the processor 801 executes the one or more programs to implement the methods provided by the above method embodiments.
[0217] According to various embodiments of this application, the computer device 800 can also be connected to a remote computer device on a network, such as the Internet. That is, the computer device 800 can be connected to a network 812 via a network interface unit 811 connected to the system bus 805, or the network interface unit 811 can be used to connect to other types of networks or remote computer device systems (not shown).
[0218] The memory further includes one or more programs stored in the memory, and the one or more programs include steps executed by the terminal device in the method provided in the embodiments of this application.
[0219] In addition, the computer device includes an atomic operation processing system 816, which is independent of the processor. The atomic operation processing system 816 includes an atomic operation processing unit, an auxiliary storage unit, and a data storage unit. The processor 801 sends atomic operation requests to the atomic operation processing unit. The access latency of the auxiliary storage unit is lower than that of the data storage unit, and it is used to cache the initial data and atomic operation results corresponding to the atomic operation requests.
[0220] When processor 801 executes a parallel computing task containing a large number of atomic reduction operations, the atomic operation processing unit receives atomic operation requests from processor 801. If an auxiliary memory location is hit, the atomic operation processing unit reads initial data from the auxiliary memory location, which is based on data reads from the data storage unit to the auxiliary memory location according to historical atomic operation requests. If an auxiliary memory location is not hit, the auxiliary memory location reads the initial data from the data storage unit and stores the initial data in the auxiliary memory location. After executing the atomic operation, the atomic operation processing unit writes the result back to the auxiliary memory location. Because the access latency of the auxiliary memory location is lower than that of the data storage unit, consecutive atomic reduction operations can be completed quickly, reducing the number of reads and writes to the data storage unit and improving the overall computing performance of the computer.
[0221] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0222] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An atomic manipulation processing system, characterized in that, The atomic operation processing system includes an atomic operation processing unit, an auxiliary storage unit, and a data storage unit; the access latency of the auxiliary storage unit is lower than the access latency of the data storage unit. The atomic operation processing unit is configured to read initial data from the auxiliary storage unit when the auxiliary storage unit is hit, wherein the initial data is read from the data storage unit to the auxiliary storage unit based on historical atomic operation requests; The atomic operation processing unit is used to execute the atomic operation based on the atomic operation request and the initial data, and obtain the atomic operation result; The atomic operation processing unit is used to write the atomic operation result back to the auxiliary storage unit.
2. The atomic manipulation processing system according to claim 1, characterized in that, The auxiliary storage unit is configured to read the initial data from the data storage unit when the auxiliary storage unit is not hit; and to store the initial data in the auxiliary storage unit.
3. The atomic manipulation processing system according to claim 1, characterized in that, The atomic operation request corresponding to the atomic operation processing unit includes multiple atomic operation requests. The multiple atomic operation requests are multiple sub-atomic operation requests corresponding to the atomic reduction operation and are all directed to a first address. The first address is the data address corresponding to each of the multiple atomic operation requests. The atomic operation processing unit is used to execute the atomic operation based on the i-th atomic operation request among the plurality of atomic operation requests, and obtain the i-th atomic operation result; And, the result of the i-th atomic operation is written back to the storage location in the auxiliary storage unit corresponding to the first address, where i is an integer greater than or equal to 1; The atomic operation processing unit is configured to read the result of the i-th atomic operation from the auxiliary storage unit based on the (i+1)-th atomic operation request among the plurality of atomic operation requests, according to the first address, wherein the result of the i-th atomic operation is the initial data corresponding to the (i+1)-th atomic operation request. Based on the (i+1)th atomic operation request and the i-th atomic operation result, the atomic operation is performed to obtain the (i+1)th atomic operation result. And, the result of the (i+1)th atomic operation is written back to the storage location in the auxiliary storage unit corresponding to the first address.
4. The atomic manipulation processing system according to any one of claims 1 to 3, characterized in that, The atomic operation processing unit is configured to write the atomic operation result back to the auxiliary storage unit and the data storage unit when the initial data is read from the data storage unit.
5. The atomic manipulation processing system according to any one of claims 1 to 3, characterized in that, The auxiliary storage unit includes at least one cache line; The auxiliary storage unit is used to write back the data in the first cache line to be replaced in the at least one cache line to the data storage unit when the at least one cache line stores data; and to write the read initial data to the first cache line to be replaced.
6. The atomic manipulation processing system according to any one of claims 1 to 3, characterized in that, The auxiliary storage unit includes an access index, which is used to locate data in the auxiliary storage unit according to the data address corresponding to the atomic operation request. The access index is determined based on all or part of the addresses of the data storage unit.
7. The atomic manipulation processing system according to any one of claims 1 to 3, characterized in that, The atomic operation processing unit is configured to send query requests for the data address corresponding to the atomic operation request to the auxiliary storage unit and the data storage unit respectively, so as to query the initial data in parallel; If the auxiliary storage unit is hit, the initial data is retrieved from the auxiliary storage unit; If the auxiliary storage unit is not hit, the initial data is obtained from the data storage unit.
8. The atomic manipulation processing system according to any one of claims 1 to 3, characterized in that, The atomic operation request includes a first atomic operation request and a second atomic operation request. The auxiliary storage unit is used to query the initial data corresponding to the second atomic operation request based on the second atomic operation request after the atomic operation result corresponding to the first atomic operation request has been written back. or, The auxiliary storage unit is used to query the initial data corresponding to the second atomic operation request based on the second atomic operation request after the initial data corresponding to the first atomic operation request is sent to the atomic operation processing unit. The first atomic operation request is an atomic operation request that is executed before the second atomic operation request.
9. The atomic manipulation processing system according to any one of claims 1 to 3, characterized in that, The atomic operation request includes a third atomic operation request and a fourth atomic operation request. The third atomic operation request is an atomic operation request executed before the fourth atomic operation request. The initial data of the fourth atomic operation request is the atomic operation result corresponding to the third atomic operation request. The atomic operation processing unit is configured to execute the fourth atomic operation request when the atomic operation result for the third atomic operation request is triggered and written back to the auxiliary storage unit.
10. The atomic manipulation processing system according to any one of claims 1 to 3, characterized in that, The atomic operation processing unit is a unit set up independently of the processor.
11. A method for atomic manipulation, characterized in that, The method is executed by an atomic operation processing system, which includes an atomic operation processing unit, an auxiliary storage unit, and a data storage unit; the access latency of the auxiliary storage unit is lower than the access latency of the data storage unit. When the atomic operation processing unit is hit in the auxiliary storage unit, it reads initial data from the auxiliary storage unit. The initial data is read from the data storage unit to the auxiliary storage unit based on historical atomic operation requests. The atomic operation processing unit executes the atomic operation based on the atomic operation request and the initial data to obtain the atomic operation result; The atomic operation processing unit writes the atomic operation result back to the auxiliary storage unit.
12. A computer device, characterized in that, The computer device includes the atomic operation processing system as described in any one of claims 1 to 10.