Data processing method, cluster, cache and system of multi-core system
By sending exclusive read and write requests to the LLC in the target processor cluster of the multi-core system, the hardware synchronization problem caused by data type or data width limitations in heterogeneous multi-core systems is solved, and more stable and efficient system performance is achieved.
Patent Information
- Application Number
- CN202510199302.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
AI Technical Summary
In heterogeneous multi-core multi-threaded computer systems, due to the differences in support of data types and data widths by different types of processor cores, the last-level cache (LLC) cannot process some atomic instructions, affecting hardware synchronization, and thus leading to system performance degradation or instability.
By sending the first exclusive read request and the first exclusive write request corresponding to the target atomic instructions to the last level cache (LLC) in the target processor cluster of the multi-core system, the modification operation of the target address data in the LLC is realized, and processing failures caused by data type or data width limitations are avoided.
It effectively avoids the problem that LLC cannot execute target atomic instructions due to data type or data width limitations, ensures hardware synchronization between different cores in multi-core systems, and improves system stability and performance.
Smart Images

Figure CN119988302A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of processor technology, and relate to but are not limited to a data processing method and cluster, cache, and system for a multi-core system. Background Art
[0002] With the development of semiconductor technology, heterogeneous multi-core and multi-threaded computer systems can achieve efficient task allocation and resource utilization by integrating different types of processor cores, such as central processing unit (CPU) cores, embedded neural network processor (NPU) cores, and graphics processing unit (GPU). According to the nature and requirements of the tasks, tasks can be dynamically assigned to the most suitable processor core, thereby significantly improving the overall performance and energy efficiency of the system.
[0003] In related technologies, to ensure the stable operation of multi-core systems, it is necessary to ensure that any address in the last level cache (LLC) shared by each core can only be accessed by one thread of one core in the system at the same time, so as to achieve hardware synchronization between different cores and ensure data consistency and atomicity of operations. However, for heterogeneous multi-core and multi-threaded computer systems, since different types of cores may support different data types and data widths, the access requests sent by some cores to the LLC cannot be processed by the LLC, affecting the hardware synchronization between different cores, and thus causing system performance degradation or instability. Summary of the invention
[0004] In view of this, the data processing method, cluster, cache, and system of the multi-core system provided in the embodiment of the present application can avoid the LLC from being unable to execute the target atomic instruction due to data type or data width restrictions, and ensure the hardware synchronization between different cores in the multi-core system. The data processing method, cluster, cache, and system of the multi-core system provided in the embodiment of the present application are implemented as follows:
[0005] A first aspect of the present application provides a data processing method for a multi-core system, the method being applied to a target processor cluster of the multi-core system, the method comprising:
[0006] Obtain a target atomic instruction, where the target atomic instruction is used to modify data at a target address to be accessed, where the target address is an address in the last level cache LLC of the multi-core system;
[0007] The modification operation is performed on the data at the target address by sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC, wherein the first exclusive read request includes the target address, and the first exclusive write request includes the target address and the modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
[0008] As an optional implementation, in the first aspect of the embodiments of the present application, the target processor cluster includes a target core and a target cache communicatively connected to the target core, the target core is used to obtain the target atomic instructions, the target cache includes a cache module and an atomic instruction module, the cache module is used to store the target atomic instructions, and the atomic instruction module is used to generate the first exclusive read request and the first exclusive write request.
[0009] As an optional implementation, in the first aspect of the embodiment of the present application, sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC to perform the modification operation on the data of the target address includes:
[0010] Sending the first exclusive read request to the LLC;
[0011] Receiving data of the target address sent by the LLC;
[0012] Obtaining the modified data according to the data of the target address and the target atomic instruction;
[0013] The first exclusive write request is sent to the LLC to modify the data at the target address to the modified data.
[0014] As an optional implementation manner, in the first aspect of the embodiment of the present application, the method further includes:
[0015] Receive execution data sent by the LLC, where the execution data is used to indicate an execution result of the modification operation performed on the data at the target address, where the execution result includes a modification success or a modification failure.
[0016] As an optional implementation manner, in the first aspect of the embodiment of the present application, the method further includes:
[0017] In the case where the execution data indicates that the execution result is a modification failure, a second exclusive read request and a second exclusive write request corresponding to the target atomic instruction are sent to the LLC to perform the modification operation on the data at the target address, wherein the second exclusive read request includes the target address, and the second exclusive write request includes the target address and the modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
[0018] As an optional implementation, in the first aspect of the embodiment of the present application, the target processor cluster is an NPU cluster, and the target atomic instruction includes a floating-point type atomic instruction, an atomic increment instruction, or an atomic decrement instruction.
[0019] A second aspect of the present application provides a data processing method for a multi-core system, the method being applied to a last level cache LLC of the multi-core system, the method comprising:
[0020] By receiving a first exclusive read request and a first exclusive write request sent by a target processor cluster of the multi-core system, a modification operation is performed on the data of the target address, the target address is an address in the LLC, the first exclusive read request and the first exclusive write request correspond to a target atomic instruction, the target atomic instruction is used to perform the modification operation on the data of the target address, the first exclusive read request includes the target address, the first exclusive write request includes the target address and modification data, and the modification data is the data after the modification operation is performed on the data of the target address.
[0021] As an optional implementation, in the second aspect of the embodiment of the present application, the modifying operation on the data of the target address by receiving the first exclusive read request and the first exclusive write request sent by the target processor cluster of the multi-core system includes:
[0022] receiving the first exclusive read request;
[0023] Sending the data of the target address to the target processor cluster;
[0024] receiving the first exclusive write request;
[0025] The data of the target address is modified according to the modification data.
[0026] As an optional implementation, in the second aspect of the embodiment of the present application, the target processor cluster includes a target core, the target core is used to obtain the target atomic instruction, the LLC includes a monitoring module, the monitoring module is used to monitor the modification status of the data of the target address, and the modification status is used to indicate whether the data of the target address is modified by other cores other than the target core. After receiving the first exclusive write request, the method further includes:
[0027] Acquire the modification status of the data at the target address;
[0028] The step of modifying the data of the target address according to the modification data comprises:
[0029] In a case where the modification status indicates that the data at the target address has not been modified by the other core, the data at the target address is modified according to the modification data.
[0030] As an optional implementation, in the second aspect of the embodiment of the present application, the method further includes:
[0031] Execution data is sent to the target processor cluster, where the execution data is used to indicate an execution result of the modification operation on the data at the target address, and when the modification status indicates that the data at the target address is modified by the other cores, the execution data indicates that the execution result is modification failure.
[0032] As an optional implementation manner, in the second aspect of the embodiment of the present application, after receiving the first exclusive read request, the method further includes:
[0033] Setting the read-write attribute of the target address to read-only;
[0034] After receiving the first exclusive write request, the method further includes:
[0035] The read / write attribute of the target address is set to be readable and writable.
[0036] A third aspect of the present application provides a target processor cluster of a multi-core system, the target processor cluster comprising a target core and a target cache in communication with the target core, wherein:
[0037] The target core is used to obtain a target atomic instruction, where the target atomic instruction is used to modify data at a target address to be accessed, where the target address is an address in the last level cache LLC of the multi-core system;
[0038] The target cache is used to perform the modification operation on the data at the target address by sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC, wherein the first exclusive read request includes the target address, and the first exclusive write request includes the target address and the modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
[0039] As an optional implementation, in the third aspect of the embodiment of the present application, the target cache includes a cache module and an atomic instruction module, the cache module is used to store the target atomic instructions, and the atomic instruction module is used to generate the first exclusive read request and the first exclusive write request.
[0040] As an optional implementation manner, in the third aspect of the embodiment of the present application, the target cache is used for:
[0041] Sending the first exclusive read request to the LLC;
[0042] Receiving data of the target address sent by the LLC;
[0043] Obtaining the modified data according to the data of the target address and the target atomic instruction;
[0044] The first exclusive write request is sent to the LLC to modify the data at the target address to the modified data.
[0045] As an optional implementation manner, in the third aspect of the embodiment of the present application, the target cache is further used for:
[0046] Receive execution data sent by the LLC, where the execution data is used to indicate an execution result of the modification operation performed on the data at the target address, where the execution result includes a modification success or a modification failure.
[0047] As an optional implementation manner, in the third aspect of the embodiment of the present application, the target cache is further used for:
[0048] In the case where the execution data indicates that the execution result is a modification failure, a second exclusive read request and a second exclusive write request corresponding to the target atomic instruction are sent to the LLC to perform the modification operation on the data at the target address, wherein the second exclusive read request includes the target address, and the second exclusive write request includes the target address and the modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
[0049] As an optional implementation, in the third aspect of the embodiments of the present application,
[0050] The target processor cluster is an NPU cluster, and the target atomic instruction includes a floating-point type atomic instruction, an atomic increment instruction, or an atomic decrement instruction.
[0051] A fourth aspect of the present application provides a last level cache LLC of a multi-core system, the LLC comprising:
[0052] An execution module is used to modify the data at the target address by receiving a first exclusive read request and a first exclusive write request sent by a target processor cluster of a multi-core system, wherein the target address is an address in the LLC, the first exclusive read request and the first exclusive write request correspond to a target atomic instruction, and the target atomic instruction is used to perform the modification operation on the data at the target address, the first exclusive read request includes the target address, and the first exclusive write request includes the target address and modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
[0053] As an optional implementation manner, in the fourth aspect of the embodiment of the present application, the execution module is used to:
[0054] receiving the first exclusive read request;
[0055] Sending the data of the target address to the target processor cluster;
[0056] receiving the first exclusive write request;
[0057] The data of the target address is modified according to the modification data.
[0058] As an optional implementation, in a fourth aspect of the embodiment of the present application, the target processor cluster includes a target core, the target core is used to obtain the target atomic instruction, the LLC also includes a monitoring module, the monitoring module is used to monitor the modification status of the data of the target address, and the modification status is used to indicate whether the data of the target address is modified by other cores other than the target core;
[0059] The monitoring module is further configured to obtain the modification status of the data at the target address after receiving the first exclusive write request;
[0060] The execution module is further configured to modify the data at the target address according to the modification data when the modification status indicates that the data at the target address has not been modified by the other cores.
[0061] As an optional implementation, in the fourth aspect of the embodiments of the present application, the LLC further includes:
[0062] An execution confirmation module is used to send execution data to the target processor cluster, wherein the execution data is used to indicate the execution result of the modification operation on the data at the target address. When the modification status indicates that the data at the target address is modified by the other cores, the execution data indicates that the execution result is a modification failure.
[0063] As an optional implementation, in the fourth aspect of the embodiments of the present application, the LLC further includes:
[0064] The attribute control module is used to set the read-write attribute of the target address to read-only after receiving the first exclusive read request; and set the read-write attribute of the target address to readable and writable after receiving the first exclusive write request.
[0065] A fifth aspect of the present application provides a multi-core system, the multi-core system comprising a target processor cluster and a last level cache LLC, wherein:
[0066] The target processor cluster is used to obtain a target atomic instruction, where the target atomic instruction is used to perform a modification operation on data at a target address to be accessed, where the target address is an address in the LLC; a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction are sent to the LLC to perform the modification operation on the data at the target address, where the first exclusive read request includes the target address, and the first exclusive write request includes the target address and modification data, where the modification data is data after the modification operation is performed on the data at the target address;
[0067] The LLC is configured to perform a modification operation on the data at the target address by receiving the first exclusive read request and the first exclusive write request sent by the target processor cluster.
[0068] Compared with the related art, the embodiments of the present application have at least the following beneficial effects:
[0069] The data processing method for a multi-core system applied to a target processor cluster provided by the present application obtains a target atomic instruction, and can implement a modification operation on the data of a target address in an LLC by sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the last-level cache of the multi-core system, wherein the first exclusive read request includes the target address, so that the target processor cluster can obtain the data of the target address, and the first exclusive write request includes the target address and the modified data after the modification operation is performed on the data of the target address, so that the LLC can update the data of the target address to the modified data after receiving the first exclusive write request, thereby avoiding the LLC from being unable to execute the target atomic instruction due to data type or data width restrictions, and ensuring hardware synchronization between different cores in the multi-core system. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.
[0071] Figure 1A A schematic diagram of the structure of a multi-core system provided in an embodiment of the present application;
[0072] Figure 1B Another structural schematic diagram of a multi-core system provided in an embodiment of the present application;
[0073] Figure 1C A schematic diagram of another structure of a multi-core system provided in an embodiment of the present application;
[0074] Figure 1D Another structural schematic diagram of a multi-core system provided in an embodiment of the present application;
[0075] Figure 2 A schematic diagram of a flow chart of a multi-core system data processing method provided in an embodiment of the present application applied to a target processor cluster;
[0076] Figure 3 Another flowchart of the data processing method for a multi-core system provided in an embodiment of the present application applied to a target processor cluster;
[0077] Figure 4 A schematic diagram of another flow chart of applying the data processing method for a multi-core system provided in an embodiment of the present application to a target processor cluster;
[0078] Figure 5 A schematic diagram of a flow chart of a multi-core system data processing method provided in an embodiment of the present application applied to LLC;
[0079] Figure 6 Another flowchart of the data processing method for a multi-core system provided in an embodiment of the present application applied to LLC;
[0080] Figure 7 A schematic diagram of a flow chart of a data processing method for a multi-core system provided in an embodiment of the present application;
[0081] Figure 8 A schematic diagram of the structure of a target processor cluster of a multi-core system provided in an embodiment of the present application;
[0082] Fig. 9 A schematic diagram of the structure of the last level cache LLC of a multi-core system provided in an embodiment of the present application;
[0083] Fig.10A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0084] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the specific technical solution of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0085] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0086] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0087] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present application are used to distinguish similar or different objects, and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0088] With the development of semiconductor technology, heterogeneous multi-core and multi-threaded computer systems have gradually become an important means to improve system performance and energy efficiency. Such systems integrate different types of processor clusters, and different processor clusters correspond to different processor cores, such as CPU cores, NPU cores, and GPU cores. They can dynamically assign tasks to the most suitable processor cores according to the nature and needs of the tasks. For example, tasks that require a large amount of parallel computing can be assigned to GPU cores or NPU cores first, while tasks that require complex logic processing can be assigned to CPU cores to ensure that all cores work under reasonable loads.
[0089] For a heterogeneous multi-core multi-threaded computer system, its various types of processor clusters, such as CPU clusters, NPU clusters, and GPU clusters, can share the last level cache (LLC) to improve data access efficiency and overall system performance. Since each address of LLC can only be accessed by a thread in a core in the processor cluster at the same time, but threads in different cores of the system may have the need to access an address in LLC at the same time, thus causing competition for the address, a strict management mechanism needs to be implemented to ensure hardware synchronization between cores in the system and maintain system stability.
[0090] For example, at a certain moment, if a thread in the NPU core of the processor cluster reads data M from address A of the LLC, then performs an addition operation of M+1, and finally writes the result M+1 back to the shared address A, then during the time interval from reading data M to writing back the result M+1, other cores and threads of the system cannot access the address A. Otherwise, the atomicity of the read, modify, and write operations of the NPU core on address A cannot be guaranteed. On the contrary, if the CPU core or threads of other cores can read or write the shared address A during the time interval, whether M or M+1 is read or written, the synchronization of the cores in the system will fail.
[0091] In the related art, in order to achieve synchronization between different cores in a heterogeneous multi-core multi-threaded computer system, such as between a CPU core and an NPU core, an atomic instruction can be issued to the LLC through the CPU core or the NPU core, so that the LLC can read, modify, and write the specified data stored therein according to the atomic instruction, and during the operation, other cores cannot access the address corresponding to the specified data, thereby ensuring the atomicity of the operation and the consistency of the data. However, for some heterogeneous multi-core multi-threaded computer systems, their LLC may have limitations in data width and supported data types, and can only support hardware synchronization of limited data width and specific types of data. When the atomic instruction issued by a core exceeds the support range of the LLC, the LLC will not be able to correctly process the instruction, which may cause the system to be unable to achieve effective hardware synchronization between the cores.
[0092] For example, if the LLC in a heterogeneous multi-core multi-threaded computer system supports a maximum of 64-byte cache line hardware synchronization, and only supports fixed-point hardware synchronization. The atomic instructions issued from the NPU core of the NPU cluster require a data width of up to 16Byte*32, where 16Byte is an element. Then, when the atomic instructions sent by the NPU core are larger than 64Byte, the LLC will not be able to process atomic instructions larger than 64Byte issued by the NPU core. In addition, the LLC is also unable to process floating-point atomic instructions, affecting the synchronization operations of each core in the heterogeneous multi-core multi-threaded computer system.
[0093] To solve the above problems, the related technology uses software methods, such as creating a new thread to modify the data in the LLC and sending it to each core to achieve synchronization between the cores in a multi-core system. However, the software implementation method is relatively complex and may affect system performance, and cannot meet the requirements of high bandwidth and low latency.
[0094] In view of this, an embodiment of the present application provides a data processing method for a multi-core system. The method can be applied to a processor cluster of a multi-core system. By sending an exclusive read and write request corresponding to the target atomic instruction to the LLC, the data modification operation on the target address can be implemented, thereby avoiding the LLC being unable to execute the target atomic instruction due to data type or data width restrictions, and ensuring hardware synchronization between different cores in the multi-core system.
[0095] The system structure of the multi-core system provided in the embodiment of the present application will be introduced below to facilitate understanding of the data processing method of the multi-core system provided in the present application.
[0096] See also Figure 1A , Figure 1A A schematic diagram of a multi-core system provided in an embodiment of the present application is shown in FIG. Figure 1A The multi-core system shown may include a CPU cluster, an NPU cluster, a GPU, other devices, an interconnect bus, and a last level cache (LLC). The CPU cluster, the NPU cluster, and other devices are connected to the last level cache via an interconnect bus, the CPU cluster includes multiple CPU cores, the NPU cluster includes multiple NPU cores, and other devices may be a memory controller, an input / output (I / O) controller, a storage device, and other hardware components, which are not limited here.
[0097] CPU clusters, NPU clusters, and other devices are connected to the last-level cache through an interconnection bus, which can increase the data transmission speed between different clusters or devices and improve the cache utilization.
[0098] Optionally, the processor cluster in the multi-core system may include multiple CPU clusters or NPU clusters to meet different performance requirements.
[0099] Optionally, the processor core of the processor cluster may be a CPU core or an NPU core.
[0100] The data processing method for a multi-core system provided in the embodiment of the present application can be applied to a target processor cluster in the multi-core system, which can be a CPU cluster or a GPU cluster, etc., so that the method provided in the embodiment of the present application can modify the data in the LLC based on the target atomic instruction obtained by the target processor cluster, thereby avoiding the situation where the target atomic instruction cannot be processed by the LLC when the data amount of the target atomic instruction is large or the data type is limited, thereby affecting the synchronization between different cores and hardware in the multi-core system.
[0101] See also Figure 1B , Figure 1B Another structural diagram of a multi-core system provided in an embodiment of the present application is as follows: Figure 1B The multi-core system shown may include a CPU cluster, an NPU cluster, a GPU, other devices, an interconnection bus, and an LLC, wherein the CPU cluster, the NPU cluster, and other devices are connected to the last level cache via the interconnection bus.
[0102] like Figure 1B As shown, any CPU core in a CPU cluster may include a level 1 instruction cache (L1 DCache), a level 1 data cache (L1ICache), and a level 2 cache (L2 Cache), and multiple CPU cores may share a level 3 cache (L3Cache). Any NPU core in an NPU cluster may include a level 1 instruction cache (L1 ICache) and a level 1 data cache (L1DCache). The level 1 instruction cache and the level 1 data cache store the instructions and data most recently used by the CPU core or NPU core, respectively.
[0103] In some possible embodiments, the target processor cluster may be an NPU cluster, and the target processor cluster includes a target core and a target cache that is communicatively connected to the target core. Figure 1B The second level cache (L2 Cache) shown.
[0104] It should be noted that multiple NPU cores can be connected to the target cache through the on-chip network bus (Network on ChipInterconnect, NoC Interconnect), and the target cache includes a cache module and an atomic instruction module. The on-chip network bus is used to achieve efficient interconnection between multiple NPU cores and the target cache to improve the response speed and processing power of the NPU cluster when processing high-concurrency tasks.
[0105] Optional, such as Figure 1B As shown, multiple NPU cores can be connected to a target cache, namely, a secondary cache, through an on-chip network bus.
[0106] See also Figure 1C , Figure 1C Another structural diagram of a multi-core system provided for an embodiment of the present application, in some possible embodiments, the secondary cache includes multiple secondary cache slices (L2 Cache Slice), and multiple NPU cores can be connected to multiple secondary cache slices in the secondary cache through an on-chip network bus, and each secondary cache slice includes a corresponding cache module and an atomic instruction module. After obtaining the atomic instruction, the NPU core can determine the target secondary cache slice among the multiple secondary cache slices through a preset allocation strategy to generate corresponding exclusive read requests and exclusive write requests.
[0107] See also Figure 1D , Figure 1D Another structural diagram of a multi-core system provided by an embodiment of the present application, in some possible embodiments, the target cache is set in the NPU core, so that the NPU core calls the target cache to implement the modification operation on the data in the LLC. Figure 1D As shown, in the case where the target processor cluster is an NPU cluster, the target cache can be a secondary cache set in the NPU core. In this way, after the target NPU core obtains the target atomic instruction, the secondary cache of the target NPU core can store the target atomic instruction through the cache module, and generate an exclusive read request and an exclusive write request through the atomic instruction module, which are then sent to the LLC via the on-chip network bus through the third-level cache (L3 Cache) to implement the data modification operation.
[0108] It is understandable that by setting up a multi-level cache structure in the processor cluster, the number of times the processor core such as the CPU or NPU accesses the main memory can be reduced. The first-level instruction cache and the first-level data cache are the cache levels closest to the core and store the most recently used instructions and data. When the CPU core or NPU core needs to read or write data, it will first try to retrieve it from the first-level cache. If the first-level cache misses, it will then try to retrieve the data from the second-level cache and even the LLC in turn. This hierarchical cache structure ensures that the processor core can access the required data as quickly as possible, thereby accelerating the execution of the program.
[0109] Exemplarily, when the processor core in the processor cluster needs to access the data in the LLC, taking the NPU core as an example, the atomic instructions obtained by the NPU can be processed by the cache module and the atomic instruction module of the target cache to modify the data in the LLC. Among them, the cache module is used to store the atomic instructions obtained by the NPU core, and the atomic instruction module is used to generate exclusive read requests and exclusive write requests corresponding to the atomic instructions. In this way, the target processor cluster sends the first exclusive read request and the first exclusive write request corresponding to the target atomic instruction to the LLC to modify the data at the target address.
[0110] It should be noted that the method provided in the embodiments of the present application can be applied to Figure 1A , Figure 1B , Figure 1C or Figure 1D The processor cluster in the heterogeneous multi-core system shown can also be applied to the processor cluster in the homogeneous multi-core system, which is not limited here.
[0111] The following will introduce the application of the data processing method of the multi-core system provided in the embodiment of the present application in the target processor cluster.
[0112] See also Figure 2 , Figure 2 A schematic diagram of a flow chart of a multi-core system data processing method provided in an embodiment of the present application applied to a target processor cluster, such as Figure 2 As shown, the method may include the following steps:
[0113] S201, obtaining a target atomic instruction, where the target atomic instruction is used to modify data at a target address to be accessed, where the target address is an address in the last level cache LLC of a multi-core system.
[0114] It should be noted that in the related art, atomic instructions are used to read data from the memory address of the memory, perform corresponding calculations on the read data, and then write the calculated results back to the memory address of the memory. Atomic instructions are indivisible, that is, the read, modify, and write operations corresponding to atomic instructions are either completed in full or not executed at all, which can prevent interference from other cores or processes to ensure data consistency among cores in a multi-core system.
[0115] There are many kinds of atomic instructions. For example, integer addition (Integer ADD) reads data at a specified address, adds a certain value to the data at a specified address, and writes the result to the specified address. This process is inseparable, that is, reading data, adding operations, and writing back data are either completed or not executed at all, thus avoiding interference with data by other processors or cores during operation.
[0116] In some possible embodiments, the LLC shared by the cores of different processor clusters in a multi-core system can support some atomic instructions, for example, integer type atomic instructions such as integer type addition (Integer ADD), integer type exclusive OR operation (Integer XOR), integer type swap (Integer Swap), etc., but for atomic instructions of floating point type or data amount exceeding the LLC support range, the LLC will not be able to process the atomic instructions after the cores of some processor clusters obtain them, affecting the synchronization between different cores and the normal data access of each core in the target processor cluster.
[0117] Optionally, LLC can execute atomic instructions of integer type and instruction size within 64 Byte. Please refer to Table 1, which is a table of examples of atomic instruction types that LLC can execute, including multiple atomic instruction types that LLC can execute.
[0118] Table 1
[0119]
[0120]
[0121]
[0122] Table 1 shows examples of atomic instructions corresponding to different atomic instruction types. Taking integer type addition as an example, in the atomic instruction example corresponding to integer type addition, Ws is the target register, which stores the result of the operation. Wt is the value register, whose value will be added to the value in the memory. Xn and SP are base registers or stack pointers, which specify the memory address. After adding an implicit offset to the memory address, it is the target memory location of the operation. Xs and Xt and Ws and Wt represent different registers to support different register combinations or operation sizes. The function of the LDADD instruction is to load the value at Xn or SP into the Wt or Xt register, and then add the value read from Xn or SP to the value read from Xs and Xt, and update the value at Xn or SP to the new value after the addition. This process is atomic. The STADD instruction has the same basic functions as the LDADD instruction, except that the Wt or Xt register is not required, and the value of the specified memory location before modification will not be returned to Xn or SP.
[0123] As shown in Table 1, the atomic instructions are all of integer type, and taking the atomic instructions of clearing specific bits of integer type and exchanging integer type as examples, the maximum instruction size of the atomic instructions is 16B (Byte), so LLC can execute the atomic instructions in Table 1.
[0124] In some possible embodiments, when the target processor cluster is a CPU cluster, that is, the core is a CPU core, the target atomic instruction sent by the CPU core to the LLC may be an atomic instruction type that the LLC can execute. Therefore, the target atomic instruction can be sent to the LLC to modify the data of the target address in the LLC. In addition, according to the method provided in the embodiment of the present application, a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction can be sent to the LLC to modify the data of the target address in the LLC, which is not limited here.
[0125] In some possible embodiments, the target processor cluster is an NPU cluster, that is, the core of the target processor cluster is an NPU core, and the target atomic instructions may include floating-point type atomic instructions, atomic increment instructions, atomic decrement instructions, and other types of atomic instructions.
[0126] Please refer to Table 2, which is an example table of atomic instruction types of target atomic instructions that the NPU core can obtain. Among them, the various atomic instruction types in Table 2 are only used as examples for reference and are not limited to these types. In actual applications, adjustments can be made according to different instruction set architectures or custom instructions in specific scenarios, etc., which are not limited here.
[0127] Table 2
[0128]
[0129]
[0130] It should be noted that the instruction size in Table 2 is used to indicate the possible instruction size of the corresponding target atomic instruction. Among them, in the integer type, U1B corresponds to an unsigned 1Byte (8bit) integer, S1B corresponds to a signed 1Byte (8bit) integer, U2B corresponds to an unsigned 2Byte (16bit) integer, S2B corresponds to a signed 2Byte (16bit) integer, U4B corresponds to an unsigned 4Byte (32bit) integer, S4B corresponds to a signed 4Byte (32bit) integer, U8B corresponds to an unsigned 8Byte (64bit) integer, S8B corresponds to a signed 8Byte (64bit) integer, U16B corresponds to an unsigned 16Byte (128bit) integer, and S16B corresponds to a signed 16Byte (128bit) integer. In the floating-point type, FADD32 corresponds to a 32-bit floating-point number in a floating-point addition operation, F16 corresponds to a 16-bit floating-point number (half-precision floating-point number), and BF16 corresponds to the 16-bit floating-point format of Brain Float 16.
[0131] As can be seen from Table 2, in some embodiments, LLC cannot directly execute floating-point target atomic instructions sent by the NPU core, nor can it execute some integer target atomic instructions, such as integer type increment or integer type decrement.
[0132] It should be noted that the instruction size of the target atomic instruction of integer type increment or integer type decrement can be the same as that of other integer type atomic instructions. However, compared with other integer type atomic instructions, the target atomic instruction of integer type increment or integer type decrement cannot be split into multiple sub-instructions corresponding to the target atomic instruction for processing, resulting in these two integer type target principle instructions cannot be directly executed by LLC.
[0133] When the target processor cluster is an NPU cluster and the target atomic instructions include floating-point type atomic instructions, atomic increment instructions or atomic decrement instructions, the target processor cluster can generate a corresponding first exclusive read request and a first exclusive write request according to the target atomic instructions through the data processing method of the multi-core system provided in the embodiment of the present application, so as to modify the data of the target address in the LLC.
[0134] Optionally, the target atomic instruction is an atomic instruction corresponding to a target thread of a target core in a target processor cluster.
[0135] It is understandable that a processor cluster may include multiple cores, and one core may run multiple threads simultaneously. Data contention and state inconsistency should also be avoided between different threads on the same core. Therefore, the target atomic instruction is corresponded to the target thread in the target processor cluster. Using the thread as the basic unit can avoid the influence of other threads other than the target thread corresponding to the target atomic instruction in the target processor cluster on the modification operation of the data at the target address.
[0136] S202, by sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC to modify the data at the target address, the first exclusive read request includes the target address, the first exclusive write request includes the target address and the modified data, and the modified data is the data after the modification operation is performed on the data at the target address.
[0137] In an embodiment of the present application, after the target processor cluster obtains the target atomic instruction, due to the limitation of data type or instruction size, the LLC may not be able to directly process the target atomic instruction to modify the data of the target address. The first exclusive read request and the first exclusive write request corresponding to the target atomic instruction may be sent to the LLC to update the data of the target address to the modified data, thereby realizing the modification operation of the target address data in the LLC through the target processor cluster.
[0138] The first exclusive read request and the first exclusive write request are both exclusive requests. It can be understood that during the period when the request is processed by LLC, the address contained in the exclusive request will be locked, and only the core or thread that initiates the request is allowed to read or write it, while other threads (or cores) are restricted from accessing the data.
[0139] Optionally, the first exclusive read condition and the first exclusive write request also include the address of the target thread of the target processor cluster, where the target thread is the thread that sends the first exclusive read request and the first exclusive write request corresponding to the target atomic instruction to the LLC, so that the target processor cluster receives data such as the target address data returned by the LLC.
[0140] Optionally, in the case where the target atomic instruction includes at least one data address other than the target address, the method provided in the embodiment of the present application also includes: based on the at least one data address, obtaining at least one instruction-related data corresponding to the at least one data address, and the modified data is obtained by performing a modification operation on the data of the target address and at least one related data.
[0141] It should be noted that for target atomic instructions of types such as addition and numerical comparison, since they involve operations on multiple data items, such target atomic instructions usually include at least one data address other than the target address to combine two or more data items stored at different addresses to produce a result. Therefore, by obtaining at least one data address other than the target address included in the target atomic instruction, the corresponding instruction-related data can be obtained to obtain the corresponding modification data, thereby realizing the modification operation on the data at the target address.
[0142] Optionally, the target atomic instruction also includes an instruction data value, so that the target processor cluster performs an operation corresponding to the target atomic instruction through the data value and the data of the target address to obtain the modified data.
[0143] In an embodiment of the present application, the modification operation performed on the data of the target address is an operation corresponding to the target atomic instruction. For example, when the target atomic instruction is a floating-point type addition, the modification operation is also a floating-point type addition. The target processor cluster can obtain the data of the target address through the first exclusive write request, and obtain another added data according to the address of another added data other than the data of the target address indicated in the target atomic instruction, and the target processor cluster performs floating-point type addition on the data of the target address and the added data to obtain the modified data.
[0144] In the data processing method for a multi-core system provided in an embodiment of the present application, a target atomic instruction is obtained, and a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction are sent to the last-level cache of the multi-core system, so that the data of the target address in the LLC can be modified. The first exclusive read request includes the target address, so that the target processor cluster can obtain the data of the target address, and the first exclusive write request includes the target address and the modification data, so that the LLC can modify the data of the target address after receiving the first exclusive write request, thereby avoiding the LLC from being unable to execute the target atomic instruction due to data type or data width restrictions, and ensuring hardware synchronization between different cores in the multi-core system.
[0145] The following will introduce the implementation process of sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC in an embodiment of the present application to modify the data of the target address.
[0146] See also Figure 3 , Figure 3 Another flowchart of the data processing method for a multi-core system provided in an embodiment of the present application applied to a target processor cluster is shown in FIG. Figure 3 As shown, the method may include the following steps:
[0147] S301, obtaining a target atomic instruction.
[0148] Optionally, when the target processor cluster is an NPU cluster, the target atomic instruction can be generated by the target NPU core in the NPU cluster and sent down step by step to the target NPU cores such as Figure 1B , Figure 1C or Figure 1D The L2 cache or L2 cache slice shown.
[0149] S302: Send a first exclusive read request to LLC.
[0150] In some possible embodiments, the target processor cluster includes a target core and a target cache communicatively connected to the target core, the target core is used to obtain target atomic instructions, the target cache includes a cache module and an atomic instruction module, the cache module is used to store the target atomic instructions, and the atomic instruction module is used to generate a first exclusive read request and a first exclusive write request.
[0151] Optional, such as Figure 1B , Figure 1C or Figure 1D As shown, when the target core is an NPU core, the cache module and the atomic instruction module can be located in the target cache corresponding to the NPU core, that is, the secondary cache in the NPU cluster.
[0152] After obtaining the target atomic instruction, the target core can store the target atomic instruction in the cache module of the target cache so that a subsequent atomic instruction module can generate a corresponding first exclusive read request and a first exclusive write request based on the target atomic instruction, or the target core can obtain the target atomic instruction based on the cached target atomic instruction, including at least one instruction-related data corresponding to at least one data address other than the target address.
[0153] Optionally, the data processing method of the multi-core system provided in the embodiment of the present application further includes:
[0154] A first exclusive read request is generated according to a target address included in the target atomic instruction.
[0155] It should be noted that in order to implement the modification operation of the data at the target address in the LLC, it is first necessary to ensure that the target processor cluster can obtain the data at the target address, so that the data at the target address can be modified accordingly based on the data at the target address according to the target atomic operation, and the modified data can be obtained and written to the target address. Among them, the first exclusive read request can also include the address corresponding to the target core or target cache, so that the LLC can accurately send the required data to the target processor cluster.
[0156] Optionally, in the case where at least one data address other than the target address included in the target atomic instruction is an address in the LLC, the first exclusive read request also includes at least one data address. So that by sending the first exclusive read request to the LLC once, the data content for the corresponding modification operation can be obtained, thereby reducing the communication frequency between the target core and the target cache and the LLC, so as to improve the efficiency of data processing and optimize the overall system performance.
[0157] Optionally, the atomic instruction module in the target processor cluster may obtain at least one instruction-related data corresponding to the at least one data address by sending a corresponding exclusive read request to at least one data address other than the target address included in the target atomic instruction.
[0158] S303, receiving data of the target address sent by LLC.
[0159] After obtaining the first exclusive read request sent by the target processor cluster, LLC will send the data of the target address to the target processor cluster. After receiving the data of the target address, the target processor cluster can generate the first exclusive write request through the atomic instruction module and continue to modify the data of the target address.
[0160] Optionally, when the first exclusive read request further includes at least one data address, the method applied to the target processor cluster further includes: receiving data of the target address and at least one instruction-related data sent by the LLC.
[0161] Optionally, when the first exclusive read request also includes the address of the target cache, data at the target address may be received through the target cache.
[0162] S304, obtaining modified data according to the data of the target address and the target atomic instruction.
[0163] After the target processor cluster generates a first exclusive read request through the atomic instruction module and sends it to the LLC, the atomic instruction module can receive the target address data required for the modification operation on the target address data and at least one instruction-related data that may exist according to the target atomic instruction type, so that the target processor cluster can process the obtained data content according to the modification operation corresponding to the target atomic instruction to obtain the modified data, and generate a first exclusive write request based on the modified data and the target address, so as to write the modified data to the target address and complete the modification operation on the target address data.
[0164] Optionally, the target processor cluster may process the obtained data content through the target cache according to the modification operation corresponding to the target atomic instruction to obtain the modified data.
[0165] S305 , sending a first exclusive write request to the LLC to modify the data at the target address as modified data.
[0166] The target processor cluster sends a first exclusive request including the modification data and the target address to the LLC, so that the LLC can write the modification data into the target address after receiving the first exclusive request, so as to perform a modification operation on the data at the target address.
[0167] By implementing the above technical solution, when LLC cannot directly execute the target atomic instruction obtained by the target processor cluster, the target processor cluster can generate the corresponding first exclusive read request and first exclusive write request through the set atomic instruction module and the target atomic instruction, thereby realizing the modification operation of the data of the target address and ensuring the hardware synchronization between different cores in the multi-core system.
[0168] The following will introduce the implementation process of confirming whether the data at the target address has been successfully modified after the target processor cluster sends the first exclusive write request to the LLC in an embodiment of the present application.
[0169] See also Figure 4 , Figure 4 Another flow chart of the data processing method for a multi-core system provided in an embodiment of the present application being applied to a target processor cluster is shown in FIG. Figure 4 As shown, the method may include the following steps:
[0170] S401, obtaining a target atomic instruction.
[0171] S402: Send a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC.
[0172] The implementation methods corresponding to S401 to S402 are similar to the implementation methods corresponding to S201 to S202 or S301 to S305 and are not described in detail here.
[0173] S403, receiving execution data sent by LLC, where the execution data is used to indicate the execution result of the modification operation on the data at the target address, and the execution result includes modification success or modification failure.
[0174] In some possible embodiments, the data processing method for a multi-core system of a target processor cluster provided by the present application further includes:
[0175] Receive execution data sent by LLC, where the execution data is used to indicate the execution result of the modification operation on the data at the target address, where the execution result includes modification success or modification failure.
[0176] In this way, the target processor cluster can receive the execution data sent by the LLC so that the target processor cluster can determine whether the data at the target address is successfully modified by the first exclusive write request.
[0177] Optionally, the target processor cluster receives the execution data sent by the LLC, including: the target cache receives the execution data sent by the LLC.
[0178] It should be noted that the method provided in the embodiment of the present application can implement the modification operation on the data of the target address in the LLC by sending a first exclusive write request and a first exclusive read request to the LLC. In the process in which the LLC receives the first exclusive request and sends the data of the target address to the target processor cluster, and the LLC receives the first exclusive write request and modifies the data of the target address to the modified data, due to the nature of the exclusive request, the target address will not be accessed by other cores or threads, but after the LLC sends the data of the target address to the target processor cluster, until the LLC receives the first exclusive write request, the target address may be accessed and modified by cores or threads other than the target core of the target processor cluster, so that the data of the target address when the LLC receives the first exclusive write request is different from the data of the target address when the LLC receives the first exclusive read request, affecting the synchronization between hardware.
[0179] For example, when the target atomic instruction indicates to add 1 to the data at the target address, after the target processor cluster sends the first exclusive read request to the target address, the data at the target address received by the target processor cluster is 3, and a modification operation of adding 1 is performed on the data at the target address, that is, the modified data included in the first exclusive write request is 4. If the target address is accessed by other cores or threads and the data at the target address is modified to 5 before LLC receives the first exclusive write request sent by the target processor cluster, when LLC receives the first exclusive write request, regardless of whether the read or written data is 4 or 5, it will cause the synchronization of the multi-core system to fail, which may lead to program errors or system instability.
[0180] Optionally, the execution data may be expressed as a numerical value, a character or other identifier, such as 0 or 1, which is not limited here.
[0181] Optionally, when the execution data indicates that the execution result is a successful modification, the target processor cluster does not need to process the execution data after receiving it.
[0182] S404, when the execution data indicates that the execution result is a modification failure, a second exclusive read request and a second exclusive write request corresponding to the target atomic instruction are sent to the LLC to modify the data of the target address, the second exclusive read request includes the target address, and the second exclusive write request includes the target address and the data after the modification operation is performed on the data of the target address.
[0183] In some possible embodiments, when the execution data indicates that the execution result is a modification failure, a second exclusive read request and a second exclusive write request corresponding to the target atomic instruction are sent to the LLC to perform a modification operation on the data at the target address, the second exclusive read request includes the target address, and the second exclusive write request includes the target address and the data after the modification operation is performed on the data at the target address.
[0184] It should be noted that if the execution result indicated by the execution data is a modification failure, it can be understood that during the period when the target processor cluster sends the first exclusive read request and the first exclusive write request, the data at the target address is modified by other cores or threads, or the modification operation on the data at the target address fails due to hardware failure, etc. The target processor cluster can send a second exclusive read request and a second exclusive write request through the atomic instruction module of the target cache to try to modify the data at the target address.
[0185] Optionally, after sending a second exclusive read request and a second exclusive write request corresponding to the target atomic instruction to the LLC to perform a modification operation on the data at the target address, the method applied to the target processor cluster further includes:
[0186] Receive execution data sent by LLC;
[0187] When the execution data indicates that the execution result is a modification failure, a third exclusive read request and a third exclusive write request corresponding to the target atomic instruction are sent to the LLC to modify the data at the target address. The third exclusive read request includes the target address, and the third exclusive write request includes the target address and the data after the modification operation is performed on the data at the target address.
[0188] It is understandable that, if after the target processor cluster sends the second exclusive write request, if the received execution data still indicates that the execution result is a modification failure, it can continuously try to send new exclusive read and write requests until the execution data indicates that the execution result is a modification success.
[0189] Optionally, an upper limit on the number of times the target processor cluster sends exclusive read and write requests can be set to avoid a long cycle of sending exclusive read and write requests due to hardware failure.
[0190] It should be noted that when the execution data indicates that the execution result is a modification failure, depending on whether other cores or threads perform modification operations on the data of the target address before the LLC receives the first exclusive write request, the data of the target address obtained by the target processor cluster through the first exclusive read request and the second exclusive read request may be different.
[0191] In some possible embodiments, in order to avoid the data of the target address being accessed by other cores or threads during the period from the first exclusive read request to the first exclusive write request, the first exclusive read request may also include a preset first identification information, and the preset first identification information is used to instruct the LLC to set the read and write attributes of the target address to read-only upon receiving the preset first identification information. The first exclusive write request may also include a preset second identification information, and the preset second identification information is used to instruct the LLC to set the read and write attributes of the target address to readable and writable upon receiving the preset second identification information, so as to ensure that the data of the target address will not be modified by other cores or threads before the first exclusive write request sent by the target processor cluster is delivered to the LLC, thereby ensuring hardware synchronization between different cores in a multi-core system.
[0192] By implementing the above technical solution, the target processor cluster can determine whether the data at the target address is successfully modified to the modified data through the first exclusive write request by receiving the execution data sent by the LLC, and when the execution data indicates that the execution result is a modification failure, the target processor cluster sends a second exclusive read request and a second exclusive write request corresponding to the target atomic instruction to the LLC, and tries the modification operation on the data at the target address again, thereby ensuring hardware synchronization between different cores in a multi-core system.
[0193] The following will introduce the implementation process of the data processing method applied to the LLC multi-core system in the embodiment of the present application.
[0194] See also Figure 5 , Figure 5 A flow chart of a multi-core system data processing method provided in an embodiment of the present application applied to LLC, as shown in FIG. Figure 5 As shown, the method may include the following steps:
[0195] S501, receiving a first exclusive read request.
[0196] In an embodiment of the present application, a data processing method for a multi-core system applied to LLC includes:
[0197] By receiving a first exclusive read request and a first exclusive write request sent by a target processor cluster of a multi-core system, a modification operation is performed on the data in the target address, the target address is an address in the LLC, the first exclusive read request and the first exclusive write request correspond to a target atomic instruction, the target atomic instruction is used to modify the data at the target address, the first exclusive read request includes the target address, the first exclusive write request includes the target address and the modified data after the modification operation is performed on the data at the target address.
[0198] In this way, LLC does not need to receive the target atomic instruction, but implements the modification operation on the data of the target address by receiving the first exclusive read request and the first exclusive write request corresponding to the target atomic instruction, so as to modify the data of the target address into the modified data.
[0199] S502, sending data of the target address to the target processor cluster.
[0200] In some possible embodiments, the LLC includes a monitoring module, which can be used to monitor the sender address of the request sent to the LLC. When the LLC receives the first exclusive request sent by the target processor cluster, the sender address corresponding to the target processor cluster can be determined through the monitoring model, so that the LLC can send the data of the target address to the target processor cluster based on the target address included in the first exclusive read request.
[0201] Optionally, the sender address of the target processor cluster refers to the address of the target thread that sends the first exclusive read request in the target processor cluster. Depending on the structure of the target processor cluster, it can be the address of the target core or target cache. The target thread can also be used to send data or request instructions such as the first exclusive write request to LLC, which is not limited here.
[0202] S503: Receive a first exclusive write request.
[0203] S504, modifying the data of the target address according to the modification data.
[0204] The LLC may update the data of the target address to the modified data according to the modified data included in the first exclusive write request.
[0205] By implementing the above technical solution, LLC receives the first exclusive read request and the first exclusive write request sent by the target processor cluster, and performs modification operations on the data in the target address, which can avoid the LLC being unable to execute the target atomic instruction due to data type or data width restrictions, thereby ensuring hardware synchronization between different cores in a multi-core system.
[0206] The following will introduce the implementation process of monitoring the modification status of the data of the target address in the data processing method of the multi-core system applied to LLC in the embodiment of the present application.
[0207] See also Figure 6 , Figure 6 Another flowchart of the data processing method of the multi-core system provided in the embodiment of the present application applied to LLC is as follows: Figure 6 As shown, the method may include the following steps:
[0208] S601, receiving a first exclusive read request.
[0209] S602, sending data of the target address to the target processor cluster.
[0210] S603: Receive a first exclusive write request.
[0211] The implementation methods corresponding to S601 to S603 are similar to those corresponding to S501 to S503 and will not be described in detail herein.
[0212] S604, obtaining the modification status of the data at the target address.
[0213] In some possible embodiments, the target processor cluster includes a target core, the target core is used to obtain a target atomic instruction, the LLC includes a monitoring module, the monitoring module is used to monitor the modification status of the data at the target address, the modification status is used to indicate whether the data at the target address is modified by other cores other than the target core, and after receiving the first exclusive write request, the method applied to the LLC also includes:
[0214] Get the modification status of the data at the target address.
[0215] It should be noted that the monitoring module can monitor the modification status of each address (cache line) in the LLC. When a core performs a write operation on a certain address (cache line), the monitoring module will update the modification status of the address.
[0216] Optionally, the modification status collected by the monitoring module may include the last modified thread address, so that the LLC can determine whether the data of the target address of the LLC is modified by other cores or threads during the period of receiving the first exclusive read request and the first exclusive write request based on the modification status of the data of the target address.
[0217] Optionally, the modification status collected by the monitoring module may include the last modification time. In this way, after the LLC receives the first exclusive write request, if the last modification time is during the period between receiving the first exclusive read request and the first exclusive write request, it means that the data of the target address is modified by other cores or threads.
[0218] S605 , when the modification status indicates that the data at the target address has not been modified by other cores, modify the data at the target address according to the modification data.
[0219] In some possible embodiments, modifying data at the target address according to the modification data includes:
[0220] When the modification status indicates that the data at the target address has not been modified by other cores, the data at the target address is modified according to the modification data.
[0221] If the modification status indicates that the data at the target address has not been modified by other cores, the LLC may modify the data at the target address according to the modification data included in the first exclusive write request to ensure hardware synchronization in the multi-core system.
[0222] In some possible embodiments, execution data is sent to a target processor cluster, the execution data is used to indicate an execution result of a modification operation on data at a target address, and when a modification status indicates that data at the target address is modified by other cores, the execution data indicates that the execution result is a modification failure.
[0223] It can be understood that by sending execution data to the target processor cluster, the target processor cluster can determine whether the modification operation is successfully performed by sending a first exclusive read request and a first exclusive write request to the LLC, so that the target processor cluster can take corresponding actions according to the execution result. For example, when the execution result indicated by the execution data is a modification failure, the target processor cluster can promptly send a second exclusive read request and a second exclusive write request to the LLC to try to modify the data at the target address again.
[0224] In some possible embodiments, after receiving the first exclusive read request, the method applied to LLC further includes:
[0225] Set the read-write attribute of the target address to read-only;
[0226] After receiving the first exclusive write request, the method applied to the LLC further includes:
[0227] Set the read-write property of the target address to read and write.
[0228] This ensures that during the period when the target processor cluster sends the first exclusive read request and the first exclusive write request, other cores or threads except the target core cannot write to the data at the target address, thereby ensuring data consistency. And after the LLC modifies the data at the target address according to the modification data included in the first exclusive write request, since the first exclusive write request is executed, the read-write attribute setting of the target address will automatically be restored to read-write, without affecting the access of other cores or threads to the target reference address.
[0229] By implementing the above technical solution, LLC ensures the consistency of the reading and modification operations of the target address data in the multi-core system by monitoring the modification status of the target address data and managing the read and write properties of the target address, thereby improving the overall performance and reliability of the system.
[0230] See also Figure 7 , Figure 7 A flow chart of a data processing method for a multi-core system provided in an embodiment of the present application is shown as follows: Figure 7 As shown, the following steps may be included:
[0231] S701, obtain the target atomic instruction.
[0232] The target processor cluster fetches the target atomic instruction.
[0233] S702: Send a first exclusive read request to the LLC.
[0234] Optionally, the target processor cluster may include a target core and a target cache, and the target cache may include a cache module and an atomic instruction module, and the atomic instruction module is used to generate the first exclusive read request.
[0235] The target processor cluster sends a first exclusive read request to the LLC.
[0236] Optionally, the target processor cluster sends a first exclusive read request to the LLC through the target cache.
[0237] S703: Receive a first exclusive read request.
[0238] The LLC receives a first exclusive read request sent by a target processor cluster.
[0239] S704, sending data of the target address to the target processor cluster.
[0240] After receiving the first exclusive read request, the LLC sends the data in the target address to the target processor cluster according to the target address included in the first exclusive read request.
[0241] S705, obtaining modified data according to the data of the target address and the target atomic instruction.
[0242] The target processor cluster can obtain the modified data according to the data of the target address and the target atomic instruction.
[0243] S706: Send a first exclusive write request to the LLC.
[0244] Optionally, the first exclusive write request is also generated by the atomic instruction module, and the first exclusive write request includes a target address and modification data for performing a modification operation on data at the target address.
[0245] S707: Receive a first exclusive write request.
[0246] LLC receives the first exclusive write request.
[0247] S708, obtaining the modification status of the data at the target address.
[0248] Optionally, the LLC includes a monitoring module, and the monitoring module is used to monitor the modification status of the data of the target address.
[0249] LLC determines whether the data at the target address is modified by other cores other than the target core of the target processor cluster by obtaining the modification status of the data at the target address.
[0250] S709 , when the modification status indicates that the data at the target address has not been modified by other cores, modify the data at the target address according to the modification data.
[0251] When the modification status indicates that the data at the target address has not been modified by other cores, the LLC can modify the data at the target address according to the modification data, thereby implementing a modification operation on the data at the target address.
[0252] Optionally, LLC can also send execution data to the target processor cluster, and when the modification status indicates that the data at the target address is modified by other cores, the execution data indicates that the execution result of the modification operation on the data at the target address is a modification failure, so that after receiving the execution data, the target processor cluster re-sends a second exclusive read request and a second exclusive write request to LLC to try to modify the data at the target address again.
[0253] By implementing the above technical solution, the target processor cluster and LLC can transmit data through the interconnection bus of the multi-core system, thereby ensuring the modification operation of the data at the target address and the data synchronization between the hardware, and improving the stability and reliability of the multi-core system.
[0254] It should be understood that, although the steps in the above-mentioned flowcharts are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flowcharts may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0255] Based on the foregoing embodiments, the embodiments of the present application provide a target processor cluster and a last-level cache LLC of a multi-core system. The target processor cluster or the last-level cache LLC includes the modules included, and the units included in the modules, which can be implemented by a processor; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0256] please Figure 8 A schematic diagram of the structure of a target processor cluster of a multi-core system provided in an embodiment of the present application, such as Figure 8As shown, the target processor cluster of the multi-core system includes a target core 801 and a target cache 802 that is in communication with the target core 801, wherein:
[0257] A target core 801 is used to obtain a target atomic instruction, where the target atomic instruction is used to modify data at a target address to be accessed, where the target address is an address in the last level cache LLC of the multi-core system;
[0258] The target cache 802 is used to modify the data at the target address by sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC, the first exclusive read request includes the target address, the first exclusive write request includes the target address and the modified data, and the modified data is the data after the modification operation is performed on the data at the target address.
[0259] In some possible embodiments, the target cache 802 includes a cache module and an atomic instruction module, the cache module is used to store the target atomic instruction, and the atomic instruction module is used to generate a first exclusive read request and a first exclusive write request.
[0260] In some possible embodiments, the target cache 802 is used to:
[0261] Send a first exclusive read request to LLC;
[0262] Receive the data of the target address sent by LLC;
[0263] According to the data of the target address and the target atomic instruction, the modified data is obtained;
[0264] A first exclusive write request is sent to the LLC to modify the data at the target address as modified data.
[0265] In some possible embodiments, the target cache 802 is further used for:
[0266] Receive execution data sent by LLC, where the execution data is used to indicate the execution result of the modification operation on the data at the target address, where the execution result includes modification success or modification failure.
[0267] In some possible embodiments, the target cache 802 is further used for:
[0268] When the execution data indicates that the execution result is a modification failure, a second exclusive read request and a second exclusive write request corresponding to the target atomic instruction are sent to the LLC to modify the data at the target address. The second exclusive read request includes the target address, and the second exclusive write request includes the target address and the modification data. The modification data is the data after the modification operation is performed on the data at the target address.
[0269] In some possible embodiments, the target processor cluster is an NPU cluster, and the target atomic instruction includes a floating-point type atomic instruction, an atomic increment instruction, or an atomic decrement instruction.
[0270] Fig. 9 A structural diagram of the last level cache LLC of a multi-core system provided in an embodiment of the present application is shown as follows: Fig. 9 As shown, the last level cache LLC of the multi-core system includes an execution module 901, a monitoring module 902 and an execution confirmation module 903, wherein:
[0271] Execution module 901 is used to modify the data of the target address by receiving the first exclusive read request and the first exclusive write request sent by the target processor cluster of the multi-core system, the target address is the address in the LLC, the first exclusive read request and the first exclusive write request correspond to the target atomic instruction, the target atomic instruction is used to modify the data of the target address, the first exclusive read request includes the target address, the first exclusive write request includes the target address and the modified data, and the modified data is the data after the modification operation is performed on the data of the target address.
[0272] In some possible embodiments, the execution module 901 is used to:
[0273] receiving a first exclusive read request;
[0274] Sending data of the target address to the target processor cluster;
[0275] receiving a first exclusive write request;
[0276] Modify the data at the target address according to the modification data.
[0277] In some possible embodiments, the target processor cluster includes a target core, the target core is used to obtain the target atomic instruction, and the LLC further includes a monitoring module 902, the monitoring module 902 is used to monitor the modification status of the data of the target address, and the modification status is used to indicate whether the data of the target address is modified by other cores other than the target core;
[0278] The monitoring module 902 is further configured to obtain a modification status of the data at the target address after receiving the first exclusive write request;
[0279] The execution module 901 is further configured to modify the data at the target address according to the modification data when the modification status indicates that the data at the target address has not been modified by other cores.
[0280] In some possible embodiments, LLC also includes an execution confirmation module 903, which is used to send execution data to the target processor cluster, and the execution data is used to indicate the execution result of the modification operation on the data of the target address. When the modification status indicates that the data of the target address is modified by other cores, the execution data indicates that the execution result is a modification failure.
[0281] In some possible embodiments, the LLC further includes an attribute control module for setting the read / write attribute of the target address to read-only after receiving the first exclusive read request; and setting the read / write attribute of the target address to readable and writable after receiving the first exclusive write request.
[0282] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0283] It should be noted that in the embodiments of this application Figure 8 or Fig. 9 The division of modules by the target processor cluster or LLC shown is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, or may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units. It may also be implemented in the form of a combination of software and hardware.
[0284] It should be noted that in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0285] The embodiment of the present application provides a multi-core system, the multi-core system includes a target processor cluster and a last level cache LLC, wherein:
[0286] A target processor cluster is used to obtain a target atomic instruction, where the target atomic instruction is used to modify the data of a target address to be accessed, where the target address is an address in the LLC; a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction are sent to the LLC to modify the data of the target address, where the first exclusive read request includes the target address, and the first exclusive write request includes the target address and modified data, where the modified data is the data after the modification operation is performed on the data of the target address;
[0287] LLC is used to modify the data of the target address by receiving a first exclusive read request and a first exclusive write request sent by the target processor cluster.
[0288] The embodiment of the present application provides a computer device, which may be a server, and its internal structure diagram may be as follows: Fig.10 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0289] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the above embodiment are implemented.
[0290] An embodiment of the present application provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.
[0291] Those skilled in the art will understand that Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0292] In one embodiment, the multi-core system provided by the present application can be implemented in the form of a computer program. The computer program can be Fig.10 Runs on the computer device shown.
[0293] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0294] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in one embodiment" or "in some embodiments" appearing throughout the specification may not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments. The above description of each embodiment tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. For the sake of brevity, this article will not repeat them.
[0295] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist at the same time, and object B exists alone.
[0296] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0297] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.
[0298] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed on multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0299] In addition, all functional modules in the embodiments of the present application may be integrated into one processing unit, or each module may be a separate unit, or two or more modules may be integrated into one unit; the above-mentioned integrated modules may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0300] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.
[0301] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0302] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0303] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0304] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0305] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A data processing method for a multi-core system, characterized in that: The method is applied to a target processor cluster of a multi-core system, and the method comprises: Obtaining a target atomic instruction, where the target atomic instruction is used to modify data at a target address to be accessed, where the target address is an address in the last level cache LLC of the multi-core system; The modification operation is performed on the data at the target address by sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC, wherein the first exclusive read request includes the target address, and the first exclusive write request includes the target address and the modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
2. The method according to claim 1, characterized in that The target processor cluster includes a target core and a target cache that is communicatively connected to the target core, the target core is used to obtain the target atomic instruction, the target cache includes a cache module and an atomic instruction module, the cache module is used to store the target atomic instruction, and the atomic instruction module is used to generate the first exclusive read request and the first exclusive write request.
3. The method according to claim 1 or 2, characterized in that: The step of sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC to perform the modification operation on the data at the target address includes: Sending the first exclusive read request to the LLC; Receiving data of the target address sent by the LLC; Obtaining the modified data according to the data of the target address and the target atomic instruction; The first exclusive write request is sent to the LLC to modify the data at the target address to the modified data.
4. The method according to claim 1 or 2, characterized in that: The method further comprises: Receive execution data sent by the LLC, where the execution data is used to indicate an execution result of the modification operation performed on the data at the target address, where the execution result includes a modification success or a modification failure.
5. The method according to claim 4, characterized in that The method further comprises: In the case where the execution data indicates that the execution result is a modification failure, a second exclusive read request and a second exclusive write request corresponding to the target atomic instruction are sent to the LLC to perform the modification operation on the data at the target address, wherein the second exclusive read request includes the target address, and the second exclusive write request includes the target address and the modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
6. The method according to claim 1, characterized in that The target processor cluster is an NPU cluster, and the target atomic instruction includes a floating-point type atomic instruction, an atomic increment instruction, or an atomic decrement instruction.
7. A data processing method for a multi-core system, characterized in that: The method is applied to the last level cache LLC of a multi-core system, and the method comprises: By receiving a first exclusive read request and a first exclusive write request sent by a target processor cluster of the multi-core system, a modification operation is performed on the data of the target address, the target address is an address in the LLC, the first exclusive read request and the first exclusive write request correspond to a target atomic instruction, the target atomic instruction is used to perform the modification operation on the data of the target address, the first exclusive read request includes the target address, the first exclusive write request includes the target address and modification data, and the modification data is the data after the modification operation is performed on the data of the target address.
8. The method according to claim 7, characterized in that The step of performing a modification operation on the data at the target address by receiving a first exclusive read request and a first exclusive write request sent by a target processor cluster of the multi-core system includes: receiving the first exclusive read request; Sending the data of the target address to the target processor cluster; receiving the first exclusive write request; The data of the target address is modified according to the modification data.
9. The method according to claim 8, characterized in that The target processor cluster includes a target core, the target core is used to obtain the target atomic instruction, the LLC includes a monitoring module, the monitoring module is used to monitor the modification status of the data at the target address, the modification status is used to indicate whether the data at the target address is modified by other cores other than the target core, and after receiving the first exclusive write request, the method further includes: Acquire the modification status of the data at the target address; The step of modifying the data of the target address according to the modification data comprises: In a case where the modification status indicates that the data at the target address has not been modified by the other core, the data at the target address is modified according to the modification data.
10. The method according to claim 9, characterized in that The method further comprises: Execution data is sent to the target processor cluster, where the execution data is used to indicate an execution result of the modification operation on the data at the target address, and when the modification status indicates that the data at the target address is modified by the other cores, the execution data indicates that the execution result is modification failure.
11. The method according to claim 8, characterized in that After receiving the first exclusive read request, the method further includes: Setting the read-write attribute of the target address to read-only; After receiving the first exclusive write request, the method further includes: The read / write attribute of the target address is set to be readable and writable.
12. A target processor cluster of a multi-core system, characterized in that: The target processor cluster includes a target core and a target cache in communication with the target core, wherein: The target core is used to obtain a target atomic instruction, where the target atomic instruction is used to modify data at a target address to be accessed, where the target address is an address in the last level cache LLC of the multi-core system; The target cache is used to perform the modification operation on the data at the target address by sending a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction to the LLC, wherein the first exclusive read request includes the target address, and the first exclusive write request includes the target address and the modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
13. The target processor cluster according to claim 12, characterized in that: The target cache includes a cache module and an atomic instruction module, the cache module is used to store the target atomic instruction, and the atomic instruction module is used to generate the first exclusive read request and the first exclusive write request.
14. The target processor cluster according to claim 12 or 13, characterized in that: The target cache is used to: Sending the first exclusive read request to the LLC; Receiving data of the target address sent by the LLC; Obtaining the modified data according to the data of the target address and the target atomic instruction; The first exclusive write request is sent to the LLC to modify the data at the target address to the modified data.
15. The target processor cluster according to claim 12 or 13, characterized in that: The target cache is also used to: Receive execution data sent by the LLC, where the execution data is used to indicate an execution result of the modification operation performed on the data at the target address, where the execution result includes a modification success or a modification failure.
16. The target processor cluster according to claim 15, characterized in that: The target cache is also used to: In the case where the execution data indicates that the execution result is a modification failure, a second exclusive read request and a second exclusive write request corresponding to the target atomic instruction are sent to the LLC to perform the modification operation on the data at the target address, wherein the second exclusive read request includes the target address, and the second exclusive write request includes the target address and the modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
17. The target processor cluster according to claim 12, characterized in that: The target processor cluster is an NPU cluster, and the target atomic instruction includes a floating-point type atomic instruction, an atomic increment instruction, or an atomic decrement instruction.
18. A last level cache LLC of a multi-core system, characterized in that: The LLC includes: An execution module is used to modify the data at the target address by receiving a first exclusive read request and a first exclusive write request sent by a target processor cluster of a multi-core system, wherein the target address is an address in the LLC, the first exclusive read request and the first exclusive write request correspond to a target atomic instruction, and the target atomic instruction is used to perform the modification operation on the data at the target address, the first exclusive read request includes the target address, and the first exclusive write request includes the target address and modification data, and the modification data is the data after the modification operation is performed on the data at the target address.
19. The LLC according to claim 18, characterized in that The execution module is used for: receiving the first exclusive read request; Sending the data of the target address to the target processor cluster; receiving the first exclusive write request; The data of the target address is modified according to the modification data.
20. The LLC according to claim 19, characterized in that The target processor cluster includes a target core, and the target core is used to obtain the target atomic instruction. The LLC also includes a monitoring module, and the monitoring module is used to monitor the modification status of the data at the target address, and the modification status is used to indicate whether the data at the target address is modified by other cores other than the target core; The monitoring module is further configured to obtain the modification status of the data at the target address after receiving the first exclusive write request; The execution module is further configured to modify the data at the target address according to the modification data when the modification status indicates that the data at the target address has not been modified by the other cores.
21. The LLC according to claim 20, characterized in that The LLC also includes: An execution confirmation module is used to send execution data to the target processor cluster, wherein the execution data is used to indicate the execution result of the modification operation on the data at the target address. When the modification status indicates that the data at the target address is modified by the other cores, the execution data indicates that the execution result is a modification failure.
22. The LLC of claim 19, wherein: The LLC also includes: The attribute control module is used to set the read-write attribute of the target address to read-only after receiving the first exclusive read request; and set the read-write attribute of the target address to readable and writable after receiving the first exclusive write request.
23. A multi-core system, characterized in that: The multi-core system includes a target processor cluster and a last level cache LLC, wherein: The target processor cluster is used to obtain a target atomic instruction, where the target atomic instruction is used to perform a modification operation on data at a target address to be accessed, where the target address is an address in the LLC; a first exclusive read request and a first exclusive write request corresponding to the target atomic instruction are sent to the LLC to perform the modification operation on the data at the target address, where the first exclusive read request includes the target address, and the first exclusive write request includes the target address and modification data, where the modification data is data after the modification operation is performed on the data at the target address; The LLC is configured to perform a modification operation on the data at the target address by receiving the first exclusive read request and the first exclusive write request sent by the target processor cluster.
Citation Information
Cited By
Data transmission system and method
CN120723314A
Data transmission system and method
CN120723314B