Method and system for data synchronization based on software and hardware collaboration

Through the coordinated design of software and hardware, registers are used to mark cache row status and batch data is eliminated to LLC, which solves the complexity of cache consistency management in heterogeneous computing systems, and improves data synchronization efficiency and system performance.

CN120353401APending Publication Date: 2025-07-22上海芯联芯智能科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510523114.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In heterogeneous computing systems, the existing technology relies on complex MESI protocols to lead to high cache consistency management costs and it is difficult to efficiently synchronize data between different computing units, especially under heterogeneous architectures such as CPU, GPU, and accelerators, resulting in low data synchronization efficiency.

Method used

Through the coordinated design of software and hardware, registers are used to mark the start and end cache addresses of data to be synchronized. The hardware automatically marks the state of cache behavior to be removed, and batches the data to LLC during write operations, avoiding operations one by one, realizing accurate identification and efficient synchronization.

Benefits of technology

Reduces invalid data transmission, saves bandwidth and power consumption, reduces hardware interrupts and software instructions, improves data synchronization efficiency and system throughput, supports dynamic adjustment of data block size, and reduces CPU computing resource usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353401A_ABST
    Figure CN120353401A_ABST
Patent Text Reader

Abstract

The invention provides a data synchronization method and system based on software and hardware collaboration, and belongs to the field of computers. When to-be-synchronized data is declared, a first arithmetic unit stores an initial cache address of the to-be-synchronized data into a first register; when the first arithmetic unit executes the write operation on the to-be-synchronized data every time, marking a cache line corresponding to the write operation in the first arithmetic unit as a to-be-rejected state; when it is determined that writing of the to-be-synchronized data is completed, the first arithmetic unit stores an ending cache address in the to-be-synchronized data to a second register; and for each cache line from the starting cache address in the first register to the ending cache address in the second register, if the cache line is marked as a to-be-removed state, the first arithmetic unit removes the data cached in the cache line to the LLC. The method does not need to depend on an MESI protocol, data synchronization of different operation units is achieved through software and hardware cooperation, and the data synchronization efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method and system for data synchronization based on software and hardware cooperation. Background Art

[0002] Some current heterogeneous computing systems include different types of computing units such as CPUs, GPUs, accelerators, etc. Each computing unit includes multiple processing cores, and each processing core has its own cache. When multiple processing cores need to access the same address, cache coherence problems may occur. Cache coherence problems can cause a core to update the data stored in a memory address, but other cores cannot obtain the latest version of the data in time, resulting in incorrect software results.

[0003] Typical multi-core CPU systems add cache coherence management logic to each layer of the multi-level cache system, that is, the Modified, Exclusive, Shared, Invalid Protocol (MESI). However, due to different requests of their own architectures, different CPU manufacturers and IP suppliers may have different state transitions. The state machine of the MESI protocol is very complex, and the hardware implementation cost is very high.

[0004] In view of this, there is an urgent need to provide a solution that does not rely on the MESI protocol and can solve data synchronization between different computing units. Summary of the Invention

[0005] Embodiments of this application provide a method and system for data synchronization based on software and hardware cooperation. Without relying on the MESI protocol, data synchronization between different computing units is achieved through software and hardware cooperation, improving the efficiency of data synchronization. The technical solutions are as follows.

[0006] In a first aspect, a method for data synchronization based on software-hardware cooperation is provided, which is applied to a computing system. The computing system includes multiple arithmetic units and a last-level cache LLC. Each arithmetic unit includes a processing core, a first-level cache, and a second-level cache. The first-level cache is located inside the processing core, and the second-level cache is located outside the processing core and connected to the processing core. The method includes: when declaring data to be synchronized, the first arithmetic unit stores the starting cache address of the data to be synchronized into a first register; each time the first arithmetic unit performs a write operation on the data to be synchronized, it marks the cache line corresponding to the write operation in the first arithmetic unit as a state to be evicted; when it is determined that the data to be synchronized has been written completely, the first arithmetic unit saves the ending cache address in the data to be synchronized into a second register; for each cache line from the starting cache address in the first register to the ending cache address in the second register, if the cache line has been marked as a state to be evicted, the data cached in the cache line is evicted to the LLC, and the eviction state of the cache line is cleared; the second arithmetic unit obtains the data to be synchronized from the LLC.

[0007] The method provided in this embodiment realizes accurate identification of data to be synchronized and batch and efficient eviction of data to the LLC through software-hardware co-design. Each time a write operation is performed, the hardware automatically marks the corresponding cache line as a state to be evicted and records the actual modified data range. The starting and ending addresses of the data to be synchronized are declared through registers to limit the scope of hardware operations. By processing the cache lines marked as to be evicted, batch eviction is performed according to the address range, avoiding full-volume or one-by-one operations.

[0008] On the one hand, it reduces invalid data transmission and optimizes bandwidth and power consumption. Specifically, only the marked cache lines are synchronized, achieving accurate eviction and avoiding evicting unmodified data, thereby improving the efficiency of data synchronization. For example, if only part of the data in a cache line is modified, the traditional scheme needs to transmit the entire cache line (such as 64B), while this scheme only transmits the actually modified part.

[0009] On the other hand, it saves bandwidth. For example, if the data to be synchronized only accounts for 50% of the cache line, the bandwidth requirement is reduced by half, reducing the probability of bus contention and improving the overall throughput of the system.

[0010] On yet another hand, by automatically recording the modification status by hardware, it eliminates the time-consuming operation of software traversing the memory to check the modification bits, reducing operation latency and power consumption.

[0011] On yet another hand, batch eviction of cache lines according to the address range reduces the number of hardware interrupts or software instructions. For example, if 10 marked cache lines need to be synchronized, the traditional one-by-one operation requires 10 instructions, while this scheme is completed through a single range scan.

[0012] On the other hand, the synchronization range is set through registers to improve flexibility. For example, it supports dynamically adjusting the data block size. State flags and batch rejection are implemented by hardware to avoid software polling or complex logic and reduce the occupancy of CPU computing resources.

[0013] In some embodiments, determining that the data to be synchronized has been written completely includes:

[0014] Determining an operation count threshold based on the program code for obtaining the data to be synchronized;

[0015] During the process of performing a write operation on the data to be synchronized, recording the execution count of the write operation;

[0016] If the execution count of the write operation reaches the operation count threshold, determining that the data to be synchronized has been written completely.

[0017] In some embodiments, the determining an operation count threshold based on the program code for obtaining the data to be synchronized includes:

[0018] Obtaining the loop count from the loop statement of the program code for obtaining the data to be synchronized as the operation count threshold.

[0019] In some embodiments, after determining the operation count threshold based on the program code for obtaining the data to be synchronized, the method further includes:

[0020] Writing the operation count threshold to the compare register;

[0021] The recording the execution count of the write operation during the process of performing a write operation on the data to be synchronized includes:

[0022] Clearing the count register, and incrementing the value stored in the count register by one whenever a write operation is performed on the data to be synchronized;

[0023] The if the execution count of the write operation reaches the operation count threshold, determining that the data to be synchronized has been written completely includes:

[0024] If the value stored in the count register is the same as the value stored in the compare register, determining that the data to be synchronized has been written completely.

[0025] In some embodiments, after the first arithmetic unit stores the starting cache address of the data to be synchronized into the first register, the method further includes:

[0026] Setting the enable bit of the first register;

[0027] After the first arithmetic unit saves the end cache address in the data to be synchronized to the second register, the method further includes:

[0028] Determine whether the enable bit of the first register is set. If the enable bit of the first register is set, perform the step of removing the data cached in the cache line to the LLC.

[0029] After removing the data cached in the cache line to the LLC, the method further includes:

[0030] If each cache line from the start cache address to the end cache address has been removed to the LLC, clear the enable bit of the first register.

[0031] In some embodiments, the first arithmetic unit and the second arithmetic unit are used to cooperate in processing deep learning tasks. The data to be synchronized is the inference result or model parameter obtained by processing the deep learning task. The first arithmetic unit or the second arithmetic unit is a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), or an accelerator.

[0032] In a second aspect, a computing system for data synchronization based on software and hardware cooperation is provided. The computing system includes a plurality of arithmetic units and a last-level cache LLC. Each arithmetic unit includes a processing core, a first-level cache, and a second-level cache. The first-level cache is inside the processing core, and the second-level cache is outside the processing core and connected to the processing core. The computing system includes:

[0033] A first arithmetic unit, configured to store the start cache address of the data to be synchronized in a first register when declaring the data to be synchronized; each time a write operation is performed on the data to be synchronized, mark the cache line corresponding to the write operation in the first arithmetic unit as a to-be-removed state; when it is determined that the data to be synchronized has been written, save the end cache address in the data to be synchronized to a second register; for each cache line from the start cache address in the first register to the end cache address in the second register, if the cache line has been marked as a to-be-removed state, remove the data cached in the cache line to the LLC, and clear the to-be-removed state of the cache line;

[0034] A second arithmetic unit, configured to obtain the data to be synchronized from the LLC.

[0035] In some embodiments, the first arithmetic unit is configured to determine an operation count threshold based on the program code for obtaining the data to be synchronized; record the execution count of the write operation during the process of performing the write operation on the data to be synchronized; if the execution count of the write operation reaches the operation count threshold, determine that the data to be synchronized has been written completely.

[0036] In some embodiments, the first arithmetic unit is configured to obtain the loop count from the loop statement of the program code for obtaining the data to be synchronized as the operation count threshold.

[0037] In some embodiments, the first arithmetic unit is further configured to write the operation count threshold into a compare register; clear the count register, and increment the value stored in the count register by one each time a write operation is performed on the data to be synchronized; if the value stored in the count register is the same as the value stored in the compare register, determine that the data to be synchronized has been written completely.

[0038] In some embodiments, the first arithmetic unit is further configured to set the enable bit of the first register; determine whether the enable bit of the first register is set, and perform the step of purging the data cached in the cache line to the LLC when the enable bit of the first register is set; if each cache line from the start cache address to the end cache address has been purged to the LLC, clear the enable bit of the first register.

[0039] In some embodiments, the first arithmetic unit and the second arithmetic unit are configured to collaboratively process deep learning tasks, the data to be synchronized is an inference result or a model parameter obtained by processing the deep learning task, and the first arithmetic unit or the second arithmetic unit is a CPU, a GPU, a GPGPU or an accelerator.

[0040] In a third aspect, a computer device is provided. The computer device includes a processor, the processor is coupled to a memory, and at least one computer program instruction is stored in the memory. The at least one computer program instruction is loaded and executed by the processor to enable the computer device to implement the method provided in the first aspect or any optional manner of the first aspect. For the specific details of the computer device provided in the third aspect, reference may be made to the first aspect or any optional manner of the first aspect, which will not be elaborated herein.

[0041] Fourthly, a computer-readable storage medium is provided, in which at least one instruction is stored. When the instruction runs on a computer, the computer is caused to execute the method provided in the first aspect or any optional implementation manner of the first aspect.

[0042] Fifthly, a computer program product is provided. The computer program product includes one or more computer program instructions. When the computer program instructions are loaded and run on a computer, the computer is caused to execute the method provided in the first aspect or any optional implementation manner of the first aspect.

[0043] Sixthly, a chip is provided, including a memory and a processor. The memory is used for storing computer instructions, and the processor is used for calling and running the computer instructions from the memory to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0044] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. Description of the Drawings

[0045] Figure 1 is a schematic diagram of a complex coherence management unit in a multi-core CPU system provided by an embodiment of the present application;

[0046] Figure 2 is a schematic architecture diagram of a heterogeneous computing system provided by an embodiment of the present application;

[0047] Figure 3 is a flowchart of a deep learning method provided by an embodiment of the present application;

[0048] Figure 4 is a flowchart of a method for data synchronization based on software and hardware cooperation provided by an embodiment of the present application;

[0049] Figure 5 is a flowchart of a method for setting registers and write operations provided by an embodiment of the present application;

[0050] Figure 6 is a flowchart of a method for performing a clearing operation on data to be synchronized provided by an embodiment of the present application. Detailed Embodiments

[0051] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.

[0052] The following gives an example of the application scenario of the embodiments of the present application.

[0053] A multi-core CPU (Central Processing Unit) system usually includes multiple cores, each of which has its own cache. When multiple cores need to access the same address, cache coherence problems may occur. Cache coherence problems can lead to a situation where after one core updates the data stored in a memory address, other cores cannot obtain the latest version of the data in a timely manner, resulting in incorrect software results. A typical multi-core CPU system adds cache coherence management logic, namely the Modified, Exclusive, Shared, Invalid Protocol (MESI), to each layer of the multi-level cache system. Due to different requests of their own architectures, different CPU manufacturers and IP suppliers may have different state transitions. The state machine of the MESI protocol is very complex, and the hardware implementation cost is very high.

[0054] The main reason for this problem is that as a general-purpose processor, the CPU faces complex and ever-changing problems. Especially after adding an operating system, virtualization, and multi-threading, the entire system may need to run hundreds or thousands of tasks (or threads) in a time-sharing multiplexing or synchronous multi-threading mode, and multiple tasks or threads may belong to the same process. When the system distributes these tasks and threads to different cores of the multi-core CPU, problems of synchronization, data exchange, and sharing will inevitably occur. In order to optimize the overall performance, the CPU chooses to use complex hardware coherence logic to solve this problem, which leads to the emergence of complex coherence management units in the multi-core CPU system, as Figure 1 shown.

[0055] Current mainstream heterogeneous architectures such as Figure 2 shown, attached Figure 2 The computing system shown includes a CPU, a GPU (Graphics Processing Unit), an accelerator, a Last Level Cache (LLC), and a memory controller.

[0056] The CPU includes multiple cores, a Level 1 cache (L1), and a Level 2 cache (L2). The Level 1 cache (L1) is inside the core. The Level 2 cache (L2) is outside the core.

[0057] The GPU contains multiple Streaming Multiprocessors (SMs), a Level 1 cache (L1), and a Level 2 cache (L2). The Level 1 cache (L1) is inside the streaming multiprocessor. The Level 2 cache (L2) is outside the streaming multiprocessor.

[0058] The accelerator contains a core, a Level 1 cache (L1), and a Level 2 cache. The Level 1 cache (L1) is inside the core. The Level 2 cache (L2) is outside the core.

[0059] The LLC provides cache services, and the memory controller manages access to the main memory (RAM).

[0060] Appendix Figure 2 The problem with the hardware architecture shown is that it is difficult to solve the coherence problem of the LLC. First, in terms of the interface protocol, the current mainstream interface protocols, whether it is ACE (AXI Coherency Extensions), OCP, Tilelink, CHI (Coherent Hub Interface) inside the SOC or UCIE, PCIE, Cxl in a multi-die system, etc., are essentially serving a system with a multi-core CPU as the main perspective and cannot be well compatible with the hardware architectures of GPUs and accelerators themselves (or these IPs need to follow the complex coherence protocol of the CPU, i.e., MESI, or some MESI states). Second, different GPGPUs and accelerators do not have a unified cache architecture, making it difficult to solve this coherence problem.

[0061] The previous text analyzed the cache hierarchy and coherence management characteristics of its multi-core CPU, the cache characteristics of GPGPUs, and the problems of their heterogeneous architectures. Next, we will try to solve such problems according to the characteristics of their own hardware architectures and software processes.

[0062] Here, first, take the process of processing natural language with a deep learning-based method as an example.

[0063] Please refer to Figure 3 , Figure 3 which is a flowchart of deep learning. Briefly introduce each process:

[0064] The main work in the data construction stage is to construct a training corpus according to the requirements of the task, also known as a corpus. With the continuous development of natural language processing research, many tasks have publicly available benchmark test sets (Benchmarks), which can be conveniently used for model training and horizontal comparison between models. For tasks without publicly available data, the method of manual annotation can also be used to construct a training corpus.

[0065] The main work in the data preprocessing stage is to use basic natural language processing algorithms to process the original input from aspects such as vocabulary, syntax, structure, and semantics to provide a basis for feature construction. Different modules and processes are adopted according to the different languages being processed and the tasks. For Chinese, word segmentation is usually required, and for English, stemming and word normalization are usually required. After that, according to the requirements of feature construction, part-of-speech tagging, syntactic analysis, semantic role tagging, etc. may also be required.

[0066] Deep learning is a subset of machine learning that transforms raw data into more abstract representations through multi - layer feature transformation. These learned representations can, to a certain extent, completely replace manually designed features, and this process is also called Representation Learning. Different from the discrete sparse representations commonly used in feature - engineering - based methods, deep - learning algorithms usually use Distributed Representation, where features are represented as low - dimensional dense vectors. Distributed representations usually start from underlying features and are obtained through multiple non - linear transformations. Since deep structures can increase the reusability of features, the representational ability increases exponentially. Therefore, the key to representation learning is to construct multi - level feature representations with a certain depth. With the continuous in - depth research of deep learning and the rapid development of computing power, the depth of the model has increased from the early 5 - 10 layers to hundreds of layers now. As the depth of the model continues to increase, its feature representational ability also continuously enhances, making the prediction part in the deep - learning model simpler and the prediction easier.

[0067] Based on the processes at each stage here, first, the data construction stage may include sensitive - character filtering, duplicate - checking, etc., which involve sequence matching and recursive operations. In the data pre - processing stage, tasks such as word segmentation need to convert continuous character sequences into word sequences. The division of words in each sentence is affected by the entire sentence and even the context, and the relationship between the front and back needs to be considered. These two types of task processes are more suitable for processing on a general - purpose CPU. The last step of the process, model learning, is more suitable for parallelized acceleration units such as NPUs and GPGPUs. It should be noted here that although each process may contain many operations, from the overall process, its essence is still a sequential execution process, and the CPU needs to schedule the corresponding acceleration unit or itself to start the next process after each process is completed.

[0068] Exemplarily, the program code for word segmentation processed on the CPU is as follows.

[0069] Program Code 1:

[0070] Input: The Chinese character sequence to be segmented x = {C1, C2,.., Cn}

[0071] Output: The word - segmentation result y = {w1, w2,.., wm}

[0072] src = [], tgt = []; / / Initialize;

[0073] for i = 1 to n do

[0074] foreach item ∈ sre do

[0075] item1 = c1; / / The current character is the start of a new word;

[0076] item2 = item[item.length] + ci; / / Append the current character to the last candidate word in item;

[0077] Add item1 and item2 to tgt;

[0078] end

[0079] Use the scoring function SCORE to score all the word segmentation results in tgt;

[0080] Sort the scoring results in tgt and keep the top K;

[0081] src = tgt;

[0082] tgt = [];

[0083] end

[0084] return src[1]; / / Return the best result in src

[0085] Model parameter learning code processed on the acceleration unit:

[0086] Input: Training data D = (xi, yi)

[0087] Output: Model parameter α

[0088] for i = 1 to T do / / T rounds of iteration;

[0089] foreach (x, y) ∈ D do

[0090] z = arg max SCORE(y):

[0091] y ∈ GEN(x)

[0092] if z ≠ y then

[0093] α = α + Φ(x, y) - Φ(x, z);

[0094] end

[0095] end

[0096] end

[0097] return a

[0098] Here, X and Y are the Chinese character sequence x and the word segmentation result y in the first code segment.

[0099] From this, it can be seen that the data flow here is (taking the Figure 1 hardware architecture as an example):

[0100] Step 1: The CPU sends a read request to read the Chinese character sequence X from the main memory or LLC, reads the Chinese character sequence X into its own L2 cache and L1 cache of the CPU, and performs a series of word segmentation operations on the Chinese character sequence X.

[0101] Step 2: After the CPU finishes processing the Chinese character sequence X, the word segmentation result Y is obtained. The CPU writes the word segmentation result Y into its own L2 cache and L1 cache of the CPU.

[0102] Step 3: The GPGPU or NPU receives the scheduling from the CPU (interrupt or other means) and reads the values of the Chinese character sequence X and the word segmentation result Y.

[0103] For the Chinese character sequence X, the Chinese character sequence X already exists in the LLC, and the Chinese character sequence X in the LLC is in a completely consistent (clean) state with the Chinese character sequence X in the main memory. The GPGPU or NPU can directly store the Chinese character sequence X in the LLC into its own L2 cache or L1 cache.

[0104] Step 4: When the GPGPU or NPU attempts to access the word segmentation result Y, the LLC proxy sends a read request to the CPU for the access request of Y, and uses the state machine of the MESI protocol to change the states of the L1 cache and L2 cache of the CPU, and returns the data to the GPGPU or NPU.

[0105] Step 5: The acceleration unit parallelizes the processing of the Chinese character sequence X and the word segmentation result Y, trains the parameter a, and the parameter a will be stored in its own L1 cache or L2 cache of the acceleration unit.

[0106] Step 6: If the CPU needs the parameter a, similarly, the CPU needs to send a request to the LLC. Here, the LLC hardware does not know the state of the parameter a in the GPGPU or NPU, and the software needs to synchronize the parameter a to the CPU. Or, the GPGPU or NPU can pre-remove the parameter a stored at its end to the LLC in advance.

[0107] Note that this is different from Step 3 because the GPGPU and NPU do not have a consistent state.

[0108] In the parallel computing GPGPU or NPU, there is no cache coherence management unit similar to the MESI protocol guaranteed by hardware, so the software needs to "remove" its internal data to the LLC or mem at the appropriate time by itself.

[0109] In summary, it can be concluded that:

[0110] 1. For the read-only data X, not much processing is required, and the current architecture can fully meet the requirements.

[0111] 2. For the dirty data Y that the CPU processes first, the LLC can also process it and use a complex protocol to fetch this data into the LLC.

[0112] 3. For the dirty data a in the GPGPU and NPU, it is difficult for the LLC to process, and only software can be borrowed for processing.

[0113] 4. For dirty data, if relying on the current coherence management component to solve the problem, it can be understood as passive, that is, only when the current IP needs data, it has to go through the LLC to proxy these data, which is a passive acquisition rather than an active push.

[0114] In this embodiment, the cache strategy for such data that needs to be frequently interacted between different computing units (such as CPU or GPGPU) is optimized. For the data that needs to be written back, that is, the data that needs to be synchronized, such as the word segmentation result y = {w1, w2, w3,..., wm} and the model parameter α in the above example, after obtaining the optimal solution several times (that is, after multiple writes), when performing the last write operation, the data needs to be removed from the L1 cache or L2 cache of the current computing unit to the LLC. For example, during the model training process, the CPU needs to load the parameter α updated by the GPU for inference or other tasks.

[0115] If only relying on software to perform the data removal operation, since it is impossible to accurately identify which data needs to be synchronized, two problems may occur. One problem is that all the data in the entire cache line in the computing unit is removed from the computing unit to the LLC, resulting in waste of bandwidth and power consumption. Another problem is that removing the data that needs to be removed one by one will undoubtedly bring a large amount of latency or power consumption to the hardware or software, thus affecting the efficiency.

[0116] In view of this, the following software and hardware collaborative unit can be designed to accurately and quickly implement the data synchronization operation.

[0117] Please refer to the appendix Figure 4 The appendix Figure 4 is a flowchart of a method for data synchronization based on software and hardware collaboration provided by an embodiment of the present application. The appendix Figure 4 The method shown is applied to, for example, Figure 2The heterogeneous computing system shown, the computing system includes multiple computing units and an LLC, each computing unit includes a processing core and a cache space, and the cache space includes, for example, a level-1 cache built into the processing core or / and a level-2 cache outside the processing core. The method is, for example, applied to data synchronization between computing unit A and computing unit B, and includes the following steps.

[0118] Step S401, when declaring the data to be synchronized, computing unit A stores the starting cache address of the data to be synchronized into the first register.

[0119] Computing unit A is, for example, a CPU, GPU, accelerator, GPGPU or NPU. The data to be synchronized is the data in the cache space of computing unit A that needs to be synchronized to the LLC. In some embodiments, computing unit A is used to execute a deep learning task, and the data to be synchronized is the model parameters or inference results obtained by computing unit A by executing the deep learning task.

[0120] The starting cache address is, for example, the address of the first data in the memory space of the data to be synchronized. For example, if the data to be synchronized is the array y representing the model parameters in the above program code, the starting cache address is the cache address of the first data w1 in the array y.

[0121] The first register is used to record the cache address of the first data to be purged. The first register is, for example, the L1 cache, L2 cache inside computing unit A or the clr_start_addr register of computing unit A (cache controller).

[0122] Exemplarily, when the CPU executes the above program code 1 for the word segmentation task, when allocating y (declaring the array y), it sets the starting cache address of the array y to the starting address start_addr, the starting cache address of the array y is the address of w1, and stores the starting cache address of the array y into the clr_start_addr register.

[0123] In some embodiments, computing unit A also sets the enable bit in the first register. The enable bit is used to indicate that the data marked as to be purged starting from the starting cache address of the data to be synchronized needs to be purged after the last write is completed. For example, set the enable bit corresponding to the clr_start_addr register to 1. Enable means that all addresses from start_addr (such as w1) to clr_end_addr (such as wm), that is, the entire memory range of the array y where the cache lines set to the wait to be cleared state need to be cleared for the entire array after the last write is completed.

[0124] Through the above-mentioned hardware registers and status bits, it is helpful to mark the cache management policy at the beginning of the data life cycle (when it is declared). On the one hand, compared with the method of traversing all cache lines through software to eliminate data, the scope of data elimination is more precisely controlled. Only the cache line where the data to be synchronized is located needs to be processed, which improves the efficiency. On the other hand, when performing a write operation, only the status bit needs to be marked, and finally all the data to be synchronized is eliminated in batches at one time, avoiding the performance jitter caused by frequent data synchronization operations.

[0125] Step S402: Each time the arithmetic unit A performs a write operation on the data to be synchronized, it marks the cache line corresponding to the write operation as the to-be-eliminated state.

[0126] For example, in each write operation on the array y, the arithmetic unit A automatically sets the wait_to_be_cleared status bit of the corresponding cache line to 1. The wait_to_be_cleared status bit being set to 1 is used to indicate that this cache line ultimately needs to be synchronized or eliminated to a higher-level cache (such as LLC). Even if the cache line is overwritten and written multiple times, the wait_to_be_cleared status remains 1 until the final batch clearing operation is completed.

[0127] Step S403: When it is determined that the writing of the data to be synchronized is completed, the arithmetic unit A saves the cache address of the last data in the data to be synchronized to the second register.

[0128] The second register is used to record the cache address of the first data to be eliminated. For example, the second register is the clr_end_addr register.

[0129] In the above program code, after the traversal in x = item[] in program code 1 is completed and the optimal solution of y = {} is solved, that is, at the last n of the for loop or the last condition that satisfies the jump-out condition of the while loop, the clr_end_addr register is set to save the address of the last data wm (the last data to be eliminated) in the array y.

[0130] By storing the cache address of the first data in the first register (clr_start_addr = w1) when declaring the data to be synchronized, and storing the cache address of the last data in the second register (clr_end_addr = wm) after the loop ends, it is convenient for the hardware to accurately lock the range of cache lines to be cleared (from w1 to wm). For example, only the cache lines within this range can be processed, avoiding the accidental clearing of other irrelevant data.

[0131] In a possible implementation for determining the completion of writing the data to be synchronized, the arithmetic unit A records the number of write operations. If the number of write operations reaches the operation count threshold, it is determined that the writing of the data to be synchronized is completed. Regarding how to determine the operation count threshold, in a possible implementation, the arithmetic unit A determines the operation count threshold based on the program code used to obtain the data to be synchronized. For example, the arithmetic unit A obtains the loop count (such as the for loop count or the while loop count) from the loop statements in the program code of the data to be synchronized as the operation count threshold. For example, a third register and a fourth register are set. The third register is used to store the write operation count threshold, and the fourth register is used to record the number of write operations that have been executed currently. The third register is, for example, a compare register, and the fourth register is, for example, a count register. When performing the write operation for the first time based on the program code corresponding to the data to be synchronized, the loop count (such as the for loop count or the while loop count) is obtained from the program code and saved in the third register. The value saved in the fourth register is cleared. Whenever a write operation is performed based on the program code corresponding to the data to be synchronized, the number of write operations recorded in the fourth register is incremented by one. It is determined whether the number saved in the fourth register is equal to the number saved in the third register. When the number saved in the fourth register is equal to the number saved in the third register, it is determined that the writing of the data to be synchronized is completed.

[0132] Step S404: Read the starting cache address of the data to be synchronized saved in the first register, and read the ending cache address of the data to be synchronized saved in the second register. Traverse each cache line from the starting cache address to the ending cache address, and detect whether the traversed cache line is in the state to be excluded. If the cache line is in the state to be excluded, the data cached in the cache line is excluded to the LLC, and then the to-be-excluded state of the cache line is cleared. When the data clearing is completed, the cache line where the data to be synchronized is located is removed from the cache space of the arithmetic unit A.

[0133] For example, when determining the final data y (when returning the word segmentation result (return src)) based on the above program code, a signal (such as a special instruction or an interrupt) is sent to the arithmetic unit A. In response to the received signal, the arithmetic unit A generates a target address range [w1, wm] according to clr_start_addr (starting address) and clr_end_addr (ending address). It traverses all cache lines within this range and checks whether the wait_to_be_cleared status bit is 1. If the wait_to_be_cleared status bit is 1, the cache line data is written back to the LLC and the wait_to_be_cleared status bit is cleared; if the wait_to_be_cleared status bit is 0, the cache line data is retained, thus achieving batch clearing. After all cache lines waiting to be purged in the cache are cleared, the enable bit in clr_start_addr is cleared.

[0134] Compared with the method of traversing each data in the data to be synchronized and separately calling the clearing instruction for each data, in this embodiment, the batch clearing operation is triggered by hardware once, and the data clearing efficiency is higher.

[0135] Registers to be added:

[0136] Clr_start_addr:

[0137] Start_addr[m:n] Enable

[0138] Clr_end_addr:

[0139] End_addr[m:n]

[0140] Add 1 bit to the Cache status bit:

[0141] Wait_to_be_cleared

[0142] When performing a write operation that requires fast synchronization or purging, the wait_to_be_cleared status bit in the cache line is set to 1, and after the data in the cache line is purged, the data in the cache line is cleared, and the wait_to_be_cleared status bit is initially set to 0. Among them, when the wait_to_be_cleared status bit of a cache line is 1, the arithmetic unit A will purge the cache line.

[0143] Step S405, the arithmetic unit B obtains the data to be synchronized from the LLC.

[0144] Through the above process, the efficiency is improved and the problem of software processing data synchronization is avoided.

[0145] For example, please refer to Appendix Figure 5 , Appendix Figure 5 FIG.

[0146] Step 1: Declare the data block that needs to be quickly synchronized.

[0147] Step 2: Set the register clr_start_addr based on the cache address of the first data in the data block.

[0148] Step 3: Determine whether the data block needs to be quickly synchronized.

[0149] Step 4: If the data block needs to be quickly synchronized, set the Wait_to_be_cleared status bit of each cache line where the data block is located to 1.

[0150] Step 5: If the data block does not need to be quickly synchronized, do not set the Wait_to_be_cleared status bit of each cache line where the data block is located.

[0151] For example, please refer to Appendix Figure 6 , Appendix Figure 6 FIG.

[0152] Step 1: Determine that the operation of the data block that needs to be quickly synchronized has ended.

[0153] By performing the subsequent steps after determining that the data block has ended its operation, it is possible to avoid triggering a synchronization operation when the data is not stable, preventing dirty data (intermediate state) from being written to the LLC.

[0154] Step 2: Determine the cache address of the last data in the data block that needs to be quickly synchronized, and write the cache address of the last data to the register clr_end_addr.

[0155] By writing the cache address of the last data to the register clr_end_addr, the end point of the synchronization range is marked. The register clr_end_addr and the register clr_start_addr jointly define an exact address range ([start, end]), and the cache controller only processes the cache lines within this range, avoiding accidentally deleting other data.

[0156] Step 3: Determine whether the enable bit in the register clr_start_addr is set to 1.

[0157] By checking whether the enable bit in the register clr_start_addr is set, the synchronization function is verified to be enabled, which is equivalent to providing a switch control that allows the flexible enabling or disabling of the synchronization mechanism to avoid the mis-triggering of data synchronization actions.

[0158] Step 4: If the enable bit in the register clr_start_addr is set to 1, starting from the address saved in the register clr_start_addr, the cache line is evicted from the L1 cache or the L2 cache to the LLC, and the Wait_to_be_cleared status bit of this cache line is cleared.

[0159] Step 5: Determine the starting address (addr + n - 1) of the next cache line.

[0160] For example, according to the cache line size n, jump to the starting address of the next cache line. For a 64B cache line, for instance, jump from address 0x1000 to address 0x1040.

[0161] Step 6: Check whether the current cache line is within the address range (addr <= clr_end_addr) saved in the register clr_end_addr.

[0162] Step 7: Check whether the Wait_to_be_cleared status bit of the cache line is 1.

[0163] By checking whether the Wait_to_be_cleared status bit of the cache line is 1, only the data marked for synchronization is cleared, achieving more precise synchronization.

[0164] Step 8: If the current cache line does not exceed the address range saved in the register clr_end_addr and the Wait_to_be_cleared status bit of this cache line is 1, the data in this cache line is evicted to the LLC, and the Wait_to_be_cleared status bit of this cache line is cleared.

[0165] Step 9: If the current cache line does not exceed the address range saved in the register clr_end_addr and the Wait_to_be_cleared status bit of this cache line is 0, return to Step 5.

[0166] Step 10: If the current cache line exceeds the address range saved in the register clr_end_addr, clear the enable bit in the register clr_start_addr and end the data synchronization operation.

[0167] Each embodiment in this specification is described in a progressive manner. Similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.

[0168] The terms "first" and "second" in the description and claims of the embodiments of this application are used to distinguish different objects, rather than to describe a specific order of the objects, nor can they be understood as indicating or implying relative importance. For example, the first arithmetic unit and the second arithmetic unit are used to distinguish different arithmetic units, rather than to describe a specific order of the arithmetic units, nor can it be understood that the first arithmetic unit is more important than the second arithmetic unit.

[0169] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive SolidState Disk (SSD)).

[0170] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for data synchronization based on software and hardware collaboration, characterized in that, Applied to a computing system, the computing system includes a plurality of arithmetic units and a last-level cache LLC. Each arithmetic unit includes a processing core, a first-level cache, and a second-level cache. The first-level cache is inside the processing core, and the second-level cache is outside the processing core and connected to the processing core. The method includes: When declaring the data to be synchronized, the first arithmetic unit stores the starting cache address of the data to be synchronized into a first register; Each time the first arithmetic unit performs a write operation on the data to be synchronized, it marks the cache line corresponding to the write operation in the first arithmetic unit as a state to be evicted; When it is determined that the writing of the data to be synchronized is completed, the first arithmetic unit saves the ending cache address in the data to be synchronized to a second register; For each cache line from the starting cache address in the first register to the ending cache address in the second register, if the cache line has been marked as a state to be evicted, the first arithmetic unit evicts the cached data in the cache line to the LLC and clears the state to be evicted of the cache line; The second arithmetic unit obtains the data to be synchronized from the LLC.

2. The method according to claim 1, characterized in that, The determining that the writing of the data to be synchronized is completed includes: Based on the program code used to obtain the data to be synchronized, determining an operation count threshold; During the process of performing a write operation on the data to be synchronized, recording the execution count of the write operation; If the execution count of the write operation reaches the operation count threshold, determining that the writing of the data to be synchronized is completed.

3. The method according to claim 2, characterized in that, The based on the program code used to obtain the data to be synchronized, determining an operation count threshold includes: Obtaining the loop count from the loop statement in the program code used to obtain the data to be synchronized as the operation count threshold.

4. The method according to claim 2 or 3, characterized in that, After the based on the program code used to obtain the data to be synchronized, determining an operation count threshold, the method further includes: Writing the operation count threshold to the compare register; The during the process of performing a write operation on the data to be synchronized, recording the execution count of the write operation includes: Clearing the count register to zero, and each time a write operation is performed on the data to be synchronized, incrementing the value saved in the count register by one; The if the execution count of the write operation reaches the operation count threshold, determining that the writing of the data to be synchronized is completed includes: If the value saved in the count register is the same as the value saved in the compare register, determining that the writing of the data to be synchronized is completed.

5. The method according to claim 1, characterized in that After the first arithmetic unit stores the starting cache address of the data to be synchronized into the first register, the method further includes: Setting the enable bit of the first register; After the first arithmetic unit saves the ending cache address in the data to be synchronized to the second register, the method further includes: Judging whether the enable bit of the first register is set, and if the enable bit of the first register is set, performing the step of evicting the cached data in the cache line to the LLC; After evicting the data cached in the cache line to the LLC, the method further includes: If each cache line from the starting cache address to the ending cache address has been evicted to the LLC, clear the enable bit of the first register.

6. The method according to claim 1, wherein The first arithmetic unit and the second arithmetic unit are used to collaboratively process deep learning tasks. The data to be synchronized is the inference result or model parameter obtained by processing the deep learning task. The first arithmetic unit or the second arithmetic unit is a CPU, GPU, GPGPU, or accelerator.

7. A computing system for data synchronization based on software and hardware cooperation, characterized in that, The computing system includes a plurality of arithmetic units and a last-level cache LLC. Each arithmetic unit includes a processing core, a first-level cache, and a second-level cache. The first-level cache is inside the processing core, and the second-level cache is outside the processing core and connected to the processing core. The computing system includes: A first arithmetic unit, configured to store the starting cache address of the data to be synchronized in a first register when declaring the data to be synchronized; mark the cache line corresponding to the write operation in the first arithmetic unit as a state to be evicted each time a write operation is performed on the data to be synchronized; when it is determined that the data to be synchronized has been written completely, save the ending cache address in the data to be synchronized to a second register; for each cache line from the starting cache address in the first register to the ending cache address in the second register, if the cache line has been marked as a state to be evicted, evict the data cached in the cache line to the LLC, and clear the state to be evicted of the cache line; A second arithmetic unit, configured to obtain the data to be synchronized from the LLC.

8. The computing system according to claim 7, wherein The first arithmetic unit is configured to determine an operation count threshold based on the program code for obtaining the data to be synchronized; record the execution count of the write operation during the process of performing the write operation on the data to be synchronized; If the execution count of the write operation reaches the operation count threshold, determine that the data to be synchronized has been written completely.

9. A computer device, characterized in that, The computer device includes: a processor, the processor is coupled to a memory, and at least one computer program instruction is stored in the memory. The at least one computer program instruction is loaded and executed by the processor to enable the computer device to implement the method according to any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, At least one instruction is stored in the storage medium. When the instruction runs on a computer, the computer is caused to execute the method according to any one of claims 1-6.

Citation Information

Cited By

  • Write operation method and electronic device

    CN122653547A