Operator Parallel Processing Method, Electronic Device, Storage Medium, and Program Product
By using multiple collaborative thread bundles to execute the processing steps of operators in parallel when processing large-scale data, and using synchronization strategies and slicing strategies, the large read and write overhead problems caused by serial execution are solved, and more efficient computing performance is achieved.
Patent Information
- Application Number
- CN202510201016.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The existing operator fusion strategy is very expensive when processing large-scale data due to serial execution methods, which affects the performance of operators and models.
By applying synchronization strategies and slicing strategies based on multiple collaborative thread bundles, each processing step of the operator is executed in parallel, and data is passed using shared register blocks to realize parallel fusion of operators.
It significantly reduces the total time of operator processing, improves the overall computing speed, and improves the performance of operators and models.
Smart Images

Figure CN119718675B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an operator parallel processing method, an electronic device, a storage medium, and a program product. Background Art
[0002] In a neural network model, operator fusion processing is usually used to reduce the overhead of intermediate data reading and writing, thereby accelerating the model. The core idea of operator fusion processing is that the subsequent operator directly reads the calculation result of the previous operator from the register, without first writing the calculation result of the previous operator to external memory and then having the subsequent operator read the data from the external memory, thus greatly reducing the overhead of data reading and writing.
[0003] However, current operator fusion strategies often adopt a serial execution method, that is, in the same thread, the subsequent operator needs to wait for the previous operator to complete data reading and calculation before it can perform its own calculation and writing operations. Due to the large overhead of reading and calculation of the previous operator, the waiting time of the subsequent operator is long, affecting the performance of the operator; and the calculation and reading / writing process of the subsequent operator often also takes a certain amount of time. Since the serial execution method cannot effectively hide the latency of the calculation and reading / writing process of the subsequent operator, this will result in the reading / writing overhead of the subsequent operator being particularly obvious when processing large-scale data, thereby making the performance of the operator and the model poor and affecting the overall calculation efficiency. Summary of the Invention
[0004] The present invention provides an operator parallel processing method, an electronic device, a storage medium, and a program product to solve the defect that the serial fusion strategy of operators in related technologies has a large reading / writing overhead when processing large-scale data, affecting the performance of operators and models.
[0005] The present invention provides an operator parallel processing method, including:
[0006] Determine steps with data dependencies based on the processing steps of a first operator and the processing steps of a second operator, where the output of the first operator is the input of the second operator;
[0007] Based on multiple cooperative warps, apply a synchronization strategy and a splitting strategy to perform parallel execution on the processing steps of the first operator and the processing steps of the second operator;
[0008] Wherein, the multiple cooperative warps share register resources, the synchronization strategy refers to inserting synchronization operations between the steps with data dependencies, and the splitting strategy refers to splitting the input data of the first operator and the input data of the second operator.
[0009] According to an operator parallel processing method provided by the present invention, the output of the first operator is passed to the second operator in the form of a shared register block.
[0010] According to an operator parallel processing method provided by the present invention, each processing step of the first operator includes a segmentation and caching step, a data loading step, a first calculation step, and a result storage step;
[0011] The segmentation and caching step is used to segment the input data of the first operator and sequentially store the segmented multiple data blocks into the cache;
[0012] The data loading step is used to segment each data block in the cache, sequentially load the segmented multiple data slices into the input buffer, and read and reuse the data slices in the input buffer;
[0013] The first calculation step is used to perform calculations based on the data slices read from the input buffer to obtain result data;
[0014] The result storage step is used to write the result data into the shared register block.
[0015] According to an operator parallel processing method provided by the present invention, each processing step of the second operator includes a data reading step, a second calculation step, and a result writing step;
[0016] The data reading step is used to read the result data from the shared register block and segment the result data based on the memory layout of the second operator to obtain multiple result slices;
[0017] The second calculation step is used to perform calculations on the multiple result slices to obtain a calculation result;
[0018] The result writing step is used to write the calculation result to a specified storage location.
[0019] According to an operator parallel processing method provided by the present invention, there is a data dependency between the segmentation and caching step, the data loading step, and the first calculation step of the first operator, and there is a data dependency between the result storage step of the first operator and the data reading step of the second operator.
[0020] According to an operator parallel processing method provided by the present invention, it further includes:
[0021] During the execution process, monitor the usage of the shared register block;
[0022] When it is detected that the usage of the shared register block exceeds the hardware limit, an automatic tuning method is adopted to determine the optimal size of the shared register block and the splitting strategy.
[0023] According to an operator parallel processing method provided by the present invention, the step of adopting an automatic tuning method to determine the optimal size of the shared register block and the splitting strategy includes:
[0024] Based on the hardware information and the shape of the operator data, determine a plurality of candidate register block sizes and a plurality of candidate splitting strategies;
[0025] Based on the plurality of candidate register block sizes and the plurality of candidate splitting strategies, determine a variety of candidate configurations;
[0026] Evaluate the performance of each candidate configuration, and based on the performance evaluation results of each candidate configuration, determine the optimal configuration, where the optimal configuration includes the optimal size of the shared register block and the splitting strategy.
[0027] The present invention also provides an operator parallel processing device, including:
[0028] A step determination unit, configured to determine steps with data dependencies based on the processing steps of the first operator and the processing steps of the second operator, where the output of the first operator is the input of the second operator;
[0029] A parallel execution unit, configured to apply a synchronization strategy and a splitting strategy based on a plurality of cooperative warps to perform parallel execution on the processing steps of the first operator and the processing steps of the second operator; wherein, register resources are shared among the plurality of cooperative warps, the synchronization strategy refers to inserting a synchronization operation between the steps with data dependencies, and the splitting strategy refers to splitting the input data of the first operator and the input data of the second operator.
[0030] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, where when the processor executes the computer program, the operator parallel processing method as described in any one of the above is implemented.
[0031] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the operator parallel processing method as described in any one of the above is implemented.
[0032] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the operator parallel processing method as described in any one of the above is implemented.
[0033] The operator parallel processing method, electronic device, storage medium, and program product provided by the present invention can, based on the processing steps of the first operator and the processing steps of the second operator, determine the steps with data dependencies, laying a solid foundation for subsequent parallel execution. By using multiple cooperative warps and applying a synchronization strategy and a splitting strategy to parallelly execute the processing steps of the first operator and the second operator, the total time for operator processing can be significantly reduced, and the overall computing speed can be improved. During the parallel execution process, by using the splitting strategy to split the input data of the operator, the large-scale data that originally needed to be serially processed can be split into multiple small pieces. Multiple cooperative warps can simultaneously process different small pieces. Moreover, while one cooperative warp is executing an operator step for a small piece of data, other cooperative warps can parallelly execute other steps of the operator for other small pieces, thereby achieving latency hiding, that is, reducing the time waiting for a certain step to complete, and thus shortening the total processing time. At the same time, by using the synchronization strategy to insert synchronization operations between steps with data dependencies, it can ensure the correct transfer of data between dependent steps, avoiding data competition and inconsistency problems.
[0034] In addition, the present invention uses multiple cooperative warps to parallelly execute the processing steps of the first operator and the second operator, and the cooperative warps can share register resources. Thus, the output of the first operator can be passed to the second operator through the shared register block, realizing the parallel fusion of operators. Compared with the traditional serial fusion method, parallel fusion can significantly improve the performance of operators and models. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 is a schematic flowchart of the operator parallel processing method provided by the present invention;
[0037] Figure 2 is a schematic diagram of the pipeline layout provided by the present invention;
[0038] Figure 3 is a schematic structural diagram of the operator parallel processing device provided by the present invention;
[0039] Figure 4 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] During the calculation process of a neural network model, different operators have different resource consumption patterns according to their characteristics. Specifically, tcore (Tensor Core) operators, such as mma (Matrix Multiply-Accumulate) and conv (convolution), etc., due to the need for large-scale numerical calculations, the computational volume of such operators is relatively large. While vcore (Vector Core) operators, such as add (addition), relu (activation function), etc., are more involved in data reading and storage operations, so the memory access overhead of such operators is relatively large. It should be understood that tcore operators and vcore operators are operators designed specifically for accelerating specific types of computing tasks. Among them, tcore operators focus on providing high-performance matrix multiplication and accumulation calculations, while vcore operators focus on the optimization of vector computing tasks. The two are mainly used to improve the efficiency of artificial intelligence chips such as GPUs (Graphics Processing Unit) in high-performance computing scenarios such as deep learning and scientific computing.
[0042] To improve the execution efficiency of the model, an optimization strategy commonly adopted is to fuse tcore operators and vcore operators. Taking the tcore operator as the previous operator and the vcore operator as the subsequent operator, the core idea of this fusion process is to directly use the result calculated by the tcore operator as the input of the vcore operator, that is, write the calculation result of the tcore operator into the register, and the vcore operator directly reads the calculation result from the register, thereby avoiding the cumbersome process of the intermediate data (i.e., the calculation result of the tcore operator) being written from the tcore to the external memory and then read by the vcore from the external memory. By directly reading the calculation result of the tcore operator from the register, the vcore operator can significantly reduce the read-write overhead.
[0043] Although the fusion process of tcore operators and vcore operators has shown certain advantages in the optimization of neural network models, there are still some obvious defects in the fusion strategies in the related technologies. Especially when dealing with large-scale data (i.e., data with a large shape), these defects are particularly prominent.
[0044] First, previous operator fusion often adopted a serial execution method. That is, in the same thread, the vcore operator had to wait for the tcore operator to finish reading and calculating before it could perform its own calculation and write operations. Due to the large overhead of reading and calculating by the tcore operator, the waiting time of the vcore operator would be relatively long, affecting the overall performance of the operator. At the same time, since the calculation, reading, and writing of the vcore operator took a certain amount of time, especially when dealing with large-scale data, that is, data with a large shape, the reading and writing overhead of the vcore operator was particularly obvious. In the serial execution method, each step of the operator was executed sequentially, and latency hiding could not be achieved, which led to poor performance of the operator and the model when dealing with large-scale data.
[0045] Second, when dealing with data with a large shape, the reading and writing overhead of the tcore operator may be greater than the calculation overhead. Therefore, during model optimization, some large-shape data input to the tcore operator is usually split into multiple small blocks for block processing. At the same time, the reading and writing overhead of the vcore operator is also relatively large. To achieve acceleration, the writing of the vcore operator is usually implemented using matrix load instructions, where one instruction corresponds to writing a group of data, that is, acceleration is achieved through the SIMD (Single Instruction Multiple Data, one instruction corresponds to executing multiple data) method. Therefore, when dealing with data with a large shape, the data in the vcore operator is also split so that one data block can be written out by one instruction each time. For different data layouts, the vcore can have different optimal splitting methods. However, since the fusion of the tcore operator and the vcore operator is executed serially in the same thread, the data splitting method of the vcore operator is often determined by the splitting method of the tcore operator, resulting in a splitting method that may not be optimal for the vcore.
[0046] In addition, the data transfer between the tcore operator and the vcore operator is often the value of a single register or the result of a single instruction execution, which leads to low data transfer efficiency between operators and cannot achieve good acceleration effects.
[0047] In response to this, the present invention provides an operator parallel processing method, which realizes the parallel execution of each step of the first operator and the second operator by means of multi-warp cooperation and the application of synchronization and segmentation strategies. Compared with serial execution, parallel execution can achieve latency hiding and reduce the waiting time between steps, thereby improving processing efficiency. During the execution process, data transfer between the first operator and the second operator is carried out through a shared register block rather than the calculation result of a single instruction. Moreover, since each step of the first operator and the second operator is executed in parallel on different warps, the second operator can determine the optimal segmentation strategy according to the data layout by itself, thereby overcoming the above-mentioned defects.
[0048] It should be noted that the operator parallel processing method provided by the present invention can realize the parallel fusion processing of multiple operators. Hereinafter, the technical solution of the present invention will be introduced mainly by taking the parallel fusion processing of the first operator and the second operator as an example. Here, in the fields of deep learning and high-performance computing, an operator usually refers to a function or process that performs specific mathematical operations or logical operations. The first operator and the second operator refer to two different mathematical operations or logical operations. For example, the first operator can be an mma operator in the tcore operator type, and the second operator can be a relu operator in the vcore operator type. The present invention does not make specific limitations in this regard.
[0049] Figure 1 is a schematic flowchart of the operator parallel processing method provided by the present invention. As Figure 1 shown, the method includes:
[0050] Step 110, based on each processing step of the first operator and each processing step of the second operator, determine the steps with data dependencies, where the output of the first operator is the input of the second operator.
[0051] Specifically, each processing step of an operator refers to a series of calculations or operations required to execute the operator. For example, when the first operator is an mma operator, each processing step of the first operator may include steps such as data segmentation, data loading, data calculation, and result storage; when the second operator is a relu operator, each processing step of the second operator may include steps such as data reading, data segmentation, data calculation, and result writing. By analyzing the data flow and dependency relationship between each processing step of the first operator and each processing step of the second operator, the steps with data dependencies can be identified. Here, the steps with data dependencies refer to that during the calculation process, the output data of one step is used as the input data of another step, and this usage method requires the latter to wait for the output data of the former before calculation.
[0052] It can be understood that the output of the first operator is the input of the second operator, indicating that the result (i.e., the output data) generated after the first operator finishes its calculation is directly used as the input data of the second operator. That is, the second operator must wait for the corresponding output data of the first operator to be ready before reading the data. Therefore, there is a data dependency between the result storage step of the first operator and the data reading step of the second operator.
[0053] In addition, by analyzing the data flow and dependency relationships between the processing steps of the first operator and the processing steps of the second operator, not only can the steps with data dependencies be identified, but also the steps without data dependencies can be identified. For example, there is no data dependency between the data splitting, data loading, etc. steps of the first operator and the data calculation, result writing, etc. steps of the second operator. Therefore, these steps can be executed concurrently to achieve the purpose of latency hiding.
[0054] It should be understood that in order to identify the steps with data dependencies and the steps without data dependencies, it can be achieved by means of automated analysis or by manual analysis. The embodiments of the present invention do not make specific limitations in this regard. For example, when using automated means to analyze the data dependencies between the processing steps of the first operator and the second operator, it can be judged according to the overlapping situation of the data indexes between the processing steps. Specifically, for the first operator and the second operator, first, the index ranges of the data blocks processed by each processing step can be determined, and then the index range of the output data block of any processing step is compared with the index ranges of the input data blocks of other steps. If there is no intersection between the two index ranges, there is no data dependency between the corresponding two processing steps; if there is an intersection, there is a data dependency between the corresponding two processing steps.
[0055] Step 120, based on multiple cooperative warps, apply a synchronization strategy and a splitting strategy to perform parallel execution on the processing steps of the first operator and the processing steps of the second operator;
[0056] Among them, the multiple cooperative warps share register resources. The synchronization strategy refers to inserting synchronization operations between the steps with data dependencies, and the splitting strategy refers to splitting the input data of the first operator and the input data of the second operator.
[0057] It should be noted that a Cooperative Warp (abbreviated as Cwarp) is a basic unit for parallel computing in a GPU. It consists of a group of threads that can cooperate with each other during execution, share certain resources (such as registers), and execute the same instructions. Different Cwarps can also share register resources, that is, multiple Cwarps can access and use the same register space to store and read data. Since register access is faster than main memory access, sharing register resources can significantly reduce data access latency and improve computing efficiency.
[0058] Different Cwarps can execute different program codes in parallel at the same time. That is, when a certain Cwarp is waiting for an operation to complete, other Cwarps can continue to execute. Thus, the latency of a single Cwarp can be hidden, achieving the effect of latency hiding. This mechanism helps to improve the overall computing efficiency of the GPU.
[0059] Specifically, after identifying the steps with data dependencies, based on multiple Cwarps, and applying a synchronization strategy and a splitting strategy, the respective processing steps of the first operator and the second operator can be executed in parallel. In other words, each processing step of the first operator and each processing step of the second operator are each executed using a different Cwarp, thereby enabling the parallel execution of each processing step.
[0060] During the parallel execution process, for steps with data dependencies, a synchronization strategy can be adopted, and corresponding synchronization operations can be inserted between these steps to ensure the production-consumption relationship between these steps, that is, to ensure the correct transfer of the data stream between these steps. Here, the synchronization strategy refers to a strategy in parallel computing of inserting synchronization operations between steps with data dependencies in order to maintain data consistency and correctness. Since each processing step is executed on a different Cwarp, the synchronization strategy can also be understood as inserting synchronization operations between different Cwarps.
[0061] Inserting a synchronization operation means forcing the relevant Cwarp to pause execution at a certain stage of parallel computing until the Cwarp reaches a certain synchronization point. For example, if the input data of the second operator depends on the output data of the first operator, when the first operator has not yet prepared the corresponding output data, a synchronization operation can be inserted between the result storage step of the first operator and the data reading step of the second operator to block the data reading operation of the second operator, so that the second operator waits until the first operator has prepared the output data and then reads the corresponding data and executes the next operation. It should be understood that inserting a synchronization operation can be achieved through some synchronization instructions on the hardware (such as a GPU), and synchronization instructions can be inserted between steps with data dependencies (that is, different Cwarps executing data-dependent steps).
[0062] In the embodiments of the present invention, after data dependencies are identified, by inserting synchronization operations in a timely manner during the execution process, it is possible to ensure the correct transfer of data between dependent steps. This synchronization mechanism is the key to maintaining data consistency in parallel processing. Although synchronization incurs some overhead, by precisely controlling the synchronization points, that is, inserting synchronization operations only between steps with data dependencies instead of between all steps, unnecessary synchronization waiting time can be minimized, thereby minimizing the synchronization overhead.
[0063] Furthermore, during the execution process, a splitting strategy can also be adopted to split the input data of the first operator and the second operator. As a result, the large-shaped data tasks that originally needed to be processed serially can be decomposed into multiple small tasks that can be processed in parallel. In this way, multiple Cwarps can simultaneously process different data chunks, and there is no data dependency between chunks, thus significantly improving the efficiency of parallel processing while reducing data dependencies between processing steps within the operator and between operators. Moreover, the splitting strategy enables data to be more evenly distributed among multiple Cwarps, avoiding the situation of a single warp being overloaded. Here, the splitting strategy refers to a strategy in parallel computing that, in order to fully utilize the resources of multiple cooperative warps, divides a computing task or a data set into multiple smaller parts and then assigns them to different Cwarps for parallel processing.
[0064] Exemplarily, for the first operator, assuming that the processing steps of the operator include a data splitting step, a data loading step, a data calculation step, and a result storage step, these steps can be executed in parallel using different Cwarps. For example, the data splitting step is executed using Cwarp0, the data loading step is executed using Cwarp1, the data calculation step is executed using Cwarp2, and the result storage step is executed using Cwarp3. During the execution process, Cwarp0 can split the input data of the first operator and allocate the split data chunks to other Cwarps for parallel processing. For example, first, Cwarp1 can load the first data chunk. After the first data chunk is loaded, Cwarp2 can calculate the first data chunk that has been loaded. At the same time, Cwarp1 can load the second data chunk. And so on. Through the parallel execution of these Cwarps, the processes of data loading, data calculation, and result storage can be overlapped with each other, thereby reducing the overall latency and improving the performance of the operator. During the execution of each Cwarp, synchronization operations can be inserted between Cwarps with data dependencies to ensure the correct transfer of data. It should be understood that the parallel execution process of each processing step of the second operator can refer to the parallel execution process of each processing step of the first operator, which will not be elaborated here.
[0065] The method provided by the embodiment of the present invention can determine the steps with data dependencies based on the processing steps of the first operator and the processing steps of the second operator, laying a solid foundation for subsequent parallel execution. By adopting multiple cooperative warps and applying a synchronization strategy and a splitting strategy to perform parallel execution on the processing steps of the first operator and the second operator, the total time for operator processing can be significantly reduced, and the overall computing speed can be improved. During the parallel execution process, by adopting the splitting strategy to split the input data of the operator, the large-scale data that originally needed to be processed serially can be split into multiple small pieces. Multiple cooperative warps can process different small pieces simultaneously. Moreover, while one cooperative warp is executing an operator step for a certain small piece of data, other cooperative warps can execute other steps of the operator for other small pieces in parallel, thereby achieving latency hiding, that is, reducing the time to wait for a certain step to complete, and thus shortening the total processing time. At the same time, by adopting the synchronization strategy to insert synchronization operations between the steps with data dependencies, it can ensure the correct transfer of data between the dependent steps and avoid data competition and inconsistency problems.
[0066] In addition, the embodiment of the present invention adopts multiple cooperative warps to perform parallel execution on the processing steps of the first operator and the second operator, and the cooperative warps can share register resources. Thus, the output of the first operator can be passed to the second operator through the shared register block, achieving parallel fusion of the operators. Compared with the traditional serial fusion method, parallel fusion can significantly improve the performance of the operators and the model.
[0067] Based on the above embodiment, the output of the first operator is passed to the second operator in the form of a shared register block.
[0068] It should be noted that considering that in the related art, the data transfer between operators is often the value of a single register or the result of a single instruction execution, which results in low data transfer efficiency between operators. To improve the data transfer efficiency between operators, the data transfer between operators in the embodiment of the present invention is implemented in the form of a shared register block, thereby improving the acceleration effect.
[0069] Specifically, a shared register block is a special memory area that is located inside the processor or computing unit and is designed to store and transfer small pieces of data. These small pieces usually have a fixed size and shape to adapt to specific computing requirements. The shared register block can be accessed by multiple Cwarps simultaneously without a complex synchronization mechanism.
[0070] Specifically, to improve the data transfer efficiency between operators, the output data of the first operator is transferred to the second operator in the form of a shared register block. This means that when the first operator (such as an mma operator) completes its computing task for a certain data block, the data block it outputs is not directly stored in ordinary memory or cache, but is organized and stored in a shared register block (i.e., registers shared among multiple Cwarps) so that the second operator (such as a relu operator) can efficiently read the data block output by the first operator from the shared register block and use it as its input data to perform the next operation. Here, the shared register block can be understood as a data organizational structure when transferring data through shared registers, which can be two-dimensional (such as a part of a matrix) or other shapes to adapt to specific computing requirements.
[0071] In the embodiments of the present invention, the data transfer between the first operator and the second operator is implemented in the form of a shared register block, rather than through the calculation results of a single instruction, which can greatly improve the data transfer efficiency between operators and thus further improve the computing efficiency.
[0072] Based on any of the above embodiments, each processing step of the first operator includes a data slicing and caching step, a data loading step, a first calculation step, and a result storage step;
[0073] The slicing and caching step is used to slice the input data of the first operator and sequentially store the sliced multiple data blocks into the cache;
[0074] The data loading step is used to slice each data block in the cache, sequentially load the sliced multiple data slices into the input buffer, and read and reuse the data slices in the input buffer;
[0075] The first calculation step is used to perform calculations based on the data slices read from the input buffer to obtain result data;
[0076] The result storage step is used to write the result data into the shared register block.
[0077] Specifically, in the split and cache step, the main purpose of splitting the input data of the first operator is to divide it into smaller and more manageable parts for subsequent processing and calculation. Moreover, there is no data dependency between the data blocks after splitting, which enables better parallel processing. Here, when splitting the input data, it can be done according to the dimensions, size, and hardware characteristics of the data, etc. Specifically, when splitting the input data of the first operator, first, based on the characteristics of the data (such as matrices), the requirements of the algorithm, and the hardware characteristics, etc., determine in which dimensions to split. For example, for matrix data, it can be split by rows or by columns, or both by rows and by columns simultaneously. Then, according to the available resources on the hardware (such as memory, computing power), etc., and the requirements of algorithm efficiency, calculate the split size (i.e., the size of each data block). Finally, split the input data into multiple data blocks according to the determined split dimensions and size. Here, multiple data blocks refer to the small parts of data obtained after the split operation for subsequent processing. These data blocks will be used as the input for the subsequent steps.
[0078] In the split and cache step, the multiple data blocks obtained after splitting will be read sequentially and written to the cache. Here, the cache refers to an efficient storage area for temporarily storing data. For example, it can be a multi-level cache. A multi-level cache is a hierarchical storage structure designed to improve the speed and efficiency of data access. For example, a multi-level cache can include global memory, shared memory, L1 cache, L2 cache, etc. Among them, the L1 cache is the cache closest to the computing core in the GPU. It is located between the core and the memory and is used to speed up data access; the L2 cache is the intermediate cache in the GPU located between the L1 cache and the global memory and is used to store a larger amount of data; the global memory and shared memory are located outside the processor and have a larger capacity. It should be understood that by splitting the input data of the first operator and reading the split data blocks into the multi-level cache, the computing efficiency of the GPU can be significantly improved.
[0079] In the data loading step, each data block on the cache can be further sliced so that the size of the sliced data slices can adapt to the size of the input buffer. Thus, the sliced data slices can be loaded into the input buffer, facilitating subsequent reading of the data slices from the input buffer for calculation. Specifically, when slicing each data block on the cache, the slicing method of the data block can be determined first according to the requirements of subsequent calculations and the limitations of hardware resources (such as the size limitation of the input buffer). This can include determining the slicing dimension, slicing size, and slicing order, etc. Subsequently, each data block in the cache is sliced according to the determined slicing method so that each data block is divided into smaller and more easily processable parts, namely data slices. The sliced data slices can be numbered, sorted, or organized so that they can be efficiently accessed and processed in subsequent calculation steps.
[0080] In the data loading step, after slicing each data block is completed, each sliced data slice can be loaded into the input buffer in a certain order. The data loading process specifically involves reading the data slices from the cache and writing them to the specified positions in the input buffer. Here, the input buffer refers to a special storage area located inside or near the processor for storing data to be processed by the processor. It usually has a relatively fast access speed and a small capacity. For example, when the first operator is the mma operator, the input buffer can include an A buffer and a B buffer. Among them, the A buffer is the buffer for storing the data slices of the left matrix, and the B buffer is the buffer for storing the data slices of the right matrix. During the execution of the mma operator, the A and B buffers serve as the input buffer to temporarily store the data slices read from the cache. After these data slices are ready, they will be sent to the next step for matrix multiplication calculation.
[0081] It should be noted that the data slices stored in the input buffer can be reused and do not need to be read repeatedly, thus greatly reducing the overhead of data reading. Here, data reuse means that when performing calculations, the data that has been loaded into the input buffer is used as many times as possible to reduce the number of times of repeated loading and reading of data.
[0082] Exemplarily, when the first operator is the mma operator, taking the input matrices A[M, K] and B[K, N] as an example, assuming that the output matrix C[M, N] needs to be calculated, in the conventional data reading and calculation process, it is usually traversed in the dimension order of M, N, and K. For each output element C[i, j], it is necessary to traverse K times to accumulate the product of A[i, k] and B[k, j]. In this way, each element in A and each element in B will be read. Secondly, this results in a relatively large data reading overhead. To reduce the data reading overhead, the order of the dimension loops can be rearranged to reuse the data elements in the input matrix A or B, thereby significantly reducing the number of data readings. For example, the loop of the k dimension can be moved to the outer layer. For each i (the row index of A), only need to traverse K once to load A[i, k] into the buffer, and then in the inner j loop, for each k, load B[k, j] and perform the calculation. In this way, each element in A only needs to be read times, instead of times, thus realizing the reuse of the data elements in A, that is, loading the data slices obtained by slicing A into the buffer, and then performing mma calculations for the k dimension, repeatedly using the data slices in this buffer, thereby greatly reducing the data reading overhead. Similarly, the data elements in the input matrix B can also be reused.
[0083] It should be understood that when reading the data slices in the input buffer, the corresponding data can be obtained according to the address offset, and the address offset here can be determined according to the start address and size of the data slice. Specifically, each data element has a fixed position in the buffer, and this position can be determined by adding an offset to a start address. By reading and reusing the data in the input buffer, it is possible to avoid repeatedly reading data from global memory or cache, which helps to reduce the number of memory accesses.
[0084] Exemplarily, when the first operator is an mma operator, during the mma calculation, due to the inherent characteristics of matrix operations (such as the repeatability and pattern of multiplication and addition operations), the data in the A and B buffers can be efficiently reused through address offset access. For example, in mma calculations, the elements in the same row may be repeatedly used when multiplying with the elements in different columns. By storing these elements in the A and B buffers and using address offsets to access them, the calculation efficiency can be significantly improved.
[0085] In the first calculation step, the computing unit (i.e., a certain Cwarp) will perform corresponding calculation operations on the data slices read from the input buffer. For example, when the first operator is an mma operator, one data slice can be read from each of the A and B buffers, and matrix multiplication calculations can be performed using the two read data slices to obtain the result data. Here, the result data refers to the output data obtained after calculating the data slices in the input cache in the first calculation step. These data will be used for subsequent storage, transmission, or further calculations.
[0086] In the result storage step, the result data can be written into a shared register block (i.e., registers shared among multiple Cwarps) so that the second operator can directly read the corresponding data from the shared register block and process it.
[0087] It can be understood that each of the above steps can be executed in parallel using one Cwarp each. For steps with data dependencies, corresponding synchronization operations can be inserted between the Cwarps executing these steps to ensure the production-consumption relationship between these steps, thereby ensuring the correct transfer of data between dependent steps.
[0088] In the embodiment of the present invention, by splitting the input data of the first operator into multiple data slices and loading these data slices into the input buffer for reuse, not only can the number of data accesses and conflicts be reduced, but also the data dependencies between the various processing steps within the first operator and the data dependencies between the first operator and the second operator can be reduced, thereby reducing the need for synchronization and further reducing the overhead of data reads and synchronization instructions, etc.
[0089] Based on any of the above embodiments, the processing steps of the second operator include a data reading step, a second calculation step, and a result writing step;
[0090] The data reading step is used to read the result data from the shared register block and split the result data based on the memory layout of the second operator to obtain multiple result slices;
[0091] The second calculation step is used to calculate the multiple result slices to obtain a calculation result;
[0092] The result writing step is used to write the calculation result to a specified storage location.
[0093] Specifically, for the second operator, the processing steps of the second operator can include a data reading step, a second calculation step, and a result writing step. Each of these steps can also be executed in parallel using one Cwarp each. For steps with data dependencies among them, corresponding synchronization operations can be inserted between the Cwarps executing these steps to ensure the production-consumption relationship between these steps, thereby ensuring the correct transfer of data between dependent steps.
[0094] In the data reading step, Cwarp can read the result data output by the first operator from the shared register block. During the data reading process, the result data can be sliced according to the memory layout of the second operator to ensure that the data can be processed by the second operator in the most efficient way. Here, the memory layout of the second operator refers to the memory layout or memory access pattern followed by the second operator (such as the relu operator) when processing data, specifically involving the storage order of data in memory, the size of data blocks, the dimensional arrangement of data (such as row-major or column-major), etc.
[0095] It can be understood that slicing the data according to the memory layout during the process of reading the result data means that the second operator will reorganize the data according to the internal memory layout of the second operator before calculating the result data of the first operator. For example, if the second operator expects to access data in row-major order, then each slice will be ensured to correspond to one row (or a part of a row) of the original data when slicing the data. Since different operators may have different memory access patterns and data processing requirements, slicing the data by matching the memory layout of the second operator can reduce unnecessary memory access overhead and optimize the computing performance of the second operator.
[0096] In the second calculation step, multiple result slices obtained by slicing can be calculated to obtain the calculation result. Here, the result slice refers to the data subset sliced from the result data of the first operator, and these slices are determined according to the memory layout of the second operator for efficient processing in the second operator. Each result slice can contain a certain number of data elements. It should be understood that when calculating the result slices, each slice can be traversed, and a specific calculation logic (such as the non-linear transformation of the relu function) can be applied to each data element in the slice.
[0097] In the result writing step, the calculation result can be written to the specified storage location. Here, the specified storage location refers to the memory address or data structure where the calculation result is stored after the second operator completes the calculation. This storage location is usually determined before the program execution, and it can correspond to a global memory area, a shared memory area, or a specific data structure (such as an array, a vector, or a matrix). Writing the calculation result to the specified storage location is to ensure that these data can be accessed and used by subsequent calculation steps or other parts of the program.
[0098] Based on any of the above embodiments, there is a data dependency between the slicing cache step, the data loading step, and the first calculation step of the first operator, and there is a data dependency between the result storage step of the first operator and the data reading step of the second operator.
[0099] Specifically, for each processing step of the first operator and each processing step of the second operator, a Cwarp can be used respectively for parallel execution. For steps with data dependencies, corresponding synchronization operations can be inserted to ensure the correct production-consumption relationship; for most steps without data dependencies, they are allowed to be executed in parallel to achieve the purpose of latency hiding.
[0100] Specifically, in the processing stage of the first operator, there are data dependencies between the cache slicing step and the data loading step and the first calculation step, that is, there is a dependency between the write-out of the cache and the data reading of the input buffer, and there is a dependency between the data reading of the input buffer and the data calculation. Therefore, corresponding synchronization operations need to be inserted between these dependent steps to ensure the correct production-consumption relationship between these steps.
[0101] Similarly, since the output data of the first operator is the input data of the second operator, that is, the input of the second operator depends on the output of the first operator, there is a data dependency between the result storage step of the first operator and the data reading step of the second operator. During parallel execution, corresponding synchronization operations need to be inserted between these two steps to ensure the correct transfer of data.
[0102] For most steps without data dependencies, such as the cache slicing step of the first operator and the calculation step of the second operator, the data loading step of the first operator and the calculation step of the second operator, the data loading step of the first operator and the result write-out step of the second operator, etc., there are no data dependencies between these steps and they can be directly assigned to different Cwarps for parallel execution to achieve latency hiding.
[0103] Figure 2 is a schematic diagram of the pipeline arrangement provided by the present invention, as Figure 2 shown, the pipeline ping-pong operation, also known as the ping-pong operation, is a processing technique for data flow control. By introducing two or more buffers, this operation enables different functional blocks to alternately read and write to the buffers, thereby achieving simultaneous operation. Exemplarily, taking the processing stage of the first operator as an example, when Cwarp1 finishes loading the ping data (i.e., a certain data slice), Cwarp2 can apply this data slice for calculation. At the same time, Cwarp1 can start loading the pong data (i.e., another data slice), that is, start reading the pong data from the input buffer. When Cwarp2 finishes calculating the ping data, Cwarp3 starts writing the calculation result into the shared register block. At the same time, Cwarp2 can apply the loaded pong data for calculation. And so on, until the calculation of all data slices is completed.
[0104] It can be understood that by adopting a data splitting strategy and multiple cooperative warps for parallel execution, multiple task pipelines are introduced, enabling better parallel execution between the read / write and calculation steps of the first operator and those of the second operator, thereby performing latency hiding and improving operator performance.
[0105] Based on any of the above embodiments, the method further includes:
[0106] Step 130, during the execution, monitor the usage of the shared register block;
[0107] Step 140, when it is detected that the usage amount of the shared register block exceeds the hardware limit, adopt an auto-tuning method to determine the optimal shared register block size and splitting strategy.
[0108] Specifically, during the execution, the usage of the shared register block can be monitored in real time. When the usage amount of the shared register block exceeds the maximum value that the GPU hardware can run (i.e., the hardware limit), it will cause register overflow and affect the performance of the operator. Here, the performance counter provided by the hardware or the monitoring tool at the software level can be used to track the usage of the shared register block in real time.
[0109] When it is detected that the usage amount of the shared register block exceeds the hardware limit, an auto-tuning method can be adopted to select the optimal shared register block size and splitting strategy. Here, in GPU parallel computing, data is usually divided into smaller blocks for processing, and the shared register block size refers to the size of these small data blocks. The optimal shared register block size and splitting strategy refer to the register configuration scheme that can maximize the computing efficiency, minimize the resource occupancy rate and conflict rate under the given hardware environment and computing task. It should be understood that the optimal shared register block size and splitting strategy can depend on multiple factors such as the characteristics of the operator, the shape of the data, the hardware limit, and the complexity of the computing task.
[0110] The auto-tuning method is a method to find the optimal configuration parameters through automated means, and these parameters can include the shared register block size and splitting strategy, etc. For example, intelligent optimization algorithms such as genetic algorithms and particle swarm optimization can be used to search for the optimal shared register block size and splitting strategy within the preset parameter space. Another example is that based on machine learning or deep learning models, the performance performance under different configurations can be predicted according to historical data or real-time monitoring information, so as to select the optimal configuration scheme. According to the determined optimal shared register block size and optimal splitting strategy, the shared register block size and data splitting method can be dynamically adjusted to adapt to different computing tasks and hardware environments.
[0111] Based on any of the above embodiments, step 140 specifically includes:
[0112] Step 141: Based on the hardware information and the shape of the operator data, determine multiple candidate register block sizes and multiple candidate partitioning strategies.
[0113] Specifically, the hardware information refers to the relevant parameters and characteristics of the physical device (such as GPU, etc.) that executes the computing task, including but not limited to the architecture of the processor, the number of registers, the memory bandwidth, the cache size, the number and type of computing units, etc. These parameters will limit the selection range of the shared register block size and affect the effectiveness of the partitioning strategy.
[0114] The shape of the operator data refers to the dimensions and layout of the data processed by the operator (i.e., the basic operation or function in the computing task). These information will have an impact on the partitioning strategy. For example, in deep learning, an mma operator processes a two-dimensional matrix (including row dimension and column dimension), and a conv operator can process a four-dimensional tensor (batch size, number of channels, height, width).
[0115] By analyzing the hardware information and the shape of the operator data, a series of candidate register block sizes and partitioning strategies can be generated. Here, the candidate register block size refers to the set of possible register block sizes under the given hardware constraints and the shape of the operator data. These sizes are obtained through the analysis of the hardware information and the data shape, aiming to find a configuration that can balance the computing efficiency and resource occupancy. The candidate partitioning strategies refer to the set of methods for splitting the data into blocks or slices of different sizes for parallel processing.
[0116] Step 142: Based on the multiple candidate register block sizes and the multiple candidate partitioning strategies, determine multiple candidate configurations.
[0117] Specifically, by combining multiple candidate register block sizes and multiple candidate partitioning strategies pairwise, a series of possible configuration combinations can be generated, that is, multiple candidate configurations are obtained. Here, the candidate configuration is a computing task configuration that includes a specific register block size and a partitioning strategy. These configurations will be used in the subsequent performance evaluation and optimization process to find the optimal configuration scheme.
[0118] Step 143: Evaluate the performance of each candidate configuration, and based on the performance evaluation results of each candidate configuration, determine the optimal configuration, where the optimal configuration includes the optimal shared register block size and partitioning strategy.
[0119] Specifically, the performance of each candidate configuration can be evaluated through the following steps: According to the requirements of the computing task and the hardware characteristics, select appropriate evaluation metrics (such as execution time, throughput, resource occupancy rate, etc.), which will be used to quantify the performance differences between different configurations. Then, in the same hardware and software environment, use the same test data set to perform benchmark tests on each candidate configuration. The benchmark tests will simulate the typical execution scenarios of the computing task and collect relevant performance data. Analyze and compare the collected performance data to determine the advantages and disadvantages of each candidate configuration.
[0120] After obtaining the performance evaluation results of each candidate configuration, the candidate configurations can be sorted according to the evaluation metrics, and the set of configurations with the best performance can be selected. For example, the candidate configuration with the shortest execution time and the highest throughput can be selected as the optimal configuration. Here, the optimal configuration refers to the computing task configuration that can maximize the computing efficiency, minimize the resource occupancy, and meet other relevant requirements under the given hardware limitations, operator data shapes, and performance evaluation metrics.
[0121] In addition, after finding the optimal configuration, it can be stored so that when the same hardware information and operator data shapes are encountered later, the optimal configuration can be directly applied without repeating the tuning and selection.
[0122] Based on any of the above embodiments, an embodiment of the present invention provides a method for fine-grained parallel fusion calculation of tcore operators and vcore operators. Here, the tcore operator takes the mma operator as an example, and the vcore operator takes the relu operator as an example. The method includes:
[0123] Step S1, the mma operator processing stage;
[0124] Step S11, split the input data of the mma operator, and read and write the split data blocks to the multi-level cache.
[0125] Step S12, split the data blocks on the cache again, and read the split data slices into the A and B buffers to prepare for matrix multiplication.
[0126] Step S13, reuse the data slices in the A and B buffers, that is, there is no need to repeatedly read the same data slices, and directly obtain the required data slices through address offset.
[0127] Step S14, perform mma calculations using the read data slices and write the results to the shared register block.
[0128] Step S15, perform a loop traversal of the intermediate dimension (i.e., the K dimension), and accumulate the results of each mma calculation to the shared register block.
[0129] Step S2, relu operator processing stage;
[0130] Step S21, read the result data of the mma calculation from the shared register block, and according to the memory layout of the relu operator, perform corresponding vcore slicing on the read data to meet the parallel requirements of relu calculation;
[0131] Step S22, perform relu calculation, that is, apply the relu activation function to each element;
[0132] Step S23, write the result of the relu calculation to the specified storage location through the matrix storage instruction.
[0133] Step S3, Cwarp parallel execution stage;
[0134] All steps in the mma operator processing stage and the relu operator processing stage (i.e., Step S11 to Step S15, Step S21 to Step S23) are respectively assigned to a Cwarp for parallel execution. For steps with data dependencies, corresponding synchronization operations can be inserted to ensure the corresponding producer-consumer relationship. Here, the steps with data dependencies mainly include the dependencies between the write-out of the cache and the data reading of A and B buffers, the dependencies between the data reading of A and B buffers and the mma calculation, and partial dependencies of the shared register block (i.e., the dependency between the storage of the mma calculation result and the relu data reading). For most steps without data dependencies (for example, there is no dependency between the data reading of mma and the calculation and write-out of relu), these steps can be executed in parallel to achieve the purpose of latency hiding.
[0135] In addition, for the part with data dependencies, a technique similar to the double buffer can be adopted to slice the data into multiple chunks for processing to reduce the data dependencies between chunks, thereby better realizing parallelization. Here, the double buffer technique can make the instruction stream more compact by introducing multiple task pipelines, achieving a certain degree of concurrency. For details, please refer to Figure 2 the process shown.
[0136] Step S4, automatic tuning stage;
[0137] During the execution process, monitor the usage of the shared register block. When it is found that the register usage exceeds the hardware limit, adopt an automatic tuning method to select the optimal shared register block size and slicing strategy.
[0138] The method provided by the embodiment of the present invention has the following advantages:
[0139] 1. By executing each processing step of the tcore operator and the vcore operator in parallel, latency hiding can be achieved, greatly improving the performance of the operator.
[0140] 2. Compared with the traditional serial fusion method, in the embodiments of the present invention, different data segmentation methods can be defined for the tcore operator and the vcore operator, better leveraging the respective performances of the tcore and the vcore.
[0141] 3. It is possible to control the pipelining and synchronization within the tcore operator, within the vcore operator, and between the tcore operator and the vcore operator at a finer granularity, enabling better parallelism between the read / write and calculation of the tcore operator and the read / write and calculation of the vcore operator for latency hiding.
[0142] 4. For scenarios where the number of available registers is insufficient, an auto-tuning method can be adopted to select the optimal shared register block size and segmentation strategy.
[0143] The operator parallel processing device provided by the present invention will be described below. The operator parallel processing device described below can be correspondingly referred to the operator parallel processing method described above.
[0144] Figure 3 is a schematic structural diagram of the operator parallel processing device provided by the present invention. As Figure 3 shown, the device includes:
[0145] A step determination unit 310, configured to determine steps with data dependencies based on each processing step of a first operator and each processing step of a second operator, where the output of the first operator is the input of the second operator.
[0146] A parallel execution unit 320, configured to apply a synchronization strategy and a segmentation strategy to parallelly execute each processing step of the first operator and each processing step of the second operator based on multiple cooperative warps; wherein, register resources are shared among the multiple cooperative warps, the synchronization strategy refers to inserting synchronization operations between the steps with data dependencies, and the segmentation strategy refers to segmenting the input data of the first operator and the input data of the second operator.
[0147] The device provided by the embodiment of the present invention can determine the steps with data dependencies through the processing steps of the first operator and the processing steps of the second operator, which can lay a solid foundation for subsequent parallel execution. By adopting multiple cooperative thread warps and applying a synchronization strategy and a splitting strategy to perform parallel execution on the processing steps of the first operator and the second operator, the total time for operator processing can be significantly reduced, and the overall computing speed can be improved. During the parallel execution process, by adopting the splitting strategy to split the input data of the operator, the large-scale data that originally needed to be processed serially can be split into multiple small pieces. Multiple cooperative thread warps can process different small pieces simultaneously. Moreover, while one cooperative thread warp is executing an operator step for a certain small piece of data, other cooperative thread warps can execute other steps of the operator for other small pieces in parallel, thereby achieving latency hiding, that is, reducing the time waiting for a certain step to complete, and thus shortening the total processing time. At the same time, by adopting the synchronization strategy to insert synchronization operations between the steps with data dependencies, it can ensure the correct transfer of data between the dependent steps and avoid data competition and inconsistency problems.
[0148] In addition, the embodiment of the present invention adopts multiple cooperative thread warps to perform parallel execution on the processing steps of the first operator and the second operator, and the cooperative thread warps can share register resources. Therefore, the output of the first operator can be passed to the second operator through the shared register block, realizing the parallel fusion of operators. Compared with the traditional serial fusion method, parallel fusion can significantly improve the performance of operators and models.
[0149] Based on any of the above embodiments, the output of the first operator is passed to the second operator in the form of a shared register block.
[0150] Based on any of the above embodiments, the processing steps of the first operator include a data splitting and caching step, a data loading step, a first calculation step, and a result storage step;
[0151] The splitting and caching step is used to split the input data of the first operator and sequentially store the multiple split data blocks into the cache;
[0152] The data loading step is used to split each data block in the cache, sequentially load the multiple split data slices into the input buffer, and read and reuse the data slices in the input buffer;
[0153] The first calculation step is used to calculate based on the data slices read from the input buffer to obtain result data;
[0154] The result storage step is used to write the result data into the shared register block.
[0155] Based on any of the above embodiments, each processing step of the second operator includes a data reading step, a second calculation step, and a result writing step;
[0156] The data reading step is used to read the result data from the shared register block, and based on the memory layout of the second operator, split the result data to obtain multiple result slices;
[0157] The second calculation step is used to calculate the multiple result slices to obtain a calculation result;
[0158] The result writing step is used to write the calculation result to a specified storage location.
[0159] Based on any of the above embodiments, there is a data dependency between the splitting cache step, the data loading step, and the first calculation step of the first operator, and there is a data dependency between the result storage step of the first operator and the data reading step of the second operator.
[0160] Based on any of the above embodiments, the device further includes a tuning unit, and the tuning unit is used for:
[0161] During the execution process, monitor the usage of the shared register block;
[0162] When it is monitored that the usage amount of the shared register block exceeds the hardware limit, adopt an automatic tuning method to determine the optimal size of the shared register block and the splitting strategy.
[0163] Based on any of the above embodiments, the tuning unit is specifically used for:
[0164] Based on the hardware information and the shape of the operator data, determine multiple candidate register block sizes and multiple candidate splitting strategies;
[0165] Based on the multiple candidate register block sizes and the multiple candidate splitting strategies, determine multiple candidate configurations;
[0166] Evaluate the performance of each candidate configuration, and based on the performance evaluation results of each candidate configuration, determine the optimal configuration, where the optimal configuration includes the optimal size of the shared register block and the splitting strategy.
[0167] Figure 4 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 4As shown in the figure, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute the operator parallel processing method, and the method includes: based on the processing steps of the first operator and the processing steps of the second operator, determining the steps with data dependencies, where the output of the first operator is the input of the second operator; based on multiple cooperative warps, applying a synchronization strategy and a splitting strategy to perform parallel execution on the processing steps of the first operator and the processing steps of the second operator; where the multiple cooperative warps share register resources, the synchronization strategy refers to inserting synchronization operations between the steps with data dependencies, and the splitting strategy refers to splitting the input data of the first operator and the input data of the second operator.
[0168] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0169] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the operator parallel processing method provided by the above-mentioned various methods. The method includes: based on the processing steps of the first operator and the processing steps of the second operator, determining the steps with data dependencies, where the output of the first operator is the input of the second operator; based on multiple cooperative warps, applying a synchronization strategy and a splitting strategy to perform parallel execution on the processing steps of the first operator and the processing steps of the second operator; where the multiple cooperative warps share register resources, the synchronization strategy refers to inserting synchronization operations between the steps with data dependencies, and the splitting strategy refers to splitting the input data of the first operator and the input data of the second operator.
[0170] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements an operator parallel processing method provided by the above-mentioned various methods. The method includes: determining steps with data dependencies based on the processing steps of a first operator and the processing steps of a second operator, where the output of the first operator is the input of the second operator; applying a synchronization strategy and a splitting strategy based on multiple cooperative warps to perform parallel execution on the processing steps of the first operator and the processing steps of the second operator; wherein, register resources are shared among the multiple cooperative warps, the synchronization strategy refers to inserting synchronization operations between the steps with data dependencies, and the splitting strategy refers to splitting the input data of the first operator and the input data of the second operator.
[0171] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0172] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An operator parallel processing method, characterized in that: include: Based on each processing step of the first operator and each processing step of the second operator, determining a step in which data dependency exists, the output of the first operator is the input of the second operator, the output of the first operator is transmitted to the second operator in the form of a shared register block, each processing step of the first operator includes a cache segmentation step, a data loading step, a first calculation step, and a result storage step, and each processing step of the second operator includes a data reading step, a second calculation step, and a result writing step; Based on the multiple cooperative thread warps, a synchronization strategy and a segmentation strategy are applied to execute each processing step of the first operator and each processing step of the second operator in parallel; Among them, register resources are shared among the multiple collaborative thread warps, the synchronization strategy refers to inserting synchronization operations between the steps with data dependencies, the splitting strategy refers to splitting the input data of the first operator and the input data of the second operator, and reusing the split data of the first operator, and the splitting of the input data of the second operator is determined based on the memory layout of the second operator.
2. The operator parallel processing method according to claim 1, characterized in that: The segmentation and caching step is used to segment the input data of the first operator and store the segmented data blocks in the cache in sequence; The data loading step is used to segment each data block in the cache, load the segmented multiple data slices into the input buffer in sequence, and read and reuse the data slices in the input buffer; The first calculation step is used to perform calculation based on the data slices read from the input buffer to obtain result data; The result storing step is used to write the result data into a shared register block.
3. The operator parallel processing method according to claim 2, characterized in that: The data reading step is used to read the result data from the shared register block, and to slice the result data based on the memory arrangement of the second operator to obtain a plurality of result slices; The second calculation step is used to calculate the multiple result slices to obtain calculation results; The result writing step is used to write the calculation result to a specified storage location.
4. The operator parallel processing method according to claim 3, characterized in that: There is data dependency between the splitting and caching step, the data loading step and the first calculation step of the first operator, and there is data dependency between the result storage step of the first operator and the data reading step of the second operator.
5. The operator parallel processing method according to any one of claims 1 to 4, characterized in that: Also includes: During execution, the usage of the shared register block is monitored; When it is detected that the usage of the shared register block exceeds the hardware limit, an automatic tuning method is adopted to determine the optimal shared register block size and segmentation strategy.
6. The operator parallel processing method according to claim 5, characterized in that: The automatic tuning method is used to determine the optimal shared register block size and segmentation strategy, including: Determine multiple candidate register block sizes and multiple candidate segmentation strategies based on hardware information and shapes of operator data; Determining a plurality of candidate configurations based on the plurality of candidate register block sizes and the plurality of candidate segmentation strategies; The performance of each candidate configuration is evaluated, and an optimal configuration is determined according to the performance evaluation result of each candidate configuration, wherein the optimal configuration includes an optimal shared register block size and a segmentation strategy.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the operator parallel processing method according to any one of claims 1 to 6 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the operator parallel processing method according to any one of claims 1 to 6 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the operator parallel processing method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method and device for optimizing tensor calculation performance
CN116775277A