Model performance optimization method, electronic equipment, storage medium and program product

By reorganizing the computational operations and communication operations of the model layer into MMA+AllReduce+MMA structures and executing them in parallel, the problem of communication operations bottlenecks in kernel fusion technology is solved, and model performance and resource utilization are improved.

CN120338052APending Publication Date: 2025-07-18SHANGHAI BIREN TECH CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510429468.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing kernel fusion technology has become a bottleneck in the artificial intelligence model, resulting in low computing resource utilization and inability to make full use of bidirectional interconnect bandwidth, affecting the performance optimization effect of model.

Method used

The calculation operations and communication operations of the model layer are reorganized into the structure of MMA+AllReduce+MMA, and the ReduceScatter, Add, LayerNorm and AllGather operations are fused into a kernel to perform, and the input data is segmented in each computing communication parallel unit to realize the parallel execution of the calculation operations and the communication operations.

Benefits of technology

It improves the parallelism of the model, reduces the waiting time, improves the utilization rate of computing resources and hardware resources, makes full use of the two-way interconnect bandwidth in modern hardware architectures, and significantly improves model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338052A_ABST
    Figure CN120338052A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides a model performance optimization method, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining a calculation operation corresponding to each model layer and a communication operation between the model layers based on a model structure; organizing all calculation operations and communication operations into a plurality of calculation and communication parallel units, wherein each unit comprises a first matrix multiply-accumulate operation, a fusion reduction operation and a second matrix multiply-accumulate operation; in each unit, input data is segmented so that a first matrix multiply-accumulate operation, a fusion reduction operation and a second matrix multiply-accumulate operation are performed in parallel based on different data blocks. According to the method, all calculation operations and communication operations are organized into a plurality of calculation and communication parallel units, and data are segmented in each unit, so that parallel execution of the calculation operations and the communication operations is realized, the degree of parallelism and the utilization rate of calculation resources and hardware resources are improved, and the model performance is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence chips, and in particular, to a method for optimizing model performance, an electronic device, a storage medium, and a program product. Background Art

[0002] In the field of optimizing the performance of artificial intelligence models, the kernel fusion technology, as an advanced means, aims to significantly improve the execution efficiency of models by reducing the context switching and data transfer overhead during the calculation process. This technology usually involves fusing multiple originally independently executed operations into a single kernel for execution. For example, fusing the Matrix Multiply Accumulate (MMA) operation with the ReduceScatter operation into a single kernel for execution, fusing the AllGather operation with the MMA operation into a single kernel for execution, etc.

[0003] However, although the kernel fusion technology helps to improve model performance, it faces certain defects in practical applications. First, communication operations (such as ReduceScatter and AllGather) may backpressure the calculation operations (such as MMA), that is, the communication delay becomes the bottleneck, restricting the full utilization of computing resources and resulting in poor model performance optimization effects. Second, when communication operations are fused with calculation operations, it may cause the tensor cores to be idle while waiting for the communication to complete, thereby reducing the effective utilization rate of hardware resources. In addition, the current fusion strategy can only utilize the unidirectional bandwidth and cannot utilize the bidirectional interconnection bandwidth, resulting in the actual performance improvement still being restricted. Summary of the Invention

[0004] The present invention provides a method for optimizing model performance, an electronic device, a storage medium, and a program product to solve the defect of poor model performance optimization effect in the related art.

[0005] The present invention provides a method for optimizing model performance, including: Based on the model structure, obtaining the calculation operations corresponding to each model layer and the communication operations between model layers; Organizing the calculation operations corresponding to all model layers and the communication operations between model layers into multiple calculation-communication parallel units, each calculation-communication parallel unit including a first Matrix Multiply Accumulate operation, a fused reduction operation, and a second Matrix Multiply Accumulate operation, and the fused reduction operation including a ReduceScatter operation, an addition operation, a normalization operation, and an AllGather operation; Within each calculation-communication parallel unit, splitting the input data so that the first Matrix Multiply Accumulate operation, the fused reduction operation, and the second Matrix Multiply Accumulate operation are executed in parallel based on different data blocks.

[0006] According to a model performance optimization method provided by the present invention, within each computing and communication parallel unit, the input data is segmented so that the first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation are executed in parallel based on different data blocks, including: Within each computing and communication parallel unit, the input data is segmented to obtain a plurality of data blocks; The first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation are sequentially executed on each data block; Wherein, the fusion reduction operation of the current data block is executed in parallel with the first matrix multiplication and accumulation operation of the next data block, and the second matrix multiplication and accumulation operation of the current data block is executed in parallel with the fusion reduction operation of the next data block.

[0007] According to a model performance optimization method provided by the present invention, the step of sequentially executing the first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation on each data block includes: Based on a first kernel, the first matrix multiplication and accumulation operation is executed on any data block to obtain a first calculation result; Based on a second kernel, the fusion reduction operation is executed on the first calculation result to obtain a fusion result; Based on a third kernel, the second matrix multiplication and accumulation operation is executed on the fusion result to obtain a second calculation result of the any data block.

[0008] According to a model performance optimization method provided by the present invention, the step of, based on the second kernel, executing the fusion reduction operation on the first calculation result to obtain a fusion result includes: Based on the number of computing devices, the first calculation result is segmented to obtain a plurality of result blocks, and a synchronization signal is sent to the second kernel of the local device and the second kernels of each remote computing device; When the second kernel detects that the synchronization signal value is equal to the number of computing devices, one result block of the local device and one result block of each remote computing device are read to obtain a first result; An addition operation and a normalization operation are sequentially executed on the first result to obtain a second result; Based on the second result of the local device and the second results of each remote computing device, the fusion result is obtained.

[0009] A model performance optimization method provided according to the present invention, the model structure includes a plurality of network modules connected in series, each network module includes a plurality of model layers, and the calculation operations corresponding to the plurality of model layers and the communication operations between the model layers include a first normalization operation, a first all-gather operation, a self-attention calculation operation, a first linear transformation operation, a first reduce-scatter operation, a first dropout operation, a first addition operation, a second normalization operation, a second all-gather operation, a second linear transformation operation, an activation operation, a third linear transformation operation, a second reduce-scatter operation, a second dropout operation, and a second addition operation.

[0010] A model performance optimization method provided according to the present invention, organizing the calculation operations corresponding to all model layers and the communication operations between the model layers into a plurality of compute-communication parallel units, includes: Based on the first linear transformation operation, the first reduce-scatter operation, the first dropout operation, the first addition operation, the second normalization operation, the second all-gather operation, the second linear transformation operation, and the activation operation in the current network module, obtain a compute-communication parallel unit; Based on the third linear transformation operation, the second reduce-scatter operation, the second dropout operation, the second addition operation in the current network module, and the first normalization operation, the first all-gather operation, and the self-attention calculation operation in the subsequent network module, obtain a compute-communication parallel unit.

[0011] A model performance optimization method provided according to the present invention, the method of obtaining a compute-communication parallel unit based on the first linear transformation operation, the first reduce-scatter operation, the first dropout operation, the first addition operation, the second normalization operation, the second all-gather operation, the second linear transformation operation, and the activation operation in the current network module, includes: Take the first linear transformation operation as a first matrix multiply-accumulate operation; Take the first reduce-scatter operation, the first dropout operation, the first addition operation, the second normalization operation, and the second all-gather operation as a fused reduction operation; Take the second linear transformation operation and the activation operation as a second matrix multiply-accumulate operation; Based on the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation, obtain the compute-communication parallel unit.

[0012] The present invention also provides a model performance optimization device, including: An acquisition unit, configured to acquire the calculation operations corresponding to each model layer and the communication operations between the model layers based on the model structure; An organizational unit for organizing the computing operations corresponding to all model layers and the communication operations between model layers into multiple compute-communication parallel units, each compute-communication parallel unit including a first matrix multiply-accumulate operation, a fused reduction operation, and a second matrix multiply-accumulate operation, the fused reduction operation including a reduce-scatter operation, an addition operation, a normalization operation, and an all-gather operation; A parallel unit for splitting input data within each compute-communication parallel unit so that the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation are executed in parallel based on different data blocks.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, where when the processor executes the computer program, the model performance optimization method as described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the model performance optimization method as described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the model performance optimization method as described in any one of the above is implemented.

[0016] The model performance optimization method, electronic device, storage medium, and program product provided by the present invention can organize the computing operations of all model layers and the communication operations between model layers into multiple compute-communication parallel units, and split the input data within each unit, so as to realize the parallel execution of the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation. Since the matrix multiply-accumulate operation is a computing operation and the fused reduction operation includes communication operations, the parallel execution of computing operations and communication operations is realized, thereby improving the parallelism and accelerating the computing process of the overall model, thus significantly improving the model performance. Within each unit, the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation are executed in parallel based on different data blocks, which means that the computing operations and communication operations are independently performed in different cores, avoiding the situation where the communication operation becomes a bottleneck and backpressures the computing operation, so the waiting time can be reduced and the utilization rate of computing resources and hardware resources can be improved.

[0017] In addition, since the fused reduction operation includes a reduce-scatter operation, an addition operation, a layer normalization operation, and an all-gather operation, this means that the reduce-scatter operation and the all-gather operation are fused into one core for execution. This fusion strategy can make full use of the bidirectional interconnection bandwidth in modern hardware architectures, reduce the latency and overhead of data transmission, and further improve the model performance. Description of the Drawings

[0018] To more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 is a schematic structural diagram of the neural network model provided by the present invention; Figure 2 is a schematic diagram of kernel fusion in related technologies; Figure 3 is a schematic flowchart of the model performance optimization method provided by the present invention; Figure 4 is a schematic structural diagram of the computing and communication parallel unit provided by the present invention; Figure 5 is a schematic diagram of the execution process of the computing device provided by the present invention; Figure 6 is a schematic diagram of the parallel execution of the first matrix multiply-accumulate operation, fusion reduction operation, and second matrix multiply-accumulate operation provided by the present invention; Figure 7 is a schematic structural diagram of the model performance optimization device provided by the present invention; Figure 8 is a schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments

[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0021] In the field of artificial intelligence (AI) computing, the overlap technique generally refers to the parallel execution between different computing tasks or data communications to reduce the total computing time. Multi-stream is a specific way to implement the overlap technique. In Graphics Processing Unit (GPU) programming, a stream can be regarded as a series of operations executed in sequence. By using multiple streams, multiple sequences of operations can be started simultaneously, and these sequences are executed in parallel on the GPU. In this way, when one stream is waiting for some operations (such as memory access) to complete, another stream can continue to execute, thus improving the overall efficiency.

[0022] For example, when performing an MMA operation on the GPU, the complete MMA can be split into multiple MMA kernels, that is, the large MMA operation originally executed as a whole is split into multiple smaller operations, and each small operation is executed on an independent kernel. Thus, the multi-stream method can be adopted to utilize the parallel processing ability of the GPU to improve the computing efficiency. It should be understood that MMA is a basic operation executed by the GPU Tensor Core (a tensor core, a hardware unit dedicated to executing deep learning operations such as MMA) to accelerate the matrix multiplication and accumulation operation in deep learning.

[0023] However, as the MMA operation is split, the workload of each independent MMA kernel or CCL (Collective Communication Library, used for data communication between GPUs) kernel will decrease accordingly, that is, the amount of data or computing volume processed by each kernel is reduced, which will lead to a decrease in the utilization rate of the GPU's Tensor Core and interconnection bandwidth. This is because starting and executing a large number of small kernels will result in higher overheads (such as thread management and scheduling), and the large-scale parallel processing ability of the Tensor Core cannot be fully utilized. At the same time, frequent data transmission (especially when there are data dependencies between kernels) may also lead to low utilization of the interconnection bandwidth.

[0024] In order to avoid the problem of Tensor Core and interconnect bandwidth utilization caused by the splitting of MMA operations, a kernel fusion technology is proposed. This technology merges multiple operations that were originally executed independently into a single kernel, and then implements parallel execution through multiple kernels. For example, the MMA operation and the ReduceScatter operation are merged into one kernel for execution, and the AllGather operation and the MMA operation are merged into another kernel for execution to achieve overlapping execution (i.e., parallel execution) of the two kernels, thereby improving the overall computing efficiency. It should be understood that in multi-GPU parallel training, different parts of the model may be assigned to different GPUs for processing. In order to maintain data consistency and synchronization, communication operations such as ReduceScatter and AllGather are required between GPUs. Among them, ReduceScatter refers to reducing the data on all GPUs (such as summing) and then distributing the results to each GPU, and each GPU only retains part of the results; AllGather refers to collecting part of the data on each GPU to all GPUs so that each GPU has complete data.

[0025] However, although kernel fusion technology has brought certain performance improvement potential, it also faces defects and challenges that cannot be ignored in practical applications. First, since communication operations (such as ReduceScatter and AllGather) and computing operations (such as MMA) are fused into one kernel for execution, if the data transmission speed is low or the amount of data output by the computing operation is large, the communication operation will become a bottleneck, resulting in the inability to output the results of the computing operation to the network in time, thus forming a communication backpressure calculation situation. In other words, the computing operation is blocked by the communication operation and cannot be calculated and executed.

[0026] Secondly, communication back pressure calculation will cause the GPU's tensor core utilization to decrease, because the tensor core cannot continue to execute new computing tasks before waiting for the communication operation to complete, and is idle, resulting in reduced tensor core utilization.

[0027] In addition, the two operations ReduceScatter and AllGather respectively utilize the unidirectional characteristics of network bandwidth. Since they are processed separately in different kernels, the current fusion strategy can only utilize unidirectional bandwidth, while ignoring the advantages of bidirectional interconnection bandwidth generally supported in modern network architecture. This means that even if the network hardware supports higher data transmission rates, the actual performance improvement is still restricted due to software-level limitations.

[0028] To overcome the above deficiencies, the present invention provides a method for optimizing model performance. By organizing the operations corresponding to all model layers into multiple compute-communication parallel units, each unit including a first matrix multiply-accumulate operation, a fused reduction operation, and a second matrix multiply-accumulate operation, parallel execution of compute operations and communication operations is achieved, thereby improving the performance of the model in multi-GPU parallel training. The technical solution provided by the present invention will be introduced in detail below.

[0029] To facilitate understanding of the technical solution provided by the present invention, the neural network model structure provided by the present invention will be described first. The model structure of the neural network model provided by the present invention includes a plurality of serially connected network modules, each network module including a plurality of model layers. The compute operations corresponding to the plurality of model layers and the communication operations between the model layers include a first normalization operation, a first all-gather operation, a self-attention calculation operation, a first linear transformation operation, a first reduce-scatter operation, a first dropout operation, a first addition operation, a second normalization operation, a second all-gather operation, a second linear transformation operation, an activation operation, a third linear transformation operation, a second reduce-scatter operation, a second dropout operation, and a second addition operation. Among them, the first all-gather operation, the first reduce-scatter operation, the second all-gather operation, and the second reduce-scatter operation are communication operations.

[0030] It should be noted that the model structure refers to the overall design framework of a deep learning model, including the components of the model, the connection methods of each component, and the path of data flow, etc. For example, Figure 1 is a schematic diagram of the structure of the neural network model provided by the present invention. As Figure 1 shown, the model structure may include a plurality of network modules (such as Transformer modules), and these modules are connected in series in sequence, that is, the output of the previous module is the input of the next module. As Figure 1 shown by "×L" in, it means that the model structure includes L (L≥1) Transformer modules. It should be understood that the neural network model provided by the present invention can receive input data and obtain a final output result after being processed by a plurality of Transformer modules. Here, the input data has different physical meanings according to different application scenarios. For example, the neural network model and the model performance optimization method provided by the present invention can be applied to fields such as speech processing, image processing, and text processing. In the field of speech processing, the input data of the model can be audio data for tasks such as speech recognition, speech synthesis, and speech segmentation; in the field of image processing, the input data of the model can be the image to be processed for tasks such as image recognition, image classification, image generation, and object detection; in the field of text processing, the input data of the model can be various text data for tasks such as text classification, sentiment analysis, text generation, and entity recognition. The present invention does not make specific limitations on this.

[0031] Each module may include multiple model layers. Here, a model layer is the basic computational unit in the model structure. Each layer receives input data, performs specific computations, and then outputs the results. Multiple layers are connected sequentially or in parallel to form a complete model. For example, each Transformer module may sequentially include Layer Normalization, an Attention layer, a residual connection, Layer Normalization, a Linear layer, a Multilayer Perceptron (MLP), and a residual connection. Among them, the Attention layer sequentially includes a Self Attention layer, a linear transformation layer, and a Dropout layer. The Multilayer Perceptron sequentially includes a linear transformation layer, an activation layer (such as GeLU), a linear transformation layer, and a Dropout layer.

[0032] It can be understood that each model layer corresponds to a specific computational operation. For example, Layer Normalization is used to normalize the input data, so its corresponding computational operation is a normalization operation; the Self Attention layer corresponds to a self-attention computational operation, and the linear transformation layer corresponds to a linear transformation operation; the residual connection layer is used to directly add the input to the output of a certain layer, so its corresponding computational operation is an Add operation. These operations are the basic steps in the model training process and will not be elaborated here. It should be understood that the first normalization operation represents the computational operation corresponding to the first Layer Normalization in the Transformer module, the second normalization operation represents the computational operation corresponding to the second Layer Normalization, and similarly, the first linear transformation operation, the second linear transformation operation, and the third linear transformation operation respectively represent the operations corresponding to different linear transformation layers.

[0033] In addition, in multi-GPU parallel training, in order to synchronize data on different GPUs, communication operations are required between GPUs. For example, a first all-gather operation is performed between the first normalization operation and the self-attention computational operation to gather data on all GPUs together; a first reduce-scatter operation is performed between the first linear transformation operation and the first Dropout operation to disperse the data from the gathered state back to each GPU; similarly, the second all-gather operation is used for data gathering, and the second reduce-scatter operation is used for the data to be dispersed again. These communication operations ensure that in a multi-GPU environment, the model can correctly synchronize and update parameters, thus achieving efficient training.

[0034] Figure 2 is a schematic diagram of kernel fusion in the related art, as Figure 2 shown, in the figure g represents an AllGather operation, g_Represents a ReduceScatter operation. A linear transformation is essentially a form of matrix multiplication. In deep learning frameworks, a linear transformation layer is typically implemented as the matrix multiplication of a weight matrix and input data, plus a bias term. Therefore, when performing a linear transformation operation on a GPU, it is actually using tensor cores to perform MMA operations. To reduce the storage and transmission requirements of intermediate data, the corresponding MMA operation and ReduceScatter operation in the left dashed box are usually fused into a single kernel for execution.

[0035] In addition, in the Transformer architecture, a linear transformation layer is followed by a non-linear activation function, such as GeLU (Gaussian Error Linear Unit). In some cases, to optimize computational efficiency and reduce memory access, these two operations are fused into a single kernel for execution. This kernel still internally utilizes tensor cores to perform MMA operations and then directly applies the GeLU activation function. Therefore, the linear transformation operation and the activation operation are essentially also performing MMA operations. Thus, the corresponding AllGather operation and MMA operation in the right dashed box can be fused into another kernel for execution.

[0036] However, during the above kernel fusion process, operations such as addition and normalization between the ReduceScatter operation and the AllGather operation are not fused. These unfused operations need to store intermediate results, increasing the number of data transfers between memory and registers, thereby increasing the data transfer overhead and resulting in limited performance improvement. Moreover, the above fusion strategy also has defects such as communication backpressure calculation, low resource utilization, and inability to fully utilize bidirectional bandwidth. In response to this, the present invention proposes a new fusion strategy that reorganizes the computational operations and communication operations corresponding to each model layer into an MMA + AllReduce + MMA structure, where the AllReduce operation fuses the ReduceScatter, Add, LayerNorm, and AllGather operations, enabling these operations to be efficiently executed in a single kernel, thereby overcoming the above defects.

[0037] Based on the above embodiments, Figure 3 is a schematic flowchart of the model performance optimization method provided by the present invention. As Figure 3 shown, the method includes: Step 310, based on the model structure, obtain the computational operations corresponding to each model layer and the communication operations between model layers.

[0038] It should be noted that the execution subject of the method provided in the embodiments of the present invention may be a computing device, such as a GPU, GPGPU (General-Purpose computing on Graphics Processing Units), TPU (Tensor Processing Unit), CPU (Central Processing Unit), etc. The embodiments of the present invention do not make specific limitations in this regard. Hereinafter, the technical solutions provided in the embodiments of the present invention will be introduced by taking the GPU as an example. Specifically, the model structure can be obtained through the interfaces provided by deep learning frameworks (such as TensorFlow, PyTorch, etc.), and the model structure can be parsed. During the parsing process, the model structure can be traversed first to identify the type of each model layer, and according to its type, the corresponding computing operation of each model layer can be determined. For example, the linear transformation layer corresponds to a linear transformation operation (i.e., matrix multiply-accumulate operation), and the residual connection layer corresponds to an addition operation, etc.

[0039] Furthermore, according to the parallel strategy adopted by each model layer, the communication operations that need to be inserted between model layers can be determined. Here, the parallel strategy is an important method in deep learning for training large neural network models, and it mainly focuses on the size of the model and the segmentation of the computing process. For example, the parallel strategy can include data parallelism, tensor parallelism, pipeline parallelism, etc. Among them, data parallelism means dividing the training data set into multiple subsets, and each subset is assigned to one or more devices (such as GPUs) for independent calculation; tensor parallelism means splitting a certain tensor operation (such as matrix multiplication) in the model along a specific dimension and distributing it to multiple devices for simultaneous execution; pipeline parallelism means dividing different layers of the neural network model into several stages, and each stage can be executed on different devices. In addition, for certain specific types of models (such as Transformer models), their parallel strategies can also include sequence parallelism, which is a method of splitting calculations in the sequence dimension. It splits a long sequence into multiple small blocks and performs parallel calculations on multiple devices.

[0040] Specifically, taking Figure 1Taking the neural network model shown as an example, each Transformer module in it sequentially includes model layers such as layer normalization, self-attention layer, linear transformation, dropout, residual connection, layer normalization, linear transformation, activation, linear transformation, dropout, and residual connection. Among them, in the order of the above model layers, layer normalization adopts a sequence parallel strategy (the layer normalization here forms a sequence parallel group), the self-attention layer and linear transformation adopt a tensor parallel strategy (the self-attention layer and linear transformation here form a tensor parallel group), dropout, residual connection, and layer normalization adopt a sequence parallel strategy (the dropout, residual connection, and layer normalization here also form a sequence parallel group), the three adjacent model layers of linear transformation, activation, and linear transformation all adopt a tensor parallel strategy, and the last two model layers of dropout and residual connection both adopt a sequence parallel strategy.

[0041] In hybrid parallel training, since the data splitting methods of sequence parallel and tensor parallel are different, when data is transferred between the sequence parallel group and the tensor parallel group, corresponding communication operations need to be performed to ensure the consistency of data transmission. For example, when the sequence parallel group and the tensor parallel group need to exchange data, an AllGather operation can be performed to meet the data requirements of the tensor parallel group; when the tensor parallel group and the sequence parallel group need to exchange data, a ReduceScatter operation can be performed to meet the data requirements of the sequence parallel group. In addition, when adjacent layers adopt the same parallel strategy, the communication in the middle can be omitted. In other words, the communication between the model layers located in the same tensor parallel group or the same sequence parallel group can be omitted.

[0042] After determining the calculation operations corresponding to the model layers and the communication operations between the model layers, these operations can be recorded in the form of code, data structure, or intermediate representation, so as to be organized into multiple computational communication parallel units later. It should be understood that step 310 can be completed in the control unit of the GPU, which is used to parse the model structure and generate corresponding operation instructions.

[0043] Step 320, organize the calculation operations corresponding to all model layers and the communication operations between the model layers into multiple computational communication parallel units. Each computational communication parallel unit includes a first matrix multiply-accumulate operation, a fused reduction operation, and a second matrix multiply-accumulate operation. The fused reduction operation includes a reduce-scatter operation, an addition operation, a normalization operation, and an all-gather operation.

[0044] It should be noted that the computational communication parallel unit is an abstraction that encapsulates the computational and communication operations in the model for efficient parallel execution on hardware. Each unit contains a series of carefully organized operations that can work together to minimize waiting time and maximize the utilization of computational resources. It should be understood that each GPU may include multiple Streaming Processor Clusters (SPCs), where one SPC is used to process one computational task, or multiple SPCs cooperate to process one computational task. Data sharing between multiple SPCs is achieved through global caches or global memory. Therefore, in the embodiments of the present invention, each computational communication unit can be assigned to one or more SPCs for processing.

[0045] In a GPU, each SPC may include multiple Compute Units (CUs, such as streaming processors). The compute units are divided into two categories, one is a vector operation unit, and the other is a tensor operation unit. Among them, the vector operation unit is used to perform arithmetic logic operations such as accumulation, conventional addition, subtraction, multiplication, division, and reduction. The tensor operation unit is specifically used to perform tensor-related operations, such as matrix multiplication and convolution operations. Each compute unit may include multiple cores (also known as compute cores or computational cores), and each core is used to perform specific computational tasks. Therefore, the first matrix multiply-accumulate operation and the second matrix multiply-accumulate operation within each computational communication unit can be assigned to the tensor operation unit within the SPC for processing, while each specific operation in the fused reduction operation can be assigned to each vector operation unit within the SPC for processing.

[0046] Specifically, considering that the model structure includes multiple network modules connected in series, and each network module includes multiple model layers, multiple computational communication parallel units can be obtained by organizing the computational and communication operations involved within each network module and the computational and communication operations of adjacent network modules. Each unit includes a first matrix multiply-accumulate operation (i.e., MMA operation), a fused reduction operation (i.e., AllReduce operation), and a second matrix multiply-accumulate operation (i.e., MMA operation), where the fused reduction operation may include ReduceScatter, Add operation, LayerNorm operation, and AllGather operation.

[0047] Figure 4 is a schematic structural diagram of the computational communication parallel unit provided by the present invention, as Figure 4As shown, taking two adjacent Transformer modules as an example, in the previous Transformer module, the operations corresponding to the structure shown by the dashed box A can be organized into a compute-communication parallel unit. Within this unit, the operations corresponding to the linear transformation layer are an MMA operation, which is executed as a kernel; g_ represents a ReduceScatter operation, the "+" represents a residual connection, which corresponds to an Add operation, and layer normalization corresponds to a LayerNorm operation, g represents an AllGather operation, and these operations are fused into an AllReduce operation and executed in a single kernel; the subsequent linear transformation layer and activation layer also correspond to an MMA operation, which is also executed in a single kernel. Therefore, the operations within each compute-communication parallel unit can be divided into 3 kernels for execution. It should be noted that the dropout layer is mainly used to randomly discard the outputs of some neurons during the training process to prevent the model from overfitting. During the inference and testing phases, this layer is turned off, so it is also fused into the AllReduce operation during the training process.

[0048] Furthermore, the operations corresponding to the structure shown by the dashed box B can also be organized into a compute-communication parallel unit. Within this unit, the previous model layer structure is the same as that within the dashed box A, which will not be elaborated here. And within the dashed box B, g the subsequent model layer is a self-attention layer. Since the self-attention layer will first perform a linear transformation on the input data after receiving the input data, this linear transformation operation can be regarded as an MMA operation, thus forming a compute-communication parallel unit, that is, an MMA+AllReduce+MMA structure is organized. The structure shown by the dashed box C is also a compute-communication parallel unit, and its structure is the same as that shown by the dashed box A, which will not be elaborated here. It should be understood that regardless of how many Transformer modules are included in the model structure, the compute operations and communication operations corresponding to the model layers can be organized into multiple compute-communication parallel units according to the above process to achieve the parallel execution of compute operations and communication operations, thereby improving the overall computing efficiency.

[0049] Step 330, within each compute-communication parallel unit, the input data is sliced so that the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation are executed in parallel based on different data blocks.

[0050] Specifically, within each compute communication parallel unit, the input data can be segmented to enable parallel execution of computation and communication using different data blocks. Specifically, first, according to the characteristics of the model and the parallel processing capabilities of the hardware, the input data can be segmented into multiple small pieces, i.e., multiple data blocks are obtained; then, different data blocks are assigned to different operations within the unit (i.e., the first matrix multiply-accumulate operation, the fused reduction operation, the second matrix multiply-accumulate operation) for processing to achieve parallel execution. Here, each operation is executed as an independent kernel. It should be understood that in different application scenarios, the physical meaning of the above input data can be different. For example, the input data can be voice data, image data, or text data, specifically depending on the scenario and field to which the method provided by the embodiments of the present invention is applied.

[0051] For example, assume that the first matrix multiply-accumulate operation is executed as kernel1, the fused reduction operation is executed as kernel2, and the second matrix multiply-accumulate operation is executed as kernel3. After the input data is segmented, the first matrix multiply-accumulate operation can be performed on the first data block in kernel1 first. After obtaining the corresponding computation result, the fused reduction operation is performed on the computation result of this data block in kernel2. At the same time, the first matrix multiply-accumulate operation is performed on the second data block in kernel1, thus achieving parallel execution of the first matrix multiply-accumulate operation and the fused reduction operation. Next, after completing the fused reduction processing of the first data block and the matrix multiply-accumulate calculation of the second data block, the first matrix multiply-accumulate operation can be performed on the third data block in kernel1, the fused reduction operation is performed on the computation result of the second data block in kernel2, and the second matrix multiply-accumulate operation is performed on the fused reduction result of the first data block in kernel3. At this time, the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation are executed in parallel. And so on, until all data blocks are processed.

[0052] The method provided by the embodiments of the present invention organizes the computational operations of all model layers and the communication operations between model layers into multiple computational communication parallel units, and slices the input data within each unit, enabling the parallel execution of the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation. Since the matrix multiply-accumulate operation is a computational operation and the fused reduction operation includes communication operations, the parallel execution of computational operations and communication operations is achieved, thereby improving the parallelism and accelerating the computational process of the overall model, and significantly enhancing the model performance. Within each unit, the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation are executed in parallel based on different data blocks, which means that the computational operations and communication operations are independently performed in different cores, avoiding the situation where the communication operation becomes a bottleneck and backpressures the computational operation, so the waiting time can be reduced and the utilization rate of computational resources and hardware resources can be improved.

[0053] In addition, since the fused reduction operation includes a reduce-scatter operation, an addition operation, a normalization operation, and an all-gather operation, this means that the reduce-scatter operation and the all-gather operation are fused and executed in one core. This fusion strategy can make full use of the bidirectional interconnection bandwidth in modern hardware architectures, reducing the latency and overhead of data transmission and further enhancing the model performance.

[0054] Based on any of the above embodiments, the neural network model provided by the embodiments of the present invention may include multiple network modules connected in series. Each network module may include multiple model layers, and each model layer corresponds to a specific computational operation, and there are also corresponding communication operations between model layers. For example, taking the Transformer module as an example, the computational operations corresponding to its model layers and the communication operations between model layers may include: a first normalization operation, a first all-gather operation, a self-attention calculation operation, a first linear transformation operation, a first reduce-scatter operation, a first dropout operation, a first addition operation, a second normalization operation, a second all-gather operation, a second linear transformation operation, an activation operation, a third linear transformation operation, a second reduce-scatter operation, a second dropout operation, and a second addition operation.

[0055] Correspondingly, step 320 specifically includes: Step 321, obtaining a computational communication parallel unit based on the first linear transformation operation, the first reduce-scatter operation, the first dropout operation, the first addition operation, the second normalization operation, the second all-gather operation, the second linear transformation operation, and the activation operation in the current network module.

[0056] It should be noted that each network module includes multiple model layers, and each model layer corresponds to a specific computing operation. To achieve parameter synchronization, communication operations are required between some model layers. By organizing the computing operations and communication operations in the network module, a computing and communication parallel unit can be obtained. Here, the network module can be a Transformer module. The model layer structure of this model, the computing operation corresponding to each model layer, and the communication operation involved between model layers can be specifically referred to the introduction in the above embodiments, and will not be elaborated here.

[0057] Further, step 321 specifically includes: Step 3211, taking the first linear transformation operation as the first matrix multiplication and accumulation operation.

[0058] Specifically, linear transformation is essentially a form of matrix multiplication. When performing a linear transformation operation on a GPU, in fact, it is using tensor cores to perform MMA operations. Therefore, the linear transformation operation can be regarded as an MMA operation.

[0059] Step 3212, taking the first reduction scatter operation, the first dropout operation, the first addition operation, the second normalization operation, and the second all-gather operation as a fused reduction operation.

[0060] Specifically, in order to achieve computing and communication parallelism while avoiding communication backpressure on computing and making full use of the interconnection bidirectional bandwidth, the embodiments of the present invention fuse the first reduction scatter operation (i.e., the ReduceScatter operation) and the second all-gather operation (i.e., the AllGather operation) into one kernel for execution.

[0061] In addition, for the Transformer architecture, in order to further improve the parallelism and overall computing efficiency, the dropout operation, addition operation, and normalization operation between the ReduceScatter operation and the AllGather operation can also be fused into this kernel for execution, thus forming a new kernel structure model, that is, the fused reduction operation. Here, the fused reduction operation includes the ReduceScatter operation, Add operation, LayerNorm operation, and AllGather operation, etc., abbreviated as the AllReduce operation.

[0062] Step 3213, taking the second linear transformation operation and the activation operation as the second matrix multiplication and accumulation operation.

[0063] Specifically, to optimize the computational efficiency and reduce memory access, the linear transformation operation and the activation operation can be fused into a single kernel for execution. Inside this kernel, tensor cores are still utilized for MMA operations, and then the GeLU activation function is directly applied. Therefore, the linear transformation operation and the activation operation are essentially also performing MMA operations.

[0064] Step 3214: Based on the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation, obtain the compute communication parallel unit.

[0065] Specifically, according to the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation organized in the above steps 3211 to 3213, a compute communication parallel unit can be formed. Specifically, reference can be made to Figure 4 the structure shown by the dashed box A or the dashed box C in

[0066] Step 322: Based on the third linear transformation operation, the second reduction-scatter operation, the second dropout operation, the second addition operation in the current network module, and the first normalization operation, the first all-gather operation, and the self-attention calculation operation in the subsequent network module, obtain the compute communication parallel unit.

[0067] Specifically, according to the computational operations and communication operations in adjacent network modules, a compute communication parallel unit can also be organized. Specifically, the third linear transformation operation in the current network module can be used as the first matrix multiply-accumulate operation; the second reduction-scatter operation, the second dropout operation, the second addition operation in the current network module, and the first normalization operation and the first all-gather operation in the subsequent network module can be fused into a single kernel for execution, that is, the fused reduction operation is obtained; the linear transformation operation involved in the self-attention calculation operation in the subsequent network module is used as the second matrix multiply-accumulate operation; thus, according to the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation, a compute communication parallel unit can be organized. It should be understood that the current network module and the subsequent network module can be two adjacent Transformer modules, where the current network module is the previous network module of the subsequent network module. Specifically, reference can be made to Figure 4 the structure shown by the dashed box B in

[0068] Based on any of the above embodiments, step 330 specifically includes: Step 331: Inside each compute communication parallel unit, split the input data to obtain multiple data blocks.

[0069] Specifically, first, based on the characteristics of the model and the parallel processing capabilities of the hardware (such as GPU), a data splitting strategy can be determined, which includes determining the size and quantity of data blocks, etc.; then, according to the determined splitting strategy, the input data is split to obtain multiple data blocks, and these data blocks will be used as the input for subsequent calculation operations. It should be understood that after the data splitting is completed, effective management of the data blocks is required, including storage, indexing, and access, etc., to ensure that they can be correctly accessed and processed during subsequent calculations.

[0070] Step 332: Sequentially perform the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation on each data block; Among them, the fused reduction operation of the current data block is executed in parallel with the first matrix multiply-accumulate operation of the next data block, and the second matrix multiply-accumulate operation of the current data block is executed in parallel with the fused reduction operation of the next data block.

[0071] Specifically, after splitting to obtain multiple data blocks, different data blocks can be assigned to different operations within the unit for parallel execution. For example, the current data block is used to execute the fused reduction operation, and the next data block is used to execute the first matrix multiply-accumulate operation. Here, the current data block refers to the data block currently being processed, and the next data block refers to the next data block after the current data block, that is, the next data block that needs to be processed.

[0072] In the embodiments of the present invention, through the implementation of Step 331 and Step 332, the input data can be effectively split into multiple data blocks, and a series of calculation and communication operations can be executed in parallel within the calculation communication parallel unit to maximize the masking of communication latency, which helps to improve the training speed of the model and make full use of the parallel processing capabilities of the hardware, thereby improving the parallelism and model performance.

[0073] Based on any of the above embodiments, Step 332 specifically includes: Step 3321: Based on the first kernel, perform the first matrix multiply-accumulate operation on any data block to obtain the first calculation result; Step 3322: Based on the second kernel, perform the fused reduction operation on the first calculation result to obtain the fused result; Step 3323: Based on the third kernel, perform the second matrix multiply-accumulate operation on the fused result to obtain the second calculation result of any data block.

[0074] Specifically, the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation for the same data block are executed serially. Below, taking any data block as an example, the processing process of each data block will be introduced.

[0075] Figure 5It is a schematic diagram of the execution process of the computing device provided by the present invention. As Figure 5 shown, it is assumed that the computing devices involved in the model distributed training include GPU0 to GPU3. On each GPU, a first kernel, a second kernel, and a third kernel are correspondingly deployed. The first kernel is used to perform the first matrix multiplication and accumulation operation (i.e., MMA operation), the second kernel is used to perform the fusion reduction operation (i.e., AllReduce operation), and the third kernel is used to perform the second matrix multiplication and accumulation operation (i.e., MMA operation).

[0076] First, the left and right matrices (such as Figure 5 the S and H matrices shown in ) are sliced. The sliced data blocks are respectively subjected to MMA calculations in the first kernels on GPU0 to GPU3 to obtain corresponding calculation results. Figure 5 The four small squares stacked in respectively represent the results obtained after MMA calculations on these 4 GPUs.

[0077] After completing the MMA operation, then perform the fusion reduction operation on the calculation results in the second kernels on GPU0 to GPU3. Specifically, first perform the ReduceScatter operation on the calculation results. These 4 GPUs will transmit the first 1 / 4 part of the calculation results to GPU0, the second 1 / 4 part to GPU1, the third 1 / 4 part to GPU2, and the fourth 1 / 4 part to GPU3, so that each GPU only retains part of the results. Subsequently, continue to perform the Add operation and the LayerNorm operation on this part of the retained results in the second kernel of each GPU in turn. These operations do not change the shape of the data. Finally, perform the AllGather operation so that each GPU has complete data.

[0078] After completing the AllReduce operation, perform the MMA operation on the complete data on each GPU in the third kernels on GPU0 to GPU3, thereby completing the processing of the same data block within one unit. And so on until all data blocks are processed.

[0079] Based on any of the above embodiments, step 3322 specifically includes: Step S1, based on the number of computing devices, slice the first calculation result to obtain multiple result blocks, and send a synchronization signal to the second kernel locally and the second kernels of each remote computing device.

[0080] Specifically, the number of computing devices refers to the total number of devices participating in the current computing task. For example, the number of computing devices can be the number of GPUs participating in the computing task. In distributed training, to make full use of each computing device, a large computing task or dataset is usually split into multiple small parts, and each part is processed by one computing device. Specifically, according to the number of computing devices, the first computing result can be evenly split into multiple parts (i.e., result chunks) equal to the number of devices, and each part is processed by one computing device.

[0081] After the computing result is split, a synchronization signal can be sent to the second kernel locally and the second kernels of each remote computing device. Here, the second kernel locally refers to the second kernel on the current computing device used to execute a specific computing task (such as a fused reduction operation). In parallel computing, each computing device can have multiple kernels to process multiple tasks simultaneously. A remote computing device refers to another computing device connected to the current computing device through a network. The remote computing device and the local computing device cooperate together to complete the computing task.

[0082] It can be understood that the synchronization signal is a signal used to coordinate the operation order among multiple computing devices. In distributed computing, the synchronization signal ensures that all computing devices start a certain computing task at the correct moment, thus maintaining data consistency and computing correctness.

[0083] Step S2, when the second kernel detects that the synchronization signal value is equal to the number of computing devices, read one result chunk locally and one result chunk from each remote computing device to obtain a first result.

[0084] Specifically, for each computing device, when its second kernel detects that the synchronization signal value is equal to the number of computing devices, it means that all computing devices are ready to start executing the next computing task. In this case, the second kernel can read one result chunk reserved locally and the corresponding result chunks on each remote computing device to complete the ReduceScatter operation, thereby obtaining a first result. Here, the synchronization signal value is a numerical value used to represent the number of computing devices that have received the synchronization signal currently. It should be understood that the first result refers to the result obtained on each device after the ReduceScatter operation is completed.

[0085] Step S3, perform an addition operation and a normalization operation on the first result in sequence to obtain a second result.

[0086] Specifically, for each computing device, after it reads the first result, it can sequentially perform an addition operation and a normalization operation on the result to obtain a second result. Here, the second result refers to the result obtained by performing the addition operation and the normalization operation on the first result on each computing device.

[0087] Subsequently, an AllGather operation can be performed on the second result. Specifically, each computing device locally retains the second result while sending the second result to each remote computing device for storage. Correspondingly, each remote computing device also transmits its respective second result to this local computing device for storage.

[0088] Step S4: Obtain the fusion result based on the local second result and the second results of the respective remote computing devices.

[0089] Specifically, for the local computing device, it can obtain the fusion result based on the second result saved locally and the second results transmitted by the respective remote computing devices received. This fusion result will continue to perform the MMA operation in the third kernel.

[0090] Figure 6 is a schematic diagram of the parallel execution of the first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation provided by the present invention. As Figure 6 (a) shows, taking the parallel execution of GPU0 and GPU1 as an example, for GPU0, assuming that the computing task corresponding to the first matrix multiplication and accumulation operation is A×B = C, where A and B are input matrices and C is the output matrix, this computing task can be split into 4 computing subtasks: A0×B0 = C0, A1×B1 = C1, A2×B2 = C2, and A3×B3 = C3, thereby obtaining 4 corresponding output submatrices (i.e., C0, C1, C2, and C3 shown in the figure). First, perform the MMA operation on data blocks A0 and B0 in the first kernel of GPU0 to obtain the corresponding computing result C0. Then, perform the AllReduce operation on the computing result C0 in the second kernel of GPU0. At the same time, the MMA operation can be performed on data blocks A1 and B1 in the first kernel of GPU0 to enable the parallel execution of the MMA operation and the AllReduce operation. Next, perform the MMA operation on the fusion result of the computing result C0 (i.e., the result obtained by the AllReduce operation) in the third kernel of GPU0, and at the same time perform the AllReduce operation on the computing result C1 in the second kernel. At this time, since the first kernel is idle, the MMA operation can be performed on data blocks A2 and B2 in the first kernel to achieve the parallel execution of the MMA operation, the AllReduce operation, and the MMA operation. The specific parallel execution process is as Figure 6 (b) shows.

[0091] The process of performing the AllReduce operation on each calculation result (i.e., C0, C1, C2, and C3) is as follows: Taking the calculation result C0 as an example, first, according to the number of computing devices, the calculation result C0 is split. For example, in the embodiments of the present invention, since there are only 2 GPUs, the calculation result C0 can be split into two small pieces (i.e., C00 and C01). Among them, the first small piece after splitting (such as Figure 6 the dark green rectangular block C00 shown in (a)) will be retained locally (i.e., GPU0), while the second small piece (such as Figure 6 the red rectangular block C01 shown in (a)) will be synchronized to the remote computing device (i.e., GPU1) for storage. At the same time, GPU1 will also perform the same AllReduce operation on the calculation results processed thereon, so that only partial results are retained on each GPU, thus completing the ReduceScatter operation. Subsequently, the Add and LayerNorm operations are sequentially performed on the results obtained after the ReduceScatter operation on each GPU. It should be understood that after the calculation result is split on GPU0, two synchronization signals need to be sent, one synchronization signal is used to inform the second kernel locally, and the other synchronization signal is used to inform the second kernel of the remote GPU1, so that when the second kernel detects that the synchronization signal value is equal to the number of computing devices, it starts to perform the AllReduce operation.

[0092] For GPU0, after performing the Add and LayerNorm operations, the obtained result will be saved locally, and at the same time, the result will be transmitted to the remote GPU1 for storage, so that each GPU has a complete result, thus completing the AllGather operation. In addition, the second kernel will send synchronization signals to the third kernel locally and the third kernel of the remote GPU1, so that when the third kernel detects that the synchronization signal value is equal to the number of computing devices, it starts to perform another MMA operation. It should be noted that Figure 6 the blue rectangular block and the orange rectangular block on the far right of (a) respectively represent the left and right matrices of the second MMA operation, which is similar to the first MMA operation and will not be elaborated here.

[0093] It can be understood that in multi-GPU parallel training, the linear transformation layer usually adopts tensor parallelism technology, the residual connection and layer normalization adopt sequence parallelism technology, and the linear transformation and activation also adopt tensor parallelism technology. A ReduceScatter operation is performed between tensor parallelism and sequence parallelism, and an AllGather operation is performed between sequence parallelism and tensor parallelism. Since tensor parallelism is executed using tensor cores (i.e., the computing cores of tensor operation units), and sequence parallelism is executed using vector cores (also known as vector cores, i.e., the computing cores of vector operation units), therefore, the ReduceScatter operation utilizes the unidirectional bandwidth from tensor cores to vector cores, and the AllGather operation utilizes the unidirectional bandwidth from vector cores to tensor cores. By fusing the ReduceScatter and AllGather operations into one kernel for execution, the tensor cores, vector cores, and bidirectional interconnection bandwidth can be fully utilized, thereby improving the overall computing efficiency.

[0094] The model performance optimization device provided by the present invention will be described below. The model performance optimization device described below can be correspondingly referred to the model performance optimization method described above.

[0095] Based on any of the above embodiments, Figure 7 is a schematic structural diagram of the model performance optimization device provided by the present invention, as Figure 7 shown. The device includes: An acquisition unit 710, configured to obtain the calculation operations corresponding to each model layer and the communication operations between model layers based on the model structure; An organization unit 720, configured to organize the calculation operations corresponding to all model layers and the communication operations between model layers into multiple calculation and communication parallel units. Each calculation and communication parallel unit includes a first matrix multiplication and accumulation operation, a fusion reduction operation, and a second matrix multiplication and accumulation operation. The fusion reduction operation includes a reduction scatter operation, an addition operation, a normalization operation, and an all-gather operation; A parallel unit 730, configured to split the input data within each calculation and communication parallel unit so that the first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation are executed in parallel based on different data blocks.

[0096] The device provided by the embodiment of the present invention organizes the calculation operations of all model layers and the communication operations between model layers into multiple calculation and communication parallel units, and divides the input data within each unit, so as to realize the parallel execution of the first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation. Since the matrix multiplication and accumulation operation is a calculation operation and the fusion reduction operation includes a communication operation, the parallel execution of the calculation operation and the communication operation is realized. Thereby, the parallelism can be improved, the calculation process of the overall model can be accelerated, and the model performance can be significantly improved. Within each unit, the first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation are executed in parallel based on different data blocks, which means that the calculation operation and the communication operation are independently performed in different cores, avoiding the situation where the communication operation becomes a bottleneck and backpressures the calculation operation. Therefore, the waiting time can be reduced, and the utilization rate of the calculation resources and the hardware resources can be improved.

[0097] In addition, since the fusion reduction operation includes a reduce-scatter operation, an addition operation, a normalization operation, and an all-gather operation, this means that the reduce-scatter operation and the all-gather operation are fused and executed in one core. This fusion strategy can make full use of the bidirectional interconnection bandwidth in the modern hardware architecture, reduce the latency and overhead of data transmission, and further improve the model performance.

[0098] Based on any of the above embodiments, the parallel unit 730 includes: A splitting subunit, configured to split the input data within each calculation and communication parallel unit to obtain a plurality of data blocks; An execution subunit, configured to sequentially execute the first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation on each data block; Wherein, the fusion reduction operation of the current data block is executed in parallel with the first matrix multiplication and accumulation operation of the next data block, and the second matrix multiplication and accumulation operation of the current data block is executed in parallel with the fusion reduction operation of the next data block.

[0099] Based on any of the above embodiments, the execution subunit includes: A first calculation subunit, configured to execute the first matrix multiplication and accumulation operation on any data block based on the first core to obtain a first calculation result; A fusion reduction subunit, configured to execute the fusion reduction operation on the first calculation result based on the second core to obtain a fusion result; A second calculation subunit, configured to execute the second matrix multiplication and accumulation operation on the fusion result based on the third core to obtain the second calculation result of the any data block.

[0100] Based on any of the above embodiments, the fusion reduction subunit is specifically configured to: Segment the first calculation result based on the number of computing devices to obtain multiple result chunks, and send a synchronization signal to the second core locally and the second cores of each remote computing device; When the second core detects that the synchronization signal value is equal to the number of computing devices, read one result chunk locally and one result chunk from each remote computing device to obtain a first result; Perform an addition operation and a normalization operation on the first result in sequence to obtain a second result; Based on the second result locally and the second results of each remote computing device, obtain the fusion result.

[0101] Based on any of the above embodiments, the model structure includes a plurality of network modules connected in series, each network module includes a plurality of model layers, and the computing operations corresponding to the plurality of model layers and the communication operations between the model layers include a first normalization operation, a first all-gather operation, a self-attention calculation operation, a first linear transformation operation, a first reduce-scatter operation, a first dropout operation, a first addition operation, a second normalization operation, a second all-gather operation, a second linear transformation operation, an activation operation, a third linear transformation operation, a second reduce-scatter operation, a second dropout operation, and a second addition operation.

[0102] Based on any of the above embodiments, the organizing unit 720 is specifically configured to: Based on the first linear transformation operation, the first reduce-scatter operation, the first dropout operation, the first addition operation, the second normalization operation, the second all-gather operation, the second linear transformation operation, and the activation operation in the current network module, obtain a computing and communication parallel unit; Based on the third linear transformation operation, the second reduce-scatter operation, the second dropout operation, the second addition operation in the current network module, and the first normalization operation, the first all-gather operation, and the self-attention calculation operation in the subsequent network module, obtain a computing and communication parallel unit.

[0103] Based on any of the above embodiments, the organizing unit 720 is specifically configured to: Regard the first linear transformation operation as a first matrix multiply-accumulate operation; Regard the first reduce-scatter operation, the first dropout operation, the first addition operation, the second normalization operation, and the second all-gather operation as a fusion reduce operation; Regard the second linear transformation operation and the activation operation as a second matrix multiply-accumulate operation; Based on the first matrix multiply-accumulate operation, the fusion reduce operation, and the second matrix multiply-accumulate operation, obtain the computing and communication parallel unit.

[0104] Figure 8 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 8 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 may call the logical instructions in the memory 830 to execute a model performance optimization method, which includes: based on the model structure, obtaining the calculation operations corresponding to each model layer and the communication operations between model layers; organizing the calculation operations corresponding to all model layers and the communication operations between model layers into multiple computation-communication parallel units, each computation-communication parallel unit including a first matrix multiply-accumulate operation, a fused reduction operation, and a second matrix multiply-accumulate operation, and the fused reduction operation including a reduce-scatter operation, an addition operation, a normalization operation, and an all-gather operation; within each computation-communication parallel unit, splitting the input data so that the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation are executed in parallel based on different data blocks.

[0105] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0106] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the model performance optimization method provided by each of the above methods. The method includes: based on the model structure, obtaining the calculation operations corresponding to each model layer and the communication operations between the model layers; organizing the calculation operations corresponding to all model layers and the communication operations between the model layers into multiple calculation-communication parallel units, each of which includes a first matrix multiplication and accumulation operation, a fusion reduction operation, and a second matrix multiplication and accumulation operation. The fusion reduction operation includes a reduction scattering operation, an addition operation, a normalization operation, and a full aggregation operation; within each calculation-communication parallel unit, the input data is segmented so that the first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation are executed in parallel based on different data blocks.

[0107] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the model performance optimization method provided by each of the above methods. The method includes: based on the model structure, obtaining the calculation operations corresponding to each model layer and the communication operations between the model layers; organizing the calculation operations corresponding to all model layers and the communication operations between the model layers into multiple calculation-communication parallel units, each of which includes a first matrix multiplication and accumulation operation, a fusion reduction operation, and a second matrix multiplication and accumulation operation. The fusion reduction operation includes a reduction scattering operation, an addition operation, a normalization operation, and a full aggregation operation; within each calculation-communication parallel unit, the input data is segmented so that the first matrix multiplication and accumulation operation, the fusion reduction operation, and the second matrix multiplication and accumulation operation are executed in parallel based on different data blocks.

[0108] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0109] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the related technologies can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for optimizing model performance, characterized in that, Including: Based on the model structure, obtain the computing operations corresponding to each model layer and the communication operations between model layers; Organize the computing operations corresponding to all model layers and the communication operations between model layers into multiple compute-communication parallel units, each compute-communication parallel unit including a first matrix multiply-accumulate operation, a fused reduction operation, and a second matrix multiply-accumulate operation, where the fused reduction operation includes a reduce-scatter operation, an addition operation, a normalization operation, and an all-gather operation; Within each compute-communication parallel unit, split the input data so that the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation are executed in parallel based on different data blocks.

2. The model performance optimization method according to claim 1, wherein The step of splitting the input data within each compute-communication parallel unit so that the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation are executed in parallel based on different data blocks includes: Within each compute-communication parallel unit, split the input data to obtain multiple data blocks; Successively perform a first matrix multiply-accumulate operation, a fused reduction operation, and a second matrix multiply-accumulate operation on each data block; Among them, the fused reduction operation of the current data block is executed in parallel with the first matrix multiply-accumulate operation of the next data block, and the second matrix multiply-accumulate operation of the current data block is executed in parallel with the fused reduction operation of the next data block.

3. The model performance optimization method according to claim 2, wherein The step of successively performing a first matrix multiply-accumulate operation, a fused reduction operation, and a second matrix multiply-accumulate operation on each data block includes: Based on a first kernel, perform a first matrix multiply-accumulate operation on any data block to obtain a first computation result; Based on a second kernel, perform a fused reduction operation on the first computation result to obtain a fused result; Based on a third kernel, perform a second matrix multiply-accumulate operation on the fused result to obtain a second computation result of the any data block.

4. The model performance optimization method according to claim 3, wherein The step of performing a fused reduction operation on the first computation result based on the second kernel to obtain a fused result includes: Based on the number of computing devices, split the first computation result to obtain multiple result chunks, and send a synchronization signal to the local second kernel and the second kernels of each remote computing device; When the second kernel detects that the synchronization signal value is equal to the number of computing devices, read one result chunk from the local and one result chunk from each remote computing device to obtain a first result; Successively perform an addition operation and a normalization operation on the first result to obtain a second result; Based on the local second result and the second results of each remote computing device, obtain the fused result.

5. The model performance optimization method according to any one of claims 1 to 4, characterized in that The model structure includes multiple network modules connected in series, each network module including multiple model layers, and the computing operations corresponding to the multiple model layers and the communication operations between model layers include a first normalization operation, a first all-gather operation, a self-attention calculation operation, a first linear transformation operation, a first reduce-scatter operation, a first dropout operation, a first addition operation, a second normalization operation, a second all-gather operation, a second linear transformation operation, an activation operation, a third linear transformation operation, a second reduce-scatter operation, a second dropout operation, and a second addition operation.

6. The model performance optimization method according to claim 5, wherein Organizing the computing operations corresponding to all model layers and the communication operations between model layers into multiple computing and communication parallel units includes: Obtaining a computing and communication parallel unit based on the first linear transformation operation, the first reduction and scattering operation, the first dropout operation, the first addition operation, the second normalization operation, the second all-gather operation, the second linear transformation operation, and the activation operation in the current network module; Obtaining a computing and communication parallel unit based on the third linear transformation operation, the second reduction and scattering operation, the second dropout operation, the second addition operation in the current network module, and the first normalization operation, the first all-gather operation, and the self-attention calculation operation in the subsequent network module.

7. The model performance optimization method according to claim 6, wherein The obtaining a computing and communication parallel unit based on the first linear transformation operation, the first reduction and scattering operation, the first dropout operation, the first addition operation, the second normalization operation, the second all-gather operation, the second linear transformation operation, and the activation operation in the current network module includes: Regarding the first linear transformation operation as the first matrix multiply-accumulate operation; Regarding the first reduction and scattering operation, the first dropout operation, the first addition operation, the second normalization operation, and the second all-gather operation as a fused reduction operation; Regarding the second linear transformation operation and the activation operation as the second matrix multiply-accumulate operation; Obtaining the computing and communication parallel unit based on the first matrix multiply-accumulate operation, the fused reduction operation, and the second matrix multiply-accumulate operation.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the model performance optimization method according to any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the model performance optimization method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the model performance optimization method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Communication calculation parallel optimization method, multiprocessor system, medium and program product

    CN121187815A

  • Multiprocessor system, data processing method, electronic device, and storage medium

    CN121188333A

  • Multiprocessor systems, data processing methods, electronic devices, storage media

    CN121188333B

  • Optimization method and device of hybrid expert model, computer equipment, readable storage medium and program product

    CN121351885A

  • Optimization method and device of mixed expert model, computer equipment, readable storage medium and program product

    CN121351885B