Neural network model optimization method and related device
By fusing and encapsulating operators in the computation graph of the Transformer neural network model, the problem of low efficiency and accuracy of the model on mobile devices is solved, and efficient deployment of neural network models is achieved.
Patent Information
- Application Number
- CN202410700521.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-12-02
AI Technical Summary
Existing technologies for deploying Transformer neural network models on mobile devices suffer from issues such as splitting network layers, inefficient fusion between layers, and misalignment of layer implementation results. These issues lead to reduced forward inference speed or decreased inference accuracy, resulting in low operational efficiency and accuracy.
By obtaining the structure and parameters of the neural network model, a first computation graph is generated. The computational complexity and memory access cost of each operator are calculated. Operators within and between sub-network modules are merged to form a custom Transformer layer. TensorRT is then used to encapsulate the layer, generating an optimized neural network model.
The optimized neural network model maintains high accuracy when deployed on mobile devices, improves operating efficiency, reduces computing resource consumption, and does not affect model accuracy.
Smart Images

Figure CN121052287A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer science, and in particular to a method and apparatus for optimizing neural network models. Background Technology
[0002] With the rapid development of the computer industry, artificial intelligence (AI) has also made tremendous progress. AI primarily relies on neural network models for decision-making. Neural network models have various structures, with the Transformer network being a current major application. The Transformer is a network structure that utilizes attention mechanisms to improve model training speed. It overcomes the limitation of recurrent neural network models in parallel computation, and the number of operations required to calculate the correlation between two locations does not increase with distance, making it parallel-friendly. Therefore, the Transformer has significantly changed the neural network design paradigm in various subfields of AI, becoming one of the fundamental structures of neural network models.
[0003] Currently, neural network models, including Transformers, are increasingly being introduced into the field of autonomous driving to achieve various perception tasks. When deploying neural network models on mobile devices, the relevant technology involves transforming and optimizing the trained neural network model to generate the runtime representation for the inference engine. This method relies entirely on the inference engine's own technical capabilities in optimizing and fusing the network structure of the neural network model. Consequently, issues such as splitting network layers, inefficient inter-layer fusion, and misalignment of layer implementation results can arise. These problems lead to reduced forward inference speed or decreased inference accuracy, resulting in lower efficiency and accuracy of the neural network model after deployment on mobile devices. Summary of the Invention
[0004] This application provides a neural network model optimization method and related apparatus, which can optimize the neural network model so that the neural network model can have high accuracy and high operating efficiency when deployed on mobile devices.
[0005] This application provides a method for optimizing a neural network model, the method comprising:
[0006] The structure and parameters of the neural network model are obtained, and a first computational graph is generated. The first computational graph includes a Transformer network module, and the Transformer module includes multiple sub-network modules, each of which includes multiple operators.
[0007] Calculate the computational complexity or memory access cost of each operator, fuse multiple operators within the sub-network module and fuse operators between multiple sub-network modules based on the computational complexity or memory access cost to obtain a custom Transformer layer, and encapsulate the custom Transformer layer using TensorRT to form a plugin;
[0008] Replace the Transformer module with a custom Transformer layer to generate a second computation graph;
[0009] Based on the second computation graph, the plugin is loaded using TensorRT to generate an engine file, which includes the optimized neural network model corresponding to the second computation graph.
[0010] Optionally, the method further includes:
[0011] Obtain the data calculation relationship between multiple operators, classify the multiple operators according to the data calculation relationship, and obtain multiple dependency relationships;
[0012] The calculation of the computational complexity or memory access cost of each operator, the fusion of multiple operators within the sub-network module based on the computational complexity or memory access cost, and the fusion of operators between multiple sub-network modules include:
[0013] Calculate the computational complexity or memory access cost of each operator and the computational complexity or memory access cost of fusing multiple operators according to the dependencies, and fuse multiple operators within the sub-network module based on the lowest computational complexity or memory access cost.
[0014] Optionally, the dependencies include serial relationships, Y-type relationships, and complex relationships;
[0015] The calculation of the computational complexity or memory access cost of each operator, and the calculation of the computational complexity or memory access cost after fusing multiple operators according to the dependency relationship, and the fusion of multiple operators within the sub-network module based on the lowest computational complexity or memory access cost, as well as the fusion of operators between multiple sub-network modules, include:
[0016] Calculate the first computational complexity and the first memory access cost for each of the operators;
[0017] The second computational complexity and the second memory access cost of the multiple operators categorized as serial relationships, Y-type relationships, and complex relationships are calculated and fused. The multiple operators within the sub-network module are fused based on the lowest of either the first computational complexity being greater than the second computational complexity or the first memory access cost being greater than the second memory access cost.
[0018] Optionally, the method further includes:
[0019] Obtain similar operators with the same structure from multiple operators in different sub-network modules;
[0020] The calculation of the computational complexity or memory access cost of each operator, the fusion of multiple operators within the sub-network module based on the computational complexity or memory access cost, and the fusion of operators between multiple sub-network modules include:
[0021] The computational complexity or memory access cost of each operator is calculated, as well as the computational complexity or memory access cost of similar operators fused in different sub-network modules. Operators among multiple sub-network modules are fused based on the lowest computational complexity or memory access cost.
[0022] Optionally, the method further includes:
[0023] Obtain multiple independent operators among the multiple operators that have no data computational relationship;
[0024] Multiple irrelevant operators are merged into a single operator.
[0025] Optionally, the irrelevance operator is an operator that processes vectors;
[0026] The step of fusing multiple irrelevant operators into a single operator includes:
[0027] When the vectors input to multiple independent operators are the same, the multiple independent operators are merged into a single operator.
[0028] Optionally, the Transformer module includes N encoding layers and N decoding layers; the encoding layer includes a multi-head self-attention (MHA) layer, a feedforward network (FFN) layer, and a normalized low-level network (LN) layer; the decoding layer includes an MHA layer, an encoder-decoder attention layer, an FFN layer, and an LN layer; and the sub-network module is an MHA layer, an FFN layer, or an LN layer.
[0029] The method further includes:
[0030] The encoding layer and the decoding layer are processed in parallel using multiple streams.
[0031] This application provides a neural network model optimization device, the device comprising:
[0032] The first generation unit is used to obtain the structure and parameters of the neural network model and generate a first computation graph. The first computation graph includes a Transformer network module, and the Transformer module includes multiple sub-network modules. Each sub-network module includes multiple operators.
[0033] A fusion unit is used to calculate the computational complexity or memory access cost of each operator, fuse multiple operators within the sub-network module and fuse operators between multiple sub-network modules according to the computational complexity or memory access cost to obtain a custom Transformer layer, and encapsulate the custom Transformer layer using TensorRT to form a plugin.
[0034] The second generation unit is used to replace the Transformer module with a custom Transformer layer to generate a second computation graph;
[0035] The loading unit is used to load the plugin using TensorRT based on the second computation graph and generate an engine file, the engine file including the optimized neural network model corresponding to the second computation graph.
[0036] This application provides a neural network model optimization device, the device comprising: a processor and a memory;
[0037] The memory is used to store instructions;
[0038] The processor is configured to execute the instructions in the memory and perform the method as described in any of the preceding descriptions.
[0039] This application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method as described in any of the preceding descriptions.
[0040] This application provides a method for optimizing a neural network model. The method includes: First, obtaining the structure and parameters of the neural network model and generating a first computational graph. The first computational graph includes a Transformer network module, which includes multiple sub-network modules. Each sub-network module includes multiple operators, i.e., the neural network model is split into multiple parts, each part being an operator. Second, calculating the computational complexity or memory access cost of each operator, and fusing multiple operators within a sub-network module and operators between multiple sub-network modules based on the computational complexity or memory access cost to obtain a custom Transformer layer. TensorRT is then used to encapsulate the custom Transformer layer into a plugin. That is, the computational complexity or memory access cost reflects the computational resources consumed by each operator. Based on the computational resources consumed by each operator, multiple operators can be easily fused, thereby reducing the computational resources consumed by multiple operators and thus reducing the overall computational resources of the neural network model. Finally, the Transformer module is replaced with a custom Transformer layer to generate a second computation graph. Then, based on the second computation graph, the TensorRT plugin is used to generate an engine file. The engine file includes the optimized neural network model corresponding to the second computation graph. The optimized neural network model consumes fewer computational resources without losing accuracy, so that the TensorRT-based neural network model can have high accuracy and high running efficiency when deployed on mobile devices. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 A flowchart illustrating a neural network model optimization method provided in this application embodiment;
[0043] Figure 2 A schematic diagram of a first computational graph provided in an embodiment of this application;
[0044] Figure 3 A schematic diagram of a multi-head self-attention layer provided in an embodiment of this application;
[0045] Figure 4 A schematic diagram of a serial relationship provided for an embodiment of this application;
[0046] Figure 5A schematic diagram of a Y-shaped relationship provided for an embodiment of this application;
[0047] Figure 6 A schematic diagram of a complex relationship provided for an embodiment of this application;
[0048] Figure 7 This application provides a schematic diagram of operator fusion for a multi-head self-attention layer.
[0049] Figure 8 A schematic diagram of operator fusion for a normalization layer provided in an embodiment of this application;
[0050] Figure 9 A schematic diagram of operator fusion in a feedforward network layer provided in an embodiment of this application;
[0051] Figure 10 A schematic diagram of operator fusion of a multi-head self-attention layer and a normalization layer provided in an embodiment of this application;
[0052] Figure 11 This application provides another schematic diagram of operator fusion for a multi-head self-attention layer.
[0053] Figure 12 This application provides another schematic diagram of operator fusion for a multi-head self-attention layer.
[0054] Figure 13 This is yet another schematic diagram of streaming parallelism provided in the embodiments of this application;
[0055] Figure 14 A schematic diagram of a second computational graph provided in an embodiment of this application;
[0056] Figure 15 This is a structural block diagram of a neural network model optimization device provided in an embodiment of this application. Detailed Implementation
[0057] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present application.
[0058] Currently, neural network models, including Transformers, are increasingly being introduced into the field of autonomous driving to achieve various perception tasks. When deploying neural network models on mobile devices, the relevant technology involves transforming and optimizing the trained neural network model to generate the runtime representation for the inference engine. This method relies entirely on the inference engine's own technical capabilities in optimizing and fusing the network structure of the neural network model. Consequently, issues such as splitting network layers, inefficient inter-layer fusion, and misalignment of layer implementation results can arise. These problems lead to reduced forward inference speed or decreased inference accuracy, resulting in lower efficiency and accuracy of the neural network model after deployment on mobile devices.
[0059] The Transformer structure contains a large number of element-wise operations, requiring element-by-element operations and reduction operations. In most mainstream inference engines, the Transformer structure is not treated as an independent layer, but rather composed of numerous basic operation layers according to its structure. Even TensorRT, a software development kit that utilizes widely used high-performance deep learning model inference and has optimized the basic operation layers commonly used in neural networks to the extreme, still suffers from the above-mentioned problems, resulting in inefficient operation of neural network models, including those with the Transformer structure.
[0060] Based on this, this application provides a method for optimizing a neural network model. The method includes: First, obtaining the structure and parameters of the neural network model and generating a first computation graph. The first computation graph includes a Transformer network module, which includes multiple sub-network modules. Each sub-network module includes multiple operators, i.e., the neural network model is split into multiple parts, each part being an operator. Second, calculating the computational complexity or memory access cost of each operator, and fusing multiple operators within a sub-network module and operators between multiple sub-network modules based on the computational complexity or memory access cost to obtain a custom Transformer layer. The custom Transformer layer is then encapsulated to form a plugin. That is, the computational complexity or memory access cost can reflect the computational resources consumed by each operator. At this point, multiple operators can be easily fused based on the computational resources consumed by each operator, thereby reducing the computational resources consumed by multiple operators and thus reducing the overall computational resources of the neural network model. Finally, the Transformer module is replaced with a custom Transformer layer to generate a second computation graph. Then, based on the second computation graph, the plugin is loaded to generate an engine file. The engine file includes the optimized neural network model corresponding to the second computation graph. The optimized neural network model consumes fewer computational resources without losing accuracy, enabling the neural network model to have high accuracy and high running efficiency when deployed on mobile devices.
[0061] To better understand the technical solution and effects of this application, the specific embodiments will be described in detail below with reference to the accompanying drawings.
[0062] See Figure 1 The figure is a flowchart of a neural network model optimization method provided in an embodiment of this application.
[0063] The neural network model optimization method provided in this embodiment includes the following steps:
[0064] S101: Obtain the structure and parameters of the neural network model, and generate the first computational graph.
[0065] In the embodiments of this application, the structure and parameters of a neural network model can be obtained, and a first computational graph can be generated based on the structure and parameters of the neural network model. The first computational graph can intuitively display the structure and computational process of the neural network model.
[0066] As an example, you can obtain the PyTorch training code of the neural network model, instantiate the training code to get the structure of the neural network model, read the corresponding floating-point model parameters, and then use the torch.onnx.export() interface to convert the neural network model into an ONNX file, i.e., the first computation graph.
[0067] Specifically, the first computation graph may include a Transformer network module, wherein the Transformer module includes multiple sub-network modules, and each sub-network module includes multiple operators. An operator refers to a data processing module in the sub-network module, that is, an operator can implement a separate data processing function, such as multiplication, addition or transpose.
[0068] refer to Figure 2 The diagram shown is a schematic of a first computational graph. This first computational graph mainly includes a Transformer module. The Transformer module includes N encoder layers and N decoder layers, where N is a natural number greater than 0. The encoder layers include a Multi-Head Self Attention (MHA) layer, a Forward Feedback Network (FFN) layer, and a LayerNorm (LN) layer. The decoder layers include an MHA layer, an Encoder-Decoder Attention layer, an FFN layer, and an LN layer; that is, the sub-network modules are MHA layers, FFN layers, or LN layers.
[0069] As an example, the first computational graph contains 6 encoding layers and 6 decoding layers. Each encoding layer contains: one MHA layer, one FFN layer, and two LN layers. Each decoding layer contains: one MHA layer, one Encoder-DecoderAttention layer, one FFN layer, and three LN layers.
[0070] In practical applications, the Encoder-Decoder Attention layer has a similar structure to the MHA layer. When the three input vectors of the Encoder-Decoder Attention layer are the same, the MHA layer is a special case of the Encoder-Decoder Attention layer.
[0071] refer to Figure 3The diagram illustrates a multi-head self-attention layer. This multi-head self-attention layer also includes a residual calculation part. The multi-head self-attention layer can receive a key vector k, a value vector v, and a query vector q. The multi-head self-attention layer includes 6 matrix multiplications (MatMul), 5 additions (Add), 4 transposes (Transpose), 4 reshapes (Reshape), 1 partition (Div), and 1 loss function (SoftMax). Each data processing corresponds to an operator, and each operator has a separate kernel function. Kernel function entry and global memory access account for a large portion of the neural network model's data computation time; kernel function fusion can effectively reduce this time consumption.
[0072] S102, calculate the computational complexity or memory access cost of each operator, and fuse multiple operators within a sub-network module and operators between multiple sub-network modules based on the computational complexity or memory access cost to obtain a custom Transformer layer. Use TensorRT to encapsulate the custom Transformer layer to form a plugin.
[0073] In the embodiments of this application, considering the neural network model, that is, the first computation graph has a large number of element-wise operations, such as activation functions (ReLU), tensor addition (Add), and residual structure addition (AddBias), element-wise operations are operations that use operators to process data.
[0074] The Element Wise operation can use the number of floating-point operations (FLOPs) and memory access cost (MAC) to reflect computational complexity and memory usage, respectively. The memory access cost can also be referred to as memory access cost. In other words, computational complexity and memory access cost can be used to quantize the computation of each operator, thereby quantifying the neural network model and facilitating its optimization.
[0075] Therefore, by calculating the computational complexity or memory access cost of each operator, it is possible to fuse multiple operators within a sub-network module and operators between multiple sub-network modules based on these computational complexity or memory access costs. In other words, computational complexity or memory access cost can quantify the computational resources consumed by each operator. Based on the computational resources consumed by each operator, a lower-cost model structure can be easily determined. Specifically, this lower-cost model structure is obtained by fusing multiple operators, thereby reducing the computational resources consumed by individual operators and ultimately reducing the overall computational resources of the neural network model. Multiple operator fusion refers to merging the kernel functions of multiple operators, thereby enabling the use of a single kernel function to perform two data processing operations, reducing data input and output, and thus lowering computational complexity and memory access costs.
[0076] Since multiple operators may belong to the same sub-network module or different sub-network modules, there are some differences in how to fuse operators based on computational complexity or memory access cost. Therefore, they will be introduced separately below:
[0077] For operators within the same sub-network module, the data computation relationships between multiple operators can be obtained. Based on these relationships, the operators can be categorized to obtain multiple dependencies. These dependencies include the multiple operators and their corresponding data computation relationships. Data computation relationships refer to the relationships in which data flows through multiple operators for data processing.
[0078] As an example, we can obtain multiple operators included in each sub-network module, confirm the input-output relationships, operator dimensions, and neighboring operator connections of these operators, thereby obtaining the data computation relationships between them. Based on these data computation relationships, we can classify the operators to obtain multiple dependencies. These dependencies can be sequential, Y-shaped, or complex. (See reference...) Figures 4-6 As shown.
[0079] refer to Figure 4 As shown in (1)-(4), in a serial relation, multiple operators are connected in series, meaning the data streams between multiple operators are serial, and the output of the previous operator is the input of the next operator. For example, the output of MatMul is the input of Add, and the output of Add is the input of ReLU. Similarly, the output of Div is the input of SoftMax. And again, the output of MatMul is the input of Transpose.
[0080] refer to Figure 5As shown in (1)-(2), in a Y-type relation, the outputs of multiple operators may be the inputs of the same operator, and the structure formed by multiple operators is similar to a Y-type structure. For example, the outputs of op1 and op2 are the inputs of Add. Similarly, the outputs of op1 and op2 are the inputs of MatMul.
[0081] refer to Figure 6 As shown, complex relationships are those other than serial and Y-shaped relationships, such as relationships involving different data processing branches. For example, the output of op1 is the input of op2 and op5, the output of op2 is the input of op3, the output of op3 is the input of op4, and the output of op4 is the input of op5.
[0082] Since multiple operators may have different dependencies, the specific operators fused when fusing multiple operators within a sub-network module based on computational complexity or memory access cost will also differ. Multiple operators can be fused according to their dependencies. This is done by calculating the computational complexity or memory access cost of each operator, as well as the computational complexity or memory access cost after fusing multiple operators based on dependencies. The operators within the sub-network module are fused based on the lowest computational complexity or memory access cost. That is, if the computational complexity or memory access cost calculated after fusing multiple operators based on dependencies is lower, then the operators are fused according to the dependency relationship; if the computational complexity or memory access cost calculated after fusing multiple operators based on dependencies is higher, then the multiple operators are not fused.
[0083] In specific fusion operations, as long as either computational complexity or memory access cost decreases, multiple operators can be fused based on their dependencies.
[0084] One possible implementation involves calculating the first computational complexity and first memory access cost for each operator, and then calculating the second computational complexity and second memory access cost after fusing multiple operators categorized into serial, Y-shaped, and complex relationships. The multiple operators within a sub-network module are then fused based on the lowest of the following: the first computational complexity is greater than the second computational complexity, or the first memory access cost is greater than the second memory access cost. In other words, if the first computational complexity is greater than the second computational complexity, or the first memory access cost is greater than the second memory access cost, then multiple operators can be directly fused based on the serial, Y-shaped, or complex relationships.
[0085] As an example, see reference Figure 7 The diagram shows an operator fusion scheme for a multi-head self-attention layer. Taking the MatMul and Add operators as examples, assume that the batch of the input matrix is B, and the number of rows and columns is equal, both being M.
[0086] The computational complexity of the MatMul operator is referenced in formula (1):
[0087] FLOPs(MatMul) = B*M 3
[0088] Among them, the matrix multiplication of 1 batch includes M 3 The multiplication-addition operation, B batch matrix multiplication contains B×M 3 This involves multiple multiply-accumulate operations.
[0089] The computational complexity of the Add operator is referenced in formula (2):
[0090] FLOPs(Add) = B*M 2
[0091] Among them, the matrix addition of 1 batch contains M 2 The addition operation, B batch matrix addition, contains B×M. 2 The next addition operation.
[0092] The memory access cost of the MatMul operator is referenced in formula (3):
[0093] MAC(MatMul) = 3 * B * M 2 *sizeof(T)
[0094] Where T represents the type of input data. A 1-batch matrix multiplication consists of 3 M... 2 A memory access of sizeof(T) bytes is required, and matrix multiplication in batch B contains 3×B×M bytes. 2 A memory access of ×sizeof(T) bytes is required, meaning a global memory access of 3×B×M is needed. 2 Second-rate.
[0095] The memory access cost of the Add operator is referenced in formula (4):
[0096] MAC(Add)=(2M+1)*B*M*sizeof(T)
[0097] Where T represents the type of input data. Matrix addition in 1 batch involves (2×M+1)×M×sizeof(T) bytes of memory access, while matrix addition in B batch involves B×(2×M+1)×M×sizeof(T) bytes of memory access, meaning it requires accessing global memory B×(2×M+1)×M times.
[0098] According to formulas (1) and (2), the total computational complexity for calculating the MatMul operator and the Add operator is (M+1)B×M. 2The total memory access cost is ((2×M+1)+3×M)×B×M×sizeof(T) bytes.
[0099] If the MatMul and Add operators are fused sequentially, that is, the kernel functions of the MatMul and Add operators are merged, the total memory access cost is reduced by (2 × B × M) because one MatMul output vector write and one Add input vector write are reduced. 2 Furthermore, it reduces the kernel function login time by one time, shortens the computation time, and greatly improves computational efficiency.
[0100] Figure 7 The diagram illustrates operator fusion based on dependencies. Operator fusion in a multi-head self-attention layer includes two types of dependencies: serial and Y-type. The serial dependency reference... Figure 7 As shown in numbers 1 and 3, the Y-type relationship is shown in number 2. That is, numbers 1 and 3 are for fusing multiple operators according to the serial relationship, and number 2 is for fusing multiple operators according to the Y-type relationship. Thus, the multi-head self-attention layer includes 5 fusions based on the serial relationship and 2 fusions based on the Y-type relationship.
[0101] As another example, see Figure 8 The diagram shown illustrates an operator fusion method for a normalization layer. The normalization layer maps the feature vectors output by the multi-head attention layer to a range of 0 to 1. The normalization layer can call the PyTorchtorch.nn.LayerNorm function, which normalizes the optimization space and accelerates convergence. The normalization layer can be calculated using formula (5):
[0102] y i =γ×(x i -μ) / σ+β
[0103] Where γ and β represent the scaling factor and displacement factor, respectively, and μ and σ represent the mean and standard deviation, respectively. i y represents the output vector of the previous layer. i This represents the normalized output vector.
[0104] The normalization layer also includes multiple operators. Complex relationships can be used to fuse these operators, combining their kernel functions into a single kernel function. (See reference...) Figure 8 As shown, the complex relationship is referenced in number 4. That is, number 4 is the fusion of multiple operators based on the complex relationship. Thus, the normalization layer includes one fusion based on the complex relationship.
[0105] As yet another example, see reference Figure 9The diagram illustrates operator fusion in a feedforward network layer. The feedforward network layer contains two linear transformations and an activation function. The linear transformations refer to matrix multiplication and matrix addition. The activation function can be Gelu or Relu; that is, the feedforward network layer also includes multiple operators. These operators can be fused using a sequential relationship, combining their kernel functions into a single kernel function. For example, matrix multiplication, matrix addition, and the activation function can be fused into a single kernel function. (See reference...) Figure 9 As shown, the serial relationship is indicated by reference numbers 1 and 5. That is, number 1 is the fusion of matrix multiplication and matrix addition based on the serial relationship, and number 5 is the fusion of matrix multiplication, matrix addition, and activation function based on the serial relationship.
[0106] Therefore, it can be seen that multiple operators are merged into one based on their data computation relationships, and the functionality of numerous operators is implemented through a kernel function. Within this kernel function, methods such as shared memory, instruction optimization, and assembly optimization are fully utilized to optimize the kernel function's implementation logic, enabling the merged kernel function to achieve optimal performance. The kernel function, as the basic function source code implementing a certain computation process, can run on the graphics processing unit (GPU) of a computer device.
[0107] For operators in different sub-network modules, similar operators with the same structure are identified from multiple operators in different sub-network modules. The computational complexity or memory access cost of each operator is calculated, as well as the computational complexity or memory access cost of fusing similar operators in different sub-network modules. Operators from multiple sub-network modules are then fused based on the lowest possible computational complexity or memory access cost. Similar operators with the same structure refer to operators that share the same data processing procedure.
[0108] As an example, the residual structures included in the MHA layer are similar to those included in the LN layer. Therefore, the operators corresponding to the residual structures included in the MHA layer and the operators corresponding to the residual structures included in the LN layer can be fused. The fused calculation formula is shown in formula (6):
[0109] y i+1 =γ×((x) i +bi)-μ)σ+β
[0110] Where γ and β represent the scaling factor and displacement factor, respectively, and μ and σ represent the mean and standard deviation, respectively. i b represents the output vector of the layer preceding the LN layer. i y represents the output vector of the layer preceding the MHA layer. i+1This represents the normalized output vector. The computational complexity or memory access cost can be calculated using the fused calculation formula, and whether to perform fusion is determined based on whether the computational complexity or memory access cost is lower than that of the unfused computational complexity or memory access cost.
[0111] refer to Figure 10 The diagram shown illustrates the operator fusion of an MHA layer and an LN layer. Figure 10 It can be seen that the operator corresponding to the residual structure included in the MHA layer is the Add operator, and the number 6 indicates that the Add operator and the operator corresponding to the residual structure of the LN layer are fused together.
[0112] Therefore, by integrating the residual addition of the previous layer with the normalization operation of the current layer into a single kernel function, the kernel function's login time and the number of global memory accesses are reduced, thus shortening memory access latency.
[0113] In the embodiments of this application, there may also be multiple irrelevant operators among the multiple operators that have no data calculation relationship. These irrelevant operators can also be merged, that is, the same kernel function is used to process the irrelevant operators, thereby greatly increasing the calculation speed and improving the calculation efficiency.
[0114] Independent operators primarily refer to operators whose data flow is independent within or between different sub-network modules. These independent operators can be optimized for parallel computing based on the hardware computing characteristics of the graphics processing unit (GPU). Parallel computing optimization mainly includes two possible implementation methods:
[0115] The first possible implementation is batch fusion, which means that multiple unrelated operators with no data computation relationship can be obtained from multiple operators, that is, multiple operators with no dependency relationship can be obtained from multiple operators, and multiple unrelated operators can be merged into a single operator. In this way, the data processing of multiple unrelated operators can be executed through the kernel function of a single operator, thereby reducing memory access costs and reducing the consumption of computing resources.
[0116] As an example, an irrelevant operator can be an operator with the same structure that processes vectors, such as a vector input to an MHA layer that is processed using multiple matrix multiplications and additions included in the MHA layer. When the input vectors to multiple irrelevant operators are the same, it means that multiple irrelevant operators perform the same data processing on the same vector, and therefore multiple irrelevant operators can be merged into a single operator.
[0117] refer to Figure 11As shown, the MHA layer contains three input vectors: key vector k, value vector v, and query vector q. When k, q, and v are the same, the three MatMul operators and three Add operators of the MHA layer corresponding to k, q, and v are merged into the kernel function of one operator and run.
[0118] refer to Figure 12 As shown, when k and v are the same input, the two MatMul operators and two Add operators in the MHA layer corresponding to k and v are merged into the kernel function of one operator and run.
[0119] The second possible implementation is stream parallelism. Since multiple independent operators with no data computational relationship may be assigned to the same stream for execution, the model's computational efficiency will be low. In this case, multiple independent operators with no data computational relationship can be assigned to different streams for parallel processing, which can greatly shorten the computation time and improve computational efficiency.
[0120] As an example, the input to the encoder-decoder attention layer in the Tansformer module comes from the multi-head self-attention layers of the Nth encoder layer and the first decoder layer. If the same stream is used to complete the forward inference of the neural network model, according to the processor scheduling rules, the output of the Nth encoder layer will be calculated before the computation related to the first decoder layer. By controlling the computation of the multi-head self-attention layers of the Nth encoder layer and the first decoder layer separately through two streams, the decoder layer does not have to wait for the encoder layer to finish its computation before starting decoding, which reduces the waiting time of the computation unit and makes full use of the stream processor and memory resources on the hardware. (Reference) Figure 13 As shown, the Nth encoding layer is executed using stream1, and the first decoding layer is executed using stream2, that is, the computation of the encoding and decoding layers is processed in parallel using multiple streams.
[0121] Therefore, by using batch fusion and stream parallelism, operators with no data flow dependencies are executed concurrently, minimizing the serial execution of control flow and reducing the time spent on model forward inference. This also involves the rational scheduling of hardware resources, making full use of the hardware's stream processors and memory bandwidth.
[0122] As can be seen from the above description, the embodiments of this application significantly simplify the model computation graph and reduce the amount of network computation by fusing multiple operators within a sub-network module, fusing operators between multiple sub-network modules, and fusing or streaming parallel processing multiple unrelated operators.
[0123] In the embodiments of this application, a custom Transformer layer is obtained by fusing multiple operators in the Transformer module. The attributes and weight parameters of the custom Transformer layer are derived from the weight file of the neural network model corresponding to the Transformer module.
[0124] Since a custom Transformer layer is a custom network layer, it is not supported by the TensorRT runtime library. Therefore, it needs to be encapsulated and registered with the TensorRT framework. The custom Transformer layer will be called by the TensorRT runtime library as a dynamic link library, existing as a plugin for the TensorRT runtime.
[0125] Specifically, the TensorRT tool is used to encapsulate the implementation code of the custom Transformer layer, register the custom Transformer layer into the TensorRT framework, and generate a dynamic link library adapted to the hardware.
[0126] S103, replace the Transformer module with a custom Transformer layer to generate a second computation graph.
[0127] In the embodiments of this application, during the process of fusing multiple operators in the Transformer module to form a custom Transformer layer, not only is the implementation code of the custom Transformer layer formed, but also the structure of the custom Transformer layer is formed. At this time, the Transformer module can be replaced with the custom Transformer layer to generate a second computation graph.
[0128] As an example, see reference Figure 14 As shown, the second computation graph can intuitively display the structure and computation process of the neural network model after fusing multiple operators.
[0129] S104, based on the second computation graph, uses TensorRT to load plugins and generate engine files.
[0130] In the embodiments of this application, based on the second computation graph, a plugin is loaded, that is, a dynamic link library of a custom Transformer layer is loaded to construct the runtime expression of the TensorRT inference engine, that is, to generate a binary engine file, which includes the optimized neural network model corresponding to the second computation graph.
[0131] Specifically, the second computation graph and dynamic link library can be copied to the domain controller. The TensorRT runtime library is installed by default; use trtexec to load the second computation graph and dynamic link library to generate the engine files.
[0132] Therefore, the neural network model optimization method provided in this application can significantly simplify the model computation graph and reduce the network computation load. By using an operator fusion strategy for different sub-network modules, the forward inference computation load of the model is further reduced, lowering the memory usage of the chip in the mobile device. Without reducing the detection accuracy of the first computation graph, the independent data streams are computed in parallel, reducing the forward inference time and meeting the frame rate requirements of mobile device deployments.
[0133] Based on the neural network model optimization method provided in the above embodiments, this application also provides a neural network model optimization device, the working principle of which will be described in detail below with reference to the accompanying drawings.
[0134] See Figure 15 The figure is a structural block diagram of a neural network model optimization device provided in an embodiment of this application.
[0135] The neural network model optimization device 200 provided in this embodiment includes:
[0136] The first generation unit 210 is used to obtain the structure and model parameters of the neural network model and generate a first computation graph. The first computation graph includes a Transformer network module, and the Transformer module includes multiple sub-network modules. Each sub-network module includes multiple operators.
[0137] The fusion unit 220 is used to calculate the computational complexity or memory access cost of each operator, fuse multiple operators within the sub-network module and fuse operators between multiple sub-network modules according to the computational complexity or memory access cost to obtain a custom Transformer layer, and encapsulate the custom Transformer layer using TensorRT to form a plugin.
[0138] The second generation unit 230 is used to replace the Transformer module with a custom Transformer layer to generate a second computation graph;
[0139] The loading unit 240 is used to load the plugin using TensorRT based on the second computation graph and generate an engine file, the engine file including the optimized neural network model corresponding to the second computation graph.
[0140] Optionally, the method further includes:
[0141] Obtain the data calculation relationship between multiple operators, classify the multiple operators according to the data calculation relationship, and obtain multiple dependency relationships;
[0142] The calculation of the computational complexity or memory access cost of each operator, the fusion of multiple operators within the sub-network module based on the computational complexity or memory access cost, and the fusion of operators between multiple sub-network modules include:
[0143] Calculate the computational complexity or memory access cost of each operator and the computational complexity or memory access cost of fusing multiple operators according to the dependencies, and fuse multiple operators within the sub-network module based on the lowest computational complexity or memory access cost.
[0144] Optionally, the dependencies include serial relationships, Y-type relationships, and complex relationships;
[0145] The calculation of the computational complexity or memory access cost of each operator, and the calculation of the computational complexity or memory access cost after fusing multiple operators according to the dependency relationship, and the fusion of multiple operators within the sub-network module based on the lowest computational complexity or memory access cost, as well as the fusion of operators between multiple sub-network modules, include:
[0146] Calculate the first computational complexity and the first memory access cost for each of the operators;
[0147] The second computational complexity and the second memory access cost of the multiple operators categorized as serial relationships, Y-type relationships, and complex relationships are calculated and fused. The multiple operators within the sub-network module are fused based on the lowest of either the first computational complexity being greater than the second computational complexity or the first memory access cost being greater than the second memory access cost.
[0148] Optionally, the method further includes:
[0149] Obtain similar operators with the same structure from multiple operators in different sub-network modules;
[0150] The calculation of the computational complexity or memory access cost of each operator, the fusion of multiple operators within the sub-network module based on the computational complexity or memory access cost, and the fusion of operators between multiple sub-network modules include:
[0151] The computational complexity or memory access cost of each operator is calculated, as well as the computational complexity or memory access cost of similar operators fused in different sub-network modules. Operators among multiple sub-network modules are fused based on the lowest computational complexity or memory access cost.
[0152] Optionally, the method further includes:
[0153] Obtain multiple independent operators among the multiple operators that have no data computational relationship;
[0154] Multiple irrelevant operators are merged into a single operator.
[0155] Optionally, the irrelevance operator is an operator that processes vectors;
[0156] The step of fusing multiple irrelevant operators into a single operator includes:
[0157] When the vectors input to multiple independent operators are the same, the multiple independent operators are merged into a single operator.
[0158] Optionally, the Transformer module includes N encoding layers and N decoding layers; the encoding layer includes a multi-head self-attention (MHA) layer, a feedforward network (FFN) layer, and a normalized low-level network (LN) layer; the decoding layer includes an MHA layer, an encoder-decoder attention layer, an FFN layer, and an LN layer; and the sub-network module is an MHA layer, an FFN layer, or an LN layer.
[0159] The method further includes:
[0160] The encoding layer and the decoding layer are processed in parallel using multiple streams.
[0161] Based on the neural network model optimization method provided in the above embodiments, this application also provides a neural network model optimization device, which includes:
[0162] The processor and memory may be present, and the number of processors may be one or more. In some embodiments of this application, the processor and memory may be connected via a bus or other means.
[0163] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory may also include NVRAM. Memory stores the operating system and operating instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. The operating instructions may include a variety of operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.
[0164] The processor controls the operation of the terminal device; the processor can also be called the CPU.
[0165] The methods disclosed in the embodiments of this application can be applied to a processor or implemented by a processor. A processor can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as execution by a hardware decoding processor, or as a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0166] This application also provides a computer-readable storage medium for storing program code that is used to execute any of the methods in the foregoing embodiments.
[0167] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0168] It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0169] When describing elements of various embodiments of this application, the articles “a,” “an,” “this,” and “described” are all intended to indicate that there are one or more elements. The words “comprising,” “including,” and “having” are inclusive and mean that there may be other elements in addition to those listed.
[0170] It should be noted that those skilled in the art will understand that all or part of the processes in the above method embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0171] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0172] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The units and modules described as separate components may or may not be physically separate. Furthermore, some or all of the units and modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0173] The above description is merely a preferred embodiment of this application. Although this application has disclosed preferred embodiments above, it is not intended to limit this application. Any person skilled in the art can make many possible variations and modifications to the technical solutions of this application using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of this application. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of this application without departing from the content of the technical solutions of this application shall still fall within the protection scope of the technical solutions of this application.
Claims
1. A method for optimizing a neural network model, characterized in that, The method includes: The structure and parameters of the neural network model are obtained, and a first computational graph is generated. The first computational graph includes a Transformer network module, and the Transformer module includes multiple sub-network modules, each of which includes multiple operators. Calculate the computational complexity or memory access cost of each operator, fuse multiple operators within the sub-network module and fuse operators between multiple sub-network modules based on the computational complexity or memory access cost to obtain a custom Transformer layer, and encapsulate the custom Transformer layer using TensorRT to form a plugin; Replace the Transformer module with a custom Transformer layer to generate a second computation graph; Based on the second computation graph, the plugin is loaded using TensorRT to generate an engine file, which includes the optimized neural network model corresponding to the second computation graph.
2. The method according to claim 1, characterized in that, The method further includes: Obtain the data calculation relationship between multiple operators, classify the multiple operators according to the data calculation relationship, and obtain multiple dependency relationships; The calculation of the computational complexity or memory access cost of each operator, the fusion of multiple operators within the sub-network module based on the computational complexity or memory access cost, and the fusion of operators between multiple sub-network modules include: Calculate the computational complexity or memory access cost of each operator and the computational complexity or memory access cost of fusing multiple operators according to the dependencies, and fuse multiple operators within the sub-network module based on the lowest computational complexity or memory access cost.
3. The method according to claim 2, characterized in that, The dependencies include serial relationships, Y-type relationships, and complex relationships; The calculation of the computational complexity or memory access cost of each operator, and the calculation of the computational complexity or memory access cost after fusing multiple operators according to the dependency relationship, and the fusion of multiple operators within the sub-network module based on the lowest computational complexity or memory access cost, as well as the fusion of operators between multiple sub-network modules, include: Calculate the first computational complexity and the first memory access cost for each of the operators; The second computational complexity and the second memory access cost of the multiple operators categorized as serial relationships, Y-type relationships, and complex relationships are calculated and fused. The multiple operators within the sub-network module are fused based on the lowest of either the first computational complexity being greater than the second computational complexity or the first memory access cost being greater than the second memory access cost.
4. The method according to claim 1, characterized in that, The method further includes: Obtain similar operators with the same structure from multiple operators in different sub-network modules; The calculation of the computational complexity or memory access cost of each operator, the fusion of multiple operators within the sub-network module based on the computational complexity or memory access cost, and the fusion of operators between multiple sub-network modules include: The computational complexity or memory access cost of each operator is calculated, as well as the computational complexity or memory access cost of similar operators fused in different sub-network modules. Operators among multiple sub-network modules are fused based on the lowest computational complexity or memory access cost.
5. The method according to claim 1, characterized in that, The method further includes: Obtain multiple independent operators from among the multiple operators that have no data computational relationship; Multiple irrelevant operators are merged into a single operator.
6. The method according to claim 5, characterized in that, The irrelevant operator is an operator that processes vectors; The step of fusing multiple irrelevant operators into a single operator includes: When the vectors input to multiple independent operators are the same, the multiple independent operators are merged into a single operator.
7. The method according to any one of claims 1-6, characterized in that, The Transformer module includes N encoding layers and N decoding layers; the encoding layer includes a multi-head self-attention (MHA) layer, a feedforward network (FFN) layer, and a normalized low-level network (LN) layer; the decoding layer includes an MHA layer, an encoder-decoder attention layer, an FFN layer, and an LN layer; and the sub-network module is an MHA layer, an FFN layer, or an LN layer. The method further includes: The encoding layer and the decoding layer are processed in parallel using multiple streams.
8. A neural network model optimization device, characterized in that, The device includes: The first generation unit is used to obtain the structure and parameters of the neural network model and generate a first computation graph. The first computation graph includes a Transformer network module, and the Transformer module includes multiple sub-network modules. Each sub-network module includes multiple operators. A fusion unit is used to calculate the computational complexity or memory access cost of each operator, fuse multiple operators within the sub-network module and fuse operators between multiple sub-network modules according to the computational complexity or memory access cost to obtain a custom Transformer layer, and encapsulate the custom Transformer layer using TensorRT to form a plugin. The second generation unit is used to replace the Transformer module with a custom Transformer layer to generate a second computation graph; The loading unit is used to load the plugin using TensorRT based on the second computation graph and generate an engine file, the engine file including the optimized neural network model corresponding to the second computation graph.
9. A neural network model optimization device, characterized in that, The device includes: a processor and a memory; The memory is used to store instructions; The processor is configured to execute the instructions in the memory and perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-7.