An intelligent model optimization method based on operator target value regression prediction

By performing operator fusion and platform matching optimization on the intelligent model, the storage and computing cost problems caused by excessive amounts of intelligent model parameters are solved, the model scale is reduced and the computing performance is improved, and the real-time needs are met.

CN119918590BActive Publication Date: 2025-07-01北京麟卓信息科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510405296.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-01
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

As the complexity of intelligent models increases, the number of model parameters has increased exponentially, resulting in a significant increase in storage demand and computing costs. It is difficult for ordinary mobile devices or embedded systems to accommodate such huge models, and the calculation speed is slow and the energy consumption is increased dramatically, which cannot meet the real-time needs.

Method used

By constructing an operator structure diagram, the optimization intelligent model is fused based on graph analysis, and the input shape and dimensional order with the highest matching degree among each operator is selected based on the constructed operator platform matching degree model, forming the first optimization model, and then dimensional rearrangement and parallelization are performed to complete the optimization of the intelligent model.

Benefits of technology

It effectively reduces the scale of the model, reduces storage requirements and computing costs, improves computing speed and energy efficiency, and meets real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918590B_ABST
    Figure CN119918590B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent model optimization method based on operator target value regression prediction. By constructing an operator structure diagram, the intelligent model to be optimized is fused based on graph analysis. Then, based on the constructed operator platform matching degree model, the input shape and dimension order of the input data with the highest matching degree in each operator are selected according to the type, number of parameters, input dimension set, and input shape set of the operator to form a first optimization model. Then, dimension rearrangement and parallelization processing are performed on the first optimization model to complete the optimization of the intelligent model, effectively reducing the scale of the model and lowering the storage requirements and computing costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer software development, and particularly relates to an intelligent model optimization method based on operator target value regression prediction. Background Art

[0002] As the complexity of intelligent models continues to increase, the number of their parameters grows exponentially. Taking large-scale deep neural networks as an example, although the increase in the number of layers and neurons improves the model performance, it also makes the volume of the model file extremely large. In the field of image recognition, some advanced models have up to billions of parameters, which poses strict requirements on storage devices, and it is difficult for ordinary mobile devices or embedded systems to accommodate such huge models. From a computational perspective, during the inference and training processes, a large number of parameter operations consume a huge amount of computing resources, resulting in slow computing speed and a sharp increase in energy consumption. For example, in a real-time video analysis scenario, the computational delay of a huge model will seriously affect the system response speed and cannot meet the real-time requirements. To alleviate these problems, model pruning technology has emerged, which aims to remove redundant connections and parameters in the model, reduce the model size without significantly reducing the model performance, thereby reducing the storage requirements and computational costs. Summary of the Invention

[0003] In view of this, the present invention provides an intelligent model optimization method based on operator target value regression prediction, which realizes the re-optimization of the fused model by evaluating the target value of the model.

[0004] An intelligent model optimization method based on operator target value regression prediction provided by the present invention specifically includes the following steps:

[0005] Step 1: Set a number extraction branch route for the operators in the intelligent model to be optimized, obtain the combinable operator combinations in each branch route, replace the operator combination with a fused operator, and complete the fusion of the intelligent model to be optimized to obtain an intermediate model;

[0006] Step 2: Use the operator type, number of parameters, input dimension set, and input shape set of each operator in the intermediate model as input data, and use the operator platform matching degree model to calculate the input data to obtain output data. The output data includes operator execution time, video memory usage, video memory bandwidth utilization rate, shared memory usage, and register usage; according to the output data, screen out the input shape and dimension order with the highest matching degree in each operator to form a first optimized model;

[0007] Step 3: Traverse each operator in the first optimized model, judge the consistency of the input shapes between it and the adjacent operators, and perform dimension rearrangement between the adjacent operators for inconsistent situations;

[0008] Step 4: The operators located at the branch nodes in the first optimization model form a set of branch operators. In the set of branch operators, for multiple branch operators with the same parent node operator, the same operator type, and the sum of the occupied video memory usage not exceeding the platform video memory, mark the corresponding branch routes as routes that can be executed in parallel, and complete the optimization of the intelligent model to be optimized.

[0009] Further, the method of obtaining the combinable operator combinations in each branch route in Step 1 and replacing the operator combinations with fused operators is as follows:

[0010] Step 1.1: Traverse the intelligent model to be optimized to obtain all nodes, determine the node types and the number of branches of the branch nodes, and establish a first node set for saving all nodes.

[0011] Step 1.2: Create a start queue and an end queue. The start queue is used to search from the starting node of the intelligent model to be optimized, and add the starting node to the start queue; the end queue is used to search from the termination node of the intelligent model to be optimized, and add the termination node to the end queue; expand layer by layer from the starting node and the termination node to the current node respectively.

[0012] Step 1.3: Check whether the current nodes in the two queues are branch nodes. If they are branch nodes, execute Step 1.4; otherwise, expand forward to the adjacent nodes of the current node, and use the adjacent nodes as the current node to execute Step 1.3.

[0013] Step 1.4: Starting from the current node, trace back to the adjacent branch node in its queue, which is recorded as the first branch node, and form a branch route by the nodes between the current node and the first branch node; delete the nodes in the branch route from the first node set, and subtract 1 from the number of branches of the first branch node. When the number of branches is 0, delete it from the first node set; if the current node is the starting node or the termination node, end the search; otherwise, expand forward to the adjacent nodes of the current node, and use the adjacent nodes as the current node to execute Step 1.3.

[0014] Step 1.5: Initialize the candidate matrix. The rows of the candidate matrix are the nodes of the fused target graph, and the columns are the nodes of the initial source graph of the intelligent model to be optimized; if the nodes of the target graph can match the nodes of the initial source graph, set the value of the corresponding point in the candidate matrix to 1, otherwise 0.

[0015] Step 1.6: According to the constraints of the adjacency matrix, recursively match each node in the target graph with the nodes in the initial source graph. When there is a complete match, eliminate and replace the matched nodes and edges in the initial source graph with the operator combination number, and retain the fused operator information until no matching items are returned and no further fusion operations can be performed.

[0016] Further, the operator platform matching degree model in step 2 is a multi-task learning model.

[0017] Further, the operator platform matching degree model includes an operator execution time prediction sub-model and a target value prediction sub-model. The input of both sub-models is the input data. The operator execution time prediction sub-model predicts the execution time of the operator by evaluating the importance of operator features, and the target value prediction sub-model is used to predict the video memory usage, video memory bandwidth utilization, shared memory usage, and register usage when the operator runs.

[0018] Further, the operator execution time prediction sub-model is a model constructed based on a random forest regression model.

[0019] Further, the target value prediction sub-model is a multi-output regression model constructed based on a baseline model.

[0020] Further, the calculation method of the matching degree is as follows: weights are assigned to the operator execution time, video memory usage, video memory bandwidth utilization, shared memory usage, and register usage respectively, and then the weighted average value is calculated according to the specific values, which is the matching degree.

[0021] Further, the determination method of the fused operator in step 1 is as follows: the operators that can be fused are represented by formulas, and then the formulas of the fused operators are formed by combining the operators according to the formulas.

[0022] Further, the implementation method of obtaining the combinable operator combinations in each branch route in step 1 is as follows: operators of the same type, convolutional operators and batch normalization operators, arithmetic operation operators and convolutional operators, arithmetic operation operators and batch normalization operators, and convolutional operators and activation function operators are all combinable operator combinations.

[0023] Further, the method of setting numbers for the operators in the intelligent model to be optimized in step 1 is as follows: the number of the operator consists of a type number and an operator node serial number. The type number is the number of the operator type of the determined type, and the operator node serial number is set in the order of the operators from input to output and from right to left; for the operators that do not belong to the determined type, the type numbers are set for them in the order of the operators from input to output and from right to left. Beneficial effects

[0024] The present invention constructs an operator structure diagram, performs a fusion process on the intelligent model to be optimized based on graph analysis, and then, based on the constructed operator platform matching degree model, screens out the input shape and dimension order of the input data with the highest matching degree among each operator according to the type of operator, the number of parameters, the input dimension set, and the input shape set to form a first optimized model. Then, dimension rearrangement and parallelization processing are performed on the first optimized model to complete the optimization of the intelligent model, effectively reducing the scale of the model, the storage requirement, and the calculation cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a schematic processing flow diagram of an intelligent model optimization method based on operator target value regression prediction provided by the present invention.

[0026] Figure 2 It is an operator structure diagram of the intelligent model to be optimized constructed by using an intelligent model optimization method based on operator target value regression prediction provided by the present invention.

[0027] Figure 3 It is an operator structure diagram after fusion by using an intelligent model optimization method based on operator target value regression prediction provided by the present invention.

[0028] Figure 4 It is an operator structure diagram after dimension rearrangement by using an intelligent model optimization method based on operator target value regression prediction provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The following examples are listed in conjunction with the accompanying drawings to describe the present invention in detail.

[0030] An intelligent model optimization method based on operator target value regression prediction provided by the present invention has a core idea that: by constructing an operator structure diagram, performing a fusion process on the intelligent model to be optimized based on graph analysis, and then, based on the constructed operator platform matching degree model, screening out the input shape and dimension order of the input data with the highest matching degree among each operator according to the type of operator, the number of parameters, the input dimension set, and the input shape set to form a first optimized model. Then, dimension rearrangement and parallelization processing are performed on the first optimized model to complete the optimization of the intelligent model.

[0031] An intelligent model optimization method based on operator target value regression prediction provided by the present invention has a processing flow as Figure 1 shown, and specifically includes the following steps:

[0032] Step 1: Set numbers for all operators in the intelligent model to be optimized, extract all branch routes, obtain the combinable operator combinations in each branch route, replace the operator combinations with fusion operators, and complete the fusion of the intelligent model to be optimized to obtain an intermediate model.

[0033] Among them, the method for setting numbers for all operators in the intelligent model to be optimized is as follows: The number of an operator consists of a type number and an operator node sequence number. Specifically, the same type number is set for operators of the same type. For example, convolution operators are all of type A, arithmetic operation operators such as addition and multiplication are all of type B, batch normalization operators are all of type C, activation function operators are all of type D, and so on; for operators that do not belong to a determined type, their type numbers are set in the order from input to output and from right to left. For example, the numbering starts from E; the operator node sequence number is set in the order from input to output and from right to left. For example, for operators at the same position in the direction from input to output, consecutive numbers are set in the order from right to left.

[0034] In the present invention, it is determined that operators of the same type can be fused. Convolution operators and batch normalization operators can be fused, arithmetic operation operators and convolution operators can be fused, arithmetic operation operators and batch normalization operators can be fused, convolution operators and activation function operators can be fused, and so on.

[0035] A branch line is a line that cannot be further decomposed. In the present invention, all branch routes in the intelligent model to be optimized are extracted, the combinable operator combinations in each branch route are obtained, and the operator combination is replaced with a fused operator. The specific method is as follows:

[0036] Step 1.1: Traverse the intelligent model to be optimized to obtain all nodes therein, determine the types of the nodes, that is, branch nodes and non-branch nodes, and determine the number of branches of the branch nodes. Among them, a branch node is a parent node with multiple child nodes, and the number of branches is the number of child nodes of the parent node; establish a first node set for storing all nodes.

[0037] Step 1.2: Create a start queue and an end queue. The start queue is used to search starting from the starting node of the intelligent model to be optimized, and add the starting node to the start queue; the end queue is used to search starting from the termination node of the intelligent model to be optimized, and add the termination node to the end queue; expand layer by layer from the starting node and the termination node to the current node respectively.

[0038] Step 1.3: Check whether the current nodes in the two queues are branch nodes. If they are branch nodes, suspend the expansion and execute Step 1.4; otherwise, expand forward to the adjacent nodes of the current node, and use the adjacent nodes as the current nodes to execute Step 1.3.

[0039] Step 1.4: Starting from the current node, trace back to the adjacent branch node in the queue where it is located, denoted as the first branch node. The nodes between the current node and the first branch node form a branch route, and this branch route is saved into the branch route set. Delete the nodes in the branch route from the first node set, and decrement the branch count of the first branch node by 1. When the branch count is 0, delete it from the first node set. If the current node is the starting node or the ending node, end the search; otherwise, expand forward to the adjacent nodes of the current node, and use the adjacent nodes as the current node to execute Step 1.3.

[0040] Construct an adjacency matrix according to the branch route to represent the connection relationships between the nodes in the graph. Among them, the branch route is represented as: a node sequence composed of all the nodes on the branch route.

[0041] The adjacency matrix is a data structure used in graph theory to represent a graph. In the matrix, the rows and columns represent the operators in the model (such as convolutional layers, fully connected layers, etc.), and the values in the matrix represent the connection relationships between the operators (such as whether there is an input-output relationship, or the weight of the relationship).

[0042] For example, for the adjacency matrix of operators A, B, and C, the first row and the first column of the matrix represent operator A, the second row and the second column represent operator B, the third row and the third column represent operator C, and the elements in the matrix represent the input-output relationships between two operators. The specific values may be as follows:

[0043] [[0, 1, 0],

[0044] [0, 0, 1],

[0045] [0, 0, 0]]

[0046] Among them, [0, 1, 0] means that the output of operator A is the input of operator B, and [0, 0, 1] means that the output of operator B is the input of operator C.

[0047] Step 1.5: Initialize the candidate matrix. The rows of the candidate matrix correspond to the nodes of the fused target graph, and the columns correspond to the nodes of the initial source graph of the intelligent model to be optimized. If a node in the target graph can match a node in the initial source graph, the value at the corresponding point in the candidate matrix is set to 1; otherwise, it is 0.

[0048] Step 1.6: According to the constraints of the adjacency matrix, for each node in the target graph, use the backtracking method to recursively match the nodes in the initial source graph. When there is a complete match (continuous operator matching is successful), replace the matched several nodes and edges in the initial source graph with the operator combination number, and retain the fused operator information until no match is returned and no further fusion operation can be performed.

[0049] Step 1.7: At the numbered points of the operator combination, obtain the fused operator information and perform the fusion operation to obtain the fused operator. Replace the nodes in the operator combination with the fused operator to obtain a new branch route. Add the marks Branch=True and cache=True to the operator at the starting position of the branch route, and then execute Step 1.5 to continue the verification.

[0050] Step 2: Obtain the types, parameter quantities, input dimension sets, and input shape sets of the operators in the intermediate model. Use these data as input data and calculate the output data using the operator platform matching degree model. The output data includes the operator execution time, video memory usage, video memory bandwidth utilization rate, shared memory usage, and register usage. Evaluate the matching degree of each operator with the target computing platform under different input shapes, and screen out the input shapes and dimension orders of the input data with the highest matching degree among the operators to form the first optimization model, which defines the input shapes and dimension orders of the operators.

[0051] Among them, the input shape set contains various input shapes that the operator can use. The input dimension set contains various dimension orders. The input shape determines the scale of the input data and the amount of data that can be processed. Different shapes will cause different amounts of computation and storage requirements during the operator's calculation.

[0052] The operator execution time refers to the execution time of the operator on the target computing platform. The shorter the time, the higher the matching degree. The video memory usage refers to the video memory capacity occupied during the operator's execution. The less the video memory usage, the higher the matching degree. The video memory bandwidth utilization rate refers to the utilization rate of the video memory bandwidth during the operator's execution. The higher the utilization rate, the higher the matching degree. The shared memory usage refers to the size of the shared memory occupied during the operator's execution. The less the shared memory usage, the higher the matching degree. The register usage refers to the number of registers occupied during the operator's execution. The less the register usage, the higher the matching degree.

[0053] The operator platform matching degree model constructed by the present invention is a multi-task learning model, including an operator execution time prediction sub-model and a target value prediction sub-model. The input of each sub-model is the input data. The operator execution time prediction sub-model is a model constructed based on the random forest regression model, which accurately predicts the execution time of the operator by evaluating the importance of operator features. The target value prediction sub-model is a multi-output regression model constructed based on the baseline model, which is used to predict the video memory usage, video memory bandwidth utilization rate, shared memory usage, and register usage during the operator's operation.

[0054] Among them, the baseline model is a model based on MultiOutputRegressor, which is a Meta Estimator in the machine learning library scikit-learn used to convert a single-output regression model into a multi-output regression model.

[0055] By collecting the types, parameter quantities, input dimensions, and input shapes of various operators on the target computing platform, as well as the operator execution time, video memory usage, video memory bandwidth utilization, shared memory usage, and register usage during the operator execution process, then transforming the input dimensions and input shapes into multiple different forms to form an input dimension set and an input shape set, calculating the matching degree between the operator and the target computing platform from the operator execution time, video memory usage, video memory bandwidth utilization, shared memory usage, and register usage, and constructing a training sample set from these data. Among them, the types, parameter quantities, input dimension set, and input shape set of the operator are input data, and the operator execution time, video memory usage, video memory bandwidth utilization, shared memory usage, and register usage are output data, and the matching degree is used as a label. The training of the operator platform matching degree model is completed using the constructed training sample set.

[0056] Specifically, the method for calculating the matching degree between the operator and the target computing platform from the operator execution time, video memory usage, video memory bandwidth utilization, shared memory usage, and register usage can be: assigning weights to the operator execution time, video memory usage, video memory bandwidth utilization, shared memory usage, and register usage respectively, and then calculating the weighted average according to the specific values of the output data, and using it as the matching degree between the operator and the target computing platform.

[0057] Step 3: Traverse each operator in the first optimized model, and judge the input shape consistency between it and adjacent operators. For the case where the input shapes are inconsistent, perform dimension rearrangement between adjacent operators.

[0058] Step 4: The branch operator set is composed of the operators located at the branch nodes in the first optimized model. In the branch operator set, for multiple branch operators with the same parent node operator, the same operator type, and the sum of the video memory usage not exceeding the platform video memory, mark the branch routes corresponding to these branch operators as routes that can be executed in parallel, and set the input data of these branch operators to be stored in the cache; for multiple branch operators with only the same parent node operator, set the input data of these branch operators to be stored in the cache to complete the optimization of the intelligent model to be optimized. Embodiment

[0059] In this embodiment, an intelligent model optimization method based on operator target value regression prediction provided by the present invention is adopted to achieve the rapid optimization of the intelligent model, including the following steps:

[0060] S1. Set the input and output nodes of the intelligent model to be optimized, and perform constant fusion and redundant operator removal on the intelligent model to be optimized.

[0061] S2. Analyze the intelligent model to be optimized to obtain the main trunk route, number all the operators in the intelligent model to be optimized, let one- and two-dimensional convolutions be class A, add and mul be class B, BN be class C, activation function classes be class D, and other operators be numbered starting from E in the order of appearance from input to output and from right to left. The serial numbers of the nodes are from input to output and from right to left. For example, the input node of the main trunk route is A1 and the output node is F1, as Figure 2 shown.

[0062] S3. Create a fusion combination, simplify the convolution or matrix multiplication, arithmetic operation class, and some activation function class operators into formulas, and then perform fusion.

[0063] The convolution operator is expressed as a formula: wx + b .

[0064] The BN operator is expressed as a formula: or wx + b .

[0065] The arithmetic operation class operator is expressed as a formula: Add = x + y , Sub = x - y , Div = x / y , Mul = xy .

[0066] The sigmoid operator in the activation function class is expressed as a formula: , The Tanh operator is expressed as a formula: , The relu operator is expressed as a formula: etc.

[0067] The operator fusion combination method is as follows:

[0068] GROUP1: Convolution + Convolution = Convolution with new parameters ( W’x + b’ ) (A1 → A2 = A1’)

[0069] GROUP2: Convolution + BN = Convolution with new parameters ( W’x + b’ ) (A1 → C1 = A1’)

[0070] GROUP3: Arithmetic operation class + Convolution = Convolution with new parameters (W’x + b’ ) (B1 → A1 = A1’)

[0071] GROUP4: Convolution + Arithmetic Operation Class = Convolution with New Parameters ( W’x + b’ ) (A1 → B1 = A1’)

[0072] GROUP5: Convolution + BN + Arithmetic Operation Class = Convolution with New Parameters ( W’x + b’ ) (A1 → C1 → B1 = A1’)

[0073] GROUP6: BN + Arithmetic Operation Class = BN with New Parameters ( W’x + b’ ) (C1 → B1 = C1’)

[0074] GROUP7: Arithmetic Operation Class + BN = BN with New Parameters ( W’x + b’ ) (B1 → C1 = C1’)

[0075] GROUP8: Convolution + Activation Function Class = Activation Function Class with Parameters (A1 → D1 = D1’), such as

[0076] relu, y (conv + relu) = max(0, W’x + b’ )

[0077] GROUP9: Convolution + BN + Activation Function Class = Activation Function Class with Parameters (A1 → C1 → D1 = D1’), such as relu, y (conv + bn + relu) = max(0, W’x + b’ )

[0078] GROUP10: Arithmetic Operation Class + Arithmetic Operation Class = New Arithmetic Operation Class (B1 → B2 = B1’), such as y 1 = x + y 、 y 2 = y 1 - y 3 Then y 2 = x + y - y 3, where y 、 y 3 are known numbers.

[0079] S4. Obtain the branch routes, and use the Ullmann subgraph isomorphism algorithm to check whether there are fusible nodes in all branch lines Lpath. Determine the starting nodes and ending nodes of all branch routes, and stop when no other branch routes can be decomposed in each branch route. For example, there are eight branch routes: A1→B1→C1, A2→B2→C2, ..., E1→B4→F1.

[0080] Let the starting node of the branch route in the main route be X and the ending node be Y. Search for branch routes for each pair of positions L(i, X, Y), including the following steps:

[0081] S4.1. Create two queues Xfront and Yrear. Xfront is used to search starting from the starting position X of the main line branch, and Yrear is used to search starting from the ending point Y of the main line branch. Add the starting node and the ending node to their respective queues, and obtain the branch route set Lpath after the search is completed.

[0082] S4.2. Construct the adjacency matrix. According to all the obtained branch routes, construct the adjacency matrix from the branch routes and operator combinations respectively. The adjacency matrix represents the connection relationship between nodes in the graph.

[0083] S4.3. Initialize the candidate matrix. According to the number of nodes in the target graph and the number of nodes in the initial source graph, construct the candidate matrix M. The rows of the candidate matrix M correspond to the nodes of the target graph, and the columns correspond to the nodes of the initial source graph. If the node u in the target graph can match the node v in the initial source graph, then M[u][v]=1, otherwise it is 0.

[0084] S4.4. Refine the candidate matrix. According to the constraints of the adjacency matrix, refine the candidate matrix M. For each edge u→v in the target graph, check whether there is a corresponding edge M[u][i]→M[v][j] in the initial source graph.

[0085] S4.5. Recursive matching. Use the backtracking method to recursively try to match nodes. For each node in the target graph, select a node in the initial source graph for matching and update the candidate matrix.

[0086] S4.6. Verify the matching. If a complete matching is found, replace the several nodes and edges matched in the initial source graph with the number of an operator combination, and retain the fusion operator information; until no matching items are returned and no further fusion operations can be performed.

[0087] S4.7. Fusion and replacement: At the numbered points of the operator combination, obtain the fusion operator information and perform the fusion operation to obtain the fusion operator. The fusion operator will replace the original operator combination to obtain a new branch route, mark the operator position of the branch route as Branch=True and cache=True. By default, all Branch and cache are False, and then return to S4.4 to continue the verification.

[0088] For example, A1 and B1 in A1→B1→C1 are fused into GROUP2. After the fusion calculation, GROUP2 is also an A operator, obtaining A1’→C2. The fused node numbers are determined according to the fusion operator type in the combination. Finally, the obtained branch routes are combined, and the result may be as shown in Figure 3 the following figure.

[0089] S5. Operator target value modeling.

[0090] S5.1. Data collection and processing

[0091] Collect the characteristics and target values of the operator on the deployment platform, and create an operator information dataset, including input data, output data, and target values.

[0092] The input data includes the following information:

[0093] Operator type, such as convolution type.

[0094] Operator parameter size, specifically: the size of the convolution kernel k is [1, 3, 5, 7], the length, width, height, and stride s are [1, 2, 3], the padding p is [0, 1, 2], the pooling kernel is [2, 3, 4], etc.

[0095] The input dimension of the operator, specifically: b is 1, c is [3, 16, 32, …, 2n, 1024], h and w are [7, 12, 14, 16, 24, 32, 64, 128, 196, 256, 512, 1024], where b is the batch size, c is the number of channels, h is the height, and w is the width.

[0096] The input shape of the operator, specifically: various shape combinations obtained by exchanging the dimensions of (b, c, h, w), (c, b, h, w), etc.

[0097] The output data includes the following information:

[0098] Operator execution time (T): The execution time of the operator on the target platform. The shorter the time, the higher the matching degree;

[0099] Video memory usage (MEM): The video memory capacity occupied when the operator is executed. The less the video memory usage, the higher the matching degree;

[0100] Video Memory Bandwidth Utilization (BW): The utilization rate of the video memory bandwidth during the execution of the operator. The higher the utilization rate, the higher the matching degree.

[0101] Shared Memory Usage (SMEM): The size of the shared memory occupied during the execution of the operator. The less the shared memory usage, the higher the matching degree.

[0102] Register Usage (REG): The number of registers occupied during the execution of the operator. The less the register usage, the higher the matching degree.

[0103] The operator matching degree is calculated based on the output data with the smallest T after multiple calculations.

[0104] S5.2. The operator platform matching degree model is a multi-task learning model. Different baseline models are used in different stages. The input is the same data set, and the output can be divided into three stages. In the first stage, a random forest regression is used to accurately predict the basic execution time (T) of the operator by evaluating the importance of operator features. In the second and third stages, a baseline model based on MultiOutputRegressor is used to build a multi-output regression model, which outputs the predicted values of each target value of the operator. Specifically, the operator target value model is divided into multiple stages, and each stage processes different features or target values.

[0105] First stage: Use the operator execution time prediction sub-model to predict the basic execution time (T) of the operator.

[0106] Second stage: Use the target value prediction sub-model to predict the video memory usage (MEM) and video memory bandwidth utilization (BW) during the operation of the operator.

[0107] Third stage: Use the target value prediction sub-model to predict the shared memory usage (SMEM) and register usage (REG) during the operation of the operator.

[0108] According to the importance of the task, different weights are assigned to each stage. First stage weight: 0.6; Second stage weight: 0.3; Third stage weight: 0.1.

[0109] S5.3. Training of the operator platform matching degree model

[0110] Train and evaluate the entire cascade model, balance the weights of multiple tasks, and the loss function is the weighted sum of multiple task loss functions: L = 0.6 * λ1L1 + 0.3 * λ2L2 + 0.3 * λ3L3 + 0.1 * λ4L4 + 0.1 * λ5L5.

[0111] Among them, L1 to L5 are the losses of each operator target value prediction respectively.

[0112] S5.4. Inference of the operator platform matching degree model

[0113] After the training of the operator platform matching degree model is completed, the operator information after the fusion of the pre-optimized model is obtained, such as conv2d (b = 1, c = 3, h = 320, w = 320, k = 3, s = 2, p = 1). These information are input into the operator target value model for inference, and according to the input dimension sizes b = 1, c = 3, h = 320, w = 320, all possible input shapes are enumerated, such as (1, 3, 320, 320), (3, 1, 320, 320), (320, 320, 3, 1), etc. It is possible to predict T, MEM, BW, SMEM, and REG of the conv2d operator under various different shape combinations, sort the Ts of all cases, and obtain the MEM, BW, SMEM, REG, and input shape Shape of the fastest Tfast in performance, that is, Tfast (MEM, BW, SMEM, REG, Shape).

[0114] S6. Allocate operator inference structure and method

[0115] S6.1. Match all the information of the fastest T of all operators in S5.4 of S5 with the model graph, judge whether the shapes of the front and rear operators are consistent. For operators with inconsistent shapes, add a permute or transpose operator after its output for dimension rearrangement, and the result is as Figure 4 shown.

[0116] S6.2. Operator alignment. Before the inference of the pre-optimized model, perform pre-inference on it. From the input Input node to the output Output node, find all operators with branch positions where Branch = True, and judge the following three points: ① Whether the numbers of its previous operator are the same; ② Whether the types of the operators at the current two branch positions are the same; ③ Whether the MEMs of the pairwise same operators are less than or equal to the video memory of the deployment platform. If all the above three points are satisfied, determine that they are parallel branches, align these operators, and calculate them simultaneously during inference, and execute S6.3; if ① is satisfied but ② and ③ are not satisfied, set its Branch to False and directly jump to S6.3.

[0117] In the example, the inputs of A3' and A2' are the same, determine that they are parallel branches, and if they are both conv2d and are of the same type of operator, and the operator video memory sum meets the alignment requirements, they can be calculated simultaneously.

[0118] S6.3. Shared memory and cache optimization. Store the common input data of the two branches that only meet ① in the cache to reduce the memory access latency. Ensure that the layout of the data in the memory is continuous to reduce the probability of cache misses.

[0119] According to the above allocation, the performance will be improved when the inference model is deployed at the deployment end.

[0120] In summary, the above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent model optimization method based on operator target value regression prediction, characterized in that: The specific steps include: Step 1: Set numbers for operators in the intelligent model to be optimized, extract branch routes, obtain operator combinations that can be fused in each branch route, replace the operator combination with the fusion operator, complete the fusion of the intelligent model to be optimized, and obtain the intermediate model; Step 2: The type, parameter quantity, input dimension set and input shape set of each operator in the intermediate model are used as input data, and the operator platform matching model is used to calculate the input data to obtain output data, where the output data includes operator execution time, video memory usage, video memory bandwidth utilization, shared memory usage and register usage; the input shape and dimension sequence with the highest matching degree in each operator are screened out according to the output data to form a first optimization model; Step 3: traverse each operator in the first optimization model, determine the consistency of the input shape between the operator and the adjacent operator, and rearrange the dimensions between the adjacent operators if there is any inconsistency; Step 4: The operators located at the branch nodes in the first optimization model constitute a branch operator set. In the branch operator set, for multiple branch operators with the same parent node operator, the same operator type, and the sum of the video memory usage not exceeding the platform video memory, the corresponding branch routes are marked as routes that can be executed in parallel, thereby completing the optimization of the intelligent model to be optimized; In step 1, the method of obtaining the operator combination that can be fused in each branch route and replacing the operator combination with the fusion operator is as follows: Step 1.1, traverse the intelligent model to be optimized to obtain all nodes, determine the node type and the number of branches of the branch node, and establish a first node set to save all nodes; Step 1.2, create a starting queue and an end queue. The starting queue is used to start searching from the starting node of the intelligent model to be optimized, and the starting node is added to the starting queue; the end queue is used to start searching from the ending node of the intelligent model to be optimized, and the ending node is added to the end queue; Start from the starting node and the ending node and expand to the current node layer by layer; Step 1.3: Check whether the current node in the two queues is a branch node. If it is a branch node, execute step 1.4; otherwise, expand forward to the adjacent node of the current node, and execute step 1.3 with the adjacent node as the current node; Step 1.4, starting from the current node, trace back to the adjacent branch node in the queue where it is located as the first branch node, and form a branch route by the nodes between the current node and the first branch node; Delete the nodes in the branch route from the first node set, and reduce the number of branches of the first branch node by 1. When the number of branches is 0, delete it from the first node set; if the current node is the starting node or the ending node, end the search, otherwise extend forward to the adjacent node of the current node, and execute step 1.3 with the adjacent node as the current node; Step 1.5, initialize the candidate matrix, the nodes of the target graph after the behavior of the candidate matrix are fused, and the nodes of the initial source graph of the intelligent model to be optimized are listed; if the node of the target graph can match the node of the initial source graph, the value of the corresponding point in the candidate matrix is ​​set to 1, otherwise it is 0; Step 1.6: According to the constraints of the adjacency matrix, recursively match each node in the target graph with the nodes of the initial source graph. When a complete match exists, eliminate the matched nodes and edges in the initial source graph and replace them with the numbers of the operator combination. Keep the fusion operator information until no match is returned and no fusion operation can be performed.

2. The intelligent model optimization method according to claim 1, characterized in that: The operator platform matching model in step 2 is a multi-task learning model.

3. The intelligent model optimization method according to claim 2, characterized in that: The operator platform matching model includes an operator execution time prediction sub-model and a target value prediction sub-model. The inputs of the sub-models are all input data. The operator execution time prediction sub-model predicts the execution time of the operator by evaluating the importance of the operator features. The target value prediction sub-model is used to predict the video memory usage, video memory bandwidth utilization, shared memory usage and register usage during the operator runtime.

4. The intelligent model optimization method according to claim 3, characterized in that: The operator execution time prediction sub-model is a model built based on a random forest regression model.

5. The intelligent model optimization method according to claim 3, characterized in that: The target value prediction sub-model is a multi-output regression model constructed based on the baseline model.

6. The intelligent model optimization method according to claim 1, characterized in that: The matching degree is calculated by assigning weights to operator execution time, video memory usage, video memory bandwidth utilization, shared memory usage and register usage, respectively, and then calculating a weighted average value based on specific values, which is the matching degree.

7. The intelligent model optimization method according to claim 1, characterized in that: The method for determining the fusion operator in step 1 is: using a formula to represent the fusionable operators, and then combining the operators according to the formula to form a formula for the fusion operator.

8. The intelligent model optimization method according to claim 1, characterized in that: The implementation method of obtaining the operator combinations that can be fused in each branch route described in step 1 is: operators of the same type, convolution operators and batch normalization operators, arithmetic operation operators and convolution operators, arithmetic operation operators and batch normalization operators, and convolution operators and activation function operators are all fusionable operator combinations.

9. The intelligent model optimization method according to claim 1, characterized in that: The method of setting numbers for operators in the intelligent model to be optimized in step 1 is as follows: the operator number consists of a type number and an operator node serial number, the type number is the number of the operator type of a certain type, and the operator node serial number is set in the order in which the operator appears from input to output and from right to left; for operators that do not belong to a certain type, the type number is set in the order in which the operator appears from input to output and from right to left.

Citation Information

Patent Citations

  • Hardware characteristic related depth model computational graph automatic optimization method

    CN115423082A