Model Optimization Method, Device, and Computer Storage Medium
The method optimizes deep learning models by decoupling and structurally reflecting operations to enhance computational speed and reduce errors, addressing the speed-accuracy tradeoff in quantized models.
Patent Information
- Application Number
- CN202111193027.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-10-13
AI Technical Summary
The quantization technology of deep learning models increases the computing speed while also reducing the computing accuracy, and the computing time-consuming after quantization increases, making it impossible to effectively balance the speed and accuracy.
By decoupling the operators in the target model, performing structural reflection and heterogeneous quantization processing, determining the optimal inference path, optimizing the model structure to improve the operation speed and reduce errors.
Without sacrificing operation accuracy, significantly improve the computing speed of the model, improve the success rate of model optimization, and reduce the error rate.
Smart Images

Figure CN114048855B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and particularly to a model optimization method, device and computer storage medium. Background Art
[0002] Deep learning has been widely applied in various branches of artificial intelligence technology, providing sufficient support for industrial productization. However, performance indicators such as speed have always been a bottleneck restricting the large-scale and wide application of deep learning technology. For this reason, the academic community has proposed a method called "quantization technology" in order to improve the operation speed of the model.
[0003] However, quantization technology has also introduced many disadvantages. On the one hand, quantization technology represents data types using fewer memory bits. Although it can improve the operation speed of the model to a certain extent, it leads to a decrease in the operation accuracy of the model. However, deep learning networks are "deep", which means that a set of data needs to go through dozens, hundreds or even thousands of operation inferences to derive the final result. Even if the requirement for numerical accuracy is not high, it is impossible to ensure that the numerical values still converge after so many operations. Once the accumulation of errors changes from quantitative to qualitative, the calculation result will be completely uncontrollable.
[0004] On the other hand, since each computing unit needs to perform independent operations, whether the operation time after quantization of each unit can completely cover the additional quantization and restoration time has become the measurement standard for the speed gain after quantization. If the time consumed by data conversion itself is very large and the time saved by the operation after quantization is small, not only will the operation accuracy be lost, but also a fruitless result will be obtained. Summary of the Invention
[0005] In view of the above problems, the present application provides a model optimization method, device and computer storage medium, which can achieve an effective balance between operation time consumption and operation accuracy.
[0006] The first aspect of the present application provides a model optimization method, which includes performing decoupling processing on each operator in a target model to obtain a decoupled model of the target model; performing structure reflection processing on each operator in the decoupled model and the inference data between the operators to obtain a reflection model of the decoupled model; performing heterogeneous quantization processing on the reflection model to obtain a heterogeneous quantization model including multiple candidate inference paths; determining an optimal inference path of the heterogeneous quantization model according to the inference time consumption corresponding to each candidate inference path in the heterogeneous quantization model; and obtaining an optimized model of the target model based on the optimal inference path.
[0007] The second aspect of the present application provides a computer storage medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor is caused to execute the method described in the first aspect above.
[0008] The third aspect of the present application provides a model optimization device, which includes a decoupling module for performing decoupling processing on each operator in a target model to obtain a decoupled model of the target model; a structure reflection module for performing structure reflection processing on each of the operators in the decoupled model and the inference data between the operators to obtain a reflected model of the decoupled model; a heterogeneous quantization module for performing heterogeneous quantization processing on the reflected model to obtain a heterogeneous quantization model including multiple candidate inference paths; and an optimization module for using each of the candidate inference paths in the heterogeneous quantization model to perform inference analysis on a preset test task respectively to determine an optimal inference path of the heterogeneous quantization model, and obtaining an optimized model of the target model based on the optimal inference path.
[0009] In summary, the model optimization method, device and computer storage medium provided by the embodiments of the present application perform structural transformation on a target model through decoupling, structure reflection and heterogeneous quantization processing, so as to use the transformed model to solve the optimal inference path, thereby obtaining an optimized model of the target model. Accordingly, the present application can not only effectively improve the operation speed of the model, but also the provided model optimization scheme has the advantages of small error rate and high model optimization success rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0011] Figure 1 It is a schematic flowchart of the model optimization method according to the first embodiment of the present application.
[0012] Figure 2 It is a schematic flowchart of the model optimization method according to the second embodiment of the present application.
[0013] Figure 3 It is a schematic flowchart of the model optimization method according to the third embodiment of the present application.
[0014] Figure 4 It is a schematic architecture diagram of the model optimization device according to the fifth embodiment of the present application.
[0015] Element reference numerals
[0016] 400: Model optimization device; 402: Decoupling module; 404: Structure reflection module; 406: Heterogeneous quantization module; 408: Optimization module. Detailed implementation manner
[0017] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.
[0018] First embodiment
[0019] Figure 1 The flowchart of the model optimization method according to the first embodiment of the present application is shown. As shown in the figure, this embodiment mainly includes the following processing steps:
[0020] Step S102, perform decoupling processing on each operator in the target model to obtain a decoupled model of the target model.
[0021] Optionally, the target model can be a trained deep learning model.
[0022] Optionally, the target network can be converted into an intermediate representation format for storage, so that the target model can communicate between different architectures and perform quantization on the target network using an inference framework (inference engine).
[0023] Optionally, for each operator in the target model, a quantization operator for performing quantization operations and an inverse quantization operator for performing inverse quantization operations can be added to obtain a decoupled model. That is, two operators, namely quantization and inverse quantization, are added at both ends of each operator to increase the complexity of the target model (deep learning neural network).
[0024] Step S104, perform structure reflection processing on each operator in the decoupled model and the inference data between the operators to obtain a reflection model of the decoupled model.
[0025] Optionally, each operator in the decoupled model can be converted from a node to a path, and each inference data between two adjacent operators can be converted from a path to a node to obtain a reflection model.
[0026] Specifically, in the target model, the operator is used as a node, and the data cache between the operators is used as a path; by performing structure reflection processing on the target model, the data cache between the operators is used as a node, and each operator is used as a path to construct a reflection model of the target model, that is, a directed acyclic graph (DAG) is constructed.
[0027] Step S106, perform heterogeneous quantization processing on the reflection model to obtain a heterogeneous quantization model including multiple candidate inference paths.
[0028] Optionally, for each operator in the reflection model, multiple preset quantization methods can be configured respectively, and the preset quantization methods corresponding to each operator can be arbitrarily combined to construct a heterogeneous quantization model including multiple candidate inference paths.
[0029] Optionally, the preset quantization methods assigned to each operator may at least include one of int8, int16, int32, fp16, and bf16.
[0030] Optionally, according to the given empirical value, a specified preset quantization method can be assigned to each operator as the alternative quantization scheme for each operator.
[0031] Optionally, if no empirical value is given, all preset quantization methods can be defaulted as the alternative quantization scheme for the operator.
[0032] Optionally, based on the preset influence factor, the weight parameters corresponding to each preset quantization method can be set.
[0033] Specifically, since the operation efficiencies of different quantization methods are different; in addition to the quantization method, a series of other influence factors will also affect the calculation speed of each operator. Therefore, by configuring the corresponding weight parameters for each preset quantization method, a directed acyclic graph carrying weight parameters can be obtained.
[0034] Optionally, the preset influence factor may include at least one of the affinity of the hardware architecture corresponding to the operator, the affinity of the CPU corresponding to the operator, the operation speed of the operator, the data quantization conversion cost, and the inference performance of the inference engine.
[0035] Among them, the hardware architecture corresponding to the nucleophilicity of the operator is, for example: the Intel x86 hardware architecture has better nucleophilicity for floating-point operation operators than the arm hardware architecture; the CPU corresponding to the affinity of the operator is, for example: the operator optimization level of arm A57 is better than other CPUs with the same computing power; the operation speed of the operator is, for example: the convolution operator can be implemented through GEMM (General Matrix Multiplication) or Winograd fast convolution algorithm; GEMM has strong generality but slightly lower efficiency; while Winograd has high sales volume but is only applicable to specific convolution kernels; the cost of data quantization conversion is, for example: when two adjacent operators both use the same quantization, without overflow, one dequantization and one quantization process can be saved, that is, two adjacent operators can share the quantized data; the inference performance of the inference engine is, for example: the self-developed mathematical operation acceleration library One-API actively promoted by Intel has the theoretical optimal solution for the 1x1 int8 convolution of the x86 architecture.
[0036] Step S108, determine the optimal inference path of the heterogeneous quantization model according to the inference time consumption corresponding to each candidate inference path in the heterogeneous quantization model.
[0037] Optionally, obtain the inference time consumption corresponding to each candidate sub-path corresponding to each pair of adjacent nodes in the candidate inference path, and determine the inference time consumption of the candidate inference path according to the sub-path inference time consumption corresponding to each candidate sub-path. By comparing the inference time consumption corresponding to each candidate inference path, determine the candidate inference path with the shortest inference time consumption as the optimal inference path of the heterogeneous quantization model.
[0038] Step S110, obtain the optimized model of the target model based on the optimal inference path.
[0039] Optionally, only retain the candidate inference path determined to be the optimal inference path, and delete other candidate inference paths in the heterogeneous quantization model to obtain the optimized model.
[0040] Optionally, the inference analysis of the preset test task can also be performed using the optimized model and the target model respectively, obtain the optimized time consumption of the optimized model and the target time consumption of the target model, and compare the optimized time consumption with the target time consumption. If the difference between the optimized time consumption and the target time consumption is less than the preset optimization threshold, it represents that the optimization of the target model is successful, and then output the optimized model.
[0041] In summary, the model optimization method of the embodiment of the present application transforms the target model by performing decoupling, structure reflection, and heterogeneous quantization, and uses the transformed model to solve the optimal inference path, so as to obtain the optimized model of the target model. Accordingly, the model optimization method provided in this embodiment can effectively improve the optimized model in both the speed and numerical error dimensions without sacrificing any conditions.
[0042] Second Embodiment
[0043] Figure 2 The flowchart shows the model optimization method according to the second embodiment of the present application. As shown in the figure, this embodiment is mainly a specific implementation of the above step S108. As shown in the figure, this embodiment mainly includes the following steps:
[0044] Step S202: Determine multiple candidate segmentation paths corresponding to each pair of adjacent nodes in the heterogeneous quantization model according to multiple preset quantization methods corresponding to each operator in the heterogeneous quantization model.
[0045] Specifically, since the reflection model of the present application uses each operator as a path of the model and each inference data between two adjacent operators as a node of the model, multiple candidate segmentation paths corresponding to each pair of adjacent nodes in the heterogeneous quantization model can be determined according to multiple preset quantization methods configured for each operator in the heterogeneous quantization model.
[0046] Step S204: Determine the shortest segmentation path corresponding to each pair of adjacent nodes according to the segmentation inference time of each candidate segmentation path corresponding to a preset test task.
[0047] Optionally, the Dijkstra algorithm can be used to predict the segmentation inference time of each candidate segmentation path corresponding to a preset test task by using a greedy strategy for each node and its adjacent nodes, and accordingly determine the shortest segmentation path corresponding to each pair of adjacent nodes.
[0048] Step S206: Determine the optimal inference path of the heterogeneous quantization model according to the shortest segmentation path corresponding to each pair of adjacent nodes.
[0049] Optionally, according to the shortest segmentation path corresponding to each pair of adjacent nodes, the shortest inference path from the input data node of the heterogeneous quantization model to each node in the heterogeneous quantization model can be gradually constructed, and finally the shortest inference path of the heterogeneous quantization model can be obtained.
[0050] Third Embodiment
[0051] Figure 3 The flowchart shows the model optimization method according to the third embodiment of the present application. This embodiment mainly shows the specific implementation of the above step S204. As shown in the figure, this embodiment mainly includes the following steps:
[0052] Step S302: Obtain a pair of adjacent nodes in the heterogeneous quantization model.
[0053] Specifically, starting from the input data node of the heterogeneous quantization model, each pair of adjacent nodes in the heterogeneous quantization model can be obtained in sequence to gradually construct the shortest inference path from the input data node of the heterogeneous quantization model to each node in the model.
[0054] Step S304: According to multiple candidate segmented paths of adjacent nodes, using Dijkstra's algorithm and greedy strategy, obtain the predicted segmented inference time corresponding to each candidate segmented path, and determine the candidate segmented path with the shortest predicted segmented inference time as the shortest segmented path.
[0055] Specifically, Dijkstra's algorithm can be used, and a greedy strategy is applied to each node and its adjacent nodes to obtain the predicted segmented inference time corresponding to each candidate segmented path, and determine the candidate segmented path with the shortest predicted segmented inference time as the shortest segmented path.
[0056] Step S306: Use the shortest segmented path to perform inference analysis on the preset test task to obtain the actual segmented inference time of the shortest segmented path.
[0057] Specifically, the actual segmented inference time of the currently determined shortest segmented path can be obtained by running the test path in real time.
[0058] Step S308: Determine whether the difference between the actual segmented inference time and the predicted segmented inference time is greater than the preset feedback threshold. If it is greater, return to step S304; if it is less, proceed to step S310.
[0059] Specifically, according to the difference between the actual segmented inference time and the predicted segmented inference time, it can be judged whether the predicted segmented inference time predicted for the currently determined shortest segmented path in step S304 is reasonable. If the difference between the actual segmented inference time and the predicted segmented inference time is greater than the preset feedback threshold, it means that the predicted segmented inference time obtained in step S304 is unreasonable, and then return to step S304 to re-predict the predicted segmented inference time corresponding to each candidate segmented path and re-determine the new shortest segmented path. If the difference between the actual segmented inference time and the predicted segmented inference time is less than the preset feedback threshold, it means that the predicted segmented inference time predicted for the currently determined shortest segmented path in step S304 is reasonable, and then proceed to step S310.
[0060] Step S310: Output the shortest segmented path of adjacent nodes.
[0061] Step S312: Determine whether there are un-inferred adjacent nodes in the heterogeneous quantization model. If there are, return to step S302 to obtain the next pair of adjacent nodes and perform inference on the shortest segmented path. If not, continue to execute step S206.
[0062] It should be noted that in the real-time solution of this embodiment, the cross-execution path prediction step (i.e., step S304) and the prediction feedback step (i.e., step S306) are adopted to sequentially determine the shortest segmented paths corresponding to each pair of adjacent nodes in the heterogeneous quantization model. That is, through a real-time solution method, after solving the shortest segmented path of the current adjacent nodes, the shortest segmented path of the next pair of connected nodes is continued to be solved until all the shortest segmented paths of all adjacent nodes in the heterogeneous quantization model are solved. It has been verified that this method can not only improve the speed but also reduce the error of inference prediction.
[0063] However, this is not the limit. In other embodiments, the loop execution path prediction step (i.e., step S304) can also be adopted to obtain the shortest segmented paths corresponding to all adjacent nodes, and then the prediction feedback step (i.e., step S306) is looped to verify the shortest segmented paths corresponding to all adjacent nodes to determine the shortest segmented paths corresponding to each pair of adjacent nodes in the heterogeneous quantization model. That is, through a pseudo-real-time solution method, all the shortest segmented paths of all adjacent nodes in the heterogeneous quantization model are first predicted, and then the shortest segmented paths corresponding to each pair of adjacent nodes predicted are verified.
[0064] Furthermore, in other embodiments, during the initial model testing stage, nodes with high time consumption can be located for targeted quantization and search, and relevant nodes with low time consumption originally can be abandoned. By this means, the model optimization efficiency can be further improved.
[0065] In summary, the model optimization method of the embodiment of the present application solves the shortest segmented paths corresponding to each adjacent node in the heterogeneous quantization model by cross-executing the path prediction step and the prediction feedback step. Using this real-time feedback adjustment method, not only can the optimal quantization path in the model be selected to improve the model optimization effect, but also it has the advantages of small error rate and high model optimization success rate.
[0066] Fourth Embodiment
[0067] The fourth embodiment of the present application provides a computer storage medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the method described in any one of the first to third embodiments above.
[0068] Fifth Embodiment
[0069] Figure 4 The architecture schematic diagram of the model optimization device according to the fifth embodiment of the present application is shown. As shown in the figure, the model optimization device 400 of this embodiment mainly includes: a decoupling module 402, a structure reflection module 404, a heterogeneous quantization module 406, and an optimization module 408.
[0070] The decoupling module 402 is used to perform decoupling processing on each operator in the target model to obtain a decoupled model of the target model.
[0071] Optionally, the decoupling module 402 is further used to add a quantization operator for performing quantization operations and an inverse quantization operator for performing inverse quantization operations for each operator in the target model, so as to obtain the decoupled model.
[0072] The structure reflection module 404 is used to perform structure reflection processing on each operator in the decoupled model and the inference data between the operators to obtain a reflection model of the decoupled model.
[0073] Optionally, the structure reflection module 404 is further used to convert each operator in the decoupled model from a node to a path, and convert each inference data between two adjacent operators from a path to a node, so as to obtain the reflection model.
[0074] The heterogeneous quantization module 406 is used to perform heterogeneous quantization processing on the reflection model to obtain a heterogeneous quantization model including multiple candidate inference paths.
[0075] Optionally, the heterogeneous quantization module 406 is further used to respectively configure multiple preset quantization methods for each operator in the reflection model; any combination of the preset quantization methods corresponding to each operator is used to construct the heterogeneous quantization model including multiple candidate inference paths.
[0076] Optionally, the preset quantization method includes at least one of int8, int16, int32, fp16, and bf16.
[0077] Optionally, the heterogeneous quantization module 406 is further used to set weight parameters corresponding to each preset quantization method based on a preset influence factor; wherein, the preset influence factor includes at least one of the nucleophilicity of the hardware architecture corresponding to the operator, the affinity of the CPU corresponding to the operator, the operation speed of the operator, the data quantization conversion cost, and the inference performance of the inference engine.
[0078] The optimization module 408 is used to respectively perform inference analysis on the preset test tasks by using each candidate inference path in the heterogeneous quantization model to determine the optimal inference path of the heterogeneous quantization model, and obtain an optimized model of the target model based on the optimal inference path.
[0079] Optionally, the optimization module 408 is further configured to determine multiple candidate segmentation paths corresponding to each pair of adjacent nodes in the heterogeneous quantization model according to the multiple preset quantization methods corresponding to each operator in the heterogeneous quantization model; determine the shortest segmentation paths corresponding to each pair of adjacent nodes according to the segmentation inference time of each candidate segmentation path corresponding to a preset test task; and determine the optimal inference path of the heterogeneous quantization model according to the shortest segmentation paths corresponding to each pair of adjacent nodes.
[0080] Optionally, the optimization module 408 is further configured to perform the following steps for each pair of adjacent nodes: perform a path prediction step, according to the multiple candidate segmentation paths of the adjacent nodes, use the Dijkstra algorithm and a greedy strategy to obtain the predicted segmentation inference time corresponding to each candidate segmentation path, and determine the candidate segmentation path with the shortest predicted segmentation inference time as the shortest segmentation path; perform a prediction feedback step, use the shortest segmentation path to perform the inference analysis of the preset test task, and obtain the actual segmentation inference time of the shortest segmentation path; wherein, if the difference between the actual segmentation inference time and the predicted segmentation inference time is greater than a preset feedback threshold, re-perform the path prediction step to determine a new shortest segmentation path until the difference between the actual segmentation inference time and the predicted segmentation inference time of the shortest segmentation path is less than the preset feedback threshold.
[0081] Optionally, the optimization module 408 is further configured to alternately execute the path prediction step and the prediction feedback step to sequentially determine the shortest segmentation paths corresponding to each pair of adjacent nodes; or loop the path prediction step to obtain the shortest segmentation paths corresponding to all the adjacent nodes, and then loop to execute the prediction feedback step to perform verification on the shortest segmentation paths corresponding to all the adjacent nodes, so as to determine the shortest segmentation paths corresponding to each pair of adjacent nodes.
[0082] Optionally, the optimization module 408 is further configured to use the optimization model and the target model to respectively perform the inference analysis of the preset test task, and obtain the optimization time of the optimization model and the target time of the target model; compare the optimization time and the target time, and if the difference between the optimization time and the target time is less than a preset optimization threshold, output the optimization model.
[0083] Optionally, the target model includes a trained deep learning model.
[0084] In addition, the model optimization device 400 according to the embodiment of the present invention can also be used to implement other steps in the foregoing embodiments of each model optimization method, and has the beneficial effects of the corresponding method step embodiments, which will not be elaborated herein.
[0085] In summary, the model optimization method, device, and computer storage medium provided in this application perform decoupling, structure reflection, and heterogeneous quantization transformation processing on the target model, and solve the optimal inference path based on the transformed model to obtain the optimized model of the target model. Thereby, the optimized model obtained in this application can achieve an effective balance between operation time consumption and operation accuracy, and has the advantages of small error rate and high model optimization success rate.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A model optimization method, characterized in that, Including: Performing decoupling processing on each operator in the target model to obtain a decoupled model of the target model; Performing structural reflection processing on each operator in the decoupled model and the inference data between the operators to obtain a reflection model of the decoupled model; Performing heterogeneous quantization processing on the reflection model to obtain a heterogeneous quantization model including multiple candidate inference paths; Determining an optimal inference path of the heterogeneous quantization model according to the inference time consumption corresponding to each candidate inference path in the heterogeneous quantization model; And Obtaining an optimized model of the target model based on the optimal inference path; The performing heterogeneous quantization processing on the reflection model to obtain a heterogeneous quantization model including multiple candidate inference paths includes: For each operator in the reflection model, respectively configuring multiple preset quantization methods; based on a preset influence factor, setting weight parameters corresponding to each of the preset quantization methods; wherein, the preset influence factor includes at least one of the nucleophilicity of the hardware architecture corresponding to the operator and the affinity of the CPU corresponding to the operator; Arbitrarily combining the preset quantization methods corresponding to each operator to construct the heterogeneous quantization model including multiple candidate inference paths.
2. The model optimization method according to claim 1, characterized in that The performing decoupling processing on each operator in the target model to obtain a decoupled model of the target model includes: For each operator in the target model, adding a quantization operator for performing quantization operations and an inverse quantization operator for performing inverse quantization operations to obtain the decoupled model.
3. The model optimization method according to claim 2, wherein The performing structural reflection processing on each operator in the decoupled model and the inference data between the operators to obtain a reflection model of the decoupled model includes: Converting each operator in the decoupled model from a node to a path, and converting the inference data between two adjacent operators from a path to a node to obtain the reflection model.
4. The model optimization method according to claim 1, characterized in that The preset quantization method includes at least one of int8, int16, int32, fp16, and bf16.
5. The model optimization method according to claim 1, wherein The determining an optimal inference path of the heterogeneous quantization model according to the inference time consumption corresponding to each candidate inference path in the heterogeneous quantization model includes: Determining multiple candidate segmented paths corresponding to each pair of adjacent nodes in the heterogeneous quantization model according to the multiple preset quantization methods corresponding to each operator in the heterogeneous quantization model; Determining the shortest segmented path corresponding to each pair of adjacent nodes according to the segmented inference time consumption corresponding to each candidate segmented path for a preset test task; and Determining the optimal inference path of the heterogeneous quantization model according to the shortest segmented paths corresponding to each pair of adjacent nodes.
6. The model optimization method according to claim 5, characterized in that The determining the shortest segmented path corresponding to each pair of adjacent nodes according to the segmented inference time consumption corresponding to each candidate segmented path for a preset test task includes: For each pair of adjacent nodes, performing the following steps: Execute the path prediction step. According to the multiple candidate segmented paths of the adjacent nodes, use the Dijkstra algorithm and the greedy strategy to obtain the predicted segmented inference time corresponding to each candidate segmented path, and determine the candidate segmented path with the shortest predicted segmented inference time as the shortest segmented path; Execute the prediction feedback step. Use the shortest segmented path to perform the inference analysis of the preset test task to obtain the actual segmented inference time of the shortest segmented path; where If the difference between the actual segmented inference time and the predicted segmented inference time is greater than the preset feedback threshold, re-execute the path prediction step to determine a new shortest segmented path until the difference between the actual segmented inference time and the predicted segmented inference time of the shortest segmented path is less than the preset feedback threshold.
7. The model optimization method according to claim 6, wherein Cross-execute the path prediction step and the prediction feedback step to sequentially determine the shortest segmented paths corresponding to each pair of adjacent nodes; Or Loop the path prediction step to obtain the shortest segmented paths corresponding to all the adjacent nodes, and then loop and execute the prediction feedback step to perform verification on the shortest segmented paths corresponding to all the adjacent nodes, so as to determine the shortest segmented paths corresponding to each pair of adjacent nodes.
8. The model optimization method according to claim 5, wherein The method further includes: Use the optimized model and the target model to respectively perform the inference analysis of the preset test task to obtain the optimization time of the optimized model and the target time of the target model; Compare the optimization time and the target time. If the difference between the optimization time and the target time is less than the preset optimization threshold, output the optimized model.
9. The model optimization method according to claim 1, wherein The target model includes a trained deep learning model.
10. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, and when the computer instructions are executed by a processor, the processor executes the method according to any one of claims 1 to 9.
11. A model optimization device, characterized in that, It includes: A decoupling module for performing decoupling processing on each operator in the target model to obtain a decoupled model of the target model; A structure reflection module for performing structure reflection processing on each operator in the decoupled model and the inference data between the operators to obtain a reflection model of the decoupled model; A heterogeneous quantization module for performing heterogeneous quantization processing on the reflection model to obtain a heterogeneous quantization model including multiple candidate inference paths; An optimization module for using each candidate inference path in the heterogeneous quantization model to respectively perform inference analysis on a preset test task to determine the optimal inference path of the heterogeneous quantization model, and based on the optimal inference path, obtain an optimized model of the target model; The heterogeneous quantization module is further used to respectively configure multiple preset quantization methods for each operator in the reflection model; Arbitrarily combine the preset quantization methods corresponding to each operator to construct the heterogeneous quantization model including multiple candidate inference paths; The heterogeneous quantization module is also used to set respective weight parameters corresponding to the respective preset quantization methods based on a preset influence factor; wherein, the preset influence factor includes at least one of the nucleophilicity of the hardware architecture corresponding to the operator and the affinity of the CPU corresponding to the operator.
Citation Information
Patent Citations
Method and system for determining mixing precision quantification strategy for deep neural network
CN112906883A
Adjusting and optimizing method and compiling method of deep learning model, and computing device
CN113269319A
Data processing method and device of terminal network model, terminal and storage medium
CN113392954A