Quantization performance evaluation method and device for model quantization, equipment and medium
Through the quantitative performance evaluation method and genetic algorithm search framework, the problem of difficulty in deploying deep learning neural networks at edge devices and difficulty in evaluating quantization performance is solved, and efficient model deployment and operation on edge devices is achieved.
Patent Information
- Application Number
- CN202510146627.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-03
AI Technical Summary
Existing deep learning neural networks are difficult to deploy at edge devices, and the quantization bit width differences at different network layers lead to difficulty in evaluating quantization performance.
Provide a quantitative performance evaluation method for model quantization. By obtaining the model to be evaluated and converting it into an optimized quantitative network structure request, determining the inference delay and network calculation amount of the model before and after quantization, and constructing quantitative performance evaluation values, and using genetic algorithms to search for the quantitative network structure with the best performance.
Systematically explore the impact of different quantization bit combinations on model performance, improve the probability of finding the most suitable model and ensuring the minimum performance loss quantization bits, and realize efficient model deployment and operation on edge devices.
Smart Images

Figure CN120087440A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model quantization, and particularly relates to a quantization performance evaluation method, device, equipment and medium for model quantization. Background Art
[0002] Since the advent of deep learning technology, deep learning has achieved remarkable breakthroughs in many fields such as image classification, speech recognition, and natural language processing with its powerful representation learning and feature extraction capabilities. The successive emergence of representative networks such as YOLO, SSD, Faster R-CNN, Transformer, BERT, and DERT has greatly improved the accuracy of relevant detections required by users. However, in order to meet the requirements of generalization performance, the structural parameters of a single deep learning network are increasing, breaking through from millions to hundreds of millions, billions, or even tens of billions of parameters. The increase in the number of parameters will inevitably lead to a huge demand for and consumption of hardware performance, thus bringing huge challenges to the deployment of deep learning networks on resource-constrained edge devices.
[0003] To solve this problem, the quantization technology of deep learning models has emerged. This technology can effectively reduce the computational amount and storage requirements of the model, enabling larger-scale models to be more conveniently deployed to edge devices with limited computing power to meet various task requirements.
[0004] In actual applications, the current model quantization technology still has certain defects, resulting in difficulties in deploying deep learning models on edge devices, and it is difficult to evaluate the quantization performance of the model due to the quantization bit-width differences existing in different network layers of the model.
[0005] To solve the problems of deploying deep learning neural network edge devices and evaluating quantization performance caused by quantization bit-width differences in different network layers, it is necessary to provide a quantization performance evaluation method for model quantization. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a quantization performance evaluation method, device, equipment and medium for model quantization to solve the existing problems of deploying deep learning neural network edge devices and evaluating quantization performance caused by quantization bit-width differences in different network layers.
[0007] According to the first aspect, embodiments of the present invention provide a quantization performance evaluation method for model quantization, and the method includes:
[0008] Obtain a model to be evaluated and a quantization request for the model to be evaluated, and convert the quantization request for the model to be evaluated into an optimized quantization network structure request; the optimized quantization network structure request is to find a quantization network structure of the model to be evaluated that meets preset requirements;
[0009] According to the request for optimizing the quantized network structure, determine the inference latency and network computational cost generated by the model to be evaluated before and after model quantization, and construct a quantization performance evaluation value based on the inference latency and network computational cost consumption; in the quantization performance evaluation value, the coefficient of determination is used to evaluate the difference between the weights and feature maps of the model to be evaluated before and after quantization.
[0010] Based on the request for optimizing the quantized network structure, construct the feasible region of the quantized network structure of the model and a search framework for searching for the quantized network structure with the optimal performance; the search framework uses a genetic algorithm constructed based on the quantization performance evaluation value to determine the quantized network structure with the optimal performance among each feasible quantized network structure.
[0011] Combined with the first aspect, in the first embodiment of the first aspect, the inference latency is determined through the following steps:
[0012] Quantize each layer of the model to be evaluated to a first preset bit width.
[0013] Quantize the first layer of the model to be evaluated from the first preset bit width to a second preset bit width; the first preset bit width is configured as 8-bit width, and the second preset bit width is configured as 16-bit width.
[0014] Execute the inference of the first layer of the network quantized to the second preset bit width, determine the layer inference latency of the first layer of the network, and quantize the first layer of the network from the second preset bit width to the first preset bit width.
[0015] Quantize the network of the next layer of the first layer of the network from the first preset bit width to the second preset bit width, execute the inference of the next layer of the network quantized to the second preset bit width, determine the layer inference latency of the next layer of the network, and quantize the next layer of the network from the second preset bit width to the first preset bit width, until the layer inference latency of each layer of the network in the model to be evaluated is determined.
[0016] Collect the layer inference latency of each layer of the network in the model to be evaluated to obtain the model inference latency of the model to be evaluated.
[0017] Combined with the first embodiment of the first aspect, in the second embodiment of the first aspect, the network computational cost is determined through the following steps:
[0018] Quantize each layer of the model to be evaluated to a first preset bit width.
[0019] Quantize the first layer of the model to be evaluated from the first preset bit width to a second preset bit width.
[0020] Determine the layer computational cost of the first layer of the network quantized to the second preset bit width, and quantize the first layer of the network from the second preset bit width to the first preset bit width.
[0021] Quantize the network of the next layer of the first-layer network from the first preset bit width to the second preset bit width, determine the layer computation amount of the next layer network, and quantize the next layer network from the second preset bit width to the first preset bit width until the layer computation amount of each layer network in the model to be evaluated is determined;
[0022] Aggregate the layer computation amounts of each layer network in the model to be evaluated to obtain the network computation amount of the model to be evaluated.
[0023] Combined with the first aspect, in the third implementation manner of the first aspect, the construction of the feasible region of the quantized network structure of the model to be evaluated and the search framework for searching for the quantized network structure with the optimal performance based on the request for optimizing the quantized network structure specifically includes:
[0024] Based on the request for optimizing the quantized network structure, obtain all feasible quantized network structures of the model to be evaluated, and encode each feasible quantized network structure to obtain the genetic algorithm coding individuals of each feasible quantized network structure. Based on all the genetic algorithm coding individuals, obtain the population and the evolutionary generation of the population;
[0025] Determine whether the preset termination condition is satisfied; the preset termination condition is that the evolutionary generation reaches the preset maximum evolutionary generation;
[0026] Determine that the preset termination condition is satisfied, use the quantized performance evaluation value to determine the fitness of each genetic algorithm coding individual in the population, and decode the genetic algorithm coding individual with the optimal fitness to obtain the quantized network structure with the optimal performance.
[0027] Combined with the third implementation manner of the first aspect, in the fourth implementation manner of the first aspect, the obtaining of the feasible region of the quantized network structure of the model to be evaluated and the construction of the search framework for searching for the quantized network structure with the optimal performance based on the request for optimizing the quantized network structure further includes:
[0028] Determine that the preset termination condition is not satisfied, and use the quantized performance evaluation value to determine the fitness of each genetic algorithm coding individual in the population;
[0029] Apply the preset selection operator to the population, and select the genetic algorithm coding individuals with fitness exceeding the preset value from the population as the parent individuals;
[0030] Randomly select two genetic algorithm coding individuals from the parent individuals, and exchange the genes of the selected genetic algorithm coding individuals with the preset crossover probability to generate two new crossed individuals. Repeat the operations of random selection and gene exchange until all the genetic algorithm coding individuals of the parent individuals have completed gene exchange to obtain a new crossed population;
[0031] Apply a preset mutation operator to the population after crossover to obtain the mutated population and the generation number of the mutated population until the preset termination condition is met.
[0032] Combined with the third and fourth embodiments of the first aspect, in the fifth embodiment of the first aspect, the search framework adopts an encoder-decoder framework. The feasible quantization network structure is encoded by an individual encoder. The feasible domain of each encoding position on the genetic algorithm encoded individual is the types of bit widths supported by the edge device. When it is necessary to determine the fitness of the genetic algorithm encoded individual, the genetic algorithm encoded individual is decoded by an individual decoder and then the fitness is determined using the quantization performance evaluation value.
[0033] Combined with the fourth embodiment of the first aspect, in the sixth embodiment of the first aspect, the preset selection operator adopts a roulette wheel selection strategy, and the gene crossover process adopts a multi-point crossover operation.
[0034] According to a third aspect, an embodiment of the present invention provides a quantization performance evaluation device for model quantization, the device includes:
[0035] A quantization conversion module, configured to obtain a model to be evaluated and a quantization request for the model to be evaluated, and convert the quantization request for the model to be evaluated into an optimized quantization network structure request; the optimized quantization network structure request is to find a quantization network structure of the model to be evaluated that meets preset requirements;
[0036] An evaluation construction module, configured to determine the inference delay and network calculation amount generated by the model to be evaluated before and after model quantization according to the optimized quantization network structure request, and construct a quantization performance evaluation value according to the inference delay and network calculation amount consumption; the determination coefficient is used in the quantization performance evaluation value to evaluate the difference between the weights and feature maps of the model to be evaluated before and after quantization;
[0037] A model deployment module, configured to construct a feasible domain of the quantization network structure of the model and a search framework for searching for the quantization network structure with the optimal performance based on the optimized quantization network structure request; the search framework is to determine the quantization network structure with the optimal performance among each feasible quantization network structure by using a genetic algorithm constructed based on the quantization performance evaluation value.
[0038] According to a fourth aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the quantization performance evaluation method for model quantization as described in any one of the above are implemented.
[0039] According to a fourth aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the quantization performance evaluation method for model quantization as described in any one of the above are implemented.
[0040] According to a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the quantization performance evaluation method for model quantization as described in any one of the above are implemented.
[0041] The quantization performance evaluation method, device, equipment, and medium for model quantization of the present invention can systematically explore the impact of different quantization bit combinations on model performance by transforming the mixed quantization problem of the deep learning model into a problem of finding the optimal mixed quantization structure. Compared with the traditional trial-and-error method or empirical selection of quantization bits, this method based on the transformation of the optimization problem greatly improves the probability of finding the quantization bits that are most suitable for the model and can ensure the minimum performance loss; by comprehensively considering the inference latency and network computational amount before and after quantization to construct a quantization performance evaluation value, it can intuitively quantify the difference degree of the network before and after quantization, which helps to ensure that when selecting the quantization bit width, not only a single accuracy or speed index is concerned, but from the perspective of overall performance, it is ensured that the selected quantization bit width can achieve balanced optimization in multiple key performance dimensions; the construction of the quantization performance evaluation value provides a clear and effective evaluation basis for the automatic search framework, enabling the search process to more accurately screen out potential individual mixed quantization network structures. The breakpoint retention mechanism of the automatic search framework ensures the stability of the entire search process, saves a large amount of search time and computational resources, and further improves the search efficiency; in view of the problem of limited resources of edge devices, from the selection of quantization bits, the optimization of the mixed quantization structure to the construction of the automatic search framework, the hardware limitations such as the computing power and storage capacity of edge devices are fully considered, which enables the finally obtained mixed quantization network structure to be well adapted to edge devices and achieve efficient model deployment and operation on edge devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as imposing any limitation on the present invention. In the drawings:
[0043] Figure 1 A schematic flowchart of the quantization performance evaluation method for model quantization provided by the present invention is shown;
[0044] Figure 2 A schematic diagram of calculating the board-side inference latency of each layer of the network when the model to be evaluated adopts the YOLO algorithm is shown;
[0045] Figure 3Shows a schematic diagram when the search framework conducts a search in the quantization performance evaluation method for model quantization provided by the present invention;
[0046] Figure 4 Shows a schematic diagram of the genetic algorithm in the quantization performance evaluation method for model quantization provided by the present invention;
[0047] Figure 5 Shows a schematic diagram of the multi - point crossover operation in the quantization performance evaluation method for model quantization provided by the present invention;
[0048] Figure 6 Shows a schematic diagram when the individual decoder decodes in the quantization performance evaluation method for model quantization provided by the present invention;
[0049] Figure 7 Shows the convergence curve of the search process of the search framework in the quantization performance evaluation method for model quantization provided by the present invention;
[0050] Figure 8 Shows a comparison chart of the output of the deployed deep - learning model and the floating - point 32 - bit model in the quantization performance evaluation method for model quantization provided by the present invention;
[0051] Figure 9 Shows a schematic diagram of the structure of the quantization performance evaluation device for model quantization provided by the present invention;
[0052] Figure 10 Is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0054] Since the advent of deep learning technology, deep learning has achieved remarkable breakthroughs in many fields such as image classification, speech recognition, and natural language processing with its powerful representation learning and feature extraction capabilities. The successive emergence of representative networks such as YOLO, SSD, Faster R-CNN, Transformer, BERT, and DERT has greatly improved the accuracy of relevant detections required by users. However, in order to meet the requirements of generalization performance, the structural parameters of a single deep learning network are increasing, breaking through from millions to hundreds of millions, billions, or even tens of billions of parameters. The increase in the number of parameters will inevitably lead to a huge demand for and consumption of hardware performance, thus posing a huge challenge to the deployment of deep learning networks on resource-constrained edge devices.
[0055] To solve this problem, the quantization technology of deep learning models has emerged. This technology can effectively reduce the computational volume and storage requirements of the model, enabling larger-scale models to be more conveniently deployed to edge devices with limited computing power to meet various task requirements.
[0056] Currently, the commonly used model quantization methods usually quantize the Float32 data type into bit-width types such as Float16, INT16, INT8, and INT4. Of course, different quantization bit widths have different impacts on the performance of the quantized model. Low-bit-width quantization methods such as INT8 and INT4 can significantly reduce the model size and computational complexity, but they will also bring relatively large accuracy losses; high-bit-width quantization methods such as Float16 and INT16 have obvious advantages in reducing the accuracy loss of the quantized model. However, high-bit-width quantization will increase the storage space of the quantized model and increase the inference latency.
[0057] To make up for the deficiencies of the above-mentioned low-bit-width and high-bit-width quantization methods, the industry has introduced a hybrid quantization method, that is, the weights and activation values of the model are represented using different quantization bit depths to achieve a better balance.
[0058] Although the hybrid quantization method can balance the advantages and disadvantages of different quantization bit widths to a certain extent, it still faces many problems in actual applications:
[0059] First of all, how to select appropriate quantization bits to ensure model performance is a difficult problem. Different tasks and application scenarios have different requirements for model performance, and a quantization bit combination that can achieve the best balance among accuracy, computational complexity, and storage space needs to be found;
[0060] Secondly, the computational complexity and optimization efficiency of the hybrid quantization method need to be improved. Due to the use of different quantization bit depths, the computational process of the hybrid quantization method is relatively complex, which may lead to an extension of the training and inference time and affect the actual application effect of the model;
[0061] Finally, under resource - constrained conditions, how to accurately evaluate the performance of a quantized model is also a challenge. Due to the limited resources of edge devices, traditional evaluation methods may not accurately reflect the performance of the quantized model in such an environment.
[0062] Based on the above, how to provide a quantization performance evaluation method for model quantization is a technical problem that the industry urgently needs to solve.
[0063] To solve the above problems, this specification provides a quantization performance evaluation method for model quantization, aiming to provide a high - accuracy quantization performance evaluation method for quantized models based on a performance sampling strategy, to ensure the quality and performance of the quantized network obtained by the search framework, solve the quantization problem when deep - learning models are deployed on edge devices, and optimize the performance of deep - learning models in resource - constrained environments. The quantization performance evaluation method for model quantization provided in this specification can be applied to electronic devices. The electronic device can include laptops, desktop computers, smartphones, smart wearable devices (such as virtual - reality glasses, smart watches, etc.), tablet computers, etc. Of course, the quantization performance evaluation method for model quantization provided in this specification can also be applied within an application running on the above - mentioned electronic devices. Figure 1 is a schematic flowchart of the quantization performance evaluation method for model quantization according to an embodiment of the present invention, as Figure 1 shown, the method may include the following steps:
[0064] S10. Obtain the model to be evaluated and the quantization request for the model to be evaluated, and convert the quantization request for the model to be evaluated into an optimized quantization network structure request. Among them, the optimized quantization network structure request is to find a quantization structure of the model to be evaluated that meets the preset requirements, and the preset requirements are usually configured by the user to find the optimal quantization structure.
[0065] In this embodiment, by converting the quantization problem of the deep - learning algorithm into a problem of finding the optimal quantization structure to ensure obtaining the minimum quantization performance evaluation value F performance , it is possible to systematically explore the impact of different quantization - bit combinations on the model performance. Compared with the traditional trial - and - error method or empirical selection of quantization bits, this method based on the transformation of the optimization problem greatly improves the probability of finding the quantization bits that are most suitable for the model and can ensure the minimum performance loss. By transforming the high - resource - consumption problem of constructing (designing) the neural - network quantization structure of the deep - learning model into a low - resource - consumption search problem, the transformation from the quantization problem to an optimization problem that can be processed by optimization means is realized, that is, converting the model quantization request sent by the user into a quantization structure request for the model.
[0066] Through such settings, the difficulty of constructing the quantization structure of the deep learning model is reduced. At the same time, it can also lay a foundation for subsequent construction of quantization performance evaluation indicators, establishment of an automatic search framework, etc.
[0067] Generally, the optimization problem for m optimization objectives can be described as:
[0068] min[f 1 (x), f 1 (x), ……, f m (x)]
[0069]
[0070] where f i (x) represents the i-th objective function to be optimized; x represents the variable to be optimized; lb represents the lower bound constraint of the variable to be optimized; ub represents the upper bound constraint of the variable to be optimized; Aeq * x = beq represents the linear equality constraint of the variable to be optimized; A * x ≤ b represents the linear inequality constraint of the variable to be optimized.
[0071] S20. According to the optimization quantization network structure request, determine the inference latency and network computational complexity generated by the model to be evaluated before and after model quantization, and construct a quantization performance evaluation value based on the inference latency and network computational complexity. Among them, the coefficient of determination is used in the quantization performance evaluation value to evaluate the difference (change) between the weights and feature maps of the model to be evaluated before and after quantization.
[0072] In this embodiment, the inference latency (precision loss) and network computational complexity before and after deep learning network quantization are considered simultaneously, and an evaluation index of quantization performance, namely the quantization performance evaluation value F performance . Specifically, the coefficient of determination (R 2 ) is used to evaluate the changes in the weights and feature maps of the model to be evaluated before and after quantization. The larger the R 2 value, the smaller the difference before and after quantization. On the contrary, the smaller the R 2 value, the smaller the difference before and after quantization. R 2 is obtained by recording the inference latency of each layer of the model to be evaluated quantized to different bit widths in the inference latency list TS and constructing the network computational complexity list FPs to record the computational complexity of each layer of the network.
[0073] Please refer to Figure 2 , taking the YOLOv5 model as an example of the model to be evaluated, the inference latency in step S20 is determined through the following steps:
[0074] S201. Quantize each layer of the model to be evaluated to the first preset bit width.
[0075] S202. Quantize the first layer of the model to be evaluated from the first preset bit width to the second preset bit width. In this embodiment, the first preset bit width is configured as 8 bits (bit), and the second preset bit width is configured as 16 bits.
[0076] S203. Perform inference on the first layer of the network after quantization to the second preset bit width, determine the layer inference latency of the first layer of the network, and quantize the first layer of the network from the second preset bit width to the first preset bit width.
[0077] S204. Quantize the network of the next layer of the first layer of the network from the first preset bit width to the second preset bit width, perform inference on the next layer of the network after quantization to the second preset bit width, determine the layer inference latency of the next layer of the network, and quantize the next layer of the network from the second preset bit width to the first preset bit width. Repeat this process until the layer inference latency of each layer of the network in the model to be evaluated is determined. Aggregate the layer inference latencies of each layer of the network in the model to be evaluated to obtain the model inference latency TS of the model to be evaluated.
[0078] The network computational amount in step S20 is determined through the following steps:
[0079] S205. Quantize each layer of the network of the model to be evaluated to the first preset bit width. The specific content is as shown in step S201.
[0080] S206. Quantize the first layer of the network in the model to be evaluated from the first preset bit width to the second preset bit width. The specific content is as shown in step S202.
[0081] S207. Determine the layer computational amount of the first layer of the network after quantization to the second preset bit width, and quantize the first layer of the network from the second preset bit width to the first preset bit width.
[0082] S208. Quantize the network of the next layer of the first layer of the network from the first preset bit width to the second preset bit width, determine the layer computational amount of the next layer of the network, and quantize the next layer of the network from the second preset bit width to the first preset bit width. Repeat this process until the layer computational amount of each layer of the network in the model to be evaluated is determined. Aggregate the layer computational amounts of each layer of the network in the model to be evaluated to obtain the network computational amount FPs of the model to be evaluated.
[0083] Finally, the constructed quantization performance evaluation value F performance has the following expression:
[0084]
[0085] where F performance represents the quantization performance evaluation value; N represents the depth of the neural network of the model to be evaluated, that is, the number of network layers of the model to be evaluated; TS nDenote the inference latency of the n-th layer network of the model to be evaluated (N≥n≥0); TS i Denote the inference latency of the i-th layer network of the model to be evaluated (N≥i≥1); FPs n Denote the computational complexity of the n-th layer network of the model to be evaluated; Denote the coefficient of determination, which is used to evaluate the changes in weights and feature maps before and after quantization of the deep learning model; w n Denote the evaluation parameter before quantization of the n-th layer network of the model to be evaluated. In this embodiment, the evaluation parameter is a weight or a feature map vector; Denote the evaluation parameter after quantization of the n-th layer network of the model to be evaluated; γ denotes the first weight factor, and ε denotes the second weight factor. The specific values of these two weight factors can be configured by the user according to requirements.
[0086] The quantization performance evaluation value F constructed in this way performance Can intuitively quantify the difference degree of the model to be evaluated before and after model quantization, which helps to, when selecting the quantization bit width, not only focus on a single accuracy or speed index, but from the perspective of the overall performance of the model, ensure that the selected quantization bit width can achieve balanced optimization in multiple key performance dimensions, and avoid sacrificing other important performances of the model due to excessive pursuit of a certain index.
[0087] In this embodiment, each layer network in the model to be evaluated is quantized into different bit widths such as 16bit or 8bit, and then the corresponding inference latency of each layer is recorded and filled into the inference latency list, so as to statistically analyze the inference latency of each quantization bit width of the model to be evaluated, which can ensure that the evaluation index provides a more objective network inference latency evaluation result.
[0088] S30. Based on the optimized quantization network structure request, construct the feasible region of the quantization network structure of the model, that is, obtain all feasible quantization network structures of the model to be evaluated, and construct a search framework for searching the quantization network structure with the optimal performance. In this embodiment, this search framework is to use the genetic algorithm based on the quantization performance evaluation value F performance Construct the genetic algorithm to determine the feasible region of the quantization network structure, that is, the quantization network structure with the optimal performance among each feasible quantization network structure.
[0089] Please refer to Figure 3 , taking the YOLOv5 model as an example of the model to be evaluated for illustration. In this embodiment, an automatic search framework is built with the genetic algorithm as the core search algorithm, and the quantization performance evaluation value F performance As the evaluation function of the genetic algorithm, use the individual decoder to encode the quantization network structure parameters and use the quantization performance evaluation value F performanceEvaluate its performance. The automated search framework adopted in this embodiment has characteristics such as parallelism and stability. By using multi-thread technology, the search task is divided into multiple sub-tasks to significantly improve the search speed.
[0090] During the search process, the search may be interrupted due to various unexpected situations (such as equipment failure, power outage, etc.). The automated search framework has a breakpoint retention mechanism by recording the key information during the search process, such as the current population state, the number of generations of evolution, etc. When the search resumes, it can continue from the breakpoint, saving a large amount of search time and computing resources, further improving the search efficiency and ensuring the stability of the entire search process.
[0091] When the automated search framework meets the termination condition of the search, it outputs all the quantized network structures of the model to be evaluated. Based on the quantized performance evaluation value F performance Select the quantized network structure with the best performance according to the pros and cons, deploy it to the development board, and test the actual performance of all network structures at the edge device end, and select the optimal neural network structure.
[0092] Specifically, please refer to Figure 4 , step S30 includes:
[0093] S301. Based on the optimized quantized network structure request, obtain all feasible quantized network structures of the model to be evaluated, and encode each feasible quantized network structure to obtain the genetic algorithm encoding individuals of each feasible quantized network structure. Based on all the genetic algorithm encoding individuals, obtain the population P(t) and the number of generations of evolution of the population P(t).
[0094] S302. Determine whether the preset termination condition is met. In this embodiment, the preset termination condition is that the number of generations of evolution reaches the preset maximum number of generations T.
[0095] S303. Determine that the preset termination condition is met, and use the quantized performance evaluation value F perfarmance Determine the fitness of each genetic algorithm encoding individual in the population P(t), and decode the genetic algorithm encoding individual with the best fitness to obtain the quantized network structure with the best performance.
[0096] S304. Determine that the preset termination condition is not met, and use the quantized performance evaluation value F performance Determine the fitness of each genetic algorithm encoding individual in the population P(t).
[0097] S305. Apply the preset selection operator to the population P(t), and select the genetic algorithm encoding individuals with fitness exceeding the preset value from the population P(t) as the parent individuals.
[0098] S306. Randomly select two genetically encoded individuals from the parental individuals, and with a preset crossover probability P c Exchange the genes of the selected genetically encoded individuals to generate two new individuals after crossover. Repeat the operations of random selection and gene exchange until all the genetically encoded individuals of the parental individuals have completed gene exchange, that is, perform times of random selection and gene exchange operations (n represents the total number of parental individuals) to obtain a new population after crossover crosspop(t + 1).
[0099] S307. Apply a preset mutation operator to the population after crossover crosspop(t + 1) to obtain the population after mutation P(t + 1) and the generation number of the population after mutation P(t + 1), that is, the next-generation population P(t + 1) and the corresponding generation number. Repeat the above steps S304 to S307 until the generation number of the population reaches the preset maximum generation number T.
[0100] Please refer to Figure 5 , in this embodiment, the preset selection operator adopts the roulette wheel selection strategy, and the gene crossover process adopts multi-point crossover operation.
[0101] In this embodiment, an encoder-decoder framework is used. For the convenience of performing calculations in the automatic search framework, the individuals in the individual encoder genetic algorithm are encoded as genetically encoded individuals. Among them, the feasible domain of each coding position on each individual is the types of bit widths supported by the edge device. For example, if the edge device supports two quantization forms of 8bit and 16bit, then the feasible domain of the individual coding position is {0, 1}, corresponding to {8bit, 16bit} respectively.
[0102] Please refer to Figure 6 , when it is necessary to evaluate the performance of an individual, the individual decoder decodes the genetically encoded individual into a neural network quantization structure, and performs a simulation inference process once to use the quantization performance evaluation value F performamce to evaluate the performance of the current individual.
[0103] Please refer to Figure 7 and Figure 8 , through actual tests, this application can take into account the influence of factors such as the hardware characteristics and operating environment of the edge device on the model performance, avoid the situation of deployment failure or poor performance caused by the difference between the theoretical performance and the actual performance, and improve the practicability and reliability of the deep learning model on the edge device.
[0104] The quantization performance evaluation method for model quantization of the present invention can systematically explore the impact of different quantization bit combinations on the model performance by transforming the quantization problem of the deep learning model into a problem of finding the optimal quantization structure. Compared with the traditional trial-and-error method or empirical selection of quantization bits, this way based on the transformation of the optimization problem greatly improves the probability of finding the quantization bits that are most suitable for the model and can ensure the minimum performance loss. By comprehensively considering the inference latency and network computational amount before and after quantization to construct the quantization performance evaluation value, it can intuitively quantify the difference degree of the network before and after quantization, which helps to not only focus on a single accuracy or speed index when selecting the quantization bit width, but from the perspective of overall performance, ensure that the selected quantization bit width can achieve balanced optimization in multiple key performance dimensions. The construction of the quantization performance evaluation value provides a clear and effective evaluation basis for the automatic search framework, enabling the search process to more accurately screen out potential quantization network structure individuals. The breakpoint retention mechanism of the automatic search framework ensures the stability of the entire search process, saves a large amount of search time and computational resources, and further improves the search efficiency. Aiming at the problem of limited resources of edge devices, from the quantization bit selection, quantization structure optimization to the construction of the automatic search framework, the computing power, storage capacity and other hardware limitations of edge devices are fully considered, so that the finally obtained quantization network structure can be well adapted to edge devices and achieve efficient model deployment and operation on edge devices.
[0105] The quantization performance evaluation device for model quantization provided by the embodiments of the present invention will be described below. The quantization performance evaluation device for model quantization described below can be correspondingly referred to the quantization performance evaluation method for model quantization described above.
[0106] To solve the above problems, a quantization performance evaluation device for model quantization is provided in this specification, aiming to provide an efficient, low-cost, highly automated and well-scalable database performance optimization solution. Figure 9 It is a schematic structural diagram of the quantization performance evaluation device for model quantization according to the embodiments of the present invention. As Figure 9 shown, the device may include:
[0107] A quantization transformation module 10, configured to obtain a model to be evaluated and a quantization request for the model to be evaluated, and transform the quantization request for the model to be evaluated into an optimized quantization network structure request, where the optimized quantization network structure request is to find a quantization structure of the model to be evaluated that meets a preset requirement, and the preset requirement is usually configured by the user to find the optimal quantization structure.
[0108] In this embodiment, by transforming the quantization problem of the deep learning algorithm into a problem of finding the optimal quantization structure to ensure obtaining the minimum quantization performance evaluation value F performance, it can systematically explore the impact of different combinations of quantization bits on the model performance. Compared with the traditional trial-and-error method or empirical selection of quantization bits, this approach based on the transformation of optimization problems greatly improves the probability of finding the quantization bits that are most suitable for the model and can ensure the minimum performance loss. It transforms the construction (design) of the neural network quantization structure of the deep learning model, a high-resource-consuming problem, into a low-resource-consuming search problem, thereby realizing the transformation from the quantization problem to an optimization problem that can be processed by optimization means, that is, converting the model quantization request sent by the user into a quantization structure request of the model,
[0109] By such settings, the difficulty of constructing the quantization structure of the deep learning model is reduced. At the same time, it can also lay a foundation for subsequent construction of quantization performance evaluation indicators, building an automatic search framework, etc.
[0110] The evaluation and construction module 20 is used to determine the inference latency and network computational volume generated by the model to be evaluated before and after model quantization according to the optimized quantization network structure request, and construct a quantization performance evaluation value based on the inference latency and network computational volume. Among them, the coefficient of determination is used in the quantization performance evaluation value to evaluate the difference (change) situation of the weights and feature maps of the model to be evaluated before and after quantization.
[0111] In this embodiment, the inference latency and network computational volume before and after quantization of the deep learning network are considered simultaneously, and a quantization performance evaluation index, that is, the quantization performance evaluation value F, is constructed based on these two parameters. performance . Specifically, the coefficient of determination (R 2 ) is used to evaluate the change situation of the weights and feature maps of the model to be evaluated before and after quantization. The larger the R 2 value, the smaller the difference before and after quantization. On the contrary, the smaller the R 2 value, the smaller the difference before and after quantization. The R 2 is obtained by recording the inference latency of each layer of the model to be evaluated quantized to different bit widths in the inference latency list TS and constructing the computational volume list FPs to record the computational volume of each layer of the network.
[0112] In this embodiment, each layer of the network in the model to be evaluated is quantized to different bit widths, such as 16bit or 8bit, and then the corresponding inference latency of each layer is recorded and filled into the inference latency list, so as to statistically analyze the inference latency situation of each quantization bit width of the model to be evaluated, which can ensure that the evaluation index provides a more objective network inference latency evaluation result.
[0113] The model deployment module 30 is used to construct the feasible region of the quantization network structure of the model based on the optimized quantization network structure request, that is, obtain all feasible quantization network structures of the model to be evaluated, and construct a search framework for searching for the quantization network structure with the optimal performance. In this embodiment, this search framework is based on the quantization performance evaluation value F performanceThe constructed genetic algorithm determines the feasible region of the quantization network structure, that is, the quantization network structure with the optimal performance among all feasible quantization network structures.
[0114] The quantization performance evaluation device for model quantization of the present invention can systematically explore the impact of different quantization bit combinations on the model performance by transforming the quantization problem of the deep learning model into a problem of finding the optimal quantization structure. Compared with the traditional trial-and-error method or empirical selection of quantization bits, this method based on the transformation of the optimization problem greatly improves the probability of finding the quantization bits that are most suitable for the model and can ensure the minimum performance loss; by comprehensively considering the inference latency and network computational amount before and after quantization to construct the quantization performance evaluation value, it can intuitively quantify the difference degree before and after network quantization, which helps to, when selecting the quantization bit width, not only focus on a single accuracy or speed index, but from the perspective of overall performance, ensure that the selected quantization bit width can achieve balanced optimization in multiple key performance dimensions; the construction of the quantization performance evaluation value provides a clear and effective evaluation basis for the automatic search framework, enabling the search process to more accurately screen out potential quantization network structure individuals, and the breakpoint retention mechanism of the automatic search framework ensures the stability of the entire search process, saving a large amount of search time and computational resources, and further improving the search efficiency; aiming at the problem of limited resources of edge devices, from the selection of quantization bits, the optimization of quantization structure to the construction of the automatic search framework, the hardware limitations such as the computing power and storage capacity of edge devices are fully considered, which enables the finally obtained quantization network structure to be well adapted to edge devices and achieve efficient model deployment and operation on edge devices.
[0115] Figure 10 An example of the physical structure diagram of an electronic device is shown as Figure 10 shown. The electronic device may include: a processor 1010 (processor), a communication interface 1020 (Communications Interface), a memory 1030 (memory), and a communication bus 1040. Among them, the processor 1010, the communication interface 1020, and the memory 1030 complete mutual communication through the communication bus 1040. The processor 1010 can call the logical commands in the memory 1030 to execute the quantization performance evaluation method for model quantization, and the method includes:
[0116] Obtain the model to be evaluated and the quantization request for the model to be evaluated, and transform the quantization request for the model to be evaluated into an optimized quantization network structure request; the optimized quantization network structure request is to find the quantization structure that satisfies the preset requirements for the model to be evaluated;
[0117] According to the request for optimizing the quantized network structure, determine the inference latency and network computational complexity generated by the model to be evaluated before and after model quantization, and construct a quantization performance evaluation value based on the inference latency and network computational complexity; in the quantization performance evaluation value, the coefficient of determination is used to evaluate the differences in weights and feature maps of the model to be evaluated before and after quantization.
[0118] Based on the request for optimizing the quantized network structure, construct a feasible region of the quantized network structure of the model and a search framework for searching for the quantized network structure with the optimal performance; the search framework uses a genetic algorithm constructed based on the quantization performance evaluation value to determine the quantized network structure with the optimal performance in the feasible region of the quantized network structure.
[0119] In addition, when the logical instructions in the above-mentioned memory 1030 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0120] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the quantization performance evaluation method for model quantization provided by the above-mentioned various methods. The method includes:
[0121] Obtain the model to be evaluated and the quantization request for the model to be evaluated, and convert the quantization request for the model to be evaluated into a request for optimizing the quantized network structure; the request for optimizing the quantized network structure is to find a quantized structure of the model to be evaluated that meets the preset requirements.
[0122] According to the request for optimizing the quantized network structure, determine the inference latency and network computational complexity generated by the model to be evaluated before and after model quantization, and construct a quantization performance evaluation value based on the inference latency and network computational complexity; in the quantization performance evaluation value, the coefficient of determination is used to evaluate the differences in weights and feature maps of the model to be evaluated before and after quantization.
[0123] Based on the request for optimizing the quantized network structure, a feasible region of the quantized network structure of the model and a search framework for searching for the quantized network structure with the optimal performance are constructed; the search framework determines the quantized network structure with the optimal performance in the feasible region of the quantized network structure by using a genetic algorithm constructed based on the quantized performance evaluation value.
[0124] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the quantized performance evaluation method for model quantization provided above. The method includes:
[0125] Obtain the model to be evaluated and the quantization request for the model to be evaluated, and convert the quantization request for the model to be evaluated into a request for optimizing the quantized network structure; the request for optimizing the quantized network structure is to find the quantized structure of the model to be evaluated that meets the preset requirements.
[0126] According to the request for optimizing the quantized network structure, determine the inference latency and network computational amount generated by the model to be evaluated before and after model quantization, and construct a quantized performance evaluation value based on the inference latency and network computational amount; the coefficient of determination is used in the quantized performance evaluation value to evaluate the difference between the weights and feature maps of the model to be evaluated before and after quantization.
[0127] Based on the request for optimizing the quantized network structure, a feasible region of the quantized network structure of the model and a search framework for searching for the quantized network structure with the optimal performance are constructed; the search framework determines the quantized network structure with the optimal performance in the feasible region of the quantized network structure by using a genetic algorithm constructed based on the quantized performance evaluation value.
[0128] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A quantitative performance evaluation method for model quantization, characterized in that: The method comprises: Obtain the model to be evaluated and the quantization request for the model to be evaluated, and convert the quantization request for the model to be evaluated into a request for optimizing the quantization network structure; the request for optimizing the quantization network structure is to find a quantization structure of the model to be evaluated that meets preset requirements; According to the request for optimizing the quantization network structure, the inference delay and network calculation amount generated by the model to be evaluated before and after the model quantization are determined, and a quantization performance evaluation value is constructed according to the inference delay and the network calculation amount; the quantization performance evaluation value uses a determination coefficient to evaluate the difference between the weights and feature maps of the model to be evaluated before and after the quantization; Based on the request for optimizing the quantitative network structure, a feasible domain of the quantitative network structure of the model and a search framework for searching the quantitative network structure with the best performance are constructed; the search framework uses a genetic algorithm constructed based on the quantitative performance evaluation value to determine the quantitative network structure with the best performance in the feasible domain of the quantitative network structure.
2. The quantitative performance evaluation method of model quantization according to claim 1, characterized in that: The inference delay is determined by the following steps: Quantize each layer of the model to be evaluated to a first preset bit width; Quantize the first layer network in the model to be evaluated from the first preset bit width to the second preset bit width; the first preset bit width is configured as 8 bits, and the second preset bit width is configured as 16 bits; Performing inference of the first layer network after quantization to the second preset bit width, determining the layer inference delay of the first layer network, and quantizing the first layer network from the second preset bit width to the first preset bit width; quantizing a network of a layer below the first layer of the network from a first preset bit width to a second preset bit width, performing inference of the next layer of the network after quantization to the second preset bit width, determining a layer inference delay of the next layer of the network, and quantizing the next layer of the network from the second preset bit width to the first preset bit width, until the layer inference delay of each layer of the network in the model to be evaluated is determined; Aggregate the layer inference delay of each layer of the network in the model to be evaluated to obtain the model inference delay of the model to be evaluated.
3. The quantitative performance evaluation method of model quantization according to claim 2, characterized in that: The network computation amount is determined by the following steps: Quantize each layer of the model to be evaluated to a first preset bit width; Quantizing the first layer network in the model to be evaluated from the first preset bit width to the second preset bit width; Determine the layer calculation amount of the first layer network after quantization to the second preset bit width, and quantize the first layer network from the second preset bit width to the first preset bit width; Quantize the network of the next layer of the first layer of the network from the first preset bit width to the second preset bit width, determine the layer calculation amount of the next layer of the network, and quantize the next layer of the network from the second preset bit width to the first preset bit width, until the layer calculation amount of each layer of the network in the model to be evaluated is determined; The layer computation amount of each network layer in the model to be evaluated is aggregated to obtain the network computation amount of the model to be evaluated.
4. The quantitative performance evaluation method of model quantization according to claim 1, characterized in that: The method of constructing a feasible domain of a quantized network structure of a model to be evaluated and a search framework for searching a quantized network structure with optimal performance based on the request for optimizing the quantized network structure specifically includes: Based on the request for optimizing the quantitative network structure, all feasible quantitative network structures of the model to be evaluated are obtained, and each feasible quantitative network structure is encoded to obtain the genetic algorithm encoded individual of each feasible quantitative network structure. Based on all the genetic algorithm encoded individuals, the population and the evolutionary generation of the population are obtained; Determine whether the preset termination condition is met; the preset termination condition is that the evolutionary generation reaches the preset maximum evolutionary generation; It is determined that the preset termination conditions are met, and the fitness of each genetic algorithm encoded individual in the population is determined using the quantitative performance evaluation value, and the genetic algorithm encoded individual with the best fitness is decoded to obtain the quantitative network structure with the best performance.
5. The quantitative performance evaluation method of model quantization according to claim 4, characterized in that: The method of obtaining the feasible domain of the quantized network structure of the model to be evaluated based on the request for optimizing the quantized network structure and constructing a search framework for searching the quantized network structure with the best performance also includes: Determining that a preset termination condition is not met, and determining the fitness of each genetic algorithm encoding individual in the population using a quantitative performance evaluation value; Apply a preset selection operator to the population, and select genetic algorithm encoded individuals whose fitness exceeds a preset value from the population as parent individuals; Randomly select two genetic algorithm coded individuals from the parent individuals, and exchange the genes of the selected genetic algorithm coded individuals with a preset crossover probability to generate two new crossover individuals. Repeat the random selection and gene exchange operations until all genetic algorithm coded individuals of the parent individuals have completed gene exchange to obtain a new crossover population. The preset mutation operator is applied to the crossover population to obtain the mutated population and the evolutionary generations of the mutated population until the preset termination condition is met.
6. The quantitative performance evaluation method of model quantization according to claims 4 and 5, characterized in that: The search framework adopts an encoder-decoder framework, and a feasible quantization network structure is encoded by an individual encoder. The feasible domain of each encoding position on the genetic algorithm encoding individual is the bit width type supported by the edge device. When it is necessary to determine the fitness of the genetic algorithm encoding individual, the genetic algorithm encoding individual is decoded by an individual decoder and then the fitness is determined by using a quantization performance evaluation value.
7. The quantitative performance evaluation method of model quantization according to claim 5, characterized in that: The preset selection operator adopts the roulette wheel selection strategy, and the gene crossover process adopts the multi-point crossover operation.
8. A quantitative performance evaluation device for model quantization, characterized in that: The device comprises: A quantization conversion module is used to obtain the model to be evaluated and the quantization request for the model to be evaluated, and convert the quantization request for the model to be evaluated into an optimized quantization network structure request; the optimized quantization network structure request is to find a quantization network structure that satisfies preset requirements for the model to be evaluated; An evaluation construction module is used to determine the inference delay and network calculation amount generated by the model to be evaluated before and after model quantization according to the optimization quantization network structure request, and to construct a quantitative performance evaluation value according to the inference delay and network calculation amount consumption; the quantitative performance evaluation value uses a determination coefficient to evaluate the difference between the weights and feature maps of the model to be evaluated before and after quantization; The model deployment module is used to construct a feasible domain of the model's quantized network structure and a search framework for searching for the quantized network structure with the best performance based on the optimization quantized network structure request; the search framework uses a genetic algorithm constructed based on the quantized performance evaluation value to determine the quantized network structure with the best performance among various feasible quantized network structures.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the quantitative performance evaluation method of model quantization according to any one of claims 1 to 7 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the quantitative performance evaluation method of model quantization as claimed in any one of claims 1 to 7 are implemented.