A method for quantifying the number of parameters of a deep learning model
By building an intelligent search framework and machine learning algorithm, the problem that quantitative methods are difficult to optimize in multi-hardware environments is solved, efficient automatic tuning of quantitative configurations is achieved, and the inference efficiency of deep learning models on resource-constrained devices is improved.
Patent Information
- Application Number
- CN202510534212.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Existing quantization methods are difficult to flexibly adjust quantization strategies in multi-hardware environments, and cannot meet the performance optimization requirements of different hardware platforms and business needs while ensuring model accuracy. The traditional search methods are complex and time-consuming.
Build an intelligent search framework, combine traditional and machine learning algorithms, quickly evaluate and find the optimal quantitative configuration in multiple iterations, and support the verification of multiple hardware platforms through the ONNX Runtime reasoning engine to achieve automatic tuning of quantitative configurations.
It significantly reduces search time, improves search efficiency, simplifies the quantization process, and ensures that the model runs efficiently on different hardware platforms and meets the balance requirements of accuracy and performance.
Smart Images

Figure CN120046759B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly to a method for quantifying the number of parameters of a deep learning model. Background Art
[0002] With the rapid development of deep learning technology, deep neural networks (DNNs) have shown excellent performance in multiple fields, such as computer vision, natural language processing, and speech recognition tasks. However, although DNNs have achieved remarkable results in inference tasks, DNNs usually rely on a large amount of computational resources and storage overhead, resulting in many challenges in their actual deployment and applications. These challenges mainly include high computational costs, storage requirements, and energy consumption, especially in edge devices or embedded systems.
[0003] To address these issues and improve the inference efficiency of DNN models on resource-constrained devices, quantization technology has become an important research direction. Quantization maps high-precision floating-point weights and activation values to low-precision integer formats, thus significantly reducing the storage requirements and computational burden of the model. Quantization can not only effectively reduce the storage size of the model, lower the memory bandwidth requirements, but also improve the computational efficiency, especially on hardware accelerators (such as TPUs, GPUs, FPGAs, etc.), where it can make full use of acceleration instructions to boost the inference speed. Although quantization technology has been proven to enhance the inference efficiency of deep learning models, how to achieve efficient quantization while ensuring model accuracy remains a challenging problem.
[0004] ONNX Runtime is a high-performance, cross-platform inference engine developed by Microsoft, aiming to accelerate the inference of models in ONNX (Open Neural Network Exchange) format. ONNX is an open-source deep learning model exchange format that supports the export of models from many popular deep learning frameworks (such as TensorFlow, PyTorch, Keras, etc.) to the ONNX format. ONNX Runtime then provides an efficient inference execution environment to accelerate the deployment and inference of these models on various platforms. It has the ability of cross-platform support, supporting running on operating systems such as Windows, Linux, MacOS, etc., and can accelerate the inference process on various hardware (such as CPUs, GPUs, FPGAs, dedicated accelerators, etc.). Through efficient optimization techniques, it improves the model inference speed and supports multiple hardware accelerators, such as CUDA (NVIDIA GPU), DirectML (Windows GPU), TensorRT, etc.
[0005] Through the quantization interface provided by ONNX Runtime, users can flexibly specify multiple quantization parameters and configurations according to specific requirements, quantize different models, and verify the accuracy and performance of the quantized models at runtime. The quantization function of ONNXRuntime supports multiple quantization methods, including dynamic quantization, static quantization, and quantization-aware training (QAT), which can be optimized for different types of models and tasks. However, due to the large number of combinations of quantization configuration information, the hardware diversity of the target inference platform (such as CPUs, GPUs, TPUs, FPGAs, etc.), and the significant differences in the degree of quantization support across different platforms, existing quantization strategies are difficult to form a unified solution. In addition, different business models (such as computer vision, natural language processing, etc.) have different requirements for the accuracy and performance after quantization, which further increases the complexity of the quantization scheme. Existing methods cannot meet the performance optimization requirements of different hardware platforms and business needs while maintaining high accuracy. Therefore, how to flexibly adjust the quantization strategy in a multi-hardware environment and ensure that the quantized model can run efficiently and accurately on various devices remains an urgent challenge to be solved.
[0006] Existing quantization methods, especially post-training quantization (PTQ), face many challenges in terms of accuracy loss, training complexity, and hardware adaptation. To reduce the accuracy loss caused by quantization, existing techniques usually need to explore different combinations of quantization parameters to optimize the model and generate model data for the target platform through these configurations. However, the process of finding the optimal parameter configuration is not only complex and time-consuming, but also in the model inference stage, different quantization configurations will directly affect the inference performance, thus affecting the scheduling and execution efficiency of underlying operators and acceleration instructions. Especially in cross-hardware platform applications, the differences in hardware characteristics and acceleration instruction sets make the optimization process more cumbersome and challenging. In addition, the quantization configuration scheme of PTQ involves parameter combinations in multiple dimensions, resulting in an extremely large configuration space. For example, quantization methods (such as integer quantization, floating-point quantization, hybrid quantization), quantization precision (such as 8-bit, 16-bit), quantization awareness (such as symmetric quantization, asymmetric quantization), the selection of quantization operators, and the adaptation of acceleration instruction sets for different hardware platforms (such as CPUs, GPUs, TPUs, FPGAs) all further increase the complexity of the configuration. These factors together make the quantization optimization process both complex and challenging.
[0007] Facing such a large configuration space, traditional manual adjustment and search methods are particularly cumbersome and inefficient. Usually, a large number of experimental verifications and repeated optimizations are required to find the most suitable quantization scheme. Different quantization configurations have different effects on model accuracy, inference speed, and energy efficiency, and there are also significant differences in the adaptability of the hardware platform and algorithm optimization, further increasing the difficulty of optimization. The selection of each quantization configuration will directly affect the computational efficiency, memory usage, bandwidth requirements, and computational accuracy of the model during the inference stage. This makes it a core problem in the research of quantization technology to find the optimal solution in multi-dimensional configurations that can meet the needs of different hardware while ensuring the efficient inference and accuracy of the model. How to balance accuracy and computational efficiency and reduce computational overhead remains a major challenge for current quantization technology. Summary of the Invention
[0008] The present invention provides a method for quantifying the number of parameters of a deep learning model to solve the problem in the prior art that the accuracy and computational efficiency of the model cannot be balanced.
[0009] In a first aspect, the present invention provides a method for quantifying the number of parameters of a deep learning model, specifically including the following steps:
[0010] Step S1: Construct a calibration test data set DATA and a deep learning model M, divide the calibration test data set into calibration data data1 and test data data2, and pre-train the deep learning model M to form a pre-trained deep learning model M1;
[0011] Step S2: Input the calibration data data1, the pre-trained deep learning model M1, and the quantization configuration into the ONNX Runtime quantization framework for quantization processing to form a quantized deep learning model M2;
[0012] Step S3: Input the quantized deep learning model M2 and the test data data2 into the quantization model test accuracy evaluation module to form model configuration information and model accuracy;
[0013] Step S4: Process according to the model configuration information and the model accuracy through a configuration search engine and a quantization configuration manager to form a new quantization configuration;
[0014] Step S5: According to the new quantization configuration, execute Step S2 - Step S3 on the pre-trained deep learning model M1 to form the current model configuration information and the current model accuracy;
[0015] Step S6: Evaluate the model based on the current model configuration information, the current model accuracy, and the structure of the current model M2 to form performance index data of the current model; wherein, the model performance index data includes model inference latency, CPU memory occupancy, and CPU utilization rate.
[0016] Step S7: Repeat steps S4 - S6 to form multiple (in this application, the "multiple" means at least two) model accuracies and model performance index data.
[0017] Step S8: Output the optimal model configuration according to the multiple model accuracies and model performance index data.
[0018] Preferably, in step S4, the configured search engine is used to quantify the search for configuration parameter data, and the quantification of the search for configuration parameter data specifically includes the following steps:
[0019] Step S401: Obtain model configuration information and select a search level according to the model configuration information.
[0020] Step S402: According to the search level, perform a search for quantified configuration parameters to form quantified configuration parameter data.
[0021] Preferably, in step S401, the search levels include three search levels: O1, O2, and O3; when the number of model configuration parameters is less than the first threshold, the search level is the O1 level; when the number of model configuration parameters is greater than or equal to the first threshold and less than the second threshold, the search level is the O2 level; when the number of model configuration parameters is greater than the second threshold, the search level is the O3 level.
[0022] Preferably, in step S402, the search method at the O1 level includes grid search and / or random search; the search method at the O2 level includes search through a heuristic search algorithm; the search method at the O3 level includes search through a heuristic search algorithm and a machine learning algorithm.
[0023] More preferably, the search method at the O3 level includes search through a heuristic search algorithm and a machine learning algorithm, and specifically includes the following steps:
[0024] Step S402a: When the number of times of forming quantified configuration parameter data is less than 500, search for model configuration parameter data through a heuristic search algorithm, and use each model configuration parameter data and its corresponding model accuracy as a training data set to train a machine learning model to form a trained machine learning model.
[0025] Step S402b: When the number of times of forming quantization configuration parameter data is greater than or equal to 500, search through the trained machine learning model to form model configuration parameter data and predict the corresponding model accuracy; when the model accuracy meets the accuracy threshold, retain the formed model configuration parameter data; otherwise, discard the model configuration parameter data. When the number of times of forming quantization configuration parameter data reaches the maximum, stop the search.
[0026] Preferably, in step S4, the quantization configuration manager is a static quantization interface based on the ONNX Runtime inference engine, used for the definition of quantization configuration parameters, including the selection of quantization paradigms, the selection of the boolean type of per-channel configuration, the selection of quantization types for activation and weights respectively (such as uint8 or int8), the selection of clip boundaries for reduce range, the selection of calibration algorithms, and the specification of the layers of the model that need to be quantized (usually a sequence composed of convolutional layers, matrix layers, and other calculation layers based on empirical configurations). In addition, other quantization configurations and the optional ranges of their configuration items can be defined through extra options.
[0027] More preferably, the selection of quantization paradigms includes insertion quantization, dequantization operators, and operator-level quantization.
[0028] More preferably, the calibration algorithm includes one or a combination of more of the min-max algorithm, entropy algorithm, and percentile algorithm.
[0029] Preferably, step S6 specifically includes the following steps:
[0030] Step S601: Input the model configuration information, the model accuracy, and the structure of model M2 into the model analysis module to form a model evaluation instruction;
[0031] Step S602: Transmit the model evaluation instruction to the ONNX Runtime target platform performance evaluation module to form model performance index data.
[0032] Preferably, in step S6, the model performance index data forms a unique ID according to its corresponding quantization configuration, quantized model structure, and model accuracy, and is stored in the cache module. The cache module provides an interface for users to read the corresponding model performance index data through the ID.
[0033] Preferably, the formation of quantization configuration and the formation of model performance index data are in an asynchronous relationship.
[0034] Preferably, in step S8, according to the multiple model accuracies and model performance index data, output the optimal model configuration, specifically including the following steps:
[0035] Step S801: Obtain the performance index data with the precision difference within T before and after quantization, and according to the performance data standardization weighting formula, output the comprehensive performance scores corresponding to all model configurations;
[0036] Step S802: Output the optimal model configuration according to the comprehensive performance score.
[0037] Preferably, in step S801, the performance data standardization weighting formula is specifically expressed as follows:
[0038] ;
[0039] where score represents the comprehensive performance score; IL represents the inference latency; mem_usage represents the CPU memory occupancy; cpu_util represents the CPU utilization rate; W1, W2, and W3 represent weights. 、 、 ; ; where ; ; ; where IL max 、IL min 、mem_usage max 、mem_usage min 、cpu_util max 、cpu_util min respectively represent the maximum and minimum inference latencies, the maximum and minimum CPU memory occupancies, and the CPU utilization rates among all performance indexes with the precision difference within T.
[0040] In the second aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for quantifying the number of parameters of a deep learning model described in any item of the first aspect of the present application.
[0041] In the third aspect, the present invention also provides an electronic device, the electronic device includes: a memory storing a computer program: a processor communicatively connected to the memory, and when the computer program is called, it executes the method for quantifying the number of parameters of a deep learning model described in any item of the first aspect of the present application.
[0042] Compared with the prior art, the present invention has the following obvious prominent substantial features and remarkable advantages:
[0043] The present invention provides a method for quantifying the number of parameters of a deep learning model, which solves the problem in the prior art that it is impossible to balance the accuracy and computational efficiency of the model. Aiming at the problem of efficient search for quantization configurations in a huge and complex configuration space, multiple search levels are set for search, and a framework for automatic tuning of quantization configurations is constructed in combination with the target platform, aiming to accelerate the post-training quantization process of deep learning models. Traditional quantization optimization methods usually need to adjust quantization parameters through complex search to optimize the model accuracy, and this process is both time-consuming and cannot guarantee finding the optimal solution. The present invention constructs an intelligent search framework, and through traditional and machine learning algorithms, quickly evaluates and finds the optimal quantization configuration in multiple rounds of iteration. This method not only greatly reduces the search time, but also significantly improves the search efficiency on the premise of ensuring or approaching the optimal model accuracy. This method not only simplifies the traditional quantization process, but also realizes efficient quantization optimization on different hardware platforms, meeting the balance requirement between model accuracy and model performance. The present invention is applicable to scenarios where machine learning models need to be deployed and quantized, especially in scenarios that require cross-platform and cross-algorithm, such as edge computing, mobile devices, embedded systems and other fields.
[0044] Compared with traditional random search and grid search, the optimization efficiency of this framework is greatly improved, shortening the search time by about 10 times, while only resulting in an accuracy loss of less than 5%, proving its efficiency and effectiveness in the field of model compression. In addition, by virtue of the design characteristics of the backend execution engine of the ONNX Runtime inference engine, it can support the rapid verification of multiple hardware platforms, including CPUs, GPUs and dedicated accelerators, providing flexibility and convenience for model deployment. In short, quantization technology improves the inference efficiency of deep learning models on resource-constrained devices by significantly reducing computational and storage overheads. However, how to optimize the quantization process while ensuring accuracy remains a long-standing challenge. The solution proposed in this paper provides an innovative solution to this problem. Through automatic tuning and efficient search, it optimizes the parameter selection in the quantization process, reduces the search calculation cost, and demonstrates good adaptability and efficiency on multiple hardware platforms. The present invention is applicable to scenarios where machine learning models need to be deployed and quantized, especially in scenarios that require cross-platform and cross-algorithm, such as edge computing, mobile devices, embedded systems and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings that form a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0046] Figure 1 is a flowchart of a method for quantifying the number of parameters of a deep learning model according to a preferred embodiment of the present invention.
[0047] Figure 2 is a schematic diagram of the quantization process of a deep learning model according to a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0048] The present invention provides a method for quantifying the number of parameters of a deep learning model. To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0050] Example 1:
[0051] As Figures 1 - 2 shown, a method for quantifying the number of parameters of a deep learning model described in this embodiment specifically includes steps S1 to S8.
[0052] Step S1: Construct a calibration test data set DATA and a deep learning model M. Divide the calibration test data set into calibration data data1 and test data data2, and pre-train the deep learning model M to form a pre-trained deep learning model M1.
[0053] Step S2: Input the calibration data data1, the pre-trained deep learning model M1, and the quantization configuration into the ONNX Runtime quantization framework for quantization processing to form a quantized deep learning model M2.
[0054] Step S3: Input the quantized deep learning model M2 and the test data data2 into the quantized model test accuracy evaluation module to form model configuration information and model accuracy.
[0055] Step S4: Process according to the model configuration information and the model accuracy through a configuration search engine and a quantization configuration manager to form a new quantization configuration.
[0056] Preferably, in step S4, the quantization configuration manager is a static quantization interface based on the ONNX Runtime inference engine for defining quantization configuration parameters, including the selection of quantization paradigms (inserting quantization, dequantization operators, and operator-level quantization, etc.), the selection of the boolean type for per-channel configuration, the selection of quantization types for activations and weights respectively (such as uint8 or int8), the selection of clip boundaries for reduce range, the selection of calibration algorithms (such as a combination of one or more of the min-max algorithm, entropy algorithm, and percentile algorithm), and specifying the layers of the model that need to be quantized (usually a sequence composed of convolutional layers, matrix layers, and other calculation layers based on empirical configurations). In addition, other quantization configurations and the optional ranges of their configuration items can be defined through extra options.
[0057] Optionally, the configuration search engine is used to search for quantization configuration parameter data. The search for quantization configuration parameter data specifically includes steps S401 - S402.
[0058] Step S401: Obtain model configuration information and select a search level according to the model configuration information. The search levels include three search levels: O1, O2, and O3. When the number of model configuration parameters is less than the first threshold, the search level is the O1 level. When the number of model configuration parameters is greater than or equal to the first threshold and less than the second threshold, the search level is the O2 level. When the number of model configuration parameters is greater than the second threshold, the search level is the O3 level.
[0059] Step S402: Search for quantization configuration parameters according to the search level to form quantization configuration parameter data.
[0060] Among them, the search method at the O1 level includes grid search and / or random search; the search method at the O2 level includes searching through a heuristic search algorithm; the search method at the O3 level includes searching through a heuristic search algorithm and a machine learning algorithm. When the number of model configuration parameters is small, directly search through grid search and / or random search at this time. This search method can obtain the optimal quantization configuration parameter data. When the number of model configuration parameters is large, directly use the heuristic algorithm to search for quantization configuration parameter data at this time to further balance the search time and search accuracy (high search accuracy means that the error between the searched quantization configuration parameter data and the optimal quantization configuration parameter data is small).
[0061] In the specific implementation of this embodiment, the search method at the O3 level includes searching through a heuristic search algorithm and a machine learning algorithm, specifically including steps S402a - S402b.
[0062] Step S402a: When the number of times of forming the quantization configuration parameter data is less than 500, search the model configuration parameter data through a heuristic search algorithm, and use each model configuration parameter data and its corresponding model accuracy as a training data set to train a machine learning model (such as an XGBoost regression model) to form a trained machine learning model.
[0063] Step S402b: When the number of times of forming the quantization configuration parameter data is greater than or equal to 500, search through the trained machine learning model to form model configuration parameter data, and predict its corresponding model accuracy; when the model accuracy meets the accuracy threshold, retain the formed model configuration parameter data; otherwise, discard the model configuration parameter data. When the number of times of forming the quantization configuration parameter data reaches the maximum, stop the search.
[0064] The O3 level uses two different types of algorithms for segmented search, including the first stage and the second stage.
[0065] 1. In the first stage, use a heuristic search algorithm with a faster search speed but lower search accuracy to form model configuration parameter data, and use all the model configuration parameter data formed by the heuristic search algorithm and their corresponding model accuracies as the training data set of the machine learning model to train the machine learning model to form a trained machine learning model.
[0066] 2. In the second stage, use a machine learning algorithm with a slower search speed but higher search accuracy to search according to the trained machine learning model obtained in the first stage to form model configuration parameter data. At this time, the trained machine learning model also predicts according to the model accuracy corresponding to the model configuration parameter data obtained by its own search. If the predicted model accuracy is not within the preset model accuracy interval, directly discard the current model configuration parameter data. In this way, only the model configuration parameter data that meets the corresponding model accuracy within the preset model accuracy interval is retained, improving the processing efficiency of subsequent selection of the optimal model configuration parameter data.
[0067] Step S5: According to the new quantization configuration, perform Step S2 - Step S3 on the pre-trained deep learning model M1 to form the current model configuration information and the current model accuracy.
[0068] Step S6: Evaluate the model according to the current model configuration information, the current model accuracy, and the structure of the current model M2 to form the performance index data of the current model.
[0069] Among them, the model performance metric data includes model inference latency, CPU memory occupancy, and CPU utilization; latency, CPU memory occupancy, and CPU utilization are performance metrics of deep learning models, but they each focus on different aspects.
[0070] 1. Latency refers to the time required from input data to the model generating an output result. In application scenarios with high real-time requirements (such as autonomous driving), low latency is very crucial.
[0071] 2. CPU memory occupancy represents the memory size required for the model to run. This includes the space occupied by the model's own parameters and the space occupied by intermediate data generated during the calculation process. Memory occupancy is particularly important for resource-constrained environments, such as mobile devices or embedded systems.
[0072] 3. CPU utilization shows the proportion of the CPU's effective working time to the total time when processing tasks. High CPU utilization means that the hardware resources are fully utilized, but if the utilization is too high, it may also imply that the system is approaching the upper limit of its processing capacity, which may lead to performance bottlenecks.
[0073] These metrics are crucial for evaluating the actual deployment efficiency of a deep learning model. In addition to the above metrics, there are other important metrics, such as power consumption, GPU utilization, etc. Depending on the application scenario, the focus will also vary.
[0074] Optionally, in step S6, the model performance metric data forms a unique ID according to its corresponding quantization configuration, quantized model structure, and model accuracy and is stored in the cache module. The cache module provides an interface for users to read the corresponding model performance metric data through the ID.
[0075] Optionally, step S6 specifically includes steps S601 - S602.
[0076] Step S601: Input the model configuration information, the model accuracy, and the structure of model M2 into the model analysis module to form a model evaluation instruction.
[0077] Step S602: Transmit the model evaluation instruction to the ONNX Runtime target platform performance evaluation module to form model performance metric data.
[0078] In the specific implementation of this embodiment, the formation of the quantization configuration and the formation of the model performance metric data are in an asynchronous relationship.
[0079] Step S7: Repeatedly execute steps S4 - S6 to form multiple model accuracies and model performance metric data.
[0080] Step S8: Output the optimal model configuration based on the multiple model accuracies and model performance metric data.
[0081] Optionally, in step S8, it specifically includes the following steps:
[0082] Step S801: Obtain the performance metric data with the accuracy difference within T (in this embodiment, T is set to 2%) before quantization (fp32 accuracy) and after quantization (8-bit). According to the performance data standardization weighting formula, output the comprehensive performance scores corresponding to all model configurations.
[0083] Among them, the performance data standardization weighting formula is specifically expressed as follows:
[0084] ;
[0085] Among them, score represents the comprehensive performance score; IL represents the inference latency; mem_usage represents the CPU memory occupancy; cpu_util represents the CPU utilization rate; W1, W2, and W3 represent weights. 、 、 ; , in this embodiment, W1, W2, and W3 take values of 0.5, 0.25, and 0.25 respectively, and can be specifically adjusted according to task requirements. For example, if the real-time system focuses on low latency, W1 can be set relatively larger; among them, ; ; ; among them, IL max 、IL min 、mem_usage max 、mem_usage min 、cpu_util max 、cpu_util min respectively represent the maximum and minimum inference latencies, the maximum and minimum CPU memory occupancies, and the CPU utilization rates among all performance metrics with the accuracy difference within T. The above comprehensive performance score data can measure the inference performance differences within the specified accuracy tolerance range.
[0086] Step S802: Output the optimal model configuration according to the comprehensive performance score.
[0087] Through the above method, the advantages and disadvantages of models in different dimensions can be systematically measured, and the actual requirements can be flexibly adapted through custom weight configuration.
[0088] Based on the ONNX Runtime inference engine, the present invention proposes a method for quantifying the parameters of a deep learning model. In response to the challenge of efficiently searching for quantization configurations in a large and complex configuration space, an intelligent algorithm that combines heuristic search and machine learning is proposed, and a framework for automatic tuning of quantization configurations is constructed in combination with the target platform, aiming to accelerate the post-training quantization process of deep learning models. Traditional quantization optimization methods usually require complex searches to adjust quantization parameters to optimize model accuracy, and this process is both time-consuming and cannot guarantee finding the optimal solution. The present invention constructs an intelligent search framework that quickly evaluates and finds the optimal quantization configuration through traditional search methods and machine learning algorithms in multiple rounds of iteration. This method not only significantly reduces the search time, but also significantly improves the search efficiency while ensuring or approaching the optimal model accuracy. This method not only simplifies the traditional quantization process, but also achieves efficient quantization optimization on different hardware platforms, meeting the balance requirements between model accuracy and model performance.
[0089] Compared with traditional random search and grid search, the optimization efficiency of this framework is greatly improved, shortening the search time by about 10 times, while only resulting in an accuracy loss of less than 5%, proving its efficiency and effectiveness in the field of model compression. In addition, by virtue of the design characteristics of the backend execution engine of the ONNX Runtime inference engine, it can support rapid verification on multiple hardware platforms, including CPUs, GPUs, and dedicated accelerators, providing flexibility and convenience for model deployment. In short, quantization technology improves the inference efficiency of deep learning models on resource-constrained devices by significantly reducing computational and storage overheads. However, how to optimize the quantization process while ensuring accuracy remains a long-standing challenge. The solution proposed in this paper provides an innovative solution to this problem. By automatically tuning and efficiently searching, it optimizes the parameter selection in the quantization process, reduces the search computational cost, and demonstrates good adaptability and efficiency on multiple hardware platforms. The present invention is applicable to scenarios where machine learning models need to be deployed and quantified, especially in scenarios that require cross-platform and cross-algorithm applications, such as edge computing, mobile devices, embedded systems, and other fields.
[0090] The specific embodiments of the present invention have been described in detail above, but they are only examples, and the present invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.
Claims
1. A method for quantifying the number of parameters of a deep learning model, characterized in that, Specifically, it includes the following steps: Step S1: Construct a calibration test dataset DATA and a deep learning model M. Divide DATA into calibration data data1 and test data data2, and pre-train the deep learning model M to form M1; Step S2: Input data1, M1, and the quantization configuration into the ONNX Runtime quantization framework for quantization processing to form M2; Step S3: Input M2 and data2 into the quantization model test accuracy evaluation module to form model configuration information and model accuracy; Step S4: According to the model configuration information and model accuracy, process through the configuration search engine and the quantization configuration manager to form a new quantization configuration; Step S5: According to the new quantization configuration, execute Step S2 - Step S3 on M1 to form the current model configuration information and the current model accuracy; Step S6: According to the current model configuration information, the current model accuracy, and the structure of the current model M2, evaluate the model to form the performance index data of the current model; among them, the model performance index data includes model inference latency, CPU memory occupancy, and CPU utilization; Step S7: Repeat Steps S4 - S6 to form multiple model accuracies and model performance index data; Step S8: Output the optimal model configuration according to the multiple model accuracies and model performance index data; Among them, in Step S4, the configuration search engine is used for searching quantization configuration parameter data. The search for quantization configuration parameter data specifically includes the following steps: Step S401: Obtain the model configuration information and select the search level according to the model configuration information; Step S402: According to the search level, search for quantization configuration parameters to form quantization configuration parameter data; Among them, in Step S401, the search levels include three search levels: O1, O2, and O3; when the number of model configuration parameters is less than the first threshold, the search level is the O1 level; when the number of model configuration parameters is greater than or equal to the first threshold and less than the second threshold, the search level is the O2 level; when the number of model configuration parameters is greater than the second threshold, the search level is the O3 level.
2. A method for quantifying the number of parameters of a deep learning model according to claim 1, characterized in that, The search method at the O3 level includes searching through a heuristic search algorithm and a machine learning algorithm, specifically including the following steps: Step S402a: When the number of times of forming quantization configuration parameter data is less than 500, search for the model configuration parameter data through the heuristic search algorithm, and use each model configuration parameter data and its corresponding model accuracy as the training dataset to train the machine learning model to form the trained machine learning model; Step S402b: When the number of times of forming quantization configuration parameter data is greater than or equal to 500, search through the trained machine learning model to form model configuration parameter data and predict its corresponding model accuracy; when the model accuracy meets the accuracy threshold, retain the formed model configuration parameter data; otherwise, discard the model configuration parameter data. When the number of times of forming quantization configuration parameter data reaches the maximum, stop the search.
3. A method for quantifying the number of parameters of a deep learning model according to claim 1, characterized in that, In step S402, the search methods at the O1 level include grid search and / or random search; the search methods at the O2 level include searching through a heuristic search algorithm; the search methods at the O3 level include searching through a heuristic search algorithm and a machine learning algorithm.
4. A method for quantifying the number of parameters of a deep learning model according to claim 1, characterized in that, In step S4, the quantization configuration manager is a static quantization interface based on the ONNX Runtime inference engine, which is used for the definition of quantization configuration parameters, including the selection of quantization paradigms, the selection of the Boolean type for per-channel configuration, the selection of quantization types for activations and weights respectively, the selection of clip boundaries for reduce range, the selection of calibration algorithms, and the specification of the layers of the model that need to be quantized. In addition, other quantization configurations and the optional ranges of their configuration items can be defined through extra options.
5. A method for quantifying the number of parameters of a deep learning model according to claim 4, characterized in that The selection of the quantization paradigm includes insertion quantization, dequantization operators, and operator-level quantization; the calibration algorithms include one or a combination of more than one of the min-max algorithm, the entropy algorithm, and the percentile algorithm.
6. A method for quantifying the number of parameters of a deep learning model according to claim 1, characterized in that, Step S6 specifically includes the following steps: Step S601: Input the model configuration information, the model accuracy, and the structure of model M2 into the model analysis module to form a model evaluation instruction. Step S602: Transmit the model evaluation instruction to the ONNX Runtime target platform performance evaluation module to form model performance metric data.
7. A method for quantifying the number of parameters of a deep learning model according to claim 1, characterized in that, In step S6, the model performance metric data forms a unique ID according to its corresponding quantization configuration, the quantized model structure, and the model accuracy, and is stored in the cache module. The cache module provides an interface for users to read the corresponding model performance metric data through the ID.
8. A method for quantifying the number of parameters of a deep learning model according to claim 1, characterized in that, In step S8, according to the multiple model accuracies and the model performance metric data, the optimal model configuration is output, specifically including the following steps: Step S801: Obtain the performance metric data with the accuracy difference within T before and after quantization. According to the performance data standardization weighting formula, output the comprehensive performance scores corresponding to all model configurations. Step S802: Output the optimal model configuration according to the comprehensive performance scores.
9. A method for quantifying the number of parameters of a deep learning model according to claim 8, characterized in that, In step S801, the performance data standardization weighting formula is specifically expressed as follows: ; Among them, score represents the comprehensive performance score; IL represents the inference latency; mem_usage represents the CPU memory occupancy; cpu_util represents the CPU utilization rate; W1, W2, and W3 represent weights. , , ; ; Among them, ; ; ; Among them, IL max , IL min , mem_usage max , mem_usage min , cpu_util max , cpu_util min respectively represent the maximum and minimum inference latencies, the maximum and minimum CPU memory occupancies, and the CPU utilization rates among all performance metrics with a precision difference within T.
10. A method for quantifying the number of parameters of a deep learning model according to claim 1, characterized in that, The formation of the quantization configuration and the formation of the model performance metric data are in an asynchronous relationship.
Citation Information
Patent Citations
Video super-division method based on quantization after training
CN117274049A