Model Quantization Method, Apparatus and Terminal Device
By optimizing the quantization function of the quantization layer in the deep learning model, the problems of the quantization model in accuracy loss and calculation error are solved, and more efficient quantitative model optimization is achieved, which improves the overall performance of the model.
Patent Information
- Application Number
- CN202111264248.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-10-28
AI Technical Summary
In practical applications, deep learning models have insufficient hardware computing power due to large data volume and high computational complexity. At the same time, the quantization model has shortcomings in accuracy loss and calculation error.
The input data is processed through the floating point model, the target output is obtained, and the quantization process is performed according to the quantization function of the node to be quantized for each layer to be quantized. Then, the first input is processed through the corresponding quantization layer of the layer to be quantized, a second output is obtained, and the quantization function is optimized based on the second output and the first output, and finally quantize the floating-point model to obtain the target quantization model.
By hierarchically refine the quantization model, the accuracy of the quantization model is improved, the calculation error is reduced, and the accuracy of data processing is improved.
Smart Images

Figure CN114065913B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of model quantization, and particularly relates to a model quantization method, apparatus, terminal device, and computer-readable storage medium. Background Art
[0002] In recent years, artificial intelligence technology has developed rapidly and continuously penetrated into various application fields represented by computer vision, natural language processing, and speech recognition. However, in actual application scenarios, the huge data volume and computational complexity of deep learning models pose a huge challenge to hardware computing power. Therefore, quantization methods for deep learning models have also emerged. Quantization technology can reduce the memory footprint of neural network models, improve data throughput, and thus reduce inference latency. However, compared with the deep learning model before quantization, the quantized model usually introduces a large accuracy loss and increases computational errors. Summary of the Invention
[0003] In view of this, the embodiments of this application provide a model quantization method, apparatus, terminal device, and computer-readable storage medium, which can improve the accuracy of the quantized model.
[0004] In a first aspect, the embodiments of this application provide a model quantization method, including:
[0005] Processing input data through a floating-point model to obtain a target output, where the target output includes the first input and the first output of each quantized layer in the floating-point model when the floating-point model processes the input data, and each quantized layer includes at least one quantized node;
[0006] For each quantized layer, performing quantization processing on the corresponding quantized layer according to the quantization function of the quantized nodes in the corresponding quantized layer to obtain a quantized layer corresponding to the corresponding quantized layer;
[0007] Processing the first input of the corresponding quantized layer through the quantized layer corresponding to the corresponding quantized layer to obtain a second output;
[0008] Optimizing the quantization function corresponding to the corresponding quantized layer according to the second output and the first output to obtain a target quantization function corresponding to the corresponding quantized layer;
[0009] Quantizing the floating-point model according to the target quantization functions corresponding to each quantized layer to obtain a target quantized model.
[0010] In a second aspect, the embodiments of this application provide a model quantization apparatus, including:
[0011] The first processing module is used to process the input data through a floating-point model to obtain a target output. The target output includes the first input and the first output of each quantized layer in the floating-point model when the floating-point model processes the input data. Each quantized layer includes at least one quantized node;
[0012] The first quantization module is used to perform quantization processing on each quantized layer according to the quantization function of the quantized nodes in the corresponding quantized layer to obtain a quantized layer corresponding to the corresponding quantized layer;
[0013] The second processing module is used to process the first input of the corresponding quantized layer through the quantized layer corresponding to the corresponding quantized layer to obtain a second output;
[0014] The optimization module is used to optimize the quantization function corresponding to the corresponding quantized layer according to the second output and the first output to obtain a target quantization function corresponding to the corresponding quantized layer;
[0015] The second quantization module is used to quantize the floating-point model according to the target quantization function corresponding to each quantized layer to obtain a target quantized model.
[0016] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the model quantization method as in the first aspect is implemented.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the model quantization method as in the first aspect is implemented.
[0018] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is enabled to execute the model quantization method in the first aspect above.
[0019] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: In the embodiments of the present application, input data can be processed through a floating-point model to obtain a target output. The target output includes the first input and the first output of each quantization layer in the floating-point model when the floating-point model processes the input data. Each quantization layer includes at least one quantization node. For each quantization layer, according to the quantization function of the quantization nodes in the corresponding quantization layer, the corresponding quantization layer is quantized to obtain a quantization layer corresponding to the corresponding quantization layer; the first input of the corresponding quantization layer is processed through the quantization layer corresponding to the corresponding quantization layer to obtain a second output; according to the second output and the first output, the quantization function corresponding to the corresponding quantization layer is optimized to obtain a target quantization function corresponding to the corresponding quantization layer; according to the target quantization functions corresponding to each quantization layer, the floating-point model is quantized to obtain a target quantization model. At this time, for each quantization layer, based on the information of the second output of the corresponding quantization layer relative to the first output of the corresponding quantization layer, the output situation of the quantization layer in the quantization model can be more comprehensively known, so as to optimize the quantization functions in each quantization layer targeted, so as to achieve refined optimization of each layer of the quantization model, obtain a target quantization model that better meets expectations, improve the accuracy of the quantization model, and reduce calculation errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 is a schematic flowchart of an implementation of a model quantization method provided by an embodiment of the present application;
[0022] Figure 2 is a schematic diagram of optimizing a quantization layer provided by an embodiment of the present application;
[0023] Figure 3 is a schematic diagram of a model quantization device provided by an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0026] Before explaining the embodiments of the present application, some terms in the embodiments of the present application are briefly introduced first.
[0027] The embodiments of the present application are specifically described below.
[0028] The model quantization method provided by the embodiments of the present application can be applied to a terminal device.
[0029] Exemplarily, the terminal device can be a server, a desktop computer, a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of the present application do not impose any restrictions on the specific type of the terminal device.
[0030] Specifically, as Figure 1 shown, the model quantization method may include:
[0031] Step S101, process the input data through a floating-point model to obtain a target output. The target output includes the first input and the first output of each quantizable layer in the floating-point model when the floating-point model processes the input data. Each quantizable layer includes at least one quantizable node.
[0032] The floating-point model is a deep learning model. The specific structure of the floating-point model is not limited herein. In some examples, the floating-point model may include at least one of a convolutional layer, a pooling layer, a fully connected layer, and an activation layer. The floating-point model can be used for image processing or text processing. For example, the floating-point model can be used for object detection or classification of images, or for text recognition, etc.
[0033] Data such as the corresponding weights, input data, and output data in the floating-point model can be floating-point data. For example, they can be data of the float type or the double type.
[0034] The input data may include content such as images, videos, and text. The specific data format of the input data is not limited herein. Exemplarily, the input data may be a pixel matrix corresponding to an image, or may be a data format that meets the input requirements of the floating-point model after preprocessing the pixel matrix through preprocessing operations corresponding to the floating-point model. The preprocessing may include data processing operations such as denoising and normalization.
[0035] It should be noted that the number of input data is not limited herein. Exemplarily, the input data may include multiple image data. At this time, the floating-point model can process the multiple image data respectively, and during the processing, the processing data of the floating-point model for each image data can be obtained.
[0036] The specific format of the floating-point model can also be various. Exemplarily, the floating-point model can be an original model constructed based on a preset deep learning framework (such as frameworks like caffe, tensorflow, pytorch, mxnet, etc.); in addition, it can also be to convert the original model into a specified format to obtain the floating-point model, where the floating-point model can describe the operators, weight values, topological structures, etc. involved in the corresponding original model. The specified format can be determined based on the requirements of the actual application scenario and is not limited herein.
[0037] Generally, the floating-point model can include multiple layers, and at least one layer to be quantized can be included in the multiple layers. Each layer to be quantized can include at least one node to be quantized. The specific type of the layer to be quantized is not limited herein. Exemplarily, the layer to be quantized can include a convolutional layer, a fully connected layer, etc.
[0038] Among them, the specific position and attributes of the node to be quantized are not limited herein. The node to be quantized can include nodes in the processing steps of the floating-point model, such as input nodes and output nodes of a specific layer in the floating-point model, or can also include parameter configuration nodes in the floating-point model.
[0039] In the embodiments of the present application, the floating-point model can process the input data through forward propagation to obtain the processing data at each node to be quantized, which may include the first input and the first output of each layer to be quantized in the floating-point model. When processing the input data through the floating-point model, the processing data at each node to be quantized can be collected and recorded.
[0040] In addition, in one example, pseudo-quantization operations can also be inserted at each node to be quantized in the floating-point model. Through the pseudo-quantization operations, each node to be quantized in the floating-point model can be marked, so that in the subsequent model quantization process, the nodes to be quantized can be quickly located, and the quantization function at the nodes to be quantized can be efficiently determined.
[0041] It should be noted that in this example, when inserting the pseudo-quantization operation at each node to be quantized, it is not necessary to actually calculate the quantization parameters in the corresponding quantization function and perform quantization. Instead, it is used to label the nodes to be quantized for subsequent model quantization.
[0042] Step S102: For each layer to be quantized, according to the quantization function of the nodes to be quantized in the corresponding layer to be quantized, perform quantization processing on the corresponding layer to be quantized to obtain the corresponding quantized layer of the corresponding layer to be quantized.
[0043] In this embodiment, each layer to be quantized can be processed separately to obtain the corresponding quantized layer of each layer to be quantized. Alternatively, according to the quantization function of each node to be quantized, the floating-point model can be quantized to obtain a quantized model, and then the corresponding quantized layer of each layer to be quantized can be obtained in the quantized model.
[0044] In the embodiments of the present application, the specific form of the quantization function and the corresponding quantization parameters can be determined in advance by the user or can be determined based on information such as the first output.
[0045] It should be noted that in the embodiments of the present application, the quantization functions corresponding to each node to be quantized can be the same or different.
[0046] The structure of the quantized model is usually the same as that of the floating-point model. Therefore, the layer to be quantized in the floating-point model has a corresponding quantized layer in the quantized model, and the node to be quantized in the floating-point model has a corresponding quantized node in the quantized model.
[0047] In some embodiments, before, for each layer to be quantized, performing quantization processing on the corresponding layer to be quantized according to the quantization function of the nodes to be quantized in the corresponding layer to be quantized to obtain the corresponding quantized layer of the corresponding layer to be quantized, it further includes:
[0048] For each node to be quantized in the corresponding layer to be quantized, determine the quantization function of the corresponding node to be quantized according to the target processing data, where the target processing data includes the processing data corresponding to the node to be quantized when the floating-point model processes the input data.
[0049] Among them, there are various specific ways to determine the quantization function according to the target processing data. Exemplarily, based on the target processing data, the specified quantization parameter in the quantization function of the corresponding node to be quantized can be determined. In one example, the quantization parameter can be a scaling factor, which is used to describe the data scaling multiple during quantization.
[0050] In the embodiments of the present application, the corresponding quantization functions can be determined separately for each node to be quantized in the layer to be quantized, so that the quantization accuracy is higher and the performance of the quantized model is better when performing corresponding quantization operations on each node to be quantized.
[0051] In some embodiments, for each quantization node in the corresponding layer to be quantized, according to the target processing data, a quantization function for the corresponding quantization node is determined. The target processing data includes the processing data corresponding to the corresponding quantization node when the floating-point model processes the input data, including:
[0052] For each quantization node in the corresponding layer to be quantized, according to the target processing data corresponding to the corresponding quantization node and the target quantization data type corresponding to the corresponding quantization node, a quantization function for the corresponding quantization node is determined.
[0053] According to the target processing data corresponding to the corresponding quantization node, the value range of the data at the corresponding quantization node can be roughly estimated. And according to the target quantization data type corresponding to the corresponding quantization node, the range of the quantized data can be determined. Therefore, according to the target processing data corresponding to the corresponding quantization node and the target quantization data type corresponding to the corresponding quantization node, specified quantization parameters such as the scaling factor in the quantization function of the corresponding quantization node can be determined, so that the determined quantization function can meet the actual quantization requirements of the quantization node.
[0054] In some embodiments, the input data includes multiple groups of input sub-data. For each quantization node, the target processing data includes multiple groups of processing data corresponding to the quantization node, and the multiple groups of processing data correspond one-to-one to the multiple groups of input sub-data;
[0055] For each quantization node in the corresponding layer to be quantized, according to the target processing data corresponding to the corresponding quantization node and the target quantization data type corresponding to the corresponding quantization node, a quantization function for the corresponding quantization node is determined, including:
[0056] For each quantization node in the corresponding layer to be quantized, according to the first data range and the second data range, a quantization function for the corresponding quantization node is determined. The first data range is determined based on the maximum value and / or minimum value in the multiple groups of processing data corresponding to the corresponding quantization node, and the second data range is determined based on the value range of the target quantization data type corresponding to the corresponding quantization node.
[0057] In this embodiment, based on multiple groups of input sub-data, multiple groups of processing data corresponding to the corresponding quantization node can be obtained, so as to roughly estimate the change range of the data at the quantization node, that is, the first data range.
[0058] And based on the target quantization data type corresponding to the corresponding quantization node, the second data range corresponding to the corresponding quantization node can be determined. For example, if the target quantization data type is int8 type, the second data range corresponding to the int8 type is 128.
[0059] The following uses a specific example to illustrate an exemplary implementation manner of this embodiment.
[0060] In a specific example, the quantization function q corresponding to each node to be quantized can be:
[0061] q = clip(round(data / Δ))
[0062] Where Δ is a scaling factor used to scale the processing data data corresponding to the corresponding node to be quantized to a suitable quantization range. round(.) is a rounding operation that can convert a floating-point number to an integer. Clip(.) is a truncation operation that can limit the range of the quantized integer.
[0063] The scaling factor Δ can be determined based on the first data range and the second data range corresponding to the corresponding node to be quantized.
[0064] Specifically, the scaling factor Δ can be determined based on the following formula:
[0065]
[0066] Where threshold is the effective data threshold of the corresponding node to be quantized and can be determined based on the first data range. In some examples, the value of threshold can be the maximum value among multiple groups of processing data corresponding to the corresponding node to be quantized, while in other examples, the value of threshold can be the difference between the maximum value and the minimum value among multiple groups of processing data corresponding to the corresponding node to be quantized. And quanted_range can be the data range of the target quantization data type and can be determined based on the second data range.
[0067] Step S103, process the first input of the corresponding layer to be quantized through the quantization layer corresponding to the corresponding layer to be quantized to obtain a second output.
[0068] In this embodiment, the input of the quantization layer is the same as the input of the corresponding layer to be quantized, so that the error of the obtained second output relative to the first output does not introduce errors due to different inputs, reducing interference terms, thereby providing a good basis for subsequent optimization of relevant quantization functions.
[0069] The quantization layer can process the corresponding first input through forward propagation to obtain information such as the processing data of each node to be quantized in the corresponding layer to be quantized at the corresponding node in the quantization layer and the second output of the corresponding quantization layer.
[0070] Step S104, optimize the quantization function corresponding to the corresponding layer to be quantized according to the second output and the first output to obtain the target quantization function corresponding to the corresponding layer to be quantized.
[0071] In this embodiment, the first output and the second output can be compared to determine the error of the output of the corresponding quantization layer relative to the output of the corresponding layer to be quantized in the floating-point model. It can be seen that through the first output and the second output, the quantization loss of the quantization layer relative to the corresponding layer to be quantized can be determined, and then the quantization parameters of the quantization function of the nodes to be quantized in the corresponding layer to be quantized can be adjusted accordingly to specifically improve the precision loss of the corresponding quantization layer, thereby improving the precision loss of the entire quantization model.
[0072] There are various ways to determine the quantization loss of the quantization layer relative to the corresponding layer to be quantized through the first output and the second output. In some examples, for any layer to be quantized, the mean square error of the second output of the corresponding quantization layer of this layer to be quantized relative to the first output of this layer to be quantized is calculated to evaluate the precision loss of the corresponding quantization layer of this layer to be quantized.
[0073] In some examples, the quantization function of the nodes to be quantized in the layer to be quantized can be iteratively optimized according to the first output and the second output until the quantization loss of the corresponding quantization layer of this layer to be quantized is minimized (for example, the mean square error of the second output of the corresponding quantization layer of the corresponding layer to be quantized relative to the first output of this layer to be quantized is minimized), or until the number of iterations reaches a preset number. After the iterative optimization is completed, the target quantization function can be obtained.
[0074] In some embodiments, optimizing the quantization function corresponding to the corresponding layer to be quantized according to the second output and the first output to obtain the target quantization function corresponding to the corresponding layer to be quantized includes:
[0075] Optimizing the quantization function corresponding to the corresponding layer to be quantized according to the first output and the second output to obtain the corresponding target quantization function, and optimizing the weight values of at least one node to be quantized in the corresponding layer to be quantized to obtain the target weight values of at least one node to be quantized in the corresponding layer to be quantized;
[0076] Quantizing the floating-point model according to the target quantization functions corresponding to each layer to be quantized to obtain the target quantization model, including:
[0077] Obtaining the target quantization model according to the target quantization functions, target weight values corresponding to each layer to be quantized, and the floating-point model.
[0078] Wherein, by way of example, if the layer to be quantized is a convolutional layer, the weight values of at least one node to be quantized in this layer to be quantized can be the weight values corresponding to the convolutional kernels in this layer to be quantized. In addition, the nodes to be quantized corresponding to the weight values can also be other nodes in the floating-point model, which is not limited herein.
[0079] In this embodiment, not only can the quantization parameters in the quantization function be optimized, but also the weight values of at least one quantization node in the layer to be quantized can be optimized. Among them, the weight value of the quantization node can be regarded as part of the model configuration parameters in the actual application process, and optimizing the weight value of the quantization node can be regarded as fine-tuning the model configuration parameters of the quantization model. At this time, the adjustable factors are not limited to the quantization parameters of the quantization function, but also some model configuration parameters of the quantization model itself, so that the quantization model can be adjusted in multiple dimensions to further improve the overall performance of the quantization model.
[0080] In some embodiments, according to the first output and the second output, optimize the quantization function corresponding to the corresponding layer to be quantized to obtain the corresponding target quantization function, and optimize the weight values of at least one quantization node in the corresponding layer to be quantized to obtain the target weight values of at least one quantization node in the corresponding layer to be quantized, including:
[0081] Calculate the mean squared error of the second output relative to the first output;
[0082] According to the mean squared error, optimize the quantization function corresponding to the corresponding layer to be quantized to obtain the corresponding target quantization function, and optimize the weight values of at least one quantization node in the corresponding layer to be quantized to obtain the target weight values of at least one quantization node in the corresponding layer to be quantized.
[0083] In the embodiments of the present application, the quantization loss of the corresponding quantization layer can be evaluated according to the mean squared error of the output of the corresponding quantization layer relative to the output of the layer to be quantized, so as to optimize the quantization function and specific weight values of the layer to be quantized. Among them, in one example, the mean squared error can be used as the corresponding precision loss value, or the mean squared error can be square-rooted to obtain the root mean squared error, and the root mean squared error can be used as the corresponding precision loss value.
[0084] The mean squared error (MSE) can be calculated based on the following formula:
[0085]
[0086] Among them, Q(·) refers to the quantization function, F(·) is the operation on the input of the corresponding layer to be quantized, W is the weight value of at least one quantization node in the layer to be quantized, and X is the processing data of other quantization nodes in the layer to be quantized. Among them, the processing data corresponding to other quantization nodes may not be weight values. For example, the processing data corresponding to other quantization nodes may be the input of the corresponding layer to be quantized or other data.
[0087] Such as Figure 2As shown, it is an exemplary schematic diagram for iterative optimization of the layer to be quantized. Among them, the iterative optimization of the quantization layer A is exemplarily described.
[0088] Exemplarily, the input of the quantization layer A is the same as the input of the corresponding quantization layer A', both of which are the first input of the quantization layer A.
[0089] After obtaining the second output by processing the first input through the quantization layer A', the mean square error of the second output of the quantization layer A' relative to the first output of the quantization layer A can be calculated. If the mean square error does not meet the preset condition, then the quantization function and the weight value corresponding to the quantization layer A are updated in reverse according to the mean square error;
[0090] After updating the quantization function and the weight value corresponding to the quantization layer A in reverse, the quantization layer A is re-quantized according to the updated quantization function and weight value, and the updated layer A' can be obtained.
[0091] Then, the precision loss of the updated layer A' can be further evaluated. Specifically, the initial input of the layer A' can be used as the input of the updated layer A' to obtain the output of the updated layer A', and then the mean square error of the updated layer A' is recalculated to determine whether the optimization is completed according to the mean square error of the updated layer A'. The subsequent steps can be analogously obtained based on the above steps. After the optimization is completed, the corresponding layer to be quantized can be processed based on the target quantization function and the target weight value to obtain the target quantization layer corresponding to the layer to be quantized.
[0092] At this time, compared with the entire quantization model, optimizing each layer to be quantized separately reduces the amount of data for a single optimization, and improves the fineness of optimizing the quantization model, improves the optimization quality, and thus can improve the precision of the finally obtained target quantization model.
[0093] In addition, since the first input of each layer to be quantized can be calculated in advance, the optimizations of each layer can be considered independent. Therefore, the optimizations of each layer to be quantized can be performed in parallel, improving the data processing efficiency.
[0094] In some embodiments, according to the mean square error, optimizing the quantization function corresponding to the corresponding layer to be quantized to obtain the corresponding target quantization function, and optimizing the weight value of at least one quantization node in the corresponding layer to be quantized to obtain the target weight value of at least one quantization node in the corresponding layer to be quantized, includes:
[0095] According to the mean square error, iteratively optimize the quantization function corresponding to the corresponding layer to be quantized and the weight value of at least one quantization node in the corresponding layer to be quantized to obtain the corresponding target quantization function and the corresponding target weight value;
[0096] Among them, in the process of each iteration;
[0097] If the mean squared error corresponding to the current iteration meets the preset condition, or the number of iterations is not less than the preset number, then the quantization function and weight value corresponding to the current iteration are used as the corresponding target quantization function and target weight value;
[0098] Otherwise, based on the mean squared error corresponding to the current iteration, update the quantization function and weight value corresponding to the current iteration, and re - execute the step of quantifying the corresponding layer to be quantified according to the quantization function of the nodes to be quantified in the corresponding layer to be quantified, and obtain the quantization layer corresponding to the corresponding layer to be quantified and subsequent steps.
[0099] In this embodiment, the preset condition can be that the mean squared error corresponding to the current iteration converges to a preset threshold, or it can make the mean squared error after iterative optimization reach the minimum, that is
[0100] (Δ′ X ,Δ′ W ,W′)=argmin(MSE)
[0101] Wherein, Δ′ X is the target quantization function corresponding to the node to be quantified such as the input node of the layer to be quantified, Δ′ W is the target quantization function of the node to be quantified corresponding to the weight value, and W′ is the target weight value. In some examples, the target quantization function and target weight value can be obtained through iterative optimization by means of back - optimization algorithms in frameworks such as tensorflow and pytoch.
[0102] Step S105, quantize the floating - point model according to the target quantization function corresponding to each layer to be quantified, and obtain the target quantization model.
[0103] In the embodiment of the present application, after obtaining the target quantization function, the corresponding layer to be quantified can be quantized through the target quantization function to obtain the target quantization layer corresponding to the layer to be quantified. After obtaining the target quantization layers corresponding to each layer to be quantified, the target quantization model can be obtained according to the non - quantization layers (i.e., the layers that do not require quantization operations) and the target quantization layers in the floating - point model.
[0104] The function of the target quantization model is the same as that of the corresponding floating-point model. That is to say, if the floating-point model is used for image processing (for example, object detection or classification of images), then the target quantization model can also be used for image processing; and if the floating-point model is used for text processing (such as text recognition), then the target quantization model can also be used for text recognition. In the embodiments of the present application, compared with the terminal device deployed with the floating-point model, the terminal device deployed with the target quantization model has a faster data processing speed for image processing or text recognition through the deployed target quantization model, consumes less storage resources and computing resources, and at the same time, can also ensure a high data accuracy and a small calculation error.
[0105] In some embodiments, after obtaining the target quantization model, the specified data to be processed can be processed by the target quantization model to obtain a data processing result.
[0106] Exemplarily, the data to be processed can be image data. According to the function of the target quantization model, the data processing result can be the classification result or object detection result of the image data, etc.; or, the data to be processed can be text data, and the data processing result can be the text recognition result, etc. Compared with the prior art, the target quantization model optimization in this solution reduces the accuracy loss and calculation error of the data, thereby improving the accuracy of the data processing result. When the data to be processed is image data and the data processing result is the classification result or object detection result of the image data, the accuracy of the classification result or object result of the image data can be improved due to the reduction of the accuracy loss and calculation error in data processing in the model. When the data to be processed is text data and the data processing result is the text recognition result, the accuracy of the text recognition result can be improved due to the reduction of the accuracy loss and calculation error in data processing in the model.
[0107] It can be seen that in the embodiments of the present application, for each quantization layer, based on the information of the second output of the corresponding quantization layer relative to the first output of the corresponding layer to be quantized, the output situation of the quantization layer in the quantization model can be comprehensively known, so as to optimize the quantization function in each quantization layer in a targeted manner, so as to achieve fine-grained optimization of the quantization model layer by layer, obtain a target quantization model that better meets the expectations, improve the accuracy of the quantization model, and reduce the calculation error.
[0108] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0109] Corresponding to the model quantization method described above in the foregoing embodiments, Figure 3The structure block diagram of a model quantization device provided by an embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.
[0110] Referring to Figure 3 , the model quantization device 3 includes:
[0111] A first processing module 301, configured to process input data through a floating-point model to obtain a target output. The target output includes the first input and the first output of each quantization layer in the floating-point model when the floating-point model processes the input data. Each quantization layer includes at least one quantization node;
[0112] A first quantization module 302, configured to perform quantization processing on each quantization layer according to the quantization function of the quantization nodes in the corresponding quantization layer to obtain a quantization layer corresponding to the corresponding quantization layer;
[0113] A second processing module 303, configured to process the first input of the corresponding quantization layer through the quantization layer corresponding to the corresponding quantization layer to obtain a second output;
[0114] An optimization module 304, configured to optimize the quantization function corresponding to the corresponding quantization layer according to the second output and the first output to obtain a target quantization function corresponding to the corresponding quantization layer;
[0115] A second quantization module 305, configured to quantize the floating-point model according to the target quantization functions corresponding to each quantization layer to obtain a target quantization model.
[0116] Optionally, the optimization module 304 is specifically configured to:
[0117] Optimize the quantization function corresponding to the corresponding quantization layer according to the first output and the second output to obtain a corresponding target quantization function, and optimize the weight values of at least one quantization node in the corresponding quantization layer to obtain the target weight values of at least one quantization node in the corresponding quantization layer;
[0118] The second quantization module 305 is specifically configured to:
[0119] Obtain a target quantization model according to the target quantization functions, target weight values corresponding to each quantization layer, and the floating-point model.
[0120] Optionally, the optimization module 304 specifically includes:
[0121] A calculation unit, configured to calculate the mean square error of the second output relative to the first output;
[0122] An optimization unit is configured to optimize the quantization function corresponding to a corresponding layer to be quantized according to the mean square error to obtain a corresponding target quantization function, and optimize the weight values of at least one quantized node in the corresponding layer to be quantized to obtain the target weight values of at least one quantized node in the corresponding layer to be quantized.
[0123] Optionally, the optimization unit is specifically configured to:
[0124] Iteratively optimize the quantization function corresponding to the corresponding layer to be quantized and the weight values of at least one quantized node in the corresponding layer to be quantized according to the mean square error to obtain a corresponding target quantization function and a corresponding target weight value;
[0125] Wherein, in the process of each iteration;
[0126] If the mean square error corresponding to this iteration meets a preset condition, or the number of iterations is not less than a preset number, then use the quantization function and weight value corresponding to this iteration as the corresponding target quantization function and target weight value;
[0127] Otherwise, based on the mean square error corresponding to this iteration, update the quantization function and weight value corresponding to this iteration, and re - execute the steps of quantizing the corresponding layer to be quantized according to the quantization function of the quantized nodes in the corresponding layer to be quantized to obtain the quantization layer corresponding to the corresponding layer to be quantized and subsequent steps.
[0128] Optionally, the model quantization device 3 further includes:
[0129] A determination module, configured to determine the quantization function of each quantized node in the corresponding layer to be quantized according to target processing data, where the target processing data includes the processing data corresponding to the quantized node when the floating - point model processes the input data.
[0130] Optionally, the determination module is specifically configured to:
[0131] For each quantized node in the corresponding layer to be quantized, determine the quantization function of the corresponding quantized node according to the target processing data corresponding to the corresponding quantized node and the target quantization data type corresponding to the corresponding quantized node.
[0132] Optionally, the input data includes multiple groups of input sub - data. For each quantized node, the target processing data includes multiple groups of processing data corresponding to the quantized node, and the multiple groups of processing data correspond one - to - one with the multiple groups of input sub - data;
[0133] The determination module is specifically configured to:
[0134] For each quantization node in the corresponding layer to be quantized, a quantization function for the corresponding quantization node is determined according to a first data range and a second data range. The first data range is determined based on the maximum value and / or minimum value in multiple sets of processed data corresponding to the corresponding quantization node, and the second data range is determined based on the value range of the target quantization data type corresponding to the corresponding quantization node.
[0135] In an embodiment of the present application, the input data can be processed by a floating-point model to obtain a target output. The target output includes the first input and the first output of each layer to be quantized in the floating-point model when the floating-point model processes the input data. Each layer to be quantized includes at least one quantization node. For each layer to be quantized, the corresponding layer to be quantized is quantized according to the quantization function of the quantization node in the corresponding layer to be quantized, and a quantized layer corresponding to the corresponding layer to be quantized is obtained; the first input of the corresponding layer to be quantized is processed through the quantized layer corresponding to the corresponding layer to be quantized to obtain a second output; according to the second output and the first output, the quantization function corresponding to the corresponding layer to be quantized is optimized to obtain a target quantization function corresponding to the corresponding layer to be quantized; the floating-point model is quantized according to the target quantization functions corresponding to each layer to be quantized to obtain a target quantization model. At this time, for each quantized layer, based on the information of the second output of the corresponding quantized layer relative to the first output of the corresponding layer to be quantized, the output situation of the quantized layer in the quantization model can be comprehensively known, so as to optimize the quantization functions in each quantized layer in a targeted manner, so as to achieve refined optimization of the layers of the quantization model, obtain a target quantization model that better meets expectations, improve the accuracy of the quantization model, and reduce calculation errors.
[0136] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units, due to being based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part for details, and will not be elaborated here.
[0137] Figure 4 is a schematic diagram of a terminal device provided by an embodiment of the present application. As Figure 4 shown, the terminal device 4 in this embodiment includes: a processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the processor 40. When the processor 40 executes the computer program 42, the steps in the above-mentioned various model quantization method embodiments are implemented, such as Figure 1 the steps S101 to S105 shown. Alternatively, when the processor 40 executes the computer program 42, the functions of each module / unit in the above-mentioned various device embodiments are implemented, such as Figure 3 the functions of the modules 301 to 303 shown.
[0138] Exemplarily, the computer program 42 can be divided into one or more modules / units. One or more modules / units are stored in the memory 41 and executed by the processor 40 to complete the embodiments of this application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 42 in the terminal device 4. For example, the computer program 42 can be divided into a first processing module, a first quantization module, a second processing module, an optimization module, and a second quantization module. The specific functions of each module are as follows:
[0139] The first processing module is used to process the input data through a floating-point model to obtain a target output. The target output includes the first input and the first output of each quantizable layer in the floating-point model when the floating-point model processes the input data. Each quantizable layer includes at least one quantizable node;
[0140] The first quantization module is used to perform quantization processing on each quantizable layer according to the quantization function of the quantizable nodes in the corresponding quantizable layer to obtain the quantized layer corresponding to the corresponding quantizable layer;
[0141] The second processing module is used to process the first input of the corresponding quantizable layer through the quantized layer corresponding to the corresponding quantizable layer to obtain a second output;
[0142] The optimization module is used to optimize the quantization function corresponding to the corresponding quantizable layer according to the second output and the first output to obtain the target quantization function corresponding to the corresponding quantizable layer;
[0143] The second quantization module is used to quantize the floating-point model according to the target quantization functions corresponding to each quantizable layer to obtain a target quantized model.
[0144] The terminal device 4 can be a computing device such as a wearable device, a desktop computer, a notebook, a handheld computer, and a cloud server. The terminal device may include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art can understand that Figure 4 merely examples of the terminal device 4, which do not constitute a limitation on the terminal device 4. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.
[0145] The so-called processor 40 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0146] The memory 41 may be an internal storage unit of the terminal device 4, such as the hard disk or memory of the terminal device 4. The memory 41 may also be an external storage device of the terminal device 4, such as a plug-in hard disk equipped on the terminal device 4, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 41 may also include both the internal storage unit and the external storage device of the terminal device 4. The memory 41 is used to store computer programs and other programs and data required by the terminal device. The memory 41 may also be used to temporarily store data that has been output or is to be output.
[0147] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0148] The embodiment of the present application provides a computer program product. When the computer program product runs on the terminal device, the terminal device can implement the steps in the above-mentioned various method embodiments when executed.
[0149] When the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The above computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the above computer program includes computer program code, and the above computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The above computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, USB flash drive, mobile hard disk, magnetic disk or optical disc, etc. In some jurisdictions, according to legislation and patent practice, computer-readable media may not be electrical carrier signals and telecommunication signals.
[0150] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0151] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0152] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the above division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0153] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0154] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A model quantization method, where the model is used for object detection or classification of images, and is characterized in that, Including: Processing input data through a floating-point model to obtain a target output, where the target output includes the first input and the first output of each quantizable layer in the floating-point model when the floating-point model processes the input data, and each quantizable layer includes at least one quantizable node; wherein, the input data is image data; For each quantizable layer, performing quantization processing on the corresponding quantizable layer according to the quantization function of the quantizable nodes in the corresponding quantizable layer to obtain a quantized layer corresponding to the corresponding quantizable layer; Processing the first input of the corresponding quantizable layer through the quantized layer corresponding to the corresponding quantizable layer to obtain a second output; Optimizing the quantization function corresponding to the corresponding quantizable layer according to the second output and the first output to obtain a target quantization function corresponding to the corresponding quantizable layer; wherein, according to the first output and the second output, iteratively optimizing the quantization function of the quantizable nodes in the quantizable layer until the quantization loss of the corresponding quantized layer of the quantizable layer is minimized; Quantizing the floating-point model according to the target quantization functions corresponding to each quantizable layer to obtain a target quantized model; Wherein, at each of the quantizable nodes of the floating-point model, a pseudo-quantization operation is inserted. When the pseudo-quantization operation is inserted, it is not necessary to calculate the quantization parameters in the corresponding quantization function and perform quantization, but it is used to label the quantizable node.
2. The model quantization method according to claim 1, characterized in that, The optimizing the quantization function corresponding to the corresponding quantizable layer according to the second output and the first output to obtain a target quantization function corresponding to the corresponding quantizable layer includes: Optimizing the quantization function corresponding to the corresponding quantizable layer according to the first output and the second output to obtain a corresponding target quantization function, and optimizing the weight values of at least one quantizable node in the corresponding quantizable layer to obtain the target weight values of at least one quantizable node in the corresponding quantizable layer; The quantizing the floating-point model according to the target quantization functions corresponding to each quantizable layer to obtain a target quantized model includes: Obtaining a target quantized model according to the target quantization functions, target weight values corresponding to each quantizable layer and the floating-point model.
3. The model quantization method according to claim 2, characterized in that, The optimizing the quantization function corresponding to the corresponding quantizable layer according to the first output and the second output to obtain a corresponding target quantization function, and optimizing the weight values of at least one quantizable node in the corresponding quantizable layer to obtain the target weight values of at least one quantizable node in the corresponding quantizable layer includes: Calculating the mean square error of the second output relative to the first output; Optimizing the quantization function corresponding to the corresponding quantizable layer according to the mean square error to obtain a corresponding target quantization function, and optimizing the weight values of at least one quantizable node in the corresponding quantizable layer to obtain the target weight values of at least one quantizable node in the corresponding quantizable layer.
4. The model quantization method according to claim 3, characterized in that, Optimizing the quantization function corresponding to the corresponding quantization layer according to the mean square error to obtain the corresponding target quantization function, and optimizing the weight values of at least one quantization node in the corresponding quantization layer to obtain the target weight values of at least one quantization node in the corresponding quantization layer, including: Iteratively optimizing the quantization function corresponding to the corresponding quantization layer and the weight values of at least one quantization node in the corresponding quantization layer according to the mean square error to obtain the corresponding target quantization function and the corresponding target weight values; Wherein, in the process of each iteration; If the mean square error corresponding to the current iteration meets the preset condition, or the number of iterations is not less than the preset number of times, then use the quantization function and weight values corresponding to the current iteration as the corresponding target quantization function and target weight values; Otherwise, based on the mean square error corresponding to the current iteration, update the quantization function and weight values corresponding to the current iteration, and re-execute the steps of quantizing the corresponding quantization layer according to the quantization function of the quantization node in the corresponding quantization layer to obtain the quantization layer corresponding to the corresponding quantization layer and subsequent steps.
5. The model quantization method according to any one of claims 1 to 4, characterized in that, Before quantizing the corresponding quantization layer according to the quantization function of the quantization node in the corresponding quantization layer to obtain the quantization layer corresponding to the corresponding quantization layer for each quantization layer, it further includes: For each quantization node in the corresponding quantization layer, determining the quantization function of the corresponding quantization node according to the target processing data, where the target processing data includes the processing data corresponding to the corresponding quantization node when the floating-point model processes the input data.
6. The model quantization method according to claim 5, characterized in that, The step of determining the quantization function of the corresponding quantization node according to the target processing data for each quantization node in the corresponding quantization layer, where the target processing data includes the processing data corresponding to the corresponding quantization node when the floating-point model processes the input data, includes: For each quantization node in the corresponding quantization layer, determining the quantization function of the corresponding quantization node according to the target processing data corresponding to the corresponding quantization node and the target quantization data type corresponding to the corresponding quantization node.
7. The model quantization method according to claim 6, characterized in that, The input data includes multiple groups of input sub-data. For each quantization node, the target processing data includes multiple groups of processing data corresponding to the quantization node, and the multiple groups of processing data correspond one-to-one with the multiple groups of input sub-data; The step of determining the quantization function of the corresponding quantization node according to the target processing data corresponding to the corresponding quantization node and the target quantization data type corresponding to the corresponding quantization node for each quantization node in the corresponding quantization layer includes: For each quantization node in the corresponding quantization layer, determining the quantization function of the corresponding quantization node according to the first data range and the second data range, where the first data range is determined based on the maximum value and / or minimum value in the multiple groups of processing data corresponding to the corresponding quantization node, and the second data range is determined based on the value range of the target quantization data type corresponding to the corresponding quantization node.
8. A model quantization device, where the model is used for object detection or classification of images, and is characterized in that, Including: The first processing module is configured to process input data through a floating-point model to obtain a target output, where the target output includes the first input and the first output of each quantizable layer in the floating-point model when the floating-point model processes the input data, and each quantizable layer includes at least one quantizable node; wherein, the input data is image data; The first quantization module is configured to, for each quantizable layer, perform quantization processing on the corresponding quantizable layer according to the quantization function of the quantizable nodes in the corresponding quantizable layer to obtain a quantized layer corresponding to the corresponding quantizable layer; The second processing module is configured to process the first input of the corresponding quantizable layer through the quantized layer corresponding to the corresponding quantizable layer to obtain a second output; The optimization module is configured to optimize the quantization function corresponding to the corresponding quantizable layer according to the second output and the first output to obtain a target quantization function corresponding to the corresponding quantizable layer; wherein, according to the first output and the second output, the quantization function of the quantizable nodes in the quantizable layer is iteratively optimized until the quantization loss of the corresponding quantized layer of the quantizable layer is minimized; The second quantization module is configured to quantize the floating-point model according to the target quantization function corresponding to each quantizable layer to obtain a target quantized model; Wherein, at each of the quantizable nodes of the floating-point model, a pseudo-quantization operation is inserted. When the pseudo-quantization operation is inserted, it is not necessary to calculate the quantization parameters in the corresponding quantization function and perform quantization, but is used to label the quantizable nodes.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Neural network model quantification method and device based on label-free data
CN110969251A