Neural network model compression method and device and electronic equipment
By establishing a combination of an exponential attenuation model with offset and a network size loss model, the problem of difficulty in combining neural network quantization and pruning technology is solved, and a better model compression effect is achieved.
Patent Information
- Application Number
- CN202510334267.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-01
AI Technical Summary
It is currently difficult to organically combine the two model compression technologies of neural network quantization and pruning, resulting in poor model compression effect.
By establishing an exponential attenuation model with offset, we describe the relationship between quantization error variance and bit width, and combining the network size loss model, we obtain the bit width size joint loss model to uniformly perform model quantization and pruning loss evaluation.
The organic combination of two different model compression methods is realized, the model compression effect is optimized, and the best compression solution can be found while meeting the limitations of computing resources.
Smart Images

Figure CN120235196A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, and electronic device for compressing a neural network model. Background Art
[0002] Neural network quantization is a method for compressing a neural network. Its principle is to convert the neural network weights and activation values from the original floating-point representation to a lower integer representation (e.g., converting 32-bit floating-point numbers to 8-bit integers) to achieve the purpose of compressing the network model size and accelerating the model operation. Neural network pruning is another technical route for model compression. Such methods structurally remove some unimportant parameters in the neural network model to achieve the reduction of the model size.
[0003] Since the two perform model compression based on different model characteristics, it is difficult to organically combine them currently. Summary of the Invention
[0004] In view of the above problems, embodiments of the present disclosure provide a method, apparatus, and electronic device for compressing a neural network model.
[0005] One aspect of the present disclosure provides a method for compressing a neural network model, including: determining a bit-width loss model corresponding to performing mixed-precision quantization on an initial neural network model, where the bit-width loss model is an exponential decay model with an offset. Determining a network size loss model corresponding to pruning the quantized neural network model. Merging the bit-width loss model and the network size loss model to obtain a bit-width size joint loss model. Based on a preset computing resource, operating on the bit-width size joint loss model to determine a target quantization bit-width and a target network scaling ratio. Compressing the initial neural network model based on the target quantization bit-width and the target network scaling ratio to obtain a target neural network model.
[0006] According to an embodiment of the present disclosure, determining a bit-width loss model corresponding to performing mixed-precision quantization on an initial neural network model includes: performing mixed-precision quantization on the initial neural network model to obtain multiple groups of first data samples, where the first data samples represent the average bit operation bit-width and the corresponding first loss function values. Based on the multiple groups of first data samples, fitting a relationship curve between the loss function corresponding to the quantized neural network model and the quantization bit-width to obtain multiple first fitting parameters, where the multiple first fitting parameters include the original loss corresponding to the initial neural network model, the exponential decay rate, and the output characteristics of the quantized neural network model. Determining the bit-width loss model based on the multiple first fitting parameters.
[0007] According to an embodiment of the present disclosure, hybrid-precision quantization is performed on an initial neural network model to obtain multiple groups of first data samples, including: determining the quantization bit width of weights and the quantization bit width of activation values of the initial neural network model layer by layer. Based on the operands of each layer, a weighted average is performed on the quantization bit width of weights and the quantization bit width of activation values of the corresponding layer to obtain an average bit operation bit width. Multiple groups of average bit operation bit widths and corresponding first loss function values are respectively determined to obtain multiple groups of first data samples.
[0008] According to an embodiment of the present disclosure, curve fitting is performed on the relationship curve between the loss function corresponding to the quantized neural network model and the quantization bit width to obtain multiple first fitting parameters, including: using the non-linear least squares method to perform curve fitting on the relationship curve between the loss function corresponding to the quantized neural network model and the quantization bit width to obtain multiple first fitting parameters.
[0009] According to an embodiment of the present disclosure, determining a network size loss model corresponding to pruning the quantized neural network model, including: scaling the number of channels of the quantized neural network model to obtain multiple groups of second data samples, where the second data samples represent the size of the scaled neural network model and the corresponding second loss function value. Based on multiple groups of second data samples, curve fitting is performed on the relationship curve between the loss function corresponding to the scaled neural network model and the total number of operations of the scaled neural network model to obtain multiple second fitting parameters, and the multiple second fitting parameters include a power exponent, a scaling factor, and a noise threshold that characterize the non-linear relationship between the size of the scaled neural network model and the corresponding second loss function value. A network size loss model is determined based on multiple second fitting parameters.
[0010] According to an embodiment of the present disclosure, combining the bit width loss model and the network size loss model to obtain a bit width-size joint loss model, including: determining an exponential relationship independent variable based on the exponential decay rate and the average bit operation bit width. Determining a power law relationship independent variable based on the total number of operations and the power exponent. Performing symbolic regression fitting on the exponential relationship independent variable and the power law relationship independent variable to obtain a bit width-size joint loss model.
[0011] According to an embodiment of the present disclosure, based on a preset computing resource, operating on the bit width-size joint loss model to determine a target quantization bit width and a target network scaling ratio, including: determining the number of bit operations corresponding to the preset computing resource. Based on the number of bit operations, using the exponential relationship independent variable as the independent variable, determining the corresponding minimum point as the target quantization bit width. Based on the number of bit operations, using the power law relationship independent variable as the independent variable, determining the corresponding minimum point as the target network scaling ratio.
[0012] According to an embodiment of the present disclosure, the operators corresponding to the symbolic regression fitting include an addition operator and a multiplication operator.
[0013] Another aspect of the present disclosure provides a neural network model compression device, including: a first determination module, configured to determine a bit-width loss model corresponding to performing mixed-precision quantization on an initial neural network model, where the bit-width loss model is an exponential decay model with an offset; a second determination module, configured to determine a network size loss model corresponding to pruning the quantized neural network model; a merging module, configured to merge the bit-width loss model and the network size loss model to obtain a bit-width size joint loss model; an operation module, configured to perform operations on the bit-width size joint loss model based on preset computing resources to determine a target quantization bit-width and a target network scaling ratio; and a compression module, configured to compress the initial neural network model based on the target quantization bit-width and the target network scaling ratio to obtain a target neural network model.
[0014] The third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory, configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the neural network model compression method according to any one of the above embodiments.
[0015] By establishing an exponential decay model with an offset, the present disclosure can organically combine the exponential decay characteristic of the quantization error variance and the bit-width, accurately capture the composite influence of the quantization bit-width of different layers on the model loss. At the same time, by combining with the network size loss model to obtain the bit-width size joint loss model, it is possible to uniformly perform loss evaluation based on model quantization and model pruning, realizing the organic combination of two different model compression methods and optimizing the model compression effect. Description of the Drawings
[0016] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features and advantages of the present disclosure will become clearer. In the drawings:
[0017] Figure 1 Schematically shows a flowchart of the neural network model compression method according to an embodiment of the present disclosure;
[0018] Figure 2 Schematically shows the fitting effect and residual analysis result of the bit-width loss model according to an embodiment of the present disclosure;
[0019] Figure 3 Schematically shows the fitting effect and residual analysis result of the network size loss model according to an embodiment of the present disclosure;
[0020] Figure 4 Schematically shows the fitting result of the bit-width size joint loss model according to an embodiment of the present disclosure;
[0021] Figure 5Schematically shows a structural block diagram of a neural network model compression device according to an embodiment of the present disclosure;
[0022] Figure 6 Schematically shows a block diagram of an electronic device suitable for implementing a neural network model compression method according to an embodiment of the present disclosure. Detailed implementation manners
[0023] To make the objectives, technical solutions and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0024] It should be noted that in the accompanying drawings or the description of the specification, similar or identical parts are all denoted by the same reference numerals. The technical features in the various embodiments exemplified in the specification can be freely combined to form a new solution on the premise of no conflict. In addition, each claim can be used as an independent embodiment, or the technical features in each claim can be combined to form a new embodiment. In the accompanying drawings, the shape or thickness of the embodiment can be enlarged and simplified or conveniently marked. Furthermore, the elements or implementation manners not shown or described in the accompanying drawings are in the forms known to those of ordinary skill in the art. In addition, although this document may provide examples containing parameters with specific values, it should be understood that the parameters do not necessarily exactly equal the corresponding values, but may approximate the corresponding values within an acceptable error tolerance or design constraint.
[0025] Unless there are technical obstacles or contradictions, the above various embodiments of the present disclosure can be freely combined to form additional embodiments, and these additional embodiments are all within the protection scope of the present disclosure.
[0026] Although the present disclosure has been described with reference to the accompanying drawings, the embodiments disclosed in the accompanying drawings are intended to exemplarily illustrate the preferred embodiments of the present disclosure and should not be construed as a limitation to the present disclosure. The dimensional ratios in the accompanying drawings are merely schematic and should not be construed as a limitation to the present disclosure.
[0027] Although some embodiments of the general concept of the present disclosure have been shown and described, those of ordinary skill in the art will understand that changes can be made to these embodiments without departing from the principles and spirit of the general concept of the present disclosure, and the scope of the present disclosure is defined by the claims and their equivalents.
[0028] Figure 1 Schematically shows a flowchart of a neural network model compression method according to an embodiment of the present disclosure.
[0029] According to an embodiment of the present disclosure, as Figure 1 shown, the present disclosure provides a neural network model compression method, for example, including operations S110 to S150.
[0030] In operation S110, determine the bit-width loss model corresponding to performing mixed-precision quantization on the initial neural network model. The bit-width loss model is an exponential decay model with an offset.
[0031] For example, the initial neural network model can be an uncompressed original neural network, which usually has high precision and large computational resource requirements.
[0032] Mixed-precision quantization is a quantization technique that reduces the computational and storage requirements of the model by converting the weights and activation values in the neural network from high precision (such as 32-bit floating-point numbers) to low precision (such as 8-bit integers). Mixed precision means that different layers or different parameters may use different quantization bit-widths.
[0033] For example, the bit-width loss model is a mathematical model used to quantify the precision loss due to reducing the bit-width. In this embodiment, the bit-width loss model can be an exponential decay model with an offset, indicating that as the bit-width decreases, the precision loss increases exponentially, but the offset can adjust the starting point or speed of this decay.
[0034] In operation S120, determine the network size loss model corresponding to pruning the quantized neural network model.
[0035] Pruning is a model compression technique that reduces the number of parameters and computational amount of the model by removing unimportant connections or neurons in the neural network.
[0036] For example, the network size loss model is a mathematical model used to quantify the precision loss due to pruning. It is usually related to the pruning ratio, indicating how the precision loss changes as the pruning ratio increases.
[0037] In operation S130, merge the bit-width loss model and the network size loss model to obtain the bit-width size joint loss model.
[0038] For example, the bit-width size joint loss model is an integrated model that combines the bit-width loss model and the network size loss model, used to simultaneously consider the effects of quantization bit-width and pruning ratio on the model precision.
[0039] In operation S140, based on the preset computational resources, perform operations on the bit-width size joint loss model to determine the target quantization bit-width and the target network scaling ratio.
[0040] For example, the preset computational resources refer to the resource limitations such as the computational power and memory size provided by the hardware device (such as GPU, TPU) when compressing the model.
[0041] The target quantization bitwidth and the target network scaling ratio are the optimal quantization bitwidth and pruning ratio found by optimizing the bitwidth size joint loss model while satisfying the computing resource constraints.
[0042] In operation S150, based on the target quantization bitwidth and the target network scaling ratio, the initial neural network model is compressed to obtain the target neural network model.
[0043] For example, the target neural network model is a compressed neural network model with a smaller size and lower computing requirements, while maintaining a relatively high accuracy as much as possible.
[0044] In this embodiment, by comprehensively considering the impact of the quantization bitwidth and the pruning ratio on the model accuracy, and under the premise of satisfying the computing resource constraints, the optimal compression scheme is found, so as to reduce the model size and computing requirements while maintaining a relatively high accuracy as much as possible.
[0045] According to an embodiment of the present disclosure, for example, the bitwidth loss model corresponding to the mixed-precision quantization of the initial neural network model can be determined through operations S211~S213.
[0046] In operation S211, the initial neural network model is subjected to mixed-precision quantization to obtain multiple groups of first data samples, and the first data samples represent the average bit operation bitwidth and the corresponding first loss function value.
[0047] For example, different quantization bitwidths are adopted for different layers or parameters of the initial neural network model (such as 8 bits for some layers and 4 bits for some layers) to balance the accuracy and computing efficiency.
[0048] By experimentally quantifying different bitwidth configurations, the average bit operation bitwidth (i.e., the average bitwidth after quantization) of each configuration and the corresponding first loss function value (i.e., the accuracy loss of the quantized model) are recorded.
[0049] Among them, the average bit operation bitwidth can be the average bitwidth of all parameters in the quantized model. The first loss function value can be the accuracy loss of the quantized model on the validation set, which can be measured by indicators such as cross-entropy loss or mean squared error.
[0050] In operation S212, based on multiple groups of first data samples, the relationship curve between the loss function corresponding to the quantized neural network model and the quantization bitwidth is fitted to obtain multiple first fitting parameters, and the multiple first fitting parameters include the original loss corresponding to the initial neural network model, the exponential decay rate, and the output characteristics of the quantized neural network model.
[0051] For example, based on multiple sets of first data samples (average bit operation bit width and corresponding first loss function values), the relationship curve between the loss function and the quantization bit width is fitted using a mathematical method (such as the least squares method).
[0052] Among them, the parameters obtained through fitting may include: the original loss, representing the initial loss value of the model when not quantized; the exponential decay rate, representing the decay rate of the loss function value with the change of the bit width when the quantization bit width decreases; the output characteristics, representing the output behavior of the quantized model, such as the distribution or range change of the output value.
[0053] In operation S213, a bit width loss model is determined based on multiple first fitting parameters.
[0054] Based on the fitted parameters (original loss, exponential decay rate, output characteristics), a mathematical model can be constructed to describe the relationship between the quantization bit width and the precision loss.
[0055] In some embodiments, for example, uniform quantization can be used to map the neural network weights and activation values in floating-point form into the integer space, and this step follows the following form:
[0056] (1)
[0057] Among them, x is the quantization object, round(·) represents the rounding function, Δ represents the quantity in the floating-point space corresponding to the value 1 in the integer space (quantization step), and x q is the quantized discrete value. For b-bit quantization, assuming the quantization range is [-Q, Q), the quantization step is:
[0058] (2)
[0059] The quantization error is defined as , in uniform quantization, it is generally considered that is a random variable subject to a uniform distribution:
[0060] (3)
[0061] Therefore, the noise variance is:
[0062] (4)
[0063] Substituting into the step formula gives:
[0064] (5)
[0065] That is, ideally, the variance of quantization noise decays exponentially with the bit width. In practical situations, due to the influence of the layer-by-layer transmission characteristics in the neural network and the differences in the specific loss calculation expressions involving noise, the present disclosure uses an exponential decay model with an offset for modeling. For a fixed network structure, the relationship between the final loss function L and the average quantization bit width b is as follows:
[0066] (6)
[0067] where α, β, and γ are parameters obtained by fitting. Since the noise introduced by quantization is additive, γ represents the original loss of the unquantized model, β represents the actual exponential decay rate. When the quantization bit width is 0, the output of the neural network is a constant, and using it to calculate the loss function still results in a non-infinite positive real number, which is regulated by α. This is different from the common power law modeling method for neural network loss prediction and is more in line with the quantization noise distribution.
[0068] A corresponding curve can represent the loss function values when a certain mixed-precision quantization method is applied to a specific neural network structure and the neural network is quantized to each bit width. The independent variable is the bit width, and the dependent variable is the loss function value. To evaluate the performance of the mixed-precision quantization method, by fixing the independent variable and comparing the loss function values, multiple mixed-precision quantization methods can be compared horizontally.
[0069] In this embodiment, by experimentally quantifying different bit width configurations, a relationship curve between the quantization bit width and the accuracy loss is fitted, and a bit width loss model is constructed based on the fitting parameters, thereby providing theoretical support for mixed-precision quantization and helping to more accurately predict the accuracy loss when compressing the model.
[0070] According to an embodiment of the present disclosure, for example, the initial neural network model can be mixed-precision quantized through operations S3111~S3113 to obtain multiple groups of first data samples.
[0071] In operation S3111, the weight quantization bit width and the activation value quantization bit width of the initial neural network model are determined layer by layer.
[0072] For example, the initial neural network model can be an uncompressed original neural network model, including multiple layers (such as convolutional layers, fully connected layers, etc.).
[0073] The weight quantization bit width represents the bit width of the weight parameter after quantization and affects the storage and calculation efficiency of the model. The activation value quantization bit width represents the bit width of the activation value after quantization and affects the calculation accuracy and speed of the model.
[0074] Quantize the weight parameters of each layer to determine their quantization bit widths (such as 8 bits, 4 bits, etc.), and quantize the activation values (i.e., output values) of each layer to determine their quantization bit widths.
[0075] Among them, for each layer, appropriately select the weight quantization bit width and activation value quantization bit width independently to achieve mixed-precision quantization.
[0076] In operation S3112, based on the operands of each layer, perform a weighted average on the weight quantization bit width and activation value quantization bit width of the corresponding layer to obtain the average bit operation bit width.
[0077] For example, the operands of each layer can be the computational amount or the number of parameters of each layer, which are used to measure the contribution of this layer to the overall model's computational resources and are usually related to the input and output dimensions, the size of the convolutional kernel, etc.
[0078] Weighted average means weighting the quantization bit widths according to the operands of each layer to reflect the influence of different layers on the overall model. Based on the operands of each layer, perform a weighted average on the weight quantization bit width and activation value quantization bit width to obtain the average bit operation bit width.
[0079] For example, the average bit operation bit width here can refer to the square root of the weighted average of the product of the quantization bit widths of the weights and activations layer by layer in the neural network, with the operands of this layer as the weights, as shown below:
[0080] (7)
[0081] Among them, b w (l) , b a (l) respectively represent the quantization bit widths of the weights and activations of the l-th layer, and N l is the operand of this layer, which can be represented by the number of multiply-accumulate operations.
[0082] In operation S3113, respectively determine multiple groups of average bit operation bit widths and the corresponding first loss function values to obtain multiple groups of first data samples.
[0083] For example, by experimentally trying different combinations of weight quantization bit widths and activation value quantization bit widths, multiple groups of average bit operation bit widths are obtained.
[0084] For each group of configurations, evaluate the precision loss of the quantized model on the validation set to obtain the corresponding first loss function value.
[0085] Record the average bit operation bit width of each group of configurations and the corresponding first loss function value to form multiple groups of data samples for subsequent fitting of the bit width loss model. The data samples are, for example, no less than 3 groups to be used for fitting the 3 parameters in the L(b) curve.
[0086] In this embodiment, the weight quantization bit width and the activation value quantization bit width are determined layer by layer, and the average bit operation bit width is calculated based on the operands, finally obtaining multiple groups of data samples, providing an experimental basis for constructing the bit width loss model, thereby supporting a more accurate mixed precision quantization strategy.
[0087] According to an embodiment of the present disclosure, for example, the relationship curve between the loss function corresponding to the quantized neural network model and the quantization bit width can be fitted through operation S4121 to obtain multiple first fitting parameters.
[0088] In operation S4121, the relationship curve between the loss function corresponding to the quantized neural network model and the quantization bit width is fitted using the non - linear least squares method to obtain multiple first fitting parameters.
[0089] In some embodiments, through the method of the above - mentioned embodiment, multiple groups of average bit operation bit widths b avg and the corresponding first loss function values L(b avg ) are obtained.
[0090] For example, each group of data samples may include an average bit operation bit width b i and the corresponding loss function value L(b i ), forming a data set {(b1, L(b1)), (b2, L(b2)), …, (b n , L(b n ))}.
[0091] Select a mathematical model suitable for describing the relationship between the quantization bit width and the loss function value, such as the exponential decay model shown in formula (6).
[0092] The non - linear least squares method is an optimization method for fitting non - linear models, which determines the model parameters by minimizing the sum of the squares of the errors between the predicted values and the actual values. The relevant parameters can be fitted using the non - linear least squares method.
[0093] The error function can be defined as:
[0094] (8)
[0095] The parameters γ, α, and β are adjusted through an iterative optimization algorithm (such as the Levenberg - Marquardt algorithm) to minimize the error function E, obtaining the optimal fitting parameters γ, α, and β.
[0096] For example, by minimizing the sum of the squares of the residuals between the predicted values and the actual data through an optimization algorithm, a parametric curve of any form (linear or non - linear) can be fitted.
[0097] For example, the coefficient of determination (R2 ), or residual analysis, to evaluate the goodness of fit between the fitted model and the actual data. If the fitting effect is not ideal, a more complex model (such as a piecewise exponential model) can be tried or the data samples can be increased.
[0098] Among them, the coefficient of determination (R 2 ) is used to measure the explanatory ability of the fitted model for data variability, and its value range is from 0 to 1. The closer it is to 1, the better the fitting effect.
[0099] In this embodiment, the relationship curve between the quantization bit width and the loss function value is fitted by the non - linear least - squares method to obtain multiple first fitting parameters (such as the original loss, attenuation amplitude, and attenuation speed), thereby providing an accurate mathematical description for constructing the bit - width loss model and supporting more scientific model compression decisions.
[0100] According to an embodiment of the present disclosure, for example, the network size loss model corresponding to pruning the quantized neural network model can be determined by operating S521~S523.
[0101] In operation S521, the number of channels of the quantized neural network model is scaled to obtain multiple groups of second data samples, and the second data samples characterize the size of the scaled neural network model and the corresponding second loss function value.
[0102] For example, a neural network model after mixed - precision quantization has a lower bit width and higher computational efficiency.
[0103] Scaling the number of channels means compressing the model size by reducing the number of channels in each layer of the neural network. The number of channels in each layer of the model can be scaled (such as reducing the number of channels) to further compress the model size.
[0104] By experimentally trying different channel - scaling ratios and recording the size of the scaled neural network model (such as the number of parameters or the amount of computation) and the corresponding second loss function value (i.e., the accuracy loss of the scaled model) for each set of configurations, multiple groups of second data samples can be obtained.
[0105] In operation S522, based on multiple groups of second data samples, the relationship curve between the loss function corresponding to the scaled neural network model and the total number of operations of the scaled neural network model is fitted to obtain multiple second fitting parameters, and the multiple second fitting parameters include the power exponent, scaling factor, and noise threshold characterizing the non - linear relationship between the size of the scaled neural network model and the corresponding second loss function value.
[0106] For example, the total number of operations can be the total computational amount of the scaled neural network model, which is usually related to the number of channels, input - output dimensions, etc.
[0107] Based on multiple groups of second data samples (scaled model sizes and corresponding second loss function values), use a mathematical method (such as non-linear least squares method) to fit the relationship curve between the loss function and the total number of operations.
[0108] The parameters obtained through fitting, for example, include: the power exponent, which describes the non-linear relationship (such as power-law relationship) between the loss function and the total number of operations. The scaling factor, which represents the influence degree of the channel number scaling on the model size and the loss function value. The noise threshold, which represents the influence of the random noise introduced during the scaling process on the loss function value.
[0109] In operation S523, determine the network size loss model based on multiple second fitting parameters.
[0110] For example, the network size loss model can be to construct a mathematical model based on the parameters obtained through fitting (power exponent, scaling factor, noise threshold) to describe the relationship between the channel number scaling and the accuracy loss. For example, a power-law model in the following form can be used:
[0111] (9)
[0112] Where N represents the number of forward propagation operations of the model, , k, m, n are coefficients to be fitted.
[0113] In this embodiment, by experimentally scaling the channel number and fitting the relationship curve between the loss function and the total number of operations, multiple second fitting parameters (such as power exponent, scaling factor, and noise threshold) are obtained, thereby constructing a network size loss model, providing theoretical support for the pruning strategy, and helping to more accurately predict the accuracy loss when compressing the model.
[0114] According to an embodiment of the present disclosure, for example, the bit-width loss model and the network size loss model can be merged through operations S631~S633 to obtain a bit-width size joint loss model.
[0115] In operation S631, determine the exponential relationship independent variable based on the exponential decay rate and the average bit operation width.
[0116] In operation S632, determine the power-law relationship independent variable based on the total number of operations and the power exponent.
[0117] In operation S633, perform symbolic regression fitting on the exponential relationship independent variable and the power-law relationship independent variable to obtain a bit-width size joint loss model.
[0118] In some embodiments, the exponential decay rate can be, for example, a parameter obtained from the bit-width loss model, which represents the decay rate of the loss function value when the quantization bit-width decreases. The average bit operation width can be the average bit-width obtained through mixed-precision quantization, which represents the computational efficiency of the quantized model.
[0119] Combine the exponential decay rate with the average bit operation width to construct the independent variable of the exponential relationship. For example, it can be e in formula (6). -βb .
[0120] For example, the total number of operations can be a parameter obtained from the network size loss model, representing the total computational amount of the scaled model. The power exponent can be a parameter obtained from the network size loss model, representing the non-linear relationship between the loss function and the total number of operations.
[0121] Combine the total number of operations with the power exponent to construct the independent variable of the power-law relationship. For example, it can be N in formula (9). -m .
[0122] For example, symbolic regression is a regression method based on genetic algorithms for discovering mathematical expressions in data without pre-specifying the model form.
[0123] Take the independent variable e of the exponential relationship -βb and the independent variable N of the power-law relationship -m as inputs.
[0124] Use the symbolic regression algorithm to automatically search for the optimal mathematical expression to describe the bit-width size joint loss model.
[0125] For example, by fixing one of the above two independent variables (such as the exponential decay rate β and the power exponent b), a model of the following form may be obtained:
[0126] (10)
[0127] where e -βb , N -m are the independent variables of the formula to be fitted by symbolic regression. Subsequently, candidate operators are defined. The symbolic regression method can be implemented through genetic algorithms. For example, it includes steps such as generating candidate expressions, evaluating the goodness of fit, crossover and mutation, and selecting the optimal model, and finally screening out the function form that meets the minimum error and is structurally simple.
[0128] In this embodiment, by integrating symbolic regression technology and computational resource constraint optimization, the technical bottleneck of the joint optimization of mixed precision and pruning is broken through. By establishing the joint distribution model of the bit-width size loss, this joint optimization mechanism enables the system to automatically discover the optimal quantization bit-width combination and network scaling ratio under fixed computational resources, significantly improving the engineering practicability of model compression..
[0129] According to the embodiments of the present disclosure, for example, operations S741~S743 can be performed to calculate the bit-width size joint loss model based on the preset computational resources to determine the target quantization bit-width and the target network scaling ratio.
[0130] In operation S741, determine the bit operand corresponding to the preset computing resources.
[0131] In operation S742, based on the bit operand, with the exponential relationship independent variable as the independent variable, determine the corresponding minimum point as the target quantization bit width.
[0132] In operation S743, based on the bit operand, with the power-law relationship independent variable as the independent variable, determine the corresponding minimum point as the target network scaling ratio.
[0133] In some embodiments, the preset computing resources can be, for example, resource limitations such as the computing power and memory size of hardware devices (such as GPUs, TPUs).
[0134] The bit operand (Bit Operations, BitOps) represents the maximum amount of computation that the model can support under the given computing resources, and is usually related to the quantization bit width and the network scaling ratio. For example, it is the product of the quantization bit width and the network scaling ratio. The bit operand can be calculated by the following formula:
[0135] (11)
[0136] Fixing BitOps as the boundary condition, obtaining the minimum point of L(b, N) within the domain of b and N is the optimal compression configuration. The minimum point represents the variable value that minimizes the objective function found through the optimization algorithm.
[0137] For example, the exponential relationship independent variable can be the exponential relationship independent variable e extracted from the bit width-size joint loss model -βb .
[0138] Under the condition of satisfying the bit operand constraint, find the quantization bit width b that minimizes the bit width-size joint loss model through an optimization algorithm (such as the gradient descent method or the Newton method), that is, the target quantization bit width.
[0139] For example, the power-law relationship independent variable can be the power-law relationship independent variable N extracted from the bit width-size joint loss model -m .
[0140] Under the condition of satisfying the bit operand constraint, find the network scaling ratio N that minimizes the bit width-size joint loss model through an optimization algorithm, that is, the target network scaling ratio.
[0141] In this embodiment, by combining the preset computing resources and the bit width-size joint loss model, jointly optimize the target quantization bit width and the target network scaling ratio, so as to find the optimal model compression scheme under the premise of meeting the hardware resource limitations, and achieve efficient and high-precision model deployment.
[0142] According to an embodiment of the present disclosure, the operators corresponding to symbolic regression fitting include an addition operator and a multiplication operator.
[0143] In some embodiments, since e -βb and N -m already contain sufficient non - linear relationships, for example, it can be defined that F(·) in formula (10) only contains addition and multiplication operators.
[0144] For example, the independent variable of the exponential relationship can be the exponential - relationship independent variable e -βb obtained from the bit - width loss model. The independent variable of the power - law relationship can be the power - law - relationship independent variable N -m obtained from the network - size loss model.
[0145] In symbolic regression, define the allowed mathematical operators, such as the addition operator (+) and the multiplication operator (·). The addition operator and the multiplication operator are, for example, the basic mathematical operators used to combine variables in symbolic regression. Through the addition operator and the multiplication operator, the exponential - relationship independent variable and the power - law - relationship independent variable are combined into possible mathematical expressions.
[0146] Symbolic regression methods such as genetic algorithms can be used to automatically search for the optimal mathematical expression from the set of operators.
[0147] The fitting process, for example, includes:
[0148] 1. Initialize the population: Randomly generate a set of mathematical expressions composed of addition operators and multiplication operators.
[0149] 2. Evaluate fitness: Calculate the goodness of fit (such as the sum of squared errors) of each expression.
[0150] 3. Selection, crossover, and mutation: Generate a new population of expressions through the operations of the genetic algorithm.
[0151] 4. Iterative optimization: Repeat the evaluation and generation process until the optimal mathematical expression is found.
[0152] In this embodiment, through symbolic regression fitting, the addition operator and the multiplication operator are used to combine the exponential - relationship independent variable and the power - law - relationship independent variable into a bit - width size joint loss model, so as to consider the impact of quantization bit - width and channel - number scaling on accuracy loss simultaneously when compressing the model, and achieve more comprehensive model optimization.
[0153] Figure 2 Schematically shows the fitting effect and residual analysis result of the bit - width loss model according to an embodiment of the present disclosure. Figure 3 Schematically shows the fitting effect and residual analysis result of the network - size loss model according to an embodiment of the present disclosure. Figure 4Schematically shows the fitting results of the bit-width size joint loss model according to an embodiment of the present disclosure.
[0154] To facilitate the understanding of the neural network model compression method of the present disclosure, the following two embodiments are further enumerated.
[0155] Embodiment 1:
[0156] This embodiment uses experimental results to verify the method of modeling the mixed-precision quantization bit-width and the loss of the neural network model in the present disclosure. For example, by performing mixed-precision quantization on the convolutional neural network ResNet-18 and adjusting the quantization target bit-width multiple times, the loss under the ImageNet classification task is statistically analyzed at different quantization bit-widths, and finally the fitting performance and prediction performance of the proposed modeling method are quantitatively analyzed.
[0157] Specifically, this embodiment includes the following steps:
[0158] Step 1.A: The open-source neural mixed-precision quantization method EdMIPS can be used to quantize ResNet-18. The mixed-precision quantization candidate bit-widths include, for example, weight bit-widths {1, 2, 3, 4}, and activation value quantization bit-widths {2, 3, 4}. By adjusting the hyperparameters (parameters constraining the bit-width), finally 8 mixed-precision quantization models with an average operation bit-width b avg ∈(2,4) are obtained.
[0159] Step 1.B: Fit an exponential decay curve and perform residual analysis, with the absolute value of the fitting error less than 0.015, and the fitting effect and marginal distribution are as Figure 2 shown. As Figure 2 can be seen, the fitting effect is better between the quantization bit-widths of 2 and 4.
[0160] Embodiment 2:
[0161] This embodiment uses experimental results to verify the quantization bit-width-network size joint loss prediction method in the present disclosure. By performing channel scaling on the mixed-precision quantization models obtained in the embodiment, a series of models with different sizes are obtained, and the classification loss is statistically analyzed.
[0162] Based on the obtained bit-width-size-loss samples, a joint loss distribution is constructed, and finally its fitting performance and prediction performance, as well as the task performance of the obtained optimal network compression scheme, are quantitatively analyzed.
[0163] This embodiment includes the following steps:
[0164] Step 2. A: Select two samples with average operation bit widths of 2 and 3 from the mixed-precision quantization configurations obtained in Example 1, and train two groups of different-sized models constructed by channel number scaling. Each group contains 8 models from 1 / 8 to 1 with a step size of 1 / 8, and each model is trained separately to obtain a total of 16 pairs of operand-loss value samples.
[0165] Step 2. B: Fit the Scaling Law power-law curve to one of the groups of data, as Figure 3 shown.
[0166] Step 2.C: Fix the key parameter m≈ -0.2793 obtained in Step 2.B and the key parameter β≈ -0.8789 determined in Step 1.B, and fit F(·) in formula (10). The fitting result is as Figure 4 shown.
[0167] The obtained joint distribution formula is:
[0168] (12)
[0169] Step 2.D: Fix the computing resource constraint , and obtain the optimal configuration b = 2, N = 3.68×10 9 under the joint distribution. The corresponding model is a 2-bit mixed-precision quantization model, and the model is scaled by 1.5 times to train the corresponding model. Compare the result with the configuration sample b = 3, N = 1.61×10 9 in the samples with similar BitOps as a benchmark. The results are shown in Table 1.
[0170] Table 1 Comparison of mixed-precision quantization compression schemes
[0171]
[0172] Based on the above neural network model compression method, the present disclosure also provides a neural network model compression device. The following will combine Figure 5 to describe the neural network model compression device in detail.
[0173] Figure 5 Schematically shows a structural block diagram of a neural network model compression device according to an embodiment of the present disclosure.
[0174] As Figure 5 shown, the neural network model compression device 500 of this embodiment includes, for example: a first determination module 510, a second determination module 520, a merging module 530, an arithmetic module 540, and a compression module 550.
[0175] The first determination module 510 is configured to determine a bit-width loss model corresponding to performing mixed-precision quantization on the initial neural network model, and the bit-width loss model is an exponential decay model with an offset. In one embodiment, the first determination module 510 may be configured to perform the operation S110 described above, which will not be elaborated herein.
[0176] The second determination module 520 is configured to determine a network size loss model corresponding to pruning the quantized neural network model. In one embodiment, the second determination module 520 may be configured to perform the operation S120 described above, which will not be elaborated herein.
[0177] The merging module 530 is configured to merge the bit-width loss model and the network size loss model to obtain a bit-width and size joint loss model. In one embodiment, the merging module 530 may be configured to perform the operation S130 described above, which will not be elaborated herein.
[0178] The operation module 540 is configured to perform operations on the bit-width and size joint loss model based on the preset computing resources to determine the target quantization bit-width and the target network scaling ratio. In one embodiment, the operation module 540 may be configured to perform the operation S140 described above, which will not be elaborated herein.
[0179] The compression module 550 is configured to compress the initial neural network model based on the target quantization bit-width and the target network scaling ratio to obtain the target neural network model. In one embodiment, the compression module 550 may be configured to perform the operation S150 described above, which will not be elaborated herein.
[0180] According to embodiments of the present disclosure, any multiple of the first determination module 510, the second determination module 520, the merging module 530, the arithmetic module 540, and the compression module 550 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present disclosure, at least one of the first determination module 510, the second determination module 520, the merging module 530, the arithmetic module 540, and the compression module 550 may be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-chip, a system-on-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or in any one or a suitable combination of the three implementation manners of software, hardware, and firmware. Alternatively, at least one of the first determination module 510, the second determination module 520, the merging module 530, the arithmetic module 540, and the compression module 550 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0181] Figure 6 A block diagram of an electronic device suitable for implementing a neural network model compression method according to an embodiment of the present disclosure is schematically shown.
[0182] As Figure 6 shown, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0183] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0184] According to an embodiment of the present disclosure, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage portion 608 as needed.
[0185] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, a neural network model compression method according to the embodiments of the present disclosure is implemented.
[0186] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.
[0187] An embodiment of the present disclosure also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the neural network model compression method provided by the embodiment of the present disclosure.
[0188] When the computer program is executed by the processor 601, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0189] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 609, and / or installed from the removable medium 611. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0190] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0191] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, programming languages such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0192] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0193] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0194] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A neural network model compression method, characterized in that: include: Determine a bit width loss model corresponding to mixed precision quantization of the initial neural network model, where the bit width loss model is an exponential decay model with an offset; Determine the network size loss model corresponding to pruning the quantized neural network model; Combining the bit width loss model and the network size loss model to obtain a bit width and size joint loss model; Based on preset computing resources, the bit width size joint loss model is operated to determine a target quantization bit width and a target network scaling ratio; Based on the target quantization bit width and the target network scaling ratio, the initial neural network model is compressed to obtain a target neural network model.
2. The method according to claim 1, characterized in that The step of determining a bit width loss model corresponding to mixed precision quantization of the initial neural network model includes: Performing mixed precision quantization on the initial neural network model to obtain multiple groups of first data samples, where the first data samples represent an average bit operation width and a corresponding first loss function value; Based on the multiple groups of first data samples, fitting the relationship curve between the loss function and the quantization bit width corresponding to the quantized neural network model to obtain multiple first fitting parameters, wherein the multiple first fitting parameters include the original loss corresponding to the initial neural network model, the exponential decay speed, and the output characteristics of the quantized neural network model; The bit width loss model is determined based on the plurality of first fitting parameters.
3. The method according to claim 2, characterized in that The performing mixed precision quantization on the initial neural network model to obtain multiple groups of first data samples includes: Determining the weight quantization bit width and activation value quantization bit width of the initial neural network model layer by layer; Based on the number of operations of each layer, weighted average is performed on the weight quantization bit width and the activation value quantization bit width of the corresponding layer to obtain the average bit operation bit width; A plurality of groups of the average bit operation bit widths and the corresponding first loss function values are determined respectively to obtain the plurality of groups of first data samples.
4. The method according to claim 2, characterized in that: The step of fitting a relationship curve between the loss function and the quantization bit width corresponding to the quantized neural network model to obtain a plurality of first fitting parameters includes: The nonlinear least square method is used to fit the relationship curve between the loss function and the quantization bit width corresponding to the quantized neural network model to obtain the multiple first fitting parameters.
5. The method according to claim 2, characterized in that: The determining of the network size loss model corresponding to pruning the quantized neural network model includes: Scaling the number of channels of the quantized neural network model to obtain multiple groups of second data samples, where the second data samples represent the size of the scaled neural network model and the corresponding second loss function value; Based on the multiple groups of second data samples, fitting a relationship curve between the loss function corresponding to the scaled neural network model and the total number of operations of the scaled neural network model to obtain multiple second fitting parameters, wherein the multiple second fitting parameters include a power exponent, a scaling factor, and a noise threshold that characterize the nonlinear relationship between the size of the scaled neural network model and the corresponding second loss function value; The network size loss model is determined based on the plurality of second fitting parameters.
6. The method according to claim 5, characterized in that The bit width loss model and the network size loss model are combined to obtain a bit width size joint loss model, including: Determining an exponential relationship independent variable based on the exponential decay speed and the average bit operation bit width; Determining a power law relationship independent variable based on the total number of operands and the power exponent; The exponential relationship independent variable and the power law relationship independent variable are subjected to symbolic regression fitting to obtain the bit width size joint loss model.
7. The method according to claim 6, characterized in that The step of computing the bit width joint loss model based on preset computing resources to determine a target quantization bit width and a target network scaling ratio includes: Determining a bit operand corresponding to the preset computing resource; Based on the bit operand, taking the exponential relationship independent variable as the independent variable, determining the corresponding minimum point as the target quantization bit width; Based on the bit operand, taking the power law relationship independent variable as the independent variable, determining the corresponding minimum point as the target network scaling ratio.
8. The method according to claim 6, characterized in that Operators corresponding to the symbolic regression fitting include addition operators and multiplication operators.
9. A neural network model compression device, characterized in that: include: A first determination module is used to determine a bit width loss model corresponding to mixed precision quantization of an initial neural network model, wherein the bit width loss model is an exponential decay model with an offset; A second determination module is used to determine a network size loss model corresponding to pruning the quantized neural network model; A merging module, used for merging the bit width loss model and the network size loss model to obtain a bit width and size joint loss model; A computing module, configured to compute the bit width joint loss model based on preset computing resources to determine a target quantization bit width and a target network scaling ratio; A compression module is used to compress the initial neural network model based on the target quantization bit width and the target network scaling ratio to obtain a target neural network model.
10. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 8.