Method for obtaining multi-bit-width quantized deep convolutional neural network
By establishing a multi-bit wide-aware quantization model with weight sharing and adopting collaborative training and adaptive label softening methods, multi-bit wide quantization of neural networks is solved, and the multi-bit wide quantization problem in the existing technology is achieved, and efficient model deployment and accuracy improvement in different hardware devices and scenarios are achieved.
Patent Information
- Application Number
- CN202110923119.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-08-12
AI Technical Summary
The prior art is difficult to achieve multi-bit wide quantization in neural network quantization, resulting in the need to independently quantify and compress the model in different hardware devices and scenarios, resulting in overhead in computing resources, human resources and time.
By establishing a multi-bit wide-aware quantization model with weight sharing, and using the minimum-random-maximum bit-width collaborative training and adaptive label softening methods, the multi-bit wide-aware quantization model is quantized hypernet training. Target constraints are set according to the requirements, mixed accuracy search is performed through Monte Carlo sampling, quantization perception accuracy predictor and genetic algorithm to obtain a sub-network that satisfies the constraints, forming a multi-bit wide quantization deep convolutional neural network.
It realizes higher model accuracy under different bit width constraints, reduces the time and computing overhead of deep convolutional neural network compression, and supports multi-scenario deployment, reducing training costs.
Smart Images

Figure CN113762489B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural networks, and in particular to a method for multi-bitwidth quantization of a deep convolutional neural network. Background Art
[0002] Neural network quantization refers to compressing a neural network model in 32-bit floating-point format to an 8- to 1-bit fixed-point format to reduce storage and computational costs. Neural network quantization technology is a currently popular technology for compressing deep neural networks, used to compress neural networks so that they can be deployed on edge devices for fixed-point calculations. And the technical route of one-time quantization and multi-scenario deployment is a new quantization direction. Current technical solutions include apq, oqa, coquant, anyprecision, and robust quantization. The multi-bitwidth-aware quantization method for one-time quantization and multi-scenario deployment only requires one-time quantization training to achieve multiple deployments, solving the training costs caused by separately quantizing and training individual models for each scenario in traditional quantization methods.
[0003] Currently, the neural network compression quantization methods in the prior art all focus on quantization models with fixed bitwidth (single precision). When the model faces different hardware device characteristics (processor computing precision) and constraints (model accuracy), independent model quantization and compression are required, which easily causes large computational resources, human resources, and time overheads when facing deployment requirements in different scenarios (such as sometimes cloud computing is needed and sometimes edge computing is needed).
[0004] Moreover, there are also many defects in other technical solutions for one-time quantization and multi-scenario deployment in the prior art. Among them, the apq method cannot achieve low-bit quantization, and can only achieve mixed-precision quantization among 3 bits of 4, 6, and 8 bits without achieving quantization below 4 bits. Oqa can only achieve unified bitwidth quantization and cannot achieve mixed-bit quantization (it means that the bit precision of different neural network layers must be the same and cannot compress different layers to different bit precisions), with poor flexibility. Others such as coquant, any precision, and robust quantization have large precision losses during low-bit quantization. Summary of the Invention
[0005] Embodiments of the present invention provide a method for obtaining a deep convolutional neural network with multi-bitwidth quantization to overcome the problems of the prior art.
[0006] To achieve the above object, the present invention adopts the following technical solutions.
[0007] A method for multi-bitwidth quantization of a deep convolutional neural network, comprising:
[0008] Build a weight - shared multi - bit - width aware quantization model;
[0009] Conduct multi - bit - width aware quantization hyper - network training on the multi - bit - width aware quantization model;
[0010] Set target constraints according to requirements, perform mixed - precision search on the trained multi - bit - width aware quantization model according to the target constraints, obtain sub - networks that meet the constraints, and use each sub - network that meets the constraints to form a multi - bit - width quantized deep convolutional neural network.
[0011] Preferably, the building of the weight - shared multi - bit - width aware quantization model includes:
[0012] Build a weight - shared multi - bit - width aware quantization model, which is a hyper - network with a multi - layer structure. The sub - networks of the multi - bit - width aware quantization model include the lowest - bit - width model, the highest - bit - width model, and the random - bit - width model. Quantize and train multiple sub - networks in the multi - bit - width aware quantization model simultaneously;
[0013] Let the quantization configuration of the multi - bit - width aware quantization model be expressed as respectively represent the bit - widths of the weights and activations of layer l. Given a floating - point weight w and activation v, the set of learnable quantization step sizes and the zero - point set Then the objective function for training the multi - bit - width aware quantization model is expressed as:
[0014]
[0015] Q(·) represents the quantization function.
[0016] Preferably, the conducting of the multi - bit - width aware quantization hyper - network training on the multi - bit - width aware quantization model includes:
[0017] Adopt the minimum - random - maximum bit - width collaborative training method. In each training iteration, optimize the lowest - bit - width model, the highest - bit - width model, and M random - bit - width models (M + 2 sub - networks) in the multi - bit - width aware quantization model simultaneously. The training objective is the objective function shown in formula 1, and the M + 2 different models are represented by different Bs in formula 1;
[0018] Adaptive label softening. Given a data set containing N classes, x i represents the input image, y i represents the corresponding true label, and define as the class - level soft label for each round. A e is an N - row and N - column square matrix. A eEach column in corresponds to a soft label of a category. When an input sample (x i , y i ) is correctly judged by an arbitrary quantization model, construct {p L (x i ), p R (x i ), p H (x i )} to update the y e column in A i . M represents the number of random subnets, n represents the predicted value, p L (x i ), p R (x i ), p H (x i ) all describe the same object and are described as follows:
[0019]
[0020] Then the Adaptive Soft Label Loss is expressed as:
[0021]
[0022] represents the value of matrix A at coordinate (n, y i ) at the e-th round, and the balance coefficient ζ is set to 0.5;
[0023] p L (x i ), p R (x i ), p H (x i ) are the logit outputs of the highest bit-width model, the random bit-width model, and the lowest bit-width model respectively;
[0024] Update according to formula 3 once in each iteration, and normalize A e after each epoch. Use it in formula 4 in the next epoch until the multi-bit-width aware quantization model converges or reaches the set number of training times, then the training process of the multi-bit-width aware quantization model ends.
[0025] Preferably, set the target constraint according to the requirement, perform mixed-precision search on the trained multi-bit-width aware quantization model according to the target constraint to obtain subnets that meet the constraint, and use each subnet that meets the constraint to form a multi-bit-width quantized deep convolutional neural network, including:
[0026] Regarding the trained multi-bitwidth aware quantization model as a model pool containing many sub-networks, set the target constraints according to the required multi-bitwidth quantized deep convolutional neural network. The target constraints include the average bit constraint. According to the target constraints, use three methods, namely Monte Carlo sampling, quantization-aware accuracy predictor, and genetic algorithm, to perform mixed-precision search on the trained multi-bitwidth aware quantization model, and search for sub-networks that meet the constraints.
[0027] According to the target sub-networks that meet the constraints, form the required multi-bitwidth quantized deep convolutional neural network, and each target sub-network is separately used as an independent unit in the multi-bitwidth quantized deep convolutional neural network.
[0028] Preferably, the step of using three methods, namely Monte Carlo sampling, quantization-aware accuracy predictor, and genetic algorithm, to perform mixed-precision search on the trained multi-bitwidth aware quantization model according to the target constraints and search for sub-networks that meet the constraints includes:
[0029] Use Monte Carlo sampling to construct the training dataset of the quantization-aware accuracy predictor, construct the initial population that meets the constraints in the sampling of the genetic algorithm for mixed-precision search, and use the quantization-aware accuracy predictor to estimate the accuracy of the mixed-precision search.
[0030] Generate a number of chromosomes using Monte Carlo sampling according to the configuration of the sub-network and the number of bits in different layers. Use the number of chromosomes as the initial Pareto solution set. Generate structure-accuracy data pairs using Monte Carlo sampling. For different chromosomes, use the predicted output of the quantization-aware accuracy predictor as the fitness score of the chromosome. Save the chromosome with the highest fitness score and add it to the elite set. Select elites for mutation and crossover with a predetermined probability to obtain a new population. The process of selection-mutation-crossover is repeated until the algorithm reaches the Pareto solution that meets the weight and activation average bitwidth target.
[0031] It can be seen from the technical solutions provided by the embodiments of the present invention described above that the embodiments of the present invention solve the problem of competitive training under different bit sub-networks through minimum-random-maximum bitwidth collaborative training and adaptive label softening, and achieve higher model accuracy under different average bitwidth constraints.
[0032] Additional aspects and advantages of the present invention will be given in part in the following description, and these will become apparent from the following description, or can be understood through the practice of the present invention. Description of the Drawings
[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0034] Figure 1 It is a processing flow chart of a method for multi-bit width quantization of a deep convolutional neural network provided by an embodiment of the present invention. Specific embodiments
[0035] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation of the present invention.
[0036] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used here may include wireless connection or coupling. The phrase "and / or" used here includes any unit and all combinations of one or more related listed items.
[0037] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.
[0038] For the convenience of understanding the embodiments of the present invention, the following will further explain with several specific embodiments as examples in combination with the accompanying drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.
[0039] An embodiment of the present invention provides a multi-bitwidth-aware quantization method for multi-scenario deployment (where each application scenario has different requirements for the computing accuracy of the neural network). By training the quantized deep convolutional neural network only once, a multi-bitwidth-aware quantization model all-in-once network that meets the requirements of any number of deployments can be obtained, greatly reducing the time and computational expenses of deep convolutional neural network compression, achieving a high model accuracy under different average bit constraints, forming a better Pareto optimal frontier, and making the neural network deployment lighter and better.
[0040] On the premise of weight sharing, multi-bitwidth awareness of the model is achieved through minimum-random-maximum bitwidth collaborative training, and a quantization model for one-time quantization and multi-scenario deployment is constructed. By adaptive label softening, the problem of malignant competition among subnets under different bitwidths is solved. The performance of the quantization-aware accuracy predictor is improved through Monte Carlo search.
[0041] The processing flow of a method for multi-bitwidth quantization of a deep convolutional neural network provided by an embodiment of the present invention is as Figure 1 shown, including the following processing steps:
[0042] Step S10: Establish a multi-bitwidth-aware quantization model with weight sharing.
[0043] First, model the multi-bitwidth-aware quantization training problem in this method. Different from the general quantization of a single model, this method needs to simultaneously quantize and train multiple sub-networks under the same model. Taking resnet18 as an example, a resnet18 model with a bitwidth range of 2-8 contains 742 sub-network models. Training multiple sub-network models simultaneously requires re-modeling the network training problem for multi-bitwidth awareness. A supernet is the abbreviation of a supernet model, and a multi-bitwidth-aware quantization model is a description of the supernet from a functional level. The supernet, the supernet model, and the multi-bitwidth-aware quantization model are all the same object, and the supernet includes multiple layers. Resnet18 has a total of 21 layers, and the activation and weight of each layer can be set independently. The quantization bitwidth can be selected from 2-8 bits, so the sub-network model includes (21×2) 7 sub-network models.
[0044] The all-in-once quantization model supports diverse quantization bitwidth configurations. Suppose the quantization configuration of a model can be expressed as At the same time respectively represent the bitwidth of the weight and activation of layer l. Given a floating-point weight w and activation v, the set of learnable quantization steps and the zero-point set Then the objective function of supernet training can be expressed as:
[0045]
[0046] Q(·) represents the quantization function. The goal of multi-bit quantization is to learn a robust weight distribution, independent quantization step sizes, and zero-point sets under different bit-width configurations. To efficiently train a quantization model, we adopt a low-bit quantization training method LSQ (Learned Step-size Quantization, low-bit quantization based on trainable step sizes). Taking the quantization of activation v to k-bit as an example, the quantization function with weight sharing is as follows:
[0047]
[0048] Equation 1 represents the objective function for hypernetwork training. Equation 2 represents the formula for lsq quantization. It can be regarded as the specific description of Q() in Equation 1, and k represents quantization to k bits.
[0049] The multi-bit width-aware quantization model aims to construct a model structure with weight sharing and independent quantization step sizes in a multi-bit width scenario by stripping the model weights and quantization step sizes. The multi-bit width-aware quantization model pre-defines the quantization step sizes for each layer under different bit widths, and by setting the quantization bit widths of each layer of the model, the corresponding quantization step sizes and quantization boundaries can be activated. Thus, the model can be flexibly adjusted to the unified quantization and mixed-precision quantization forms in different bit width scenarios.
[0050] Step S20: Perform multi-bit width-aware quantization hypernetwork training on the multi-bit width-aware quantization model.
[0051] This method proposes a minimum-random-maximum bit width collaborative training and an adaptive label softening method to iteratively train the multi-bit width-aware quantization model.
[0052] The training of the multi-bit width-aware quantization model includes optimizing the lowest bit width model, the highest bit width model, and M random bit width models, a total of M + 2 seed networks simultaneously. The training objective is the objective function shown in Equation 1, and the M + 2 different models are represented by different in Equation 1. Adopt the minimum-random-maximum bit width collaborative training method. In each training iteration, train the lowest bit width model (for example, 2 bits fixed for each layer), the highest bit width model (for example, 8 bits fixed for each layer), and two random bit width models simultaneously to improve the overall performance of the hypernetwork model.
[0053] Adaptive label softening. Given a dataset containing N classes, x i represents the input image, and y i represents the corresponding true label. Define As the class-level soft label for each round, A e is an N-by-N square matrix, and each column in A e corresponds to the soft label of a category. When an input sample (x i , y i ) is correctly judged by any quantization model, we construct {p L (x i ), p R (x i ), p H (x i )} to update the y e column in A i . M represents the number of random subnets, and n represents the predicted value. p L (x i ), p R (x i ), p H (x i ) all describe the same object.
[0054] can be described as follows:
[0055]
[0056] Then the Adaptive Soft Label Loss can be expressed as:
[0057]
[0058] represents the value of matrix A at the coordinate (n, y i ) in the e-th round. The balance coefficient ζ is generally set to 0.5.
[0059] p L (x i ), p R (x i ), p H (x i ) are the logit outputs of the highest bit-width model, the random bit-width model, and the lowest bit-width model mentioned above, respectively.
[0060] The update of formula 3 is performed in each iteration, and A e is normalized after each epoch. It is used in formula 4 in the next epoch. The total number of epochs is set manually. The training process of the multi-bit-width aware quantization model ends until the multi-bit-width aware quantization model converges or reaches the set number of training rounds. The conditions for judging the convergence of the multi-bit-width aware quantization model include that the accuracy no longer improves as the number of training rounds increases.
[0061] Step S30: Consider the trained multi-bitwidth perceptual quantization model as a large model pool, which contains many sub-networks. Sub-networks that meet the requirements can be selected from it according to needs. For example, if a quantization depth convolutional neural network with an average bitwidth of 4 is required, set the target constraint to 4. According to the target constraint, use three methods: Monte Carlo sampling, quantization-aware accuracy predictor, and genetic algorithm to perform mixed-precision search on the trained multi-bitwidth perceptual quantization model to search for the target sub-network.
[0062] The target constraint includes the average bit constraint. The average bit constraint means that the activations and weights of each layer have different bitwidth representations. The value obtained by multiplying the activations and weights of all layers by their proportional weights is the average bit.
[0063] According to the target sub-networks that meet the constraints, form a multi-bitwidth quantized depth convolutional neural network, and each target sub-network is separately used as an independent unit in the multi-bitwidth quantized depth convolutional neural network.
[0064] Monte Carlo sampling. First, explain Monte Carlo sampling. In a super-network, a sampling pool of (sub-network architecture, average bit) is obtained through random uniform sampling. For example, randomly sample 500,000 sub-network models and calculate the corresponding average bit numbers, and an empirical distribution of different layer bit numbers under each average bit number can be obtained. Sampling from this empirical distribution can obtain results that meet the target distribution with a higher probability.
[0065] Monte Carlo sampling is applied in two aspects: constructing a quantization accuracy prediction training data set in the quantization-aware accuracy predictor, and sampling an initial population that meets the constraints in the genetic algorithm for mixed-precision search.
[0066] The technical details are as follows:
[0067] Respectively given the average bit constraints τ w and τ a , The empirical approximation of is For the convenience of statistics, It is calculated in the following way:
[0068]
[0069] To construct the above distribution, we randomly sample a large number of structure-average bit data pairs in the sampling space to construct a sampling pool. Let #(τ w = τ0) represent the total number of merits sub-networks with an average bitwidth of τ0 in the sampling pool. At the same time, represents the data pair The total number that appears in the sampling pool, then can be estimated as follows:
[0070]
[0071] Quantization-Aware Accuracy Predictor.
[0072] During the search process, it is very important to accelerate the evaluation process of the search model. We propose a Quantization-Aware Accuracy Predictor to accurately estimate the accuracy of the network, which can predict the accuracy of a model given a configuration. More specifically, it is a 7-layer feedforward neural network, and each embedding dimension is equal to 150. The bit-width configuration is encoded as a one-hot vector as the input, (for example, a set of weight bit-width configurations like [2, 4, 6, 4, 8], where each number represents the quantization bit-width of the weights of a certain layer, and the activation values are the same), and input into the predictor to obtain the predicted accuracy as the output.
[0073] In particular, we use Monte Carlo sampling to generate structure-accuracy data pairs, which can avoid the imbalance of the dataset and improve the prediction performance for lower and higher bit-widths, such as the accuracy prediction of models below 3 bits or models above 7 bits.
[0074] The specific approach is to uniformly and randomly sample an average number of bits, such as 5 bits, and then use the Monte Carlo sampling technique to sample from the empirical distribution at 5 bits, which can make the sampled models easily meet the 5-bit constraint. In this way, the constructed dataset can be more uniform, rather than having a large number of sampled sub-networks concentrated in the middle bit part like random uniform sampling.
[0075] The genetic algorithm for mixed-precision search first uses Monte Carlo sampling to generate several chromosomes (i.e., the configurations of sub-networks: the bit-number settings of different layers) as the initial Pareto solution set. Monte Carlo sampling can greatly accelerate the time for constructing the initial solution set.
[0076] Then, for different chromosomes, the predicted output of the Quantization-Aware Accuracy Predictor is used as the fitness score of the chromosome.
[0077] Finally, the chromosome with the highest fitness score is saved and added to the elite set, and then the elites are selected for mutation and crossover according to a predetermined probability to obtain a new population. The process of selection-mutation-crossover is repeated until the algorithm reaches a Pareto solution that meets the average bit-width targets of weights and activations.
[0078] In summary, the embodiments of the present invention solve the problem of competitive training under different bit subnets through minimum-random-maximum bitwidth collaborative training and adaptive label softening, achieving higher model accuracy under different average bitwidth constraints, enabling fast deployment of high-performance models for application scenarios with different quantization constraints without having to re-perform quantization training, and reducing a large amount of computing resources and time overhead.
[0079] The embodiments of the present invention can improve the performance of the quantization-aware accuracy predictor and greatly improve the search efficiency, reducing the time to obtain the target subnet, through an evolutionary algorithm optimized by Monte Carlo sampling.
[0080] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0081] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0082] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and the relevant parts can be referred to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0083] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for obtaining a multi-bitwidth quantized deep convolutional neural network, characterized in that, Including: Establish a multi-bitwidth aware quantization model with weight sharing; Perform multi-bitwidth aware quantization supernetwork training on the multi-bitwidth aware quantization model; Set target constraints according to requirements, perform mixed-precision search on the trained multi-bitwidth aware quantization model according to the target constraints to obtain sub-networks that meet the constraints, and use each sub-network that meets the constraints to form a multi-bitwidth quantized deep convolutional neural network; The establishment of the multi-bitwidth aware quantization model with weight sharing includes: Establish a multi-bitwidth aware quantization model with weight sharing. This multi-bitwidth aware quantization model is a supernetwork with a multi-layer structure. The sub-networks of the multi-bitwidth aware quantization model include the lowest bitwidth model, the highest bitwidth model, and the random bitwidth model. Quantize and train multiple sub-networks in the multi-bitwidth aware quantization model simultaneously; Let the quantization configuration of the multi-bitwidth aware quantization model be represented as representing the bitwidths of the weights and activations of layer l respectively. Given a floating-point weight w, activation b, the set of learnable quantization steps and the set of zero-points Then the objective function for training the multi-bitwidth aware quantization model is expressed as: Q(·) represents the quantization function; The multi-bitwidth aware quantization supernetwork training on the multi-bitwidth aware quantization model includes: In each training iteration, using the minimum-random-maximum bitwidth collaborative training method, the lowest-bitwidth model, the highest-bitwidth model, and M random-bitwidth models in the multi-bitwidth perceptually quantized model, a total of M+2 seed networks, are optimized simultaneously. The training objective is the objective function shown in formula (1), and the M+2 different models are represented by different in formula (1); Adaptive label softening, given a dataset containing N classes, x i represents the input image, y i represents the corresponding ground truth label, define as the class-level soft label for each round, A e is an N-by-N square matrix, A e each column in corresponds to the soft label of a class. When an input sample (x i , y i ) is correctly judged by an arbitrary quantization model, construct {p L (x i ), p R (x i ), p H (x i )} to update the y e column in A i , M represents the number of random subnets, n represents the predicted value, p L (x i ), p R (x i ), p H (x i ) all describe the same object and are described as follows: Then the Adaptive Soft Label Loss is expressed as: Denote the value of matrix A at coordinate (n, y i ) at the e-th round, and set the balance coefficient ζ to 0.5; p L (x i ), p R (x i ), p H (x i ) are the logit outputs of the lowest bit-width model, the random bit-width model, and the highest bit-width model, respectively; Perform the update of formula (3) in each iteration, and normalize A after each epoch. Use it in formula (4) in the next epoch until the multi-bitwidth aware quantization model converges or reaches the set number of training times, then the training process of the multi-bitwidth aware quantization model ends; e The setting of target constraints according to requirements, performing mixed-precision search on the trained multi-bitwidth aware quantization model according to the target constraints to obtain sub-networks that meet the constraints, and using each sub-network that meets the constraints to form a multi-bitwidth quantized deep convolutional neural network includes: Regard the trained multi-bitwidth aware quantization model as a model pool containing many sub-networks. Set target constraints according to the required multi-bitwidth quantized deep convolutional neural network. This target constraint includes the average bit constraint. Use Monte Carlo sampling, quantization-aware accuracy predictor, and genetic algorithm to perform mixed-precision search on the trained multi-bitwidth aware quantization model according to the target constraints to search for sub-networks that meet the constraints; Form the required multi-bitwidth quantized deep convolutional neural network according to the target sub-networks that meet the constraints. Each target sub-network is separately used as an independent unit in the multi-bitwidth quantized deep convolutional neural network.
2. The method according to claim 1, wherein The use of Monte Carlo sampling, quantization-aware accuracy predictor, and genetic algorithm to perform mixed-precision search on the trained multi-bitwidth aware quantization model according to the target constraints to search for sub-networks that meet the constraints includes: Use Monte Carlo sampling to construct the training dataset of the quantization-aware accuracy predictor, construct the initial population that meets the constraints for sampling in the genetic algorithm for mixed-precision search, and use the quantization-aware accuracy predictor to estimate the accuracy of the mixed-precision search; Generate a number of chromosomes using Monte Carlo sampling according to the configuration of the sub-network and the number of bits in different layers. Use the number of chromosomes as the initial Pareto solution set. Generate structure-accuracy data pairs using Monte Carlo sampling. For different chromosomes, use the predicted output of the quantization-aware accuracy predictor as the fitness score of the chromosome. Save the chromosome with the highest fitness score and add it to the elite set. Select elites for mutation and crossover with a predetermined probability to obtain a new population. Repeat the process of selection-mutation-crossover until the algorithm reaches the Pareto solution that meets the weight and activation average bitwidth target.
Citation Information
Patent Citations
Neural network hybrid quantization method based on progressive quantization and Hessian information
CN112183742A
Differentiable search method and device for hybrid precision neural network
CN112364981A