Model compression method and device, equipment, storage medium and program product
By determining the sensitivity type of each layer of a deep neural network and using a differentiable neural structure search method, a suitable compression strategy is automatically found for the deep neural network. This solves the problems of high memory consumption and computational complexity of deep neural networks in computer vision and speech recognition, achieving efficient compression and accuracy preservation.
Patent Information
- Application Number
- CN202511117591.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-14
AI Technical Summary
The widespread deployment of deep neural networks in computer vision and speech recognition is hampered by high memory consumption and computational complexity. Existing model compression methods struggle to find suitable compression strategies to achieve a balance between efficiency and accuracy.
By determining the sensitivity type of each network layer in a deep neural network, and based on the highest pruning rate and the lowest quantization bit width, combined with the training and validation sets, a suitable compression strategy for each network layer is automatically found. The probability of candidate pruning rate and quantization bit width is updated using a differentiable neural structure search method, thereby achieving accurate compression.
It achieves efficient compression of deep neural networks, maintains or improves model accuracy, and achieves a higher balance between compression efficiency and accuracy.
Smart Images

Figure CN120952085A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a model compression method, apparatus, device, storage medium, and program product. Background Technology
[0002] In recent years, deep neural networks have achieved remarkable performance breakthroughs in artificial intelligence fields such as computer vision and speech recognition due to their powerful feature learning capabilities. However, deep neural networks suffer from large memory consumption and high computational complexity, requiring significant storage and computing resources for training and operation, which severely hinders their widespread deployment in artificial intelligence systems. Therefore, model compression methods, such as network pruning and model quantization, have emerged. These methods can significantly reduce model resource consumption at the cost of limited performance loss, but finding suitable compression strategies for each network layer remains extremely difficult. Summary of the Invention
[0003] This application provides a model compression method, apparatus, device, storage medium, and program product that can automatically find suitable compression strategies for each network layer of a deep neural network, achieving a higher balance between compression efficiency and accuracy.
[0004] In a first aspect, embodiments of this application provide a model compression method, comprising: determining the sensitivity type of each network layer based on the highest pruning rate and the lowest quantization bit width of each network layer of a basic deep neural network; determining the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected for each network layer based on the architecture parameters of a first predetermined number of candidate pruning rates and a second predetermined number of candidate quantization bit widths corresponding to the sensitivity type of each network layer; updating the probability of each candidate pruning rate being selected for each network layer based on a first training set and a validation set, in combination with the highest pruning rate, and updating the probability of each candidate quantization bit width being selected for each network layer based on the lowest quantization bit width; and compressing each network layer according to the candidate pruning rate with the highest probability after the update and the candidate quantization bit width with the highest probability after the update, to obtain a compressed model.
[0005] Secondly, embodiments of this application also provide a model compression apparatus, comprising: a sensitivity type determination module, configured to determine the sensitivity type of each network layer based on the highest pruning rate and the lowest quantization bit width of each network layer of a basic deep neural network; a probability determination module, configured to determine the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected for each network layer based on the architecture parameters of a first predetermined number of candidate pruning rates and a second predetermined number of candidate quantization bit widths corresponding to the sensitivity type of each network layer; a probability update module, configured to update the probability of each candidate pruning rate being selected for each network layer based on a first training set and a validation set, in combination with the highest pruning rate, and update the probability of each candidate quantization bit width being selected for each network layer in combination with the lowest quantization bit width; and a compression module, configured to compress each network layer according to the candidate pruning rate with the highest probability after the update and the candidate quantization bit width with the highest probability after the update, to obtain a compressed model.
[0006] Thirdly, embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the model compression method as described in embodiments of this application.
[0007] Fourthly, embodiments of this application also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the model compression method as described in embodiments of this application.
[0008] Fifthly, embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the model compression method as described in embodiments of this application.
[0009] The technical solution of this application embodiment determines the sensitivity type of each network layer based on the highest pruning rate and the lowest quantization bit width of each network layer of the basic deep neural network; based on the architecture parameters of a first predetermined number of candidate pruning rates and a second predetermined number of candidate quantization bit widths corresponding to the sensitivity type of each network layer, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected of each network layer are determined respectively; based on a first training set and a validation set, the probability of each candidate pruning rate being selected of each network layer is updated by combining the highest pruning rate and the probability of each candidate quantization bit width being selected of each network layer is updated by combining the lowest quantization bit width; and each network layer is compressed according to the candidate pruning rate with the highest probability after the update and the candidate quantization bit width after the update to obtain a compressed model. In this embodiment, the sensitivity type of each network layer is determined by the highest pruning rate and the lowest quantization bit width. Based on the architectural parameters of the candidate pruning rates and candidate quantization bit widths corresponding to the sensitivity types of each network layer, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected are determined. Based on the first training set and validation set, and by combining the highest pruning rate and the lowest quantization bit width, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected are updated. This method can automatically find suitable compression strategies for each network layer of the deep neural network, achieving a higher balance between compression efficiency and accuracy. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 This is a schematic diagram of a model compression method provided in an embodiment of this application;
[0012] Figure 2 This is a schematic diagram of a model compression device provided in an embodiment of this application;
[0013] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". It should be noted that the concepts of "first," "second," etc., mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should be noted that the modifications "a" and "a plurality" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless explicitly indicated otherwise in the context, they should be understood as "one or more". It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) shall comply with the requirements of applicable laws, regulations, and relevant provisions.
[0016] Figure 1 This is a schematic flowchart of a model compression method provided in an embodiment of this application, applicable to model compression scenarios. The method can be executed by a model compression device, which can be implemented in software and / or hardware, optionally through an electronic device such as a mobile terminal, PC, or server. Figure 1 As shown, the method includes:
[0017] S110. Based on the highest pruning rate and lowest quantization bit width of each network layer of the basic deep neural network, determine the sensitivity type of each network layer.
[0018] The basic deep neural network is an untrained neural network. The maximum pruning rate and minimum quantization bit width are the highest pruning rate and the lowest quantization bit width that can be compressed for each layer of the basic deep neural network at a given accuracy.
[0019] Among them, the sensitivity types include quantification sensitivity and pruning sensitivity.
[0020] In this embodiment, for each network layer, a first result value is determined based on the highest pruning rate and a set pruning rate threshold, and a second result value is determined based on the lowest quantization bit width and a set quantization bit width threshold. By comparing the magnitude of the first result value and the second result value, the sensitivity type of each network layer is determined.
[0021] Optionally, the sensitivity type of each network layer is determined based on the highest pruning rate and lowest quantization bit width of each network layer of the basic deep neural network, including: determining the highest pruning rate and lowest quantization bit width of each network layer of the basic deep neural network based on a second training set and a set precision; determining a first result value based on the second training set and for each network layer, based on the highest pruning rate of the corresponding network layer and a set pruning rate threshold; determining a second result value based on the lowest quantization bit width of the corresponding network layer and a set quantization bit width threshold; if the first result value is greater than the second result value, the sensitivity type of the corresponding network layer is quantization sensitive; if the first result value is less than or equal to the second result value, the sensitivity type of the corresponding network layer is pruning sensitive.
[0022] In this embodiment, there is no limitation on the setting accuracy; for example, it can be 98%, 95%, etc.
[0023] For example, based on historical experimental experience, the pruning rate and quantization bit width of each network layer can be set, and each network layer can be compressed according to the set pruning rate and quantization bit width. Then, the compressed deep neural network is trained using a second training set. If the accuracy of the compressed deep neural network reaches the set accuracy, the set pruning rate can be used as the highest pruning rate, and the set quantization bit width as the lowest quantization bit width. Otherwise, the pruning rate and quantization bit width are reset until the accuracy of the compressed deep neural network reaches the set accuracy. If the accuracy of the compressed deep neural network still does not reach the set accuracy after a set time, the set accuracy can be adjusted.
[0024] Optionally, determining the highest pruning rate and lowest quantization bit width of each network layer of the basic deep neural network based on the second training set and a set precision includes: compressing the basic deep neural network based on the set pruning rate and set quantization bit width of each network layer of the basic deep neural network to obtain a first target deep neural network; training the first target deep neural network based on the second training set to obtain training precision; wherein, the second training set includes the first training set and the validation set; if the training precision is greater than or equal to the set precision, then the set pruning rate of each network layer is respectively determined as the highest pruning rate of each network layer and the set quantization bit width of each network layer is respectively determined as the lowest quantization bit width of each network layer; if the training precision is less than the set precision, then the set pruning rate and set quantization bit width of each network layer are adjusted, and the process is returned to and the step of compressing the basic deep neural network based on the set pruning rate and set quantization bit width of each network layer of the basic deep neural network is executed again to obtain the compressed basic deep neural network.
[0025] In this embodiment, there are no restrictions on the setting of the pruning rate and the setting of the quantization bit width. The specific settings can be determined according to the actual situation. For example, the pruning rate can be set to 60%, and the quantization bit width can be set to 2. The first training set and the validation set can be combined to form the second training set.
[0026] Specifically, the basic deep neural network is first compressed according to the set pruning rate and set quantization bit width of each network layer to obtain a first target deep neural network. Then, the first target deep neural network is trained using a second training set to obtain training accuracy. If the training accuracy is greater than or equal to the set accuracy, the set pruning rate of each network layer is determined as the highest pruning rate of each network layer, and the set quantization bit width of each network layer is determined as the lowest quantization bit width of each network layer. If the training accuracy is less than the set accuracy, the set pruning rate and set quantization bit width of each network layer are adjusted, and the process is repeated to compress the basic deep neural network based on the set pruning rate and set quantization bit width of each network layer of the basic deep neural network to obtain the compressed basic deep neural network.
[0027] In this embodiment, by combining the second training set with a set precision to determine the maximum pruning rate and minimum quantization bit width of each network layer, the accuracy of determining the maximum pruning rate and minimum quantization bit width can be improved.
[0028] For example, the formula for determining the first result value can be:
[0029]
[0030] Where Result1 is the first result value. For the highest pruning rate of the i-th network layer, pr threshold To set the pruning rate threshold.
[0031] For example, the formula for determining the second result value can be:
[0032]
[0033] Result2 is the second result value. qb is the minimum quantization bit width of the i-th network layer. threshold To set the quantization bit width threshold.
[0034] In this embodiment, by quantitatively evaluating the sensitivity of each network layer to pruning and quantization, the sensitivity type of each network layer can be accurately distinguished, which facilitates the precise layered customization of the model compression strategy and maximizes compression efficiency while ensuring accuracy.
[0035] S120. Based on the architecture parameters of the first set number of candidate pruning rates and the architecture parameters of the second set number of candidate quantization bit widths corresponding to the sensitivity type of each network layer, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected for each network layer are determined respectively.
[0036] In this embodiment, if the sensitivity type is pruning-sensitive, it indicates that under the same compression ratio, pruning the network layer results in a higher performance loss than quantization, suggesting that the network layer is more suitable for compression using quantization. Therefore, when setting the initial search space for the network layer, more candidate quantization bit widths are assigned (the total number of candidate quantization bit widths is greater than the total number of candidate pruning rates, i.e., the second set number is greater than the first set number). When the sensitivity type is quantization-sensitive, the total number of candidate pruning rates is greater than the total number of candidate quantization bit widths, i.e., the first set number is greater than the second set number.
[0037] In this embodiment, for each candidate pruning rate of each network layer, the probability of the corresponding candidate pruning rate of the corresponding network layer being selected is determined based on the architecture parameters of the corresponding candidate pruning rate and the first set quantity; for each candidate quantization bit width of each network layer, the probability of the corresponding candidate quantization bit width of the corresponding network layer being selected is determined based on the architecture parameters of the corresponding candidate quantization bit width and the second set quantity.
[0038] For example, determine the j-th candidate pruning rate in the i-th layer. The method is as follows:
[0039]
[0040] in, Let be the architecture parameter for the j-th candidate pruning rate in the i-th layer. C is the total number of candidate pruning rates (i.e., the first set number). Let be the architecture parameters corresponding to the k-th candidate pruning rate of the i-th layer.
[0041] For example, determine the probability that the j-th candidate quantization bit width in the i-th layer is selected. The method is as follows:
[0042]
[0043] in, Let B be the architecture parameter corresponding to the j-th candidate quantization bit width in the i-th layer. B is the total number of candidate quantization bit widths (i.e., the second set number). Let be the architecture parameter corresponding to the a-th candidate quantization bit width of the i-th layer.
[0044] S130. Based on the first training set and validation set, update the probability of each candidate pruning rate of each network layer being selected by combining the highest pruning rate, and update the probability of each candidate quantization bit width of each network layer being selected by combining the lowest quantization bit width.
[0045] In this embodiment, the basic neural network can be compressed according to the candidate pruning rate and the candidate quantization bit width with the highest probability in each network layer in the current round to obtain the second target deep neural network. Using the idea of differentiable neural structure search, the weight parameters and architecture parameters of each network layer of the second target deep neural network are updated based on the first training set and validation set, thereby updating the probability of the candidate pruning rate and the probability of the candidate quantization bit width being selected in each network layer. During the search process, a new candidate pruning rate is generated based on the highest pruning rate and a new candidate quantization bit width is generated based on the lowest quantization bit width, and the candidate pruning rate and candidate quantization bit width with the lowest probability are replaced accordingly.
[0046] Optionally, based on the first training set and validation set, the probability of each candidate pruning rate of each network layer being selected is updated by combining the highest pruning rate, and the probability of each candidate quantization bit width of each network layer being selected is updated by combining the lowest quantization bit width, including: each iteration process is as follows: compressing the basic neural network according to the candidate pruning rate with the highest probability and the candidate quantization bit width with the highest probability in the current round to obtain a second target deep neural network; determining the training loss of the second target deep neural network based on the first training set; wherein, the training loss is used to update the weight parameters of the second target deep neural network; determining the validation loss of the second target deep neural network based on the validation set; wherein, the The verification loss is used to update the weight parameters and the architecture parameters of each network layer of the second target deep neural network, so as to update the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected in each network layer; when the verification loss is minimized and the training loss is minimized, the current iteration ends; based on the highest pruning rate and the lowest quantization bit width of each network layer, new candidate pruning rates and new candidate quantization bit widths are generated for each network layer; the target candidate pruning rates and target candidate quantization bit widths of each network layer are replaced with the new candidate pruning rates and new candidate quantization bit widths; when the number of iterations is greater than or equal to the set number of rounds, all iteration processes end.
[0047] In this embodiment, the idea of differentiable neural structure search can be used to update the weight parameters and architecture parameters of each network layer of the second target deep neural network based on the first training set and validation set, thereby updating the probability of each candidate pruning rate and the probability of each candidate quantization bit width being selected in each network layer. During the search process, a new candidate pruning rate is generated based on the highest pruning rate and a new candidate quantization bit width is generated based on the lowest quantization bit width, and the candidate pruning rate and candidate quantization bit width with the lowest probability are replaced accordingly.
[0048] The specific process for each iteration is as follows:
[0049] The base neural network is compressed according to the candidate pruning rate and the candidate quantization bit width with the highest probability in each network layer of the current round to obtain the second target deep neural network. The training loss of the second target deep neural network is calculated on the first training set to update the weight parameters of the second target deep neural network. The validation loss of the second target deep neural network is calculated on the validation set to update the weight parameters and the architecture parameters of each network layer of the second target deep neural network, thereby updating the probability of each candidate pruning rate and the probability of each candidate quantization bit width being selected in each network layer. The validation loss is jointly determined by the weight parameters and the network architecture (second target deep neural network) obtained by pruning according to the candidate pruning rate with the highest probability and quantizing according to the candidate quantization bit width with the highest probability in each network layer during the search process. The final search goal is to obtain an ideal second target deep neural network with the lowest validation loss and that minimizes the training loss. That is, the current iteration ends when both the validation loss and the training loss are minimized.
[0050]
[0051] in, Let A be the ideal weight parameters in the ideal second-target deep neural network, and ω be the weight parameters of the second-target deep neural network. val To verify the loss, L train This is due to training losses.
[0052] After each round of search and training, new candidate pruning rates and new candidate quantization bit widths are randomly generated for each network layer. The new candidate pruning rates are limited to being less than or equal to the highest pruning rate of the corresponding network layer, and the new candidate quantization bit widths are limited to being greater than or equal to the lowest quantization bit width of the corresponding network layer. If a new candidate pruning rate is greater than the highest pruning rate of the corresponding network layer, a new candidate quantization bit width is greater than or equal to the lowest quantization bit width, or a new candidate quantization bit width is less than the lowest quantization bit width of the corresponding network layer, then it should be discarded and regenerated. The target candidate pruning rates and target candidate quantization bit widths of each network layer are replaced with the new candidate pruning rates and new candidate quantization bit widths. When the number of iterations is greater than or equal to a set number of rounds, all iterations end. In this embodiment, the set number of rounds is not limited; for example, it can be 100 rounds.
[0053] In this embodiment, the new candidate pruning rate and new candidate quantization bit width generated by the corresponding constraints of the highest pruning rate and the lowest quantization bit width can effectively reduce the search space.
[0054] Specifically, for each network layer, among all the candidate pruning rates with the lowest probability corresponding to multiple iterations in each round, the candidate pruning rate that appears most frequently is taken as the target candidate pruning rate for the corresponding network layer. For each network layer, among all the candidate quantization bit widths with the lowest probability corresponding to multiple iterations in each round, the candidate quantization bit width that appears most frequently is taken as the target candidate quantization bit width for the corresponding network layer.
[0055] In this embodiment, each iteration includes multiple iterations (the forward propagation and backward propagation processes count as one iteration). In each iteration, the probability of selecting a candidate pruning rate and the probability of selecting a candidate quantization bit width for each network layer are updated. Therefore, each network layer will have a candidate pruning rate and a candidate quantization bit width with the lowest probability in each iteration. For each network layer, among all the candidate pruning rates with the lowest probability corresponding to multiple iterations in each round, the candidate pruning rate that appears most frequently is taken as the target candidate pruning rate for the corresponding network layer. For each network layer, among all the candidate quantization bit widths with the lowest probability corresponding to multiple iterations in each round, the candidate quantization bit width that appears most frequently is taken as the target candidate quantization bit width for the corresponding network layer.
[0056] In this embodiment, the candidate pruning rate that appears most frequently among all the candidate pruning rates with the lowest probability in each round of multiple iterations is taken as the target candidate pruning rate of the corresponding network layer. Similarly, the candidate quantization bit width that appears most frequently among all the candidate quantization bit widths with the lowest probability in each round of multiple iterations is taken as the target candidate quantization bit width of the corresponding network layer. This approach ensures that the candidate pruning rate and candidate quantization bit width with the worst quality are replaced in each round.
[0057] In this embodiment, during the search process, by continuously replacing the candidate pruning rate and candidate quantization bit width with poor quality, the search space is updated and evolved, making it easier to find a simplified network architecture with better quality; the search and update are carried out alternately, taking into account both search efficiency and search quality.
[0058] S140. Compress each network layer according to the candidate pruning rate with the highest probability after each network layer is updated and the candidate quantization bit width with the highest probability after each network layer is updated, to obtain the compressed model.
[0059] In this embodiment, after obtaining the candidate pruning rate and the candidate quantization bit width with the highest probability after the final update of each network layer, pruning and quantization are performed on each network layer according to these parameters to obtain a compressed model. The compressed model is a simplified network. Compression includes pruning and quantization. Pruning refers to network pruning, and quantization refers to model quantization; both are model compression methods.
[0060] Optionally, after compressing each network layer according to the candidate pruning rate with the highest probability after each network layer is updated and the candidate quantization bit width with the highest probability after each network layer is updated, to obtain the compressed model, the method further includes: fine-tuning the compressed model based on the second training set.
[0061] In this embodiment, the compressed model is fine-tuned using the second training set to obtain a simplified network, thereby restoring the performance of the simplified network.
[0062] In this embodiment, the sensitivity of different network layers in the basic deep neural network to pruning and quantization is considered. Compared with a single compression strategy, the hybrid search space composed of candidate pruning rate and candidate quantization bit width helps to further improve the compression rate and performance. It can automatically select the appropriate pruning rate for pruning and the appropriate quantization bit width for quantization for each network layer.
[0063] The technical solution of this application embodiment determines the sensitivity type of each network layer based on the highest pruning rate and the lowest quantization bit width of each network layer of the basic deep neural network; based on the architecture parameters of a first predetermined number of candidate pruning rates and a second predetermined number of candidate quantization bit widths corresponding to the sensitivity type of each network layer, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected of each network layer are determined respectively; based on a first training set and a validation set, the probability of each candidate pruning rate being selected of each network layer is updated by combining the highest pruning rate and the probability of each candidate quantization bit width being selected of each network layer is updated by combining the lowest quantization bit width; and each network layer is compressed according to the candidate pruning rate with the highest probability after the update and the candidate quantization bit width after the update to obtain a compressed model. In this embodiment, the sensitivity type of each network layer is determined by the highest pruning rate and the lowest quantization bit width. Based on the architectural parameters of the candidate pruning rates and candidate quantization bit widths corresponding to the sensitivity types of each network layer, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected are determined. Based on the first training set and validation set, and by combining the highest pruning rate and the lowest quantization bit width, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected are updated. This method can automatically find suitable compression strategies for each network layer of the deep neural network, achieving a higher balance between compression efficiency and accuracy.
[0064] Figure 2 This is a schematic diagram of a model compression device provided in an embodiment of this application, as shown below. Figure 2 As shown, the device includes: a sensitivity type determination module 210, a probability determination module 220, a probability update module 230, and a compression module 240;
[0065] Sensitivity type determination module 210 is used to determine the sensitivity type of each network layer based on the highest pruning rate and lowest quantization bit width of each network layer of the basic deep neural network.
[0066] The probability determination module 220 is used to determine the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected for each network layer based on the architecture parameters of a first set number of candidate pruning rates and a second set number of candidate quantization bit widths corresponding to the sensitivity type of each network layer, respectively.
[0067] The probability update module 230 is used to update the probability of each candidate pruning rate of each network layer being selected based on the first training set and validation set, combined with the highest pruning rate, and to update the probability of each candidate quantization bit width of each network layer being selected based on the lowest quantization bit width.
[0068] Compression module 240 is used to compress each network layer according to the candidate pruning rate with the highest probability after each network layer is updated and the candidate quantization bit width with the highest probability after each network layer is updated, so as to obtain a compressed model.
[0069] The technical solution of this application embodiment determines the sensitivity type of each network layer based on the highest pruning rate and lowest quantization bit width of each network layer in the basic deep neural network through a sensitivity type determination module; determines the probability of each candidate pruning rate and each candidate quantization bit width being selected for each network layer based on the architecture parameters of a first predetermined number of candidate pruning rates and a second predetermined number of candidate quantization bit widths corresponding to the sensitivity type of each network layer; updates the probability of each candidate pruning rate being selected for each network layer based on a first training set and a validation set, combined with the highest pruning rate, and updates the probability of each candidate quantization bit width being selected for each network layer, combined with the lowest quantization bit width; and compresses each network layer according to the candidate pruning rate and the candidate quantization bit width with the highest probability after the update, using a compression module to compress the respective network layers to obtain a compressed model. In this embodiment, the sensitivity type of each network layer is determined by the highest pruning rate and the lowest quantization bit width. Based on the architectural parameters of the candidate pruning rates and candidate quantization bit widths corresponding to the sensitivity types of each network layer, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected are determined. Based on the first training set and validation set, and by combining the highest pruning rate and the lowest quantization bit width, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected are updated. This method can automatically find suitable compression strategies for each network layer of the deep neural network, achieving a higher balance between compression efficiency and accuracy.
[0070] Optionally, the sensitivity type determination module is specifically used for: determining the highest pruning rate and lowest quantization bit width of each network layer of the basic deep neural network based on the second training set and a set precision; determining a first result value based on the second training set and for each network layer, based on the highest pruning rate of the corresponding network layer and a set pruning rate threshold; determining a second result value based on the lowest quantization bit width of the corresponding network layer and a set quantization bit width threshold; if the first result value is greater than the second result value, then the sensitivity type of the corresponding network layer is quantization sensitive; if the first result value is less than or equal to the second result value, then the sensitivity type of the corresponding network layer is pruning sensitive.
[0071] Optionally, the sensitive type determination module is further configured to: compress the basic deep neural network based on the set pruning rate and set quantization bit width of each network layer of the basic deep neural network to obtain a first target deep neural network; train the first target deep neural network based on a second training set to obtain training accuracy; wherein, the second training set includes the first training set and the validation set; if the training accuracy is greater than or equal to the set accuracy, then the set pruning rate of each network layer is respectively determined as the highest pruning rate of each network layer and the set quantization bit width of each network layer is respectively determined as the lowest quantization bit width of each network layer; if the training accuracy is less than the set accuracy, then the set pruning rate and set quantization bit width of each network layer are adjusted, and the process is returned to and the step of compressing the basic deep neural network based on the set pruning rate and set quantization bit width of each network layer of the basic deep neural network to obtain the compressed basic deep neural network is executed again.
[0072] Optionally, the probability update module is specifically used for the following iteration process in each round: compressing the basic neural network according to the candidate pruning rate and the candidate quantization bit width with the highest probability in each network layer of the current round to obtain a second target deep neural network; determining the training loss of the second target deep neural network based on the first training set; wherein the training loss is used to update the weight parameters of the second target deep neural network; determining the validation loss of the second target deep neural network based on the validation set; wherein the validation loss is used to update the weight parameters and the architecture parameters of each network layer of the second target deep neural network to update the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected in each network layer; when the validation loss is minimized and the training loss is minimized, the current round of iteration ends; generating new candidate pruning rates and new candidate quantization bit widths for each network layer based on the highest pruning rate and the lowest quantization bit width of each network layer; replacing the target candidate pruning rate and target candidate quantization bit width of each network layer with the new candidate pruning rate and the new candidate quantization bit width; ending all iteration processes when the number of iterations is greater than or equal to the set number of rounds.
[0073] Specifically, for each network layer, among all the candidate pruning rates with the lowest probability corresponding to multiple iterations in each round, the candidate pruning rate that appears most frequently is taken as the target candidate pruning rate for the corresponding network layer; for each network layer, among all the candidate quantization bit widths with the lowest probability corresponding to multiple iterations in each round, the candidate quantization bit width that appears most frequently is taken as the target candidate quantization bit width for the corresponding network layer.
[0074] Optionally, the above-mentioned device further includes a fine-tuning module for fine-tuning the compressed model based on the second training set.
[0075] The model compression apparatus provided in this application embodiment can execute the model compression method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method execution.
[0076] Figure 3 A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0077] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0078] Multiple components in electronic device 10 are connected to input / output (I / O) interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0079] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as method model compression.
[0080] In some embodiments, method model compression may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of the method model compression described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform method model compression by any other suitable means (e.g., by means of firmware).
[0081] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0082] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0083] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0084] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0085] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0086] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0087] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the model compression method provided in any embodiment of this application.
[0088] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0089] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.
Claims
1. A model compression method, characterized in that, include: Based on the highest pruning rate and lowest quantization bit width of each layer of the basic deep neural network, the sensitivity type of each layer is determined. Based on the architecture parameters of the first set number of candidate pruning rates and the second set number of candidate quantization bit widths corresponding to the sensitivity type of each network layer, the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected for each network layer are determined respectively. Based on the first training set and validation set, the probability of each candidate pruning rate being selected for each network layer is updated by combining the highest pruning rate, and the probability of each candidate quantization bit width being selected for each network layer is updated by combining the lowest quantization bit width. The network layers are compressed according to the candidate pruning rate with the highest probability after each network layer is updated and the candidate quantization bit width with the highest probability after each network layer is updated, so as to obtain the compressed model.
2. The method according to claim 1, characterized in that, Based on the highest pruning rate and lowest quantization bit width of each layer in the basic deep neural network, the sensitivity type of each layer is determined, including: Based on the second training set and the set precision, the highest pruning rate and the lowest quantization bit width of each network layer of the basic deep neural network are determined. Based on the second training set and for each network layer, the first result value is determined based on the highest pruning rate of the corresponding network layer and a set pruning rate threshold. The second result value is determined based on the minimum quantization bit width of the corresponding network layer and the set quantization bit width threshold. If the first result value is greater than the second result value, then the sensitivity type of the corresponding network layer is quantization sensitive. If the first result value is less than or equal to the second result value, then the sensitivity type of the corresponding network layer is pruning sensitive.
3. The method according to claim 2, characterized in that, Based on the second training set and a set precision, the highest pruning rate and lowest quantization bit width of each layer of the basic deep neural network are determined, including: Based on the set pruning rate and set quantization bit width of each network layer of the basic deep neural network, the basic deep neural network is compressed to obtain the first target deep neural network. The first target deep neural network is trained based on the second training set to obtain training accuracy; wherein, the second training set includes the first training set and the validation set; If the training accuracy is greater than or equal to the set accuracy, then the set pruning rate of each network layer is respectively determined as the highest pruning rate of each network layer and the set quantization bit width of each network layer is respectively determined as the lowest quantization bit width of each network layer. If the training accuracy is less than the set accuracy, the set pruning rate and set quantization bit width of each network layer are adjusted, and the process is repeated to compress the basic deep neural network based on the set pruning rate and set quantization bit width of each network layer of the basic deep neural network to obtain the compressed basic deep neural network.
4. The method according to claim 1, characterized in that, Based on the first training set and validation set, the probability of each candidate pruning rate being selected for each network layer is updated by combining the highest pruning rate, and the probability of each candidate quantization bit width being selected for each network layer is updated by combining the lowest quantization bit width, including: The iteration process for each round is as follows: The basic neural network is compressed according to the candidate pruning rate with the highest probability and the candidate quantization bit width with the highest probability of each network layer in the current round to obtain the second target deep neural network. The training loss of the second target deep neural network is determined based on the first training set; wherein the training loss is used to update the weight parameters of the second target deep neural network. The validation loss of the second target deep neural network is determined based on the validation set; wherein the validation loss is used to update the weight parameters and the architecture parameters of each network layer of the second target deep neural network, so as to update the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected in each network layer. The current iteration ends when both the validation loss and the training loss are minimized. Based on the highest pruning rate and the lowest quantization bit width of each network layer, new candidate pruning rates and new candidate quantization bit widths are generated for each network layer. Replace the target candidate pruning rate and target candidate quantization bit width of each network layer with the new candidate pruning rate and the new candidate quantization bit width; When the number of iterations is greater than or equal to the set number of rounds, the entire iteration process ends.
5. The method according to claim 4, characterized in that, in, For each network layer, among all the candidate pruning rates with the lowest probability corresponding to multiple iterations in each round, the candidate pruning rate that appears most frequently is taken as the target candidate pruning rate of the corresponding network layer; for each network layer, among all the candidate quantization bit widths with the lowest probability corresponding to multiple iterations in each round, the candidate quantization bit width that appears most frequently is taken as the target candidate quantization bit width of the corresponding network layer.
6. The method according to claim 1, characterized in that, After compressing each network layer according to the candidate pruning rate with the highest probability after the update of each network layer and the candidate quantization bit width with the highest probability after the update of each network layer, to obtain the compressed model, the process further includes: The compressed model was fine-tuned based on the second training set.
7. A model compression device, characterized in that, include: The sensitivity type determination module is used to determine the sensitivity type of each network layer based on the highest pruning rate and lowest quantization bit width of each network layer of the basic deep neural network. The probability determination module is used to determine the probability of each candidate pruning rate being selected and the probability of each candidate quantization bit width being selected for each network layer based on the architecture parameters of a first set number of candidate pruning rates and a second set number of candidate quantization bit widths corresponding to the sensitivity type of each network layer. The probability update module is used to update the probability of each candidate pruning rate of each network layer being selected based on the first training set and validation set, combined with the highest pruning rate, and to update the probability of each candidate quantization bit width of each network layer being selected based on the lowest quantization bit width. The compression module is used to compress each network layer according to the candidate pruning rate with the highest probability after the update of each network layer and the candidate quantization bit width with the highest probability after the update of each network layer, so as to obtain the compressed model.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the model compression method as described in any one of claims 1-6.
9. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the model compression method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the model compression method as described in any one of claims 1-6.