Self-defined widening method for low-precision neural network

By performing pruning sensitivity analysis and quantization perception training on the neural network, the degree of expansion of the outstanding contribution layer is selectively increased when the width is expanded, and the problem of low-precision neural network widening method in the existing technology is solved, which significantly improves the model's inference accuracy.

CN120031092APending Publication Date: 2025-05-23UNIV OF CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311554596.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing low-precision neural network widening methods uniformly expand each layer when width is expanded, resulting in the inability to achieve peak performance.

Method used

By performing pruning sensitivity analysis on the target neural network, the impact of each layer on the final performance at different sparsity is obtained, and the impact of each layer in sparseness is quantified according to the contribution of each layer, and then selectively increase the extent of expansion of the layer with outstanding contributions when expanding the width.

Benefits of technology

The peak performance when the accuracy is reduced is improved, and the inference accuracy of the model can be significantly improved while maintaining a low quantization bit width.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031092A_ABST
    Figure CN120031092A_ABST
Patent Text Reader

Abstract

The invention discloses a low-precision neural network custom widening method, and the method comprises the steps: carrying out the pruning sensitivity analysis of a target neural network, and obtaining the overall performance of the network when each layer of the target neural network is at different sparseness; calculating the contribution of each layer to the performance of the target neural network according to the overall performance of the network when each layer of the target neural network is at different sparseness; according to the contribution of each layer to the performance of the target neural network, the widening proportion of the layer is calculated, and width expansion is carried out on the neural network of the layer; and performing quantitative perception training on the target neural network after width expansion to obtain a final model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computers and deep learning, and more specifically, to a low-precision neural network custom widening method based on network pruning sensitivity. Background Art

[0002] Generally speaking, in recent years, the development of deep learning and artificial intelligence technology has provided great convenience for people's lives and created huge productivity. Deep learning models have been applied to various fields of various industries, creating greater scientific and economic value than traditional non-artificial intelligence methods.

[0003] Model compression is an important but challenging task in deep learning. Model compression includes typical methods such as pruning, quantization, and knowledge distillation, which have been widely used in the deployment of deep learning model applications. Existing solutions, such as the wide reduced-precision network (WRPN: Wide Reduced-Precision Networks[C]. International Conference on Learning Representations, 2018.) proposed by Asit Mishra et al. in 2018, reduce the computational and storage requirements of deep neural networks by using width expansion and reduced precision while maintaining high accuracy.

[0004] However, the above method uniformly expands each layer of the deep neural network when expanding the width. Uniform expansion expands each layer of the network in equal proportion. Each layer of the neural network contributes differently to the final result, which means that the above method cannot bring peak performance to the neural network.

[0005] In summary, considering the limitations of the above methods, the present invention provides a low-precision neural network custom widening method based on neural network pruning sensitivity analysis to solve the above problems. The method obtains the impact of each layer on the final performance when the sparsity is increased based on the neural network pruning sensitivity analysis, and quantifies the impact of each layer on the final result when it is sparse to obtain the contribution of each layer to the final result, and improves the peak performance of the network when the precision is reduced by selectively increasing the expansion degree of the layer with outstanding contribution when the width is expanded. The core algorithm of the present invention has good scalability and can be fully applied to all types of current convolutional neural networks; the method described in the present invention can greatly improve the inference accuracy of the model while maintaining a low quantization bit width; the method can be easily applied to the deployment and implementation of almost all deep learning fields, and has industrial application prospects. Summary of the invention

[0006] The purpose of the present invention is to provide a low-precision neural network custom widening method, which can overcome the problem of low peak performance of existing low-precision neural network widening methods and can be effectively applied to the deployment of large-scale deep learning applications.

[0007] In order to achieve the above-mentioned purpose of the invention, according to one aspect of the present invention, a method for custom widening a low-precision neural network is provided, the method comprising: performing pruning sensitivity analysis on the target neural network to obtain the overall performance of each layer of the target neural network at different sparsities; calculating the contribution of each layer to the performance of the target neural network according to the overall performance of each layer of the target neural network at different sparsities; calculating the widening ratio of each layer according to the contribution of the layer to the performance of the target neural network, and performing width expansion on the neural network layer; performing quantization perception training on the target neural network after width expansion to obtain the final model.

[0008] The steps of the method include:

[0009] Perform pruning sensitivity analysis on the target neural network to obtain the overall performance of the network at different sparsities for each layer of the target neural network. Use Top1 as the performance metric;

[0010] According to the performance of each layer of the target neural network at different sparsity, the contribution of each layer to the performance of the target neural network is calculated. Specifically, the arithmetic mean L of the Top1 index of each layer of the target neural network at all sparsity is calculated. average_top1 ;

[0011] The widening ratio of each layer is calculated based on its contribution to the performance of the target neural network, and the width of the neural network layer is expanded. First, all layers of the target network structure are divided into several groups according to the connection topology of the target neural network structure. The layers whose input channels need to be kept consistent in the target network structure are grouped into the same group, and the layers whose input channels are not constrained are grouped independently. A group number is assigned to each group. The group number assigned to each layer is denoted as L. group_constraints . Secondly, the number of convolution kernels in each layer is L original_filter_num and the group number L assigned to the layer group_constraints According to the arithmetic mean L of the Top1 index at all sparsities of this layer average_top1 Then, given the basic widening ratio α, assign a subscript l to each layer after sorting, and calculate the widening ratio B[L group_constraints [l]]=max(B[L group_constraints [l]],α-0.5l / L average_top1 [l]). Finally, the number of convolution kernels L in each layer after widening according to the widening ratio of the group is calculated.resize_filter_num , L resize_filter_num [l] = Round (B [L group_constraints [l]]*L original_filter_num [l],0). Use the number of convolution kernels L after each layer is proportionally widened resize_filter_num Modify the width of each layer of the target neural network structure to perform width expansion;

[0012] The target neural network after width expansion is trained with quantization awareness to obtain the final model. The modified target neural network is used to perform neural network quantization awareness training while maintaining the quantization bit width of 1 bit for weight and 1 bit for activation. After the training, a low-precision custom widened neural network is obtained, which has better performance and smaller parameters than the non-custom widened low-precision neural network.

[0013] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-readable instructions, and wherein the computer-readable instructions implement the above-mentioned low-precision neural network custom widening method when executed by a processor.

[0014] The core algorithm of the present invention has good scalability and can be fully applied to all current types of convolutional neural networks; the method of the present invention can greatly improve the inference accuracy of the model while maintaining a low quantization bit width; the method can be conveniently applied to the deployment and implementation of almost all deep learning fields, and has industrial application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flow chart of a method for customizing widening a low-precision neural network according to an embodiment of the present invention;

[0016] Figure 2 is a schematic diagram of a process of performing pruning sensitivity analysis on a target neural network according to an embodiment of the present invention;

[0017] Figure 3 is a schematic diagram of a process for calculating the contribution of each layer to the performance of a target neural network according to an embodiment of the present invention;

[0018] Figure 4 is a schematic diagram of a process of calculating a widening ratio of each layer and expanding the width of the neural network layer according to an embodiment of the present invention; and

[0019] Figure 5 It is a flowchart of performing quantization-aware training and obtaining a final model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The following description is made with reference to the accompanying drawings and is provided to assist in a comprehensive understanding of the exemplary embodiments of the present invention as defined by the claims and their equivalents. Various details are disclosed to assist in understanding, but these are considered as examples only. Therefore, it will be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, for clarity and brevity, descriptions of well-known functions and structures may be omitted.

[0021] The terms and words used in the following description and claims are not limited to the written meanings, but are only used to enable a clear and consistent understanding of the present disclosure. Therefore, it should be clear to those skilled in the art that the following description of exemplary embodiments of the present disclosure is provided for illustrative purposes only, and not for the purpose of limiting the present disclosure as defined by the attached claims and their equivalents.

[0022] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] Figure 1 It is a flowchart of a low-precision neural network custom widening method according to an embodiment of the present invention.

[0024] The target neural network can be subjected to pruning sensitivity analysis by a norm-based deep neural network pruning algorithm. The deep neural network pruning algorithm can be any pruning algorithm based on the neural network weight amplitude, for example, it can be a deep convolutional neural network structured pruning algorithm based on the L1 norm. Of course, any pruning algorithm based on the neural network weight amplitude can be used.

[0025] Reference Figure 1 The low-precision neural network custom widening method based on network pruning sensitivity analysis proposed in the present invention comprises the following steps:

[0026] In step S1, a pruning sensitivity analysis may be performed on the target neural network to obtain the overall performance of the network at different sparsities for each layer of the target neural network.

[0027] Specifically, refer to Figure 2 Describes performing pruning sensitivity analysis on a target neural network.

[0028] Specifically, Figure 2 As shown, in step S101, the first layer of the neural network is selected.

[0029] Then, proceed to step S102. In step S102, the L1 norm of the filter in the weight tensor of the layer is calculated.

[0030] Then, the process proceeds to step S103, in which x% is set to 0 (ie, no pruning is performed).

[0031] Then, the process proceeds to step S104, in which x% of filters with the smallest L1 norm are set to zero.

[0032] Then, proceed to step S105, in which the performance of the current model is evaluated.

[0033] Next, proceed to step S106 to determine whether x% reaches 90%. If not, return to S104 and increase x% by 5%; if yes, proceed to the next step.

[0034] Next, proceed to step S107 to determine whether there are any unprocessed neural network layers. If yes, return to S102 and select the next layer; if no, proceed to the next step.

[0035] Next, return Figure 1 , in the low-precision neural network custom widening method based on network pruning sensitivity analysis, step S2 is performed. In step S2, based on the overall performance of each layer of the target neural network at different sparsities obtained in step S1, the contribution of each layer to the performance of the target neural network is calculated.

[0036] Specifically, refer to Figure 3 Calculate the contribution of each layer to the performance of the target neural network.

[0037] Specifically, Figure 3 As shown, in step S201, the Top1 indicators of all layers of the target neural network at all sparsities are extracted.

[0038] Then, proceed to step S202. In step S202, for each layer, calculate the arithmetic mean L of its Top1 index. average_top1 , and derive the contribution of this layer to the network performance.

[0039] Next, return Figure 1 , step S3 is performed in the low-precision neural network custom widening method based on network pruning sensitivity analysis. In step S3, the widening ratio of each layer is calculated based on the contribution of each layer to the target neural network performance obtained in step S2, and the width of the neural network layer is expanded.

[0040] Specifically, refer to Figure 4 Calculate the widening ratio of each layer and expand the width of the neural network layer.

[0041] Specifically, Figure 4As shown, in step S301, according to the connection topology of the target neural network structure, all layers of the target network structure are divided into several groups, the layers whose input channel numbers need to remain consistent are grouped into the same group, and the layers whose input channel numbers are not constrained are grouped independently.

[0042] Then, proceed to step S302. In step S302, a group number is assigned to each group. The group number assigned to each layer is recorded as L. group_constraints .

[0043] Then, proceed to step S303, in which the number of convolution kernels of each layer is L original_filter_num and the group number L assigned to the layer group_constraints According to the arithmetic mean L of the Top1 index at all sparsities of this layer average_top1 Sort in ascending order.

[0044] Then, the process proceeds to step S304, in which a subscript l is assigned to each of the sorted layers given a basic widening ratio α.

[0045] Then, the process proceeds to step S305, in which the widening ratio B[L group_constraints [l]]=max(B[L group_constraints [l]],α-0.5l / L average_top1 [l]).

[0046] Then, proceed to step S306, in which the number of convolution kernels L of each layer after widening according to the widening ratio of the group is calculated. resize_filter_num , L resize_filter_num [l] = Round (B [L group_constraints [l]]*L original_filter_num [l],0).

[0047] Then, proceed to step S307, in which the number of convolution kernels L after each layer is proportionally widened is used. resize_filter_num Modify the width of each layer of the target neural network structure to perform width expansion.

[0048] Next, return Figure 1 , step S4 is performed in the low-precision neural network custom widening method based on network pruning sensitivity analysis. In step S4, quantization-aware training is performed based on the target neural network after width expansion obtained in step 3 to obtain the final model.

[0049] Specifically, refer to Figure 5 Perform quantization-aware training to obtain the final model.

[0050] In step S401, we use the target neural network that has been width-expanded as the network to be trained. This network has been designed and optimized to adapt to more complex tasks and larger amounts of data.

[0051] Next, we proceed to step S402. At this stage, we begin preparations for quantization-aware training. First, we need to prepare the training data, including collecting and cleaning the data, and dividing the data into training and validation sets. Then, we need to set the training parameters, such as learning rate, batch size, optimizer type, etc.

[0052] Then, we go to step S403 and start quantization-aware training. In this step, we quantize the weights of the neural network from floating point numbers to low-precision integers. Specifically, we quantize the weights and activations to 1 bit. The quantization scheme we use is symmetric quantization, which is a common quantization method. It can ensure that the numerical distribution after quantization still maintains the original symmetry, thereby reducing the quantization error. After determining the quantization parameters, we can quantize the weights and activations of the neural network to 1 bit. During the training process, we need to consider both the quantization error and the task error to achieve the best training effect.

[0053] Finally, in step S404, the training is completed and we obtain a low-precision custom widened neural network. This neural network has better performance and smaller parameters than the non-custom widened low-precision neural network while maintaining the quantization bit width of 1 bit for weights and 1 bit for activations. This means that through this training process, we have successfully created a neural network that is both efficient and resource-saving. It has a smaller number of parameters while completing the task, and is more suitable for running on resource-constrained devices.

[0054] While the inventive concept has been described with reference to exemplary embodiments thereof, it will be apparent to those skilled in the art that various changes and modifications may be made thereto without departing from the scope of the inventive concept as set forth in the appended claims.

Claims

1. A method for custom widening a low-precision neural network, the method include: Perform pruning sensitivity analysis on the target neural network to obtain the overall performance of the network at different sparsities for each layer of the target neural network; Calculate the contribution of each layer to the performance of the target neural network based on the overall performance of the network at different sparsities of each layer of the target neural network; The widening ratio of each layer is calculated based on the contribution of each layer to the performance of the target neural network, and the width of the neural network of this layer is expanded; as well as The target neural network after width expansion is trained with quantization perception to obtain the final model.

2. The method according to claim 1, in, Top1 is used as the performance measurement indicator.

3. The method according to claim 2, in, The contribution of each layer to the performance of the target neural network is calculated based on the overall performance of the network at different sparsities, including: Extract the top 1 indicators of all layers of the target neural network at all sparsities; For each layer, calculate the arithmetic mean L of its Top1 index average_top1 , and derive the contribution of this layer to the network performance.

4. The method according to claim 1, in, The widening ratio of each layer is calculated based on its contribution to the performance of the target neural network, and the width of the neural network is expanded including: According to the connection topology of the target neural network structure, all layers of the target network structure are divided into several groups. The layers whose input channel numbers need to be consistent in the target network structure are grouped into the same group, and the layers whose input channel numbers are not constrained are grouped independently. A group number is assigned to each group. The group number assigned to each layer is denoted as L. group_constraints ; The number of convolution kernels in each layer is L original_filter_num and the group number L assigned to the layer group_constraints According to the arithmetic mean L of the Top1 index at all sparsities of this layer average_top1 Then, given the basic widening ratio α, assign a subscript l to each layer after sorting, and calculate the widening ratio B[L group_constraints [l]]=max(B[L group_constraints [l]],α-0.5l / L average_top1 [l]); Calculate the number of convolution kernels L after each layer is widened according to the widening ratio of the group to which it belongs resize_filter_num , L resize_filter_num [l] = Round (B [L group_constraints [l]]*L original_filter_num [l],0), using the number of convolution kernels L after each layer is proportionally widened resize_filter_num Modify the width of each layer of the target neural network structure to perform width expansion.

5. The method according to claim 1, in, The target neural network after width expansion is subjected to quantization-aware training to obtain the final model, including: using the modified target neural network to perform neural network quantization-aware training while maintaining the quantization bit width of 1 bit for weight and 1 bit for activation, and obtaining a low-precision custom widened neural network after the training.

6. A computer-readable storage medium storing computer-readable instructions. It is characterized in that When the computer-readable instructions are executed by a processor, the low-precision neural network custom widening method of claim 1 is implemented.