Deep learning model pruning technology based on FPGM

Through the deep learning model pruning technology based on FPGM, redundant parameters and connections are eliminated, and compact and efficient model structure is generated, which solves the problem of excessive storage and computing resources of deep learning models on mobile devices and embedded systems, and achieves efficient deployment and rapid inference.

CN120471130APending Publication Date: 2025-08-12SMIC FUTURE (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410159967.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-04
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The storage and computing resources demands of deep learning models on mobile devices and embedded systems are too high, resulting in low deployment efficiency, excessive inference latency and power consumption, and it is difficult for the existing technology to effectively solve these problems.

Method used

Using FPGM-based deep learning model pruning technology, by eliminating redundant parameters and connections, using importance evaluation, soft pruning, pruning mask and weight-removing modules, a compact and efficient model structure is generated, retaining the main features and performance.

Benefits of technology

Significantly reduce storage and transmission costs, improve deployment efficiency, reduce computing resource requirements and power consumption, accelerate inference processes, and support online updates and rapid deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471130A_ABST
    Figure CN120471130A_ABST
Patent Text Reader

Abstract

The invention relates to a deep learning model pruning technology based on an FPGM (Field Programmable Gate Array), which is a method for optimizing a deep learning model and realizes simplification and compression of the model by removing redundant and unnecessary parameters and connections in the model. The storage space and the transmission cost of the model are reduced, the deployment efficiency of the model on the mobile equipment and the embedded system is improved, the computing resource requirement can be remarkably reduced, the reasoning process is accelerated, and then the power consumption and the delay are reduced. Meanwhile, the model pruning maintains the main characteristics and performance of the model, provides convenience for online updating and rapid deployment, and is a technical means for effectively optimizing the deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a deep learning model pruning technology based on FPGM, which is a method for optimizing deep learning models by removing redundant and unnecessary parameters and connections in the model, thereby streamlining and compressing the model. This not only helps to reduce the storage space and transmission cost of the model, improve the deployment efficiency of the model on mobile devices and embedded systems, but also significantly reduces the computing resource requirements, accelerates the inference process, and thus reduces power consumption and latency. At the same time, model pruning maintains the model while retaining its main features and performance, providing convenience for online updates and rapid deployment, and is an effective technical means to optimize deep learning models. Background Art

[0002] With the rapid development of deep learning, deep learning models have achieved remarkable results in various fields. However, this has been accompanied by a dramatic increase in model complexity and size, posing significant challenges in terms of storage, computing resources, and power consumption for deployment on mobile devices and embedded systems. Traditional model compression methods face difficulties in balancing model performance with practical application requirements. Therefore, there is an urgent need for new technologies that can effectively compress model size, improve inference efficiency, and reduce resource consumption.

[0003] However, core challenges facing current technologies include unreasonably high storage and transmission costs due to model complexity, performance bottlenecks on mobile devices, and inference latency and excessive power consumption caused by high computing resource requirements. These issues severely limit the widespread deployment of deep learning models in practical applications, necessitating an innovative and efficient technology to address this complex and resource-efficient trade-off.

[0004] This application combines model compression technology with FPGM-based deep learning model pruning. This technology achieves a highly streamlined and efficient model by removing redundant parameters and connections from the model, significantly reducing storage and transmission costs and improving deployment efficiency on mobile devices and embedded systems. Its unique ability to simultaneously reduce computing resource requirements, accelerate the inference process, and minimize power consumption and latency provides a much-needed solution for the rapid and efficient application of large-scale deep learning models. Summary of the Invention

[0005] The purpose of this application is to propose a deep learning model pruning technology based on FPGM, which achieves compactness and efficiency of the model, reduces storage and transmission costs, and improves deployment efficiency on mobile devices and embedded systems.

[0006] In order to achieve the above objectives, this application combines model pruning technology and proposes a deep learning model pruning technology based on FPGM. Figure 1 As shown, we first use the VGG16 architecture (such as Figure 2 The deep learning model is trained on the CIFAR-10 dataset (as shown in the figure) and CIFAR-10 datasets. This includes preprocessing the input images, defining the model structure, selecting the loss function and optimizer, and optimizing the model parameters through multiple iterations of forward propagation and backpropagation to make it suitable for the classification task of the CIFAR-10 dataset. The trained VGG16 model is then evaluated for importance to identify the contribution of each filter or layer, and a soft pruning strategy is adopted to reduce the influence of some weights based on the evaluation results instead of completely pruning them to retain a certain degree of model flexibility. The generated pruning mask identifies which weights need to be retained and which need to be pruned. Subsequently, the mask is applied to reset the weights corresponding to zero to achieve the sparsity of the model. Through this process, the pruned network structure is derived, including the removal of unnecessary filters or layers. Finally, the pruned model is fine-tuned and the parameters are adjusted through additional training iterations to ensure that the performance level of the model is maintained while pruning.

[0007] The method of this application mainly consists of 4 modules, such as Figure 3 As shown, it is divided into importance assessment module, soft pruning module, pruning mask module and weight zero removal module. The following describes the four modules in the method of this application respectively.

[0008] (1) Importance Assessment Module

[0009] The importance evaluation module uses the principle of geometric median to calculate the importance score of each filter by calculating the filter output feature map of the trained model. Specifically, the geometric median of the output feature map of each filter on the validation set is calculated, such as Figure 4 As shown in , the resulting comprehensive score is used as a measure of the importance of the filter. This process uses a subset of the validation set to robustly identify the filters in the model that contribute more to performance. The purpose of doing so is to more comprehensively evaluate the contribution of the filter to the entire feature map by considering the output of each position, thereby determining the importance of the filter.

[0010] (2) Soft pruning module

[0011] In order to prune more flexibly in engineering implementation, the soft pruning method of this application is usually simulated by resetting the weights to zero, such as Figure 5As shown in Figure 2. For example, in a convolutional layer, there are five filters numbered 0-4. If filters 1 and 3 need to be pruned, soft pruning will set all weights for these two filters to 0. This is achieved by introducing a pruning coefficient, where a pruning coefficient of 0 means that the corresponding weights will be set to zero, while a non-zero pruning coefficient means that the corresponding weights will be retained. This soft pruning method effectively simulates the filter pruning process by resetting the weights to zero, providing engineering flexibility for model compactness.

[0012] (3) Pruning mask module

[0013] The pruning mask module plays a key role in the method of this application. It records the information of weight retention or pruning. Taking the convolution layer as an example, if the dimension of the convolution layer is nchw, where n represents the number of convolution kernels, c represents the number of channels, h and w represent the height and width, then the total number of parameters of the convolution layer is total = nchw. The corresponding pruning mask is composed of a total list of 1 or 0, where 1 represents that the corresponding parameter is retained and 0 represents that the parameter is pruned. According to the FPGM pruning principle, the smallest unit of pruning is the filter, that is, the dimension chw. Therefore, in the pruning mask, the total number of 0s or the total number of 1s must be an integer multiple of ch*w.

[0014] The pruning mask module generates a pruning mask based on the importance assessment and pruning coefficients. This mask indicates which parameters should be retained and which should be pruned. By applying the pruning mask to each parameter, the model parameters can be precisely pruned, ensuring that the integrity of the filter is maintained, thereby effectively reducing the size of the model.

[0015] (4) Weight zero removal module

[0016] The weight removal module in this application's method plays a key role after soft pruning, as soft pruning does not directly reduce model size or speed up inference. To obtain the final streamlined weight file, the soft-pruned model file requires further processing. Specifically, this module's task is to remove zero-weight channels in the model to further optimize the model's scale.

[0017] In practice, zero removal is achieved by reconstructing the pruned network. First, based on the pruning mask and pruning coefficients generated by soft pruning, a new small network model is reconstructed, containing only the non-zero weights retained after pruning. Next, the non-zero weights from the original model file that were not pruned are copied to the new small network model. This process effectively eliminates zero-weight channels in the soft-pruned model, resulting in the final streamlined weight file.

[0018] By implementing the weight zero removal module, the resulting model file is smaller in size, helping to improve the model's efficiency during inference while maintaining model performance. This step is a key step in the soft pruning process and provides important support for optimizing the overall effectiveness of deep learning models. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The overall block diagram of the deep learning model pruning technology based on FPGM

[0020] Figure 2 This is the VGG16 algorithm architecture diagram

[0021] Figure 3 This is a module diagram of the application method

[0022] Figure 4 Schematic diagram for filter pruning

[0023] Figure 5 Pruning for specification-based standards Specific implementation methods

[0024] In order to better describe the compression technology based on the abnormal event early warning detection model, the specific implementation of this application is given below.

[0025] This project uses a Windows system as the development environment. First, in the first step, the VGG16 model is trained using the CIFAR-10 dataset. The model input is CIFAR-10 image data, and the output is the image class label. During training, the model parameters are gradually adjusted to improve classification performance on CIFAR-10 using the optimization algorithm stochastic gradient descent and the cross-entropy loss function. Next, the second step, importance evaluation, is performed. On the trained VGG16 model, the output feature map of each filter on the validation set is calculated through forward propagation. These outputs are used to calculate gradients or other relevant metrics to obtain an importance score for each filter. This score takes the model weights as input and outputs the importance score for each filter. Third, based on the results of the importance evaluation, the VGG16 model is fine-tuned using a soft pruning strategy. This step takes the model weights and importance scores as input, and outputs the soft-pruned VGG16 model. The fourth step involves generating a pruning mask. The input of this mask is the weights of the fine-tuned VGG16 model and the set pruning threshold. The output is a list of 0s and 1s, where 1 indicates that the corresponding parameters are retained and 0 indicates that they are pruned. The fifth step is to perform a weight zero removal operation. The input of this step is the pruning mask and the weights of the original VGG16 model. The output is a new small network model that only contains non-zero weights retained after pruning. The sixth step is to derive the pruned self-network structure. Its input is the new small network model, and the output is the pruned network structure, in which the pruned filters are deleted. The seventh step is to fine-tune and optimize the model. Its input is the pruned network structure, and the output is the final model after fine-tuning. The last step is to evaluate the performance. The fine-tuned model is evaluated using the CIFAR-10 test set, including indicators such as classification accuracy, to verify the impact of pruning on model performance.

Claims

1. FPGM-based deep learning model pruning technology, characterized by: Through FPGM-based deep learning model pruning technology, we achieve compact and efficient models, reduce storage and transmission costs, and improve deployment efficiency on mobile devices and embedded systems. It has the characteristics of significantly reducing computing resource requirements, accelerating the inference process, and reducing power consumption and latency. The specific steps are as follows: Step 1) Model training; Step 2) Importance assessment; Step 3) soft pruning; Step 4) pruning mask; Step 5) remove the weights to zero; Step 6) Export the pruned sub-network structure; Step 7) Fine-tune and optimize.

2. The deep learning model pruning technology based on FPGM according to claim 1 is characterized in that Step 1) Through the traditional supervised learning training process, the training set is used for forward propagation and loss calculation, and then the model parameters are adjusted through backpropagation and optimization algorithms to fit the given training data.

3. The deep learning model pruning technology based on FPGM according to claim 1 is characterized in that Step 2) Use the model’s gradient or other relevant metrics to evaluate the weights or filters in the model to quantify their impact on the overall performance of the model and identify their importance.

4. The deep learning model pruning technology based on FPGM according to claim 1 is characterized in that Step 3) Fine-tune the model parameters by introducing sparsity regularization terms or thresholds to reduce the weight influence of some parameters instead of directly pruning them.

5. The deep learning model pruning technology based on FPGM according to claim 1 is characterized in that Step 4) Generate a binary mask through a threshold or other criteria to indicate which parameters need to be retained (value 1) and which need to be pruned (value 0).

6. The deep learning model pruning technology based on FPGM according to claim 1 is characterized in that Step 5) According to the generated pruning mask, the parameters corresponding to zero are set to zero, thereby pruning the weights and generating sparsity.

7. The deep learning model pruning technology based on FPGM according to claim 1 is characterized in that Step 6) Based on the obtained sparse weights, derive the pruned network structure, including deleting unnecessary neurons or filters.

8. The deep learning model pruning technology based on FPGM according to claim 1 is characterized in that Step 7) Fine-tune and optimize the pruned model, and then conduct a comprehensive evaluation, including performance indicators such as accuracy, precision, and recall, to ensure that the model can still meet application requirements while being compressed.