Pruning method for visual neural network model based on pre-training model

By introducing the pruning matrix w' into the pre-trained model, combining the classification frequency of user data during the updating and screening process, personalized pruning of the pre-trained model is achieved, solving the problem of high computing resource consumption in the prior art, and maintaining the accuracy of the model.

CN119990230APending Publication Date: 2025-05-13ANHUI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411880174.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to personalize the pre-trained model without losing accuracy, reduce computing resource consumption, and adapt to the needs of resource-constrained devices.

Method used

By loading the pretrained model and initializing a pruning matrix w' of the same size as the model weight matrix, the w' matrix is ​​updated during the training process, the classification frequency is counted according to the user data, the model is filtered and pruned, and finally fine-tuning is performed to ensure accuracy.

Benefits of technology

It realizes that the pre-trained model is pruned without losing accuracy, reduces computing resource consumption, and adapts to the needs of resource-constrained devices to achieve the effect of model compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990230A_ABST
    Figure CN119990230A_ABST
Patent Text Reader

Abstract

The invention provides an optic neural network model pruning method based on a pre-training model, and relates to the technical field of optic neural network ViT model pruning, and the method comprises the steps: firstly loading a pre-training model, and initializing an auxiliary matrix w'with the same size as a model weight matrix W for recording the importance of each parameter in a training process; in the training process, a w'matrix is updated by using a data set and a soft target label, and the contribution of each parameter to a model output classification task is reflected. By counting the classification frequency in the user data, the most common classification type of the user is found out, and unimportant parameters are screened out according to the information for pruning. And finally, performing fine adjustment on the model by using the residual w'matrix to ensure that the precision of the pruned model is not lower than a preset requirement. According to the method, the pruning strategy can be adaptively adjusted, the method is suitable for mobile equipment or equipment with limited computing resources, and the computing efficiency and the personalized adaptation capability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual neural network ViT model pruning, and in particular to a visual neural network model pruning method based on a pre-training model. Background Art

[0002] In recent years, the Transformer model has become a hot topic of research and application due to its outstanding performance in fields such as natural language processing and computer vision. In particular, the Vision Transformer (ViT) model, as a Transformer-based visual neural network model, has demonstrated strong performance in tasks such as image classification. However, the ViT model usually requires a lot of computing resources and storage space. Although its huge number of parameters brings high accuracy, it also makes its application on devices with limited computing and storage (such as mobile devices, embedded devices, etc.) challenging.

[0003] In order to reduce computing resource consumption and accelerate the reasoning process, model pruning technology has become an effective means to solve this problem. Existing pruning technologies reduce model size and accelerate reasoning by removing redundant network parameters. Common pruning methods include weight-based pruning, structural pruning, quantization, and knowledge distillation. Although these methods can effectively reduce the computational workload and storage requirements of the model, there are still some problems: First, existing pruning methods often do not fully consider personalized needs and cannot optimize the model according to the data or application scenarios of different users. Second, most pruning technologies only focus on general data sets and ignore the accuracy optimization under user-specific needs. Finally, existing methods often find it difficult to balance the relationship between computational efficiency and model accuracy after pruning.

[0004] Therefore, in view of the shortcomings of existing technologies, how to personalize the pruning of pre-trained models without losing accuracy, reduce the consumption of computing resources, and adapt to the needs of resource-constrained devices has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] The purpose of the present invention is to solve the deficiencies in the prior art and provide a method for pruning a visual neural network model based on a pre-trained model. The present invention aims to prune the pre-trained model to make the pre-trained parameter model more targeted at personalized data, while removing unimportant parameters to maintain the accuracy of the personalized model. Ultimately, personalized model compression is achieved.

[0006] The present invention provides a visual neural network model pruning method based on a pre-trained model, comprising:

[0007] S1: Load the pre-trained model and initialize a w' matrix of the same size as the model weight matrix W, with the initial value set to zero;

[0008] S2: Use the data_train dataset to train the model. During the training process, update the w' matrix according to the softtarget label of the model;

[0009] S3: Collect user data and count the frequency of occurrence of each category in the user data to find out the most commonly used category type by users;

[0010] S4: According to the most commonly used classification types by users, the corresponding parameters in the w' matrix are screened out, and the uncommon classification parameters are deleted using the pruning model;

[0011] S5: Use the remaining w' matrix to fine-tune the model, and the accuracy of the pruned model meets the requirements.

[0012] Optionally, the data_train dataset in S2 is a training set for training a pre-trained model, which contains a large number of images or data samples for training a visual neural network. The soft target label is the basis for calculating the classification importance and model performance during the training process. By analyzing the soft target label, the model can identify which parameters are the most important in the classification process.

[0013] Optionally, the step of updating the w' matrix in S2 includes: for each training sample, calculating the classification importance of the model output, and accumulating it to the corresponding position of the w' matrix.

[0014] Optionally, the collection of user data described in S3 includes: collecting actual application data related to specific users through various means, including but not limited to user interaction records, device usage logs, real-time classification task feedback, user preference settings, historical classification task results, etc. The collection of user data can be achieved by obtaining data generated by the user in actual use from the user terminal device, or by collecting data through user online behavior analysis, user feedback system, etc.

[0015] Optionally, the counting of occurrence frequencies of various categories in the user data in S3 includes: performing frequency statistics on each category in the collected data related to the specific user, determining the category types most frequently used by the user by analyzing the frequency information, and retaining these important category types during the pruning process.

[0016] Optionally, the step of screening and pruning the model in S4 includes: screening out the most important parameters in the matrix w' according to the most commonly used classification types in the user data, and pruning classification parameters that are irrelevant to user needs or are not commonly used.

[0017] Optionally, the remaining w' matrix in S5 includes: a pruned parameter set that has a high contribution to the user's personalized needs, which is used to optimize the model performance during the fine-tuning process.

[0018] Optionally, the method is applicable to a Vision Transformer model.

[0019] Beneficial effects: The present invention can adaptively obtain a personalized compression model of the current data set according to the data set; compared with the prior art, the present invention has the following advantages:

[0020] 1. The present invention proposes a personalized model compression method, which achieves the effect of model pruning by introducing the w' matrix into the model and recording important parameters.

[0021] 2. The present invention achieves a balance between pruning and accuracy by continuously optimizing the w' matrix and fine-tuning it. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a model framework flow chart of the present invention. DETAILED DESCRIPTION

[0023] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.

[0024] This implementation provides a visual neural network model pruning method based on a pre-trained model, which is particularly suitable for the ViT model. It aims to optimize the model structure and reduce the amount of calculation through the processing and fine-tuning of personalized data sets, so as to adapt to the deployment requirements of resource-constrained devices. Figure 1 As shown, the specific implementation steps are as follows:

[0025] Step 1: Load the pre-trained model and initialize the w' matrix

[0026] First, load a pre-trained visual neural network model. While loading the model, initialize a pruning matrix w' with the same size as the model weight matrix W, and set the initial value to zero. This matrix is ​​used to record the contribution of each parameter to the classification task during the training process and provide a basis for subsequent pruning operations.

[0027] Initialize the model and w' matrix code:

[0028] model=load_pretrained_model()

[0029] w_prime=zeros_like(model.weights).

[0030] Step 2: Use the training data set to train the model and update the w' matrix

[0031] The model is trained using the data_train dataset. During each training process, the output of the model is compared with the soft target label to calculate the importance of each parameter for the final classification result. The matrix w' is updated based on the calculated classification importance. The purpose of this step is to identify which parameters are critical to the model output based on the performance of the model during training, thereby providing information for subsequent pruning operations.

[0032] Update w' matrix implementation code:

[0033] for epoch in range(num_epochs):

[0034] for data,soft_target in data_train:

[0035] output = model.forward(data)

[0036] importance=calculate_importance(output,soft_target)

[0037] w_prime+=update_w_prime(importance).

[0038] Step 3: Collect user data and count classification frequency

[0039] Next, we collect data sets related to users and count the frequency of each category appearing in the user data. By counting these frequencies, we can determine the classification tasks that users are most concerned about. The purpose of this process is to ensure that the classification functions that users need most are retained during the pruning process, so as to avoid affecting the personalized needs of the model due to pruning.

[0040] The code for implementing statistical user data includes:

[0041] user_data=collect_user_data()

[0042] classification_frequency=

[0043] calculate_classification_frequency(user_data).

[0044] Step 4: Filter and prune the model

[0045] According to the most common classification tasks in the user data, the most important parameters in the pruning matrix w' are selected. Then, pruning is performed based on these important parameters, and infrequently used classification parameters are deleted. This step ensures that the classification functions most relevant to the user are retained after pruning, while reducing unnecessary computing resource consumption.

[0046] The pruning model code includes:

[0047] common_classes=get_common_classes(classification_frequency)

[0048] w_prime_pruned=w_prime[common_classes].

[0049] Step 5: Fine-tune the model

[0050] The pruned model needs to be fine-tuned to ensure that the accuracy is not less than the preset threshold. During the fine-tuning process, the model is trained using the remaining matrix w'. During fine-tuning, the performance of the model is continuously monitored to ensure that the model meets the accuracy requirements while effectively reducing the computational overhead. During the fine-tuning process, the loss and gradient updates of each training are adjusted based on the output of the model and the user labels.

[0051] The fine-tuning model code includes:

[0052] for epoch in range(num_epochs_finetune):

[0053] for data,label in user_data:

[0054] output = model.forward(data)

[0055] loss=calculate_loss(output,label)

[0056] w_prime_grad=calculate_gradient(loss,w_prime_pruned)

[0057] w_prime_pruned-=learning_rate*w_prime_grad

[0058] monitor_model_performance(model,w_prime_pruned).

[0059] After fine-tuning, the performance of the model is verified to ensure that the accuracy meets the requirements. In this process, the model has been personalized and optimized to significantly reduce the computational overhead without sacrificing accuracy.

[0060] The above-described embodiments merely express the implementation methods of the present invention, but they cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A visual neural network model pruning method based on a pre-trained model, characterized in that: The following steps are involved: S1: Load the pre-trained model and initialize a w' matrix of the same size as the model weight matrix W, with the initial value set to zero; S2: Use the data_train dataset to train the model. During the training process, update the w' matrix according to the soft target label of the model; S3: Collect user data and count the frequency of occurrence of each category in the user data to find out the most commonly used category type by users; S4: According to the most commonly used classification types by users, the corresponding parameters in the w' matrix are screened out, and the uncommon classification parameters are deleted using the pruning model; S5: Use the remaining w' matrix to fine-tune the model, and the accuracy of the pruned model meets the requirements.

2. The method for pruning a visual neural network model based on a pre-trained model according to claim 1, characterized in that: The data_train dataset in S2 is a training set for training the pre-trained model. It contains a large number of images or data samples for training the visual neural network. The soft target label is the basis for calculating the classification importance and model performance during the training process. By analyzing the soft target label, the model can identify which parameters are the most important in the classification process.

3. The method for pruning a visual neural network model based on a pre-trained model according to claim 1, characterized in that: The step of updating the w' matrix in S2 includes: for each training sample, calculating the classification importance of the model output, and accumulating it to the corresponding position of the w' matrix.

4. The method for pruning a visual neural network model based on a pre-trained model according to claim 1, characterized in that: The collection of user data described in S3 includes: collecting actual application data related to specific users through various means, including but not limited to user interaction records, device usage logs, real-time classification task feedback, user preference settings, historical classification task results, etc. The collection of user data can be achieved by obtaining data generated by users in actual use from user terminal devices, or by collecting data through user online behavior analysis, user feedback systems, etc.

5. The method for pruning a visual neural network model based on a pre-trained model according to claim 1, characterized in that: The frequency of occurrence of each category in the user data described in S3 includes: performing frequency statistics on each category in the collected data related to the specific user, determining the category types most frequently used by the user by analyzing the frequency information, and retaining these important category types during the pruning process.

6. The method for pruning a visual neural network model based on a pre-trained model according to claim 1, characterized in that: The step of screening and pruning the model in S4 includes: screening out the most important parameters in the matrix w' according to the most commonly used classification types in the user data, and pruning classification parameters that are irrelevant to user needs or are not commonly used.

7. The method for pruning a visual neural network model based on a pre-trained model according to claim 1, characterized in that: The remaining w' matrix in S5 includes: a pruned parameter set that has a high contribution to the user's personalized needs, which is used to optimize the model performance during the fine-tuning process.

8. The method for pruning a visual neural network model based on a pre-trained model according to claim 1, characterized in that: The fine-tuning step described in S5 includes fine-tuning the model: during the fine-tuning process, the model is trained using the pruned w' matrix.

9. The method for pruning a visual neural network model based on a pre-trained model according to claim 1, characterized in that: The method is applicable to the Vision Transformer model.