Neural network selective forgetting learning method based on neuron importance
By evaluating the memory value of the convolutional neural network model for the forgotten data set, quantifying the importance of each neuron and generating parameter masks, the problem of insufficient research on neurons in the existing technology is solved, efficient selective forgetting is achieved, and the performance of the model on the original task is restored.
Patent Information
- Application Number
- CN202510496214.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing selective forgetting learning methods are insufficiently studied in neuronal granularity, mainly focusing on the significance of individual parameters, neglecting the correlation and synergy between parameters, and lacking in-depth analysis of the relationship between neuronal hierarchical characteristic response and forgotten data.
By evaluating the memory value of the convolutional neural network model for the forgotten data set, the importance of each neuron is quantified and the parameter mask is generated, and only some of the parameters that are significant to the forgotten data set are updated to achieve selective forgetting.
Effectively identify and update neuronal parameters that have significant contributions to the forgetting process, significantly reduce the scale of parameter updates, improve the forgetting efficiency, and restore the model's classification performance on the original task while ensuring the forgetting effect.
Smart Images

Figure CN120012871A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning image classification and forgetting learning, and in particular to a neural network selective forgetting learning method based on neuron importance. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, the amount of data used in model training has increased exponentially, and data privacy and security issues have become increasingly prominent. In many practical application scenarios, according to the provisions of privacy protection regulations such as the General Data Protection Regulation (GDPR), users have the right to request the deletion of their data and revoke their authorization to use the data. However, simply deleting the target sample from the dataset is not enough to achieve true data forgetting, because the model parameters may have memorized the sensitive information of the sample (Feldman, V. (2019). Does learning require memorization? ashort tale about a long tail. Proceedings of the 52nd Annual ACM SIGACTSymposium on Theory of Computing.). This memory effect is particularly significant in deep neural networks. Even if the training data is deleted, the attacker may still be able to recover the original data through methods such as model reverse engineering (Yang, Z., Chang, E., & Liang, Z. (2019). Adversarial Neural Network Inversion via AuxiliaryKnowledge Alignment. ArXiv, abs / 1902.08552.). How to efficiently and effectively "forget" specific data from the model without affecting the performance of the model has become a problem that needs to be solved urgently, and related research on machine unlearning has emerged.
[0003] Existing forgetting learning methods can be divided into two categories: precise forgetting and approximate forgetting. Precise forgetting is retraining, which removes the data sample set requested by the user from the original training dataset, and then trains a reinitialized model from scratch with the remaining dataset. However, precise learning methods based on retraining require a lot of computing resources, which is especially challenging for large-scale models (such as diffusion-based generative models). Approximate forgetting aims to remove the influence of the forgotten dataset in the model by updating the model parameters, so that the difference between the updated forgotten model and the model that was not trained is statistically indistinguishable. Common approximate forgetting methods include: fine-tuning the model using the remaining dataset (Warnecke, A., Pirch, L., Wressnegger, C., & Rieck, K. (2021). Machine Unlearning of Features and Labels. ArXiv, abs / 2108.11577.) Compared with the retraining method, fine-tuning only requires a small number of training rounds to achieve the desired effect, significantly reducing the computational overhead; randomly perturbing the sample labels in the forgetting dataset and then fine-tuning the model using the dataset (Golatkar, A., Achille, A., & Soatto, S. (2019). Eternal Sunshine of the Spotless Net: Selective Forgetting in Deep Networks. 2020 IEEE / CVF Conferenceon Computer Vision and Pattern Recognition (CVPR), 9301-9309.); based on the influence function method (Koh, P., & Liang, P. (2017). Understanding Black-box Predictions via Influence Functions. International Conference on Machine Learning.), by calculating the contribution of data points to model parameters and adjusting the parameters inversely to remove the influence of specific data sample points in the model parameters (Guo, C., Goldstein, T., Hannun, AY, & Maaten, LV (2019). Certified Data Removalfrom Machine Learning Models. International Conference on Machine Learning.); adding normal distribution noise to the model parameters as a whole (Golatkar, A., Achille, A.,&Soatto, S. (2020).Forgetting Outside the Box: Scrubbing Deep Networks of Information Accessiblefrom Input-Output Observations. European Conference on Computer Vision.), weakening the model's dependence on specific data by randomly perturbing parameters; by adjusting the decision boundary of the classification model, the model's prediction results on the forgotten dataset tend to be randomized (Chen, M., Gao, W., Liu, G., Peng, K.,&Wang, C.(2023). Boundary Unlearning: Rapid Forgetting of Deep Networks via Shiftingthe Decision Boundary. 2023 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7766-7775.); based on the gradient ascent of the forgotten dataset, maximize the model's loss in , thereby weakening the model's memory of the forgotten dataset (Thudi, A., Deza, G., Chandrasekaran,V.,&Papernot, N. (2021). Unrolling SGD: Understanding Factors Influencing Machine Unlearning. 2022 IEEE 7th European Symposium on Security and Privacy (EuroS&P), 303-319.). The above forgetting learning method needs to update all parameters of the model, and it takes a lot of computing resources and storage resources to calculate the Hessian matrix and its inverse and the Fisher information matrix of all parameters. .
[0004] The forgetting method of "pruning first and forgetting later" (Jia, J., Liu, J., Ram, P., Yao, Y., Liu, G., Liu, Y., Sharma, P., & Liu, S. (2023). Model Sparsity Can Simplify Machine Unlearning. Neural Information Processing Systems.) points out that model sparsity helps forgetting, and incorporates model sparsity as a regular term in the loss function into the forgetting process, but pruning will destroy the model structure. "Pruning first and forgetting later" provides an idea for selective forgetting, that is, selectively updating some parameters that are significant to the samples in the forgetting data set. The key lies in selecting significant parameters. Common parameter significance evaluation methods for forgotten datasets include: calculating the Fisher information matrix and on the forgotten dataset and the retained dataset respectively, selecting parameters that are sensitive to changes and insensitive to changes (Liu, Y., Sun, C., Wu, Y.,&Zhou, A. (2023). Unlearning with Fisher Masking.ArXiv, abs / 2310.05331.); generating a parameter significance map based on the model parameter gradient for a specific loss function on the forgotten dataset (Fan, C., Liu, J., Zhang, Y., Wei, D., Wong, E.,&Liu, S. (2023).SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency inBoth Image Classification and Generation. ArXiv, abs / 2310.12508.), and updating the significant parameters therein; based on the connection sensitivity analysis method, the change value of the loss function on the forgotten dataset when a single parameter is set to 0 is used as the importance evaluation criterion of the corresponding parameter (Wu, J.,&Harandi,M. (2024). Scissorhands: Scrub Data Influence via Connection Sensitivity in Networks. European Conference on Computer Vision.). The above method of selecting significance parameters only considers a single parameter and does not consider the correlation between parameters.Therefore, the existing selective forgetting learning methods based on parameter significance have the following limitations: they mainly focus on the significant impact of a single parameter on the forgotten data, ignoring the correlation and synergy between parameters; lack of in-depth analysis of the relationship between the characteristic response of the neuron level and the forgotten data; and lack of interpretability. Summary of the invention
[0005] In view of the deficiencies in the prior art, the purpose of the present invention is to fill the research deficiencies in selective forgetting learning at the neuron granularity, and to provide a neural network selective forgetting learning method based on neuron importance, which screens important neurons and updates significant partial parameters to meet the requirements of forgetting learning.
[0006] The objective of the present invention is achieved through the following technical solutions:
[0007] The first aspect of the present invention is a neural network selective forgetting learning method based on neuron importance, comprising the following steps:
[0008] (1) Determine the forgotten data set: In the random forgetting scenario, randomly select part of the training data as the forgotten data set; in the category forgetting scenario, randomly select a category of data as the forgotten data set; the remaining data is used as the retained data set;
[0009] (2) Evaluating the importance of neurons: assigning the memory value of the convolutional neural network (CNN) model to the samples to the neurons to evaluate the importance of the neurons to the data set; the samples are forgotten samples, and the data set is a forgotten data set, where the forgotten data set consists of a certain number of forgotten samples;
[0010] (3) Generate parameter mask: set the mask of neurons with memory values greater than a certain threshold to 1, otherwise set the mask to 0;
[0011] (4) Random label forgetting process: Randomly set the labels of the samples in the forgetting dataset and use the stochastic gradient descent method to update the parameters with the mask set to 1;
[0012] (5) Fine-tune the CNN model: Use the reserved dataset to fine-tune the entire model to obtain the final forgotten model;
[0013] (6) Evaluate the forgetting effect: Test the accuracy of the model after forgetting on the forgotten dataset and the accuracy of the member reasoning attack to evaluate the effect of the forgetting method; test the accuracy of the model after forgetting on the retained dataset and the test dataset to evaluate the performance and generalization of the model after forgetting on the original task.
[0014] Furthermore, the step (1) includes the following two forgetting scenarios:
[0015] (1.1) Random forgetting scenario: complete training dataset ,include data sample points, among which For images, For labels; in the random forgetting scenario, according to the user's actual forgetting request, some data sample points are randomly selected from the complete data set as the forgetting data set , the rest is the reserved data set ;
[0016] (1.2) Category forgetting scenario: The complete training dataset includes category images, namely In the case of category forgetting, all categories in the training data set are set to the user-specified forgetting categories. The data samples are divided into the forget data set , the remaining categories of data are divided into the reserved data set .
[0017] Furthermore, the specific calculation steps for evaluating the importance of neurons are:
[0018] (2.1) Calculate the memory value of the entire CNN model for all samples in the forgotten data set, and take the average as the memory value of the entire model for the forgotten data set. The specific calculation formula is:
[0019]
[0020] in: Represents an indicator function, which takes a value of 1 when the condition in the brackets is true, otherwise it takes a value of 0. Indicates the algorithm used to train the model, Representation dataset The trained model, Representation dataset The trained model;
[0021] (2.2) For each neuron in the model Ablate one by one and calculate the ablation of a certain neuron The memory value of the incomplete model structure for the forgotten data set is calculated as follows:
[0022]
[0023] in: Representation Model Ablate a neuron After the incomplete model, Representation Model Ablate a neuron The incomplete model after
[0024] (2.3) Subtract the result of step (2.1) from the result of step (2.2) to get the neuron Forgotten dataset The memory value of the neuron, i.e. the neuron importance score for:
[0025]
[0026] Among them, due to the calculation formulas in steps (2.1) and (2.2), the model and incomplete models Forgotten dataset The memory values of are approximately equal, and the result of subtracting the two is close to 0, so the corresponding two terms can be omitted.
[0027] Furthermore, the specific process of generating the parameter mask in step (3) is as follows: Neurons By importance score Sort in descending order; according to actual needs, The neuron importance score of the position is used as the threshold , the importance score is less than The neuron mask value of is set to 0, and the mask values of the remaining neurons are set to 1, that is, the neuron mask for:
[0028]
[0029] in is a binary vector with dimension The number of neurons Consistent.
[0030] Furthermore, the step (4) includes the following sub-steps:
[0031] (4.1) Randomize labels: Assume that the dataset is forgotten , for which Generate random labels for each sample , construct a new data set :
[0032]
[0033] in is the number of data set categories;
[0034] (4.2) Update some parameters: Dataset after randomizing labels On, combined with the mask vector , use the stochastic gradient descent algorithm to update the model parameters Specifically, only the parameters corresponding to the neurons with mask values 1 are updated, and the updated model parameters It is expressed as:
[0035]
[0036] in represents element-wise multiplication of vectors, Represents the parameter update amount; the optimization goal during the parameter update process is to minimize the loss function :
[0037]
[0038] Among them, in order to enhance the forgetting effect, the loss function is introduced for the original forgotten data The regular term of .
[0039] Furthermore, the specific process of step (5) is as follows: using the reserved data set The updated model in step (4) Perform global update to obtain the final forgetting model parameters ; By optimizing the loss function , using the stochastic gradient descent algorithm to update the model parameters, thereby ensuring that the model retains the data set The performance on the dataset is not affected, while the forgotten dataset is implemented effective forgetting.
[0040] Furthermore, the step (6) includes the following sub-steps:
[0041] (6.1) Accuracy test: Update the model parameters in step (5) , respectively, in the forgotten data set , retain the dataset and test dataset The accuracy test was performed on the model to evaluate the forgetting effect, original task performance, and generalization of the model. In the dataset The accuracy of the above is:
[0042]
[0043] According to actual needs, use , , replace That's it; among them The model is The prediction results;
[0044] (6.2) Membership Inference Attack Test: Training a binary membership inference attack model , used to distinguish data samples Is it the target model? From the training data set; from the test data set Select the quantity The samples are used as non-member test sets , computational attack model exist and Attack success rate ASR:
[0045]
[0046] The first sum represents the prediction is the number of members, and the second sum represents the prediction For the number of non-members, ASR has a good forgetting effect when it is close to 50%.
[0047] The second aspect of the present invention is a neural network selective forgetting learning device based on neuron importance, comprising the following modules:
[0048] Determine the forgotten data set module: In the random forgetting scenario, randomly select part of the training data as the forgotten data set; in the category forgetting scenario, randomly select a category of data as the forgotten data set; the remaining data is used as the reserved data set;
[0049] Neuron importance evaluation module: assign the memory value of the convolutional neural network model to the forgotten samples to the neurons to evaluate the importance of the neurons to the forgotten data set; parameter mask generation module: set the neuron mask with a memory value greater than a specific threshold to 1, otherwise the mask is set to 0;
[0050] Random label forgetting module: randomly set the labels of samples in the forgetting dataset and use the stochastic gradient descent method to update the parameters with the mask set to 1;
[0051] Fine-tune the CNN model module: Use the reserved dataset to fine-tune the entire model to obtain the final forgotten model;
[0052] Evaluation module for forgetting effect: Test the accuracy of the model after forgetting on the forgotten dataset and the accuracy of member reasoning attack to evaluate the effect of the forgetting method; test the accuracy of the model after forgetting on the retained dataset and test dataset to evaluate the performance and generalization of the model after forgetting on the original task.
[0053] The beneficial effects of the present invention are as follows: in the image classification task, by comparing the convolutional neural network model before and after the ablation of a single neuron, the forgetting data set The change in classification accuracy can effectively identify neurons that have made significant contributions to the forgetting process, thereby achieving accurate updates of key parameters. This method significantly reduces the scale of parameter updates and improves forgetting efficiency while ensuring the forgetting effect. In addition, by retaining the data set Fine-tuning the model for a small number of epochs can effectively restore the classification performance of the model on the original task, ensuring that the impact of the forgetting process on the overall performance of the model is minimized. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a schematic diagram of the application of a neural network selective forgetting learning method based on neuron importance provided by the present invention;
[0055] Figure 2 It is a flow chart of a neural network selective forgetting learning method based on neuron importance provided by the present invention. DETAILED DESCRIPTION
[0056] The present invention will be further described below in conjunction with the accompanying drawings.
[0057] The core technology of the present invention is to distribute the memory value of the convolutional neural network model for forgotten data samples to neurons, and evaluate the importance of neurons to the forgotten data set by comparing the changes in the model's prediction accuracy for samples in the forgotten data set before and after the ablation of a single neuron, thereby achieving selective forgetting at the neuron granularity.
[0058] like Figure 1 As shown in the figure, the application of the neural network selective forgetting learning method based on neuron importance proposed in the present invention in deep learning image classification tasks, especially for convolutional neural network (CNN). The framework mainly includes the following parts: the data requested to be deleted by the user is divided into the forgotten data set , and from the original training dataset Remove , the remaining data constitute the retained dataset ; Construct a random label forgetting dataset by randomly disturbing the labels of samples in the forgetting dataset Based on the ablation analysis of single neurons, the original model Forgotten dataset The memory values are assigned to neurons and ranked according to their importance scores. The neurons are sorted by memory value, and the masks of neurons with significant memory value are set to 1, and the masks of other neurons are set to 0. Fine-tune the saliency parameter and introduce the forgotten dataset The loss function regularization term on , in order to strengthen forgetting; further on the reserved data set The model is globally fine-tuned on the retained dataset to restore the performance of the model on the original task; finally, , test data set , and forget datasets Accuracy tests on The member reasoning attack test is used to evaluate the forgetting effect of this method and the performance of the model after forgetting.
[0059] More specifically, Figure 2 As shown, the neural network selective forgetting learning method based on neuron importance proposed by the present invention comprises the following steps:
[0060] (1) Determine the forgotten data set: In the random forgetting scenario, randomly select part of the training data as the forgotten data set; in the category forgetting scenario, randomly select a category of data as the forgotten data set; the remaining data is used as the retained data set; specifically, the following sub-steps are included:
[0061] (1.1) Random forgetting scenario: Assuming the complete training dataset include data sample points, among which is the image data, is the corresponding label. In the random forgetting scenario, according to the user’s actual forgetting request, Randomly select some data samples to form a forgotten data set , the remaining data sample points constitute the reserved data set . Forget Dataset and the reserved dataset The division reflects the user's specific needs for data deletion while retaining the data set It will be used for subsequent fine-tuning of the model to restore the performance of the forgotten model on the original task.
[0062] (1.2) Category forgetting scenario: Assume that the complete training dataset includes categories of image data, namely In the category forgetting scenario, the user specifies the category to be forgotten. , all labels in the training data set are The data samples are divided into the forget data set , the remaining categories of data samples are divided into the reserved data set Taking face recognition as an example, the training data set contains photos of multiple users, each of which corresponds to an independent category. When a user requests to delete all photos related to him / her, these photos belong to the same category.
[0063] (2) Evaluate the importance of neurons: Assign the model's memory value of the forgotten samples to neurons to evaluate the importance of neurons to the forgotten data set; specifically, this includes the following sub-steps:
[0064] (2.1) Calculate the memory value of the entire model for all samples in the forgotten data set, and take the mean as the memory value of the entire model for the forgotten data set. The specific calculation formula is:
[0065]
[0066] in Represents an indicator function, which takes a value of 1 when the condition in the brackets is true, otherwise it takes a value of 0. Indicates the algorithm used to train the model, Representation dataset The trained model, Representation dataset The calculation formula refers to the calculation method of the memory value of a data sample point proposed by Feldman et al. (Feldman, V., & Zhang, C. (2020). What neural networks memorize and why: Discovering the long tail via influence estimation. Advances in Neural Information Processing Systems, 33, 2881-2891.), and the memory value of the model for all sample points in the forgotten data set is averaged as the memory value of the model for the entire forgotten data set.
[0067] (2.2) For each neuron in the model Ablate one by one and calculate the ablation of a certain neuron The memory value of the incomplete model structure for the forgotten data set is calculated as follows:
[0068]
[0069] in Representation Model Ablation of a neuron After the incomplete model, Representation Model Ablation of a neuron The incomplete model.
[0070] Specifically, during model inference, for a specific neuron The ablation operation is to suppress the function of the neuron by setting the output feature map generated by the neuron in the forward propagation to zero. This operation does not modify the weight parameters of the neuron. and bias parameters , nor does it involve pruning or adjustment of the network structure.
[0071] (2.3) Subtract the result of step (2.1) from the result of step (2.2) to get the neuron Forgotten dataset The memory value, i.e. the importance score for:
[0072]
[0073] Among them, in the forgetting learning scenario, due to the model Using the holdout dataset Retrained Model Therefore, this method simplifies the result of subtracting the calculation formula of step (2.1) from that of step (2.2). Specifically, the model in step (2.1) Forgotten dataset The prediction accuracy of right The prediction accuracy of is approximately equal, and the result of subtracting the two approaches 0. Based on this approximate relationship, this method omits these two items when calculating the neuron importance score, thereby obtaining a simplified calculation formula.
[0074] (3) Generate parameter mask: Set the mask of neurons with memory values greater than a specific threshold to 1, otherwise set the mask to 0; the specific threshold is a relative threshold. If the top 20% of significant neurons are selected, the neurons are sorted in descending order by memory value, and the memory value at the 20th percentile is taken as the threshold. The specific process of generating parameter mask is as follows: Based on the neurons’ forgetting data set Memory salience score , for the model Neurons Sort in descending order; select the ratio according to the preset neuron (For example ), the first The importance score of neurons at the percentile is used as the threshold ; For each neuron , if its memory salience score , then the corresponding mask value is set to 1, otherwise it is set to 0, that is, the neuron mask for:
[0075]
[0076] is a binary vector with dimension The number of neurons , where each element corresponds to the mask state of a neuron.
[0077] (4) Random label forgetting process: Randomly set the labels of the samples in the forgetting dataset and use the stochastic gradient descent method to update the parameters with the mask set to 1. Specifically, it includes the following sub-steps:
[0078] (4.1) Randomize labels: Assume that the dataset is forgotten , for which Generate random labels for each sample , construct a new data set :
[0079]
[0080] in is the number of categories in the dataset. It should be noted that in the category forgetting scenario, for all samples of the target forgotten category, their labels will be randomly reallocated to any of the remaining categories instead of being uniformly assigned to a specific category.
[0081] (4.2) Update some parameters: Dataset after randomizing labels On, combined with the mask vector , use the Stochastic Gradient Descent (SGD) algorithm to update the model parameters Specifically, only the parameters corresponding to the neurons with mask values 1 are updated, and the updated model parameters It is expressed as:
[0082]
[0083] in represents element-wise multiplication of vectors, Represents the parameter update amount; the optimization goal during the parameter update process is to minimize the loss function :
[0084]
[0085] Among them, based on the dataset after minimizing random perturbation labels The loss function on the original forgotten dataset is the same as maximizing The loss function on the optimization target has a consistent basis. In order to enhance the forgetting effect, this method introduces the original forgotten data into the loss function The regular term of .
[0086] In image classification tasks, especially multi-class classification problems, the cross-entropy loss function is widely used because it can effectively measure the difference between the classification model output and the true label. Specifically, this method uses the cross-entropy loss function. The original CNN model Forgotten dataset The cross entropy loss function :
[0087]
[0088] in is the number of categories, Indicates the data set The one-hot encoding vector of the true label of the sample components, where when the sample belongs to the category The value is 1 when it is, otherwise it is 0; Representation Model For The sample prediction The probability value of each category satisfies and ; Dataset after randomly perturbing labels The calculation form of the cross entropy loss function remains unchanged, only the forgotten dataset needs to be Replace with That's it.
[0089] (5) Fine-tune the model: Use the reserved dataset to fine-tune the entire model to obtain the final forgotten model; the specific process is: use the reserved dataset The updated model in step (4) Perform global update to obtain the final forgetting model parameters ; By optimizing the loss function , using the stochastic gradient descent algorithm to update the model parameters, thereby ensuring that the model retains the data set The performance on the dataset is not affected, while the forgotten dataset is implemented effective forgetting.
[0090] Among them, the cross entropy loss function is still used in the fine-tuning stage using the reserved dataset, and the dataset in the above cross entropy loss function calculation formula is replaced by the reserved dataset That's it.
[0091] (6) Evaluate the forgetting effect: Test the accuracy of the model after forgetting on the forgotten dataset and the accuracy of the member reasoning attack to evaluate the effect of the forgetting method; test the accuracy of the model after forgetting on the retained dataset and the test dataset to evaluate the performance and generalization of the model after forgetting on the original task; specifically, it includes the following sub-steps:
[0092] (6.1) Accuracy test: Update the model parameters in step (5) , respectively, in the forgotten data set , retain the dataset and test dataset The accuracy test was performed on the model to evaluate the forgetting effect, original task performance, and generalization of the model. In the dataset The accuracy of the above is:
[0093]
[0094] According to actual needs, use , , replace That's it; among them The model is The prediction results;
[0095] (6.2) Membership Inference Attack Test: Training a binary membership inference attack model , the model is used to discriminate the given sample Is it the target model? The training data set outputs the prediction results as binary labels , respectively indicating that a given sample belongs to or does not belong to the training data set; from the test data set Select the quantity The samples are used as non-member test sets , computational attack model exist and Attack success rate ASR:
[0096]
[0097] The first sum represents the prediction is the number of members, and the second sum represents the prediction For the number of non-members, ASR has a good forgetting effect when it is close to 50%.
[0098] The present invention also proposes a neural network selective forgetting learning device based on neuron importance, comprising the following modules:
[0099] Determine the forgotten data set module: In the random forgetting scenario, randomly select part of the training data as the forgotten data set; in the category forgetting scenario, randomly select a category of data as the forgotten data set; the remaining data is used as the reserved data set;
[0100] Neuron importance evaluation module: Assign the memory value of the convolutional neural network model to the forgotten samples to the neurons to evaluate the importance of the neurons to the forgotten data set;
[0101] Generate parameter mask module: set the mask of neurons with memory values greater than a certain threshold to 1, otherwise set the mask to 0;
[0102] Random label forgetting module: randomly set the labels of samples in the forgetting dataset and use the stochastic gradient descent method to update the parameters with the mask set to 1;
[0103] Fine-tune the CNN model module: Use the reserved dataset to fine-tune the entire model to obtain the final forgotten model;
[0104] Evaluation module for forgetting effect: Test the accuracy of the model after forgetting on the forgotten dataset and the accuracy of member reasoning attack to evaluate the effect of the forgetting method; test the accuracy of the model after forgetting on the retained dataset and test dataset to evaluate the performance and generalization of the model after forgetting on the original task.
[0105] In summary, the present invention aims at the current lack of research on selective forgetting learning at the neuron granularity, and proposes a neural network selective forgetting learning method based on neuron importance. Aiming at the selective forgetting requirements of convolutional neural network models in image classification tasks, the present invention proposes a significance evaluation method based on the change in prediction accuracy before and after ablation of a single neuron. By comparing the convolutional neural network model on the forgetting dataset before and after ablation of a single neuron, the prediction accuracy of the image classification task is evaluated. The classification accuracy changes are quantitatively evaluated to evaluate the importance of each neuron to the forgotten data, and then the neuron parameters with significant memory values are screened out to achieve selective forgetting at the neuron granularity. This invention innovatively constructs a multi-stage, multi-method fusion forgetting learning framework: the first stage adopts a coordinated optimization strategy of random label perturbation and regularization constraints to eliminate the influence of samples to be forgotten at the neuron granularity; the second stage uses the remaining data set to Global fine-tuning is performed on the original classification task, and the model performance can be restored with only a small number of training rounds. This method can effectively achieve the forgetting of target samples while better maintaining the performance of the model on the original classification task, providing a new solution for the controllable forgetting of deep neural networks.
[0106] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. A neural network selective forgetting learning method based on neuron importance, characterized in that: The steps include: (1) Determine the forgotten data set: In the random forgetting scenario, randomly select part of the training data as the forgotten data set; in the category forgetting scenario, randomly select a category of data as the forgotten data set; the remaining data is used as the retained data set; (2) Evaluate the importance of neurons: Assign the memory value of the convolutional neural network (CNN) model to the forgotten samples to the neurons to evaluate the importance of the neurons to the forgotten dataset; the forgotten dataset consists of a certain number of forgotten samples; (3) Generate parameter mask: set the mask of neurons with memory values greater than a certain threshold to 1, otherwise set the mask to 0; (4) Random label forgetting process: Randomly set the labels of the samples in the forgetting dataset and use the stochastic gradient descent method to update the parameters with the mask set to 1; (5) Fine-tune the CNN model: Use the reserved dataset to fine-tune the entire CNN model to obtain the final forgotten model; (6) Evaluate the forgetting effect: Test the accuracy of the model after forgetting on the forgotten dataset and the accuracy of the member reasoning attack to evaluate the effect of the forgetting method; test the accuracy of the model after forgetting on the retained dataset and the test dataset to evaluate the performance and generalization of the model after forgetting on the original task.
2. The neural network selective forgetting learning method based on neuron importance according to claim 1, characterized in that: The step (1) is specifically as follows: (1.1) Random forgetting scenario: complete training dataset ,include data sample points, among which For images, For labels; in the random forgetting scenario, according to the user's actual forgetting request, some data sample points are randomly selected from the complete data set as the forgetting data set , the rest is the reserved data set ; (1.2) Category forgetting scenario: The complete training dataset includes category images, namely In the case of category forgetting, all categories in the training data set are set to the user-specified forgetting categories. The data samples are divided into the forget data set , the remaining categories of data are divided into the reserved data set .
3. The neural network selective forgetting learning method based on neuron importance according to claim 1, characterized in that: Evaluate the importance of neurons for the forgetting dataset. The specific steps are: (2.1) Calculate the memory value of the entire CNN model for all samples in the forgotten data set, and take the average as the memory value of the entire model for the forgotten data set. The specific calculation formula is: ;in: Represents an indicator function, which takes a value of 1 when the condition in the brackets is true, otherwise it takes a value of 0. Indicates the algorithm used to train the model, Representation dataset The trained model, Representation dataset The trained model; (2.2) For each neuron in the model Ablate one by one and calculate the ablation of a certain neuron The memory value of the incomplete model structure for the forgotten data set is calculated as follows: ;in: Representation Model Ablate a neuron After the incomplete model, Representation Model Ablate a neuron The incomplete model after (2.3) Subtract the result of step (2.1) from the result of step (2.2) to get the neuron Forgotten dataset The memory value of the neuron, i.e. the neuron importance score for: ; Among them, due to the calculation formulas in steps (2.1) and (2.2), the model and incomplete models Forgotten dataset The memory values of are approximately equal, and the result of subtracting the two is close to 0, so the corresponding two terms can be omitted.
4. The neural network selective forgetting learning method based on neuron importance according to claim 1, characterized in that: The specific process of generating the parameter mask in step (3) is as follows: Neurons By importance score Sort in descending order; according to actual needs, The neuron importance score of the position is used as the threshold , the importance score is less than The neuron mask value of is set to 0, and the mask values of the remaining neurons are set to 1, that is, the neuron mask for: ;in is a binary vector with dimension The number of neurons Consistent.
5. The neural network selective forgetting learning method based on neuron importance according to claim 1, characterized in that: The step (4) includes the following sub-steps: (4.1) Randomize labels: Assume that the dataset is forgotten , for which Generate random labels for each sample , construct a new data set : ;in is the number of data set categories; (4.2) Update some parameters: Dataset after randomizing labels On, combined with the mask vector , use the stochastic gradient descent algorithm to update the model parameters Specifically, only the parameters corresponding to the neurons with mask values 1 are updated, and the updated model parameters It is expressed as: ;in represents element-wise multiplication of vectors, Represents the parameter update amount; the optimization goal during the parameter update process is to minimize the loss function : ; Among them, in order to enhance the forgetting effect, the loss function is introduced for the original forgotten data The regular term of .
6. The neural network selective forgetting learning method based on neuron importance according to claim 1, characterized in that: The step (5) is specifically as follows: using the reserved data set The updated model in step (4) Perform global update to obtain the final forgetting model parameters ; By optimizing the loss function , using the stochastic gradient descent algorithm to update the model parameters, thereby ensuring that the model retains the data set The performance on the dataset is not affected, while the forgotten dataset is implemented effective forgetting.
7. The neural network selective forgetting learning method based on neuron importance according to claim 1, characterized in that: The step (6) includes the following sub-steps: (6.1) Accuracy test: Update the model parameters in step (5) , respectively, in the forgotten data set , retain the dataset and test dataset The accuracy test was performed on the model to evaluate the forgetting effect, original task performance, and generalization of the model. In the dataset The accuracy of the above is: ; According to actual needs, use , , replace That's it; among them The model is The prediction results; (6.2) Membership Inference Attack Test: Training a binary membership inference attack model , used to distinguish data samples Is it the target model? From the training data set; from the test data set Select the quantity The samples are used as non-member test sets , Computational attack model exist and Attack success rate ASR: ; The first sum represents the prediction is the number of members, and the second sum represents the prediction For the number of non-members, ASR has a good forgetting effect when it is close to 50%.
8. A neural network selective forgetting learning device based on neuron importance, characterized in that: Includes the following modules: Determine the forgotten data set module: In the random forgetting scenario, randomly select part of the training data as the forgotten data set; in the category forgetting scenario, randomly select a category of data as the forgotten data set; the remaining data is used as the reserved data set; Neuron importance evaluation module: assign the memory value of the convolutional neural network model to the forgotten samples to the neurons to evaluate the importance of the neurons to the forgotten data set; parameter mask generation module: set the neuron mask with a memory value greater than a specific threshold to 1, otherwise the mask is set to 0; Random label forgetting module: randomly set the labels of samples in the forgetting dataset and use the stochastic gradient descent method to update the parameters with the mask set to 1; Fine-tune the CNN model module: Use the reserved dataset to fine-tune the entire model to obtain the final forgotten model; Evaluation module for forgetting effect: Test the accuracy of the model after forgetting on the forgotten dataset and the accuracy of member reasoning attack to evaluate the effect of the forgetting method; test the accuracy of the model after forgetting on the retained dataset and test dataset to evaluate the performance and generalization of the model after forgetting on the original task.
Citation Information
Patent Citations
Behavior recognition method for relieving old class forgetting based on new class feature space
CN117292294A
Variable forgetting rate circuit based on memristor
CN118278478A
Method for overcoming catastrophic forgetting by neuron-level plasticity control and computing system performing the same
KR102471514B1
Artificial neural networks having attention-based selective plasticity and methods of training the same
US11210559B1
Neuromorphic memory circuit and method of neurogenesis for an artificial neural network
WO2022235789A1
Cited By
Method, device and equipment for machine learning model and storage medium
CN121072634A
Machine learning forgetting method for resisting model inversion attack
CN122153938A
Fish school ingestion desire assessment method and device coupling adaptive forgetting learning and deep learning
CN122368869A