Model forgetting method and device, electronic equipment, storage medium and computer program product
By orthogonally projecting the gradients of forgetting loss and retention loss, the optimized gradient direction is constructed, and the balance problem between forgetting and retention goals in machine forgetting is solved, and the stability and learning effect of the model in forgetting tasks are achieved.
Patent Information
- Application Number
- CN202510547060.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to balance the two optimization goals forgotten by machines, resulting in destroying the original learning effect of the model when specific information is forgotten.
By orthogonally projecting the gradients of forgetting loss and retention loss, the optimized gradient direction of the final model is constructed to ensure that the forgetting task is consistent with the optimization direction of the retention task, and a fine-grained learning vector aggregation method is used to adjust the model parameters.
It realizes the maximum protection of the model's learning effect while forgetting specific information, improves the stability and adaptability of the model, and ensures accurate and consistent performance in different forgetting scenarios.
Smart Images

Figure CN120354971A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a model forgetting method, apparatus, electronic device, storage medium, and computer program product. Background Art
[0002] Machine learning technologies can provide users with intelligent and personalized services, bringing great convenience to users' work and life. However, with the emergence of threats such as membership inference attacks and model inversion attacks, more and more users have begun to pay attention to the security and privacy of personal data. These attack methods can infer sensitive information in the training data by analyzing the output and behavior of the model, which will pose a serious threat to user privacy. Therefore, the machine forgetting technology has emerged.
[0003] "Machine forgetting" is a machine learning process aimed at deleting the influence of specific data from a trained model to protect user privacy. Specifically, machine forgetting focuses on how to make the trained model forget specific information related to user privacy while maintaining the original learning effectiveness of the model. It can be seen that there are mainly two optimization objectives in the machine forgetting process: one is to make the model forget specific information as much as possible; the other is to maintain the original training effect of the model as much as possible. However, the optimization directions of these two optimization objectives themselves are in conflict, and it is very likely that in the actual training process, success is achieved in one optimization objective while the other optimization objective is damaged, that is, it is difficult to balance between the two optimization objectives of machine forgetting in the related technologies. Summary of the Invention
[0004] The present disclosure provides a model forgetting method, apparatus, electronic device, storage medium, and computer program product to at least solve the problem in the above related technologies that it is difficult to balance between the two optimization objectives of machine forgetting.
[0005] According to the first aspect of the embodiments of the present disclosure, a model forgetting method is provided, including: performing a first training on the original model based on forgetting data to obtain a forgetting model, and performing a second training on the original model based on retention data to obtain a retention model, where the forgetting data and the retention data are data included in the original training dataset, the original training dataset includes multiple image samples and is used to train the original model for classifying images; constructing a final model based on the original model, the forgetting model, and the retention model; inputting the forgetting data and the retention data into the final model respectively to obtain a forgetting output and a retention output; calculating a forgetting loss based on the forgetting output and the forgetting label of the forgetting data; calculating a retention loss based on the retention output and the retention label of the retention data; calculating a final loss of the final model based on the forgetting loss and the retention loss; taking partial derivatives of the trainable parameters of the final model with respect to the final loss to obtain gradients of the forgetting loss and the retention loss; projecting the gradients of the forgetting loss and the gradients of the retention loss in corresponding orthogonal directions respectively to obtain an orthogonal forgetting gradient and an orthogonal retention gradient; calculating the sum of the orthogonal forgetting gradient and the orthogonal retention gradient to obtain an optimized gradient direction; performing gradient descent on the final loss in accordance with the optimized gradient direction to adjust the trainable parameters, thereby training the final model.
[0006] Optionally, the orthogonal forgetting gradient is calculated by the following formula:
[0007] The orthogonal retention gradient is calculated by the following formula: ; where is the orthogonal forgetting gradient, is the orthogonal retention gradient, is the gradient of the forgetting loss, is the gradient of the retention loss, is the norm of the vector.
[0008] Optionally, the constructing the final model based on the original model, the forgetting model, and the retention model includes: The final model is constructed by the following formula:
[0009]
[0010]
[0011] where is the final model, is the original model, is the forgetting vector, is the retaining vector, is the forgetting model, is the retaining model, are the trainable parameters of the final model.
[0012] Optionally, calculating the final loss of the final model based on the forgetting loss and the retaining loss includes: Calculating the final loss of the final model through the following formula:
[0013] where, is the final loss of the final model, is a hyperparameter, is the forgetting loss, is the retaining loss.
[0014] Optionally, the forgetting loss is represented by the following formula:
[0015] where, is the forgetting data, is the forgetting label of the forgetting data, is the final model; The retaining loss is represented by the following formula:
[0016] where, is the retaining data, is the retaining label of the retaining data, is the final model;
[0017] where, is the input, is the output, is the number of classes, is the input corresponding label, is the score predicted by the model for the input under the parameter belonging to the i-th class, is the score predicted by the model for the input under the parameter belonging to the class.
[0018] Optionally, the first training of the original model based on the forgotten data to obtain a forgotten model includes: adjusting the parameters of a preset number of layers close to the output layer of the original model based on the forgotten data to perform the first training on the original model to obtain the forgotten model.
[0019] According to a second aspect of the embodiments of the present disclosure, there is provided a model forgetting device, including: a fine-tuning module configured to perform a first training on an original model based on forgotten data to obtain a forgotten model, and perform a second training on the original model based on retained data to obtain a retained model, where the forgotten data and the retained data are data included in an original training data set, the original training data set includes a plurality of image samples and is used to train the original model for classifying images; a construction module configured to construct a final model based on the original model, the forgotten model, and the retained model; a data input module configured to input the forgotten data and the retained data into the final model respectively to obtain a forgotten output and a retained output; a forgotten loss calculation module configured to calculate a forgotten loss based on the forgotten output and a forgotten label of the forgotten data; a retained loss calculation module configured to calculate a retained loss based on the retained output and a retained label of the retained data; a final loss calculation module configured to calculate a final loss of the final model based on the forgotten loss and the retained loss; a gradient acquisition module configured to obtain gradients of the forgotten loss and the retained loss by taking partial derivatives of the trainable parameters of the final model with respect to the final loss; a projection module configured to project the gradients of the forgotten loss and the gradients of the retained loss in corresponding orthogonal directions respectively to obtain an orthogonal forgotten gradient and an orthogonal retained gradient; an optimized gradient direction calculation module configured to calculate the sum of the orthogonal forgotten gradient and the orthogonal retained gradient to obtain an optimized gradient direction; and a gradient descent module configured to perform gradient descent on the final loss in accordance with the optimized gradient direction to adjust the trainable parameters, thereby training the final model.
[0020] Optionally, the orthogonal forgotten gradient is calculated by the following formula:
[0021] The orthogonal retained gradient is calculated by the following formula: ; where is the orthogonal forgotten gradient, is the orthogonal retained gradient, is the gradient of the forgotten loss, is the gradient of the retained loss, is the norm of the vector.
[0022] Optionally, the building block is configured to: Construct the final model through the following formula:
[0023]
[0024]
[0025] where is the final model, is the original model, is the forgetting vector, is the retaining vector, is the forgetting model, is the retaining model, are the trainable parameters of the final model.
[0026] Optionally, the final loss calculation module is configured to: Calculate the final loss of the final model through the following formula:
[0027] where is the final loss of the final model, is a hyperparameter, is the forgetting loss, is the retaining loss.
[0028] Optionally, the forgetting loss is represented by the following formula:
[0029] where is the forgetting data, is the forgetting label of the forgetting data, is the final model; The retaining loss is represented by the following formula:
[0030] where is the retaining data, is the retaining label of the retaining data, is the final model;
[0031] where is the input, is the output, is the number of classes, is the input The corresponding label, is the score predicted by the model for the input under the parameter and belonging to the i-th class, is the score predicted by the model for the input under the parameter and belonging to the class.
[0032] Optionally, the fine-tuning module is configured to: adjust the parameters of a preset number of layers close to the output layer of the original model based on the forgetting data to perform the first training on the original model, and obtain the forgetting model.
[0033] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the model forgetting method according to the present disclosure.
[0034] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the model forgetting method according to the present disclosure.
[0035] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, implementing the model forgetting method according to the present disclosure.
[0036] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: In the present disclosure, by projecting the gradients of the forgetting loss and the retention loss onto the orthogonal directions of each other respectively, it can be ensured that the conflicting parts between the two are removed. And, by adding the two obtained projections, an update direction that takes both optimization objectives into account and does not introduce conflicts can be obtained. By performing gradient descent on the final loss of the final model according to this update direction to adjust the trainable parameters of the final model, it can be ensured that the optimization direction of the forgetting task is consistent with the optimization direction of the retention task, thereby avoiding unreasonable parameter changes. That is, in the present disclosure, by constraining the gradient direction in the optimization process, a better balance can be achieved between the two optimization objectives of machine forgetting, that is, while effectively eliminating the influence of forgetting data on the model, the characteristics of the retention data can be protected to the greatest extent.
[0037] Furthermore, by performing gradient projection, the stability and adaptability of the model can be further improved, enabling the model to achieve accurate and consistent performance in different forgetting scenarios.
[0038] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation on the present disclosure.
[0040] Figure 1 is a flowchart showing a model forgetting method according to an exemplary embodiment of the present disclosure; Figure 2 is a comparison diagram showing the calculation of the gradient update direction directly based on the gradients of the forgetting loss and the retention loss in the related art and the calculation of the optimized gradient direction by gradient projection in an exemplary embodiment of the present disclosure; Figure 3 is a block diagram showing a model forgetting apparatus according to an exemplary embodiment of the present disclosure; Figure 4 is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0043] It should be noted here that "at least one of several items" in this disclosure all represents three parallel situations, including "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "executing at least one of Step 1 and Step 2" means the following three parallel situations: (1) executing Step 1; (2) executing Step 2; (3) executing Step 1 and Step 2.
[0044] As mentioned above, "machine forgetting" is a machine learning process aimed at deleting the influence of specific data from a trained model to protect user privacy. Traditionally, this requirement can be achieved by directly deleting the corresponding user data from the backend database and then retraining the model. However, with the expansion of the model and dataset scale, the feasibility and efficiency of this method have been greatly reduced. Especially for complex models that require a large amount of computing resources and time for training, each response to a user's forgetting request requires retraining, which not only consumes huge resources but also is time-consuming and laborious.
[0045] To efficiently respond to users' privacy protection requirements, academia and industry have proposed various machine forgetting methods. In related technologies, the main machine forgetting methods include: fine-tuning on the retained data, gradient ascent on the forgotten data, and using a knowledge distillation framework to transfer different knowledge to achieve machine forgetting. Specifically, the fine-tuning method can weaken the influence of the forgotten data in the model by adjusting the model on the retained data; the gradient ascent method can make the model gradually forget the forgotten data by inversely optimizing the model parameters on the forgotten data; the knowledge distillation method can exclude the influence of the forgotten data in the new model by constructing a new model and using the original model to transfer knowledge, so as to achieve the purpose of forgetting.
[0046] However, although the above methods have achieved machine forgetting to a certain extent, they cannot well respond to different forgetting requests. In addition, when the proportion of retained data is small, the effects of these methods are often unstable. Considering that enterprises cannot hold user data for a long time and the retained data may only account for a small part of the total data, this further increases the difficulty of achieving stable and efficient machine forgetting. Therefore, in related technologies, it is difficult to ensure the forgetting effect and model performance when dealing with forgetting requests for different proportions of retained data.
[0047] To solve the above problems existing in the related art, the model forgetting method, device, electronic device, storage medium, and computer program product provided by the present disclosure can ensure that the conflicting parts between the gradient of the forgetting loss and the gradient of the retention loss are removed by projecting them in the orthogonal directions of each other. And, by adding the two obtained projections, an update direction that takes into account both optimization objectives and does not introduce conflicts can be obtained. By performing gradient descent on the final loss of the final model according to this update direction to adjust the trainable parameters of the final model, it can be ensured that the optimization direction of the forgetting task is consistent with the optimization direction of the retention task, thereby avoiding unreasonable parameter changes. That is, in the present disclosure, by constraining the gradient direction in the optimization process, a better balance can be achieved between the two optimization objectives of machine forgetting, that is, while effectively eliminating the influence of forgotten data on the model, the characteristics of the retained data can be protected to the greatest extent.
[0048] Furthermore, by performing gradient projection, the stability and adaptability of the model can be improved, so that the model can obtain accurate and consistent performance in different forgetting scenarios.
[0049] Figure 1 It is a flowchart showing the model forgetting method according to an exemplary embodiment of the present disclosure.
[0050] Refer to Figure 1 , in step 101, the original model can be first trained based on the forgetting data to obtain a forgetting model, and the original model can be second trained based on the retention data to obtain a retention model. The forgetting data and the retention data can be data included in the original training dataset, and the original training dataset can include multiple image samples and can be used to train the original model for classifying images. Specifically: First, the original training dataset and the pre-trained original model can be selected. When a forgetting request for specific data is received, the original model can be independently fine-tuned twice on the forgetting data and the remaining data respectively. The purpose of this process is to expect the original model to further fit their respective datasets, so as to learn the potential data distribution characteristics of these datasets.
[0051] Exemplarily, in the present disclosure, the CIFAR10 dataset can be selected as the original training dataset, and the ResNet-18 model can be selected as the original model. Assume that there is an original model trained on the complete dataset. When a request to forget the data of class 0 is received, the original model can be independently fine-tuned twice on the forgetting data (i.e., the data of class 0) and the remaining data respectively.
[0052] According to an exemplary embodiment of the present disclosure, the parameters of a preset number of layers close to the output layer of the original model can be adjusted based on the forgetting data to perform the first training on the original model to obtain a forgetting model, that is, only the parameters of several layers close to the output layer of the original model can be fine-tuned. In this way, the computational complexity can be reduced and the processing speed can be accelerated.
[0053] In step 102, a final model can be constructed based on the original model, the forgetting model, and the retention model.
[0054] According to an exemplary embodiment of the present disclosure, the final model can be constructed by the following formula:
[0055]
[0056]
[0057] Wherein, is the final model, is the original model, is the forgetting vector, is the retention vector, is the forgetting model, is the retention model, are the trainable parameters of the final model, and, , , are the learnable weights allocated layer by layer, d represents the number of layers of the model. Specifically: When constructing the final model, the forgetting model and the retention model obtained after fine-tuning can be subtracted from the original model in the parameter space respectively, and then the forgetting vector and the retention vector can be obtained. The forgetting vector can represent the parameter adjustment direction required to forget specific information, and the retention vector can represent the parameter adjustment direction required to retain the existing learning effect.
[0058] In this way, the model forgetting method provided by the present disclosure can effectively handle various types of forgetting requests, that is, it can provide customized forgetting solutions for different data sets. Further, since this method only needs to use a small amount of retention data when calculating the vector direction, therefore, under different retention data ratios, this method can provide stable and excellent performance, that is, it can ensure the performance stability, consistency, and adaptability of the model in different data environments.
[0059] In step 103, the forgetting data and the retention data can be input into the final model respectively , obtain the forgotten output and the retained output, where both the forgotten output and the retained output can be trainable parameters of the final model function.
[0060] In step 104, the forgetting loss can be calculated based on the forgotten output and the forgetting label of the forgotten data .
[0061] In step 105, the retention loss can be calculated based on the retained output and the retention label of the retained data .
[0062] In step 106, based on the forgetting loss and the retention loss , calculate the final loss of the final model .
[0063] According to an exemplary embodiment of the present disclosure, the final loss of the final model can be calculated by the following formula: The final loss is:
[0064] where, is the final loss of the final model , is a hyperparameter, is the forgetting loss, is the retention loss.
[0065] According to an exemplary embodiment of the present disclosure, the forgetting loss can be expressed by the following formula:
[0066] where, is the forgotten data, is the forgetting label of the forgotten data, is the final model; The retention loss can be expressed by the following formula:
[0067] where, is the retained data, is the retention label of the retained data, is the final model;
[0068] where, is the input, is the output, is the number of classes, is the input corresponding label, The score predicted by the model for the input under the parameter and belonging to the i-th class The score predicted by the model for the input under the parameter and belonging to the class
[0069] In step 107, the final loss can be used to take the partial derivative of the trainable parameters of the final model to obtain the gradient of the forgetting loss and the gradient of the retention loss
[0070] It should be noted that in order to make the parameter adjustment of the model more accurate and flexible, the present disclosure provides a fine-grained learnable vector aggregation method. Specifically, independent learnable weights can be set for each layer of the model. These weights allow the model to make more detailed adjustments at different levels, so as to better adapt to the different requirements of forgetting data and retaining data. For example, as mentioned above, the trainable parameters of the final model , are learnable weights allocated layer by layer, where d represents the number of layers of the model. Exemplarily, d can take values including but not limited to: 20.
[0071] In addition, the above learnable weights can be learned and adjusted through an optimization algorithm, so as to achieve the best model performance. And because this parameter adjustment method can combine existing objective functions and optimization strategies, it has high flexibility and generality, and can thus adapt to different application scenarios and requirements. Further, this parameter adjustment method can also enable the model to provide robust results under various retention data ratios, thereby enhancing the application value of the model in data privacy protection and security.
[0072] It should be noted that in multi-objective optimization, when the gradient directions of two optimization objectives (for example, the forgetting loss and the retention loss ) are inconsistent, if the gradient descent is directly performed on the loss function , it may cause the optimization result of the model on one optimization objective to be improved while the optimization result on the other optimization objective is damaged.
[0073] Therefore, in step 108, the gradient of the forgetting loss and the gradient of the retention loss can be projected in the corresponding orthogonal directions respectively to obtain the orthogonal forgetting gradient and the orthogonal retention gradient , specifically: The gradient of the forgetting loss can be projected in the orthogonal direction of the gradient of the retention loss to obtain the orthogonal forgetting gradient , and this orthogonal forgetting gradient can ensure that the part of the gradient of the forgetting loss that conflicts with the gradient of the retention loss is removed, so as to maintain the stability in the optimization process. Similarly, the gradient of the retention loss can be projected in the orthogonal direction of the gradient of the forgetting loss to obtain the orthogonal retention gradient , and this orthogonal retention gradient can ensure that the part of the gradient of the retention loss that conflicts with the gradient of the forgetting loss is removed, so as to maintain the stability in the optimization process.
[0074] In this way, during the model parameter adjustment process, by performing gradient projection, it is equivalent to constructing a feasible region, and then the parameter update can be restricted within this feasible region, which can ensure that the optimization direction of the forgetting task is consistent with the optimization direction of the retention task, thus avoiding optimization conflicts. Specifically, the gradient projection technique can balance the optimization process between the forgetting task and the retention task by constraining the direction of the gradient during the optimization process, and then can ensure that when the model parameters are adjusted, they can effectively eliminate the influence of the forgotten data while maximizing the protection of the characteristics of the retained data.
[0075] In addition, through this gradient projection technique, the model can also provide more accurate and stable results under different data forgetting ratios, improving the stability and efficiency of the model in dealing with the tasks of data forgetting and protecting retained data. Moreover, the application of this gradient projection technique not only improves the accuracy of the model in processing forgetting requests, but also enhances its robustness in practical applications.
[0076] According to an exemplary embodiment of the present disclosure, the orthogonal forgetting gradient can be calculated by the following formula:
[0077] The orthogonal retention gradient can be calculated by the following formula: ; where is the orthogonal forgetting gradient, is the orthogonal retention gradient, is the gradient of the forgetting loss, is the gradient of the retention loss, is the norm of the vector.
[0078] Thus, in view of the possible conflict problems in the optimization process, the present disclosure provides a gradient projection technique. By projecting the gradient update onto the feasible region, this technique can ensure that the optimization directions of the forgetting task and the retention task are consistent, thereby avoiding mutual interference between the two, that is, avoiding conflicts in the optimization directions between the two, and thus improving the stability and accuracy of the model in various data forgetting tasks.
[0079] In addition, through this gradient projection technique, when processing a forgetting request, the model can not only effectively complete the forgetting task, but also maximally protect the learning effect of the retained data. Moreover, the gradient projection technique can also improve the stability and accuracy of the model under different forgetting data ratios, and thus ensure that the model can exhibit strong robustness and high efficiency when dealing with complex data forgetting and privacy protection tasks.
[0080] In step 109, the orthogonal forgetting gradient and the orthogonal retention gradient can be summed to obtain the optimized gradient direction .
[0081] Figure 2 is a comparison schematic diagram showing the gradient directly based on the forgetting loss and the gradient of the retention loss in the related art to calculate the gradient update direction and the optimized gradient direction calculated by gradient projection in the exemplary embodiments of the present disclosure.
[0082] Referring to Figure 2 , Figure 2 on the left side in is a schematic diagram of calculating the gradient update direction directly based on the gradient of the forgetting loss and the gradient Figure 2 without orthogonal projection; and the gradient of the retention loss in are projected in the corresponding orthogonal directions to obtain the orthogonal forgetting gradient and the orthogonal retention gradient . After that, the optimized gradient direction is calculated based on the orthogonal forgetting gradient is shown on the right side in
[0083] In step 1010, according to the optimized gradient direction , the final loss Perform gradient descent to adjust the trainable parameters so as to train the final model thereby.
[0084] In the present disclosure, through an accurate model parameter adjustment method and a fine-grained learnable vector aggregation method, various data forgetting requests can be effectively addressed, and excellent performance can be provided in different application scenarios. Additionally, these methods not only improve the flexibility and stability of the model but also provide reliable technical guarantees for data privacy protection and the security of machine learning models.
[0085] Figure 3 is a block diagram showing a model forgetting device 300 according to an exemplary embodiment of the present disclosure.
[0086] Referring to Figure 3 , the model forgetting device 300 may include a fine-tuning module 301, a construction module 302, a data input module 303, a forgetting loss calculation module 304, a retention loss calculation module 305, a final loss calculation module 306, a gradient acquisition module 307, a projection module 308, an optimized gradient direction calculation module 309, and a gradient descent module 3010.
[0087] The fine-tuning module 301 may perform a first training on the original model based on the forgetting data to obtain a forgetting model, and may perform a second training on the original model based on the retention data to obtain a retention model. The forgetting data and the retention data may be data included in the original training dataset, and the original training dataset may include multiple image samples and may be used to train the original model for classifying images. Specifically: First, the original training dataset and the pre-trained original model may be selected. When a forgetting request for specific data is received, the original model may be independently fine-tuned twice on the forgetting data and the remaining data respectively. The purpose of this process is to expect the original model to further fit each dataset, thereby learning the potential data distribution characteristics of these datasets.
[0088] According to an exemplary embodiment of the present disclosure, the fine-tuning module 301 may adjust the parameters of a preset number of layers close to the output layer of the original model based on the forgetting data to perform the first training on the original model to obtain a forgetting model, that is, only the parameters of several layers close to the output layer of the original model may be fine-tuned. In this way, the computational complexity can be reduced and the processing speed can be accelerated.
[0089] The construction module 302 may construct a final model based on the original model, the forgetting model, and the retention model.
[0090] According to an exemplary embodiment of the present disclosure, the construction module 302 may construct the final model through the following formula:
[0091]
[0092]
[0093] Among them, is the final model, is the original model, is the forgetting vector, is the retention vector, is the forgetting model, is the retention model, are the trainable parameters of the final model, and, , , are the learnable weights assigned layer by layer, where d represents the number of layers of the model. Specifically: When constructing the final model, the forgetting model and the retention model obtained after fine-tuning can be subtracted from the original model in the parameter space respectively, and then the forgetting vector and the retention vector can be obtained. The forgetting vector can represent the parameter adjustment direction required to forget specific information, and the retention vector can represent the parameter adjustment direction required to retain the existing learning achievements.
[0094] In this way, the model forgetting method provided by the present disclosure can effectively handle various types of forgetting requests, that is, it can provide customized forgetting solutions for different data sets. Further, since this method only needs to use a small amount of retained data when calculating the vector direction, therefore, under different retained data ratios, this method can provide stable and excellent performance, that is, it can ensure the performance stability, consistency and adaptability of the model in different data environments.
[0095] The data input module 303 can input the forgetting data and the retention data into the final model respectively, and obtain the forgetting output and the retention output. Among them, both the forgetting output and the retention output can be functions of the trainable parameters of the final model.
[0096] The forgetting loss calculation module 304 can calculate the forgetting loss based on the forgetting output and the forgetting label of the forgetting data.
[0097] The retention loss calculation module 305 can calculate the retention loss based on the retention output and the retention label of the retention data.
[0098] The final loss calculation module 306 can be based on the forgetting loss and the retention loss , calculate the final model 's final loss .
[0099] According to an exemplary embodiment of the present disclosure, the final loss calculation module 306 may calculate the final loss of the final model by the following formula:
[0100] where, is the final loss of the final model , is a hyperparameter, is the forgetting loss, is the retention loss.
[0101] According to an exemplary embodiment of the present disclosure, the forgetting loss can be expressed by the following formula:
[0102] where, is the forgetting data, is the forgetting label of the forgetting data, is the final model; The retention loss can be expressed by the following formula:
[0103] where, is the retention data, is the retention label of the retention data, is the final model;
[0104] where, is the input, is the output, is the number of classes, is the input corresponding label, is the score predicted by the model for the input under the parameter belonging to the i-th class, is the score predicted by the model for the input under the parameter belonging to the th class.
[0105] The gradient acquisition module 307 can utilize the final loss for the trainable parameters of the final model Find the partial derivative to obtain the gradient of the forgetting loss and the gradient of the retention loss.
[0106] It should be noted that in order to make the parameter adjustment of the model more precise and flexible, the present disclosure provides a fine-grained learnable vector aggregation method. Specifically, independent learnable weights can be set for each layer of the model. These weights allow the model to make more detailed adjustments at different levels, so as to better adapt to the different requirements of forgetting data and retaining data. For example, as mentioned above, the trainable parameters of the final model , are layer-wise allocated learnable weights, where d represents the number of layers of the model. Exemplarily, d can take values including but not limited to: 20.
[0107] In addition, the above-mentioned learnable weights can be learned and adjusted through an optimization algorithm, so as to achieve the best model performance. And since this parameter adjustment method can combine existing objective functions and optimization strategies, it has high flexibility and generality, and can thus adapt to different application scenarios and requirements. Further, this parameter adjustment method can also enable the model to provide robust results under various retention data ratios, thereby enhancing the application value of the model in data privacy protection and security.
[0108] It should be noted that in multi-objective optimization, when the gradient directions of two optimization objectives (for example, the forgetting loss and the retention loss ) are inconsistent, if the gradient descent is directly performed on the loss function , it may cause the optimization result of the model on one optimization objective to be optimized while the optimization result on the other optimization objective is damaged.
[0109] The projection module 308 can project the gradient of the forgetting loss and the gradient of the retention loss onto the corresponding orthogonal directions respectively to obtain the orthogonal forgetting gradient and the orthogonal retention gradient . Specifically: The gradient of the forgetting loss can be projected onto the orthogonal direction of the gradient of the retention loss to obtain the orthogonal forgetting gradient . This orthogonal forgetting gradient can ensure that the part of the gradient of the forgetting loss that conflicts with the gradient of the retention loss is removed, so as to maintain the stability during the optimization process. Similarly, the gradient of the retention loss can be projected onto the gradient of the forgetting loss Perform projection in the orthogonal direction to obtain the orthogonal retention gradient , and this orthogonal retention gradient can ensure the retention of the gradient of the loss and remove the part that conflicts with the gradient of the forgetting loss in
[0110] This way, during the process of adjusting model parameters, by performing gradient projection, it is equivalent to constructing a feasible region, and then the parameter update can be restricted within this feasible region, which can ensure that the optimization direction of the forgetting task is consistent with the optimization direction of the retention task, thus avoiding optimization conflicts. Specifically, the gradient projection technique can balance the optimization process between the forgetting task and the retention task by constraining the direction of the gradient during the optimization process, and then ensure that when adjusting model parameters, it can effectively eliminate the influence of forgotten data and protect the characteristics of the retained data to the greatest extent.
[0111] In addition, through this gradient projection technique, the model can also provide more accurate and stable results under different data forgetting ratios, improving the stability and efficiency of the model in dealing with data forgetting and protecting retained data tasks. Moreover, the application of this gradient projection technique not only improves the accuracy of the model in processing forgetting requests but also enhances its robustness in practical applications.
[0112] According to an exemplary embodiment of the present disclosure, the orthogonal forgetting gradient can be calculated by the following formula:
[0113] The orthogonal retention gradient can be calculated by the following formula: ; wherein is the orthogonal forgetting gradient, is the orthogonal retention gradient, is the gradient of the forgetting loss, is the gradient of the retention loss, is the norm of the vector.
[0114] In this way, for the possible conflict problems in the optimization process, the present disclosure provides a gradient projection technique. By projecting the gradient update into the feasible region, this technique can ensure that the optimization direction of the forgetting task is consistent with the optimization direction of the retention task, thus avoiding mutual interference between the two, that is, avoiding conflicts in the optimization direction between the two, and improving the stability and accuracy of the model in various data forgetting tasks.
[0115] In addition, through this gradient projection technique, when the model processes a forgetting request, it can not only effectively complete the forgetting task, but also maximize the learning effect of the retained data. Moreover, the gradient projection technique can also improve the stability and accuracy of the model under different forgetting data ratios, and thus can ensure that the model can exhibit strong robustness and high efficiency when dealing with complex data forgetting and privacy protection tasks.
[0116] The optimized gradient direction calculation module 309 can calculate the orthogonal forgetting gradient and the orthogonal retention gradient and obtain the optimized gradient direction .
[0117] The gradient descent module 3010 can, according to the optimized gradient direction , perform gradient descent on the final loss to adjust the trainable parameters , thereby training the final model .
[0118] Figure 4 FIG. is a block diagram showing an electronic device 400 according to an exemplary embodiment of the present disclosure.
[0119] Referring to Figure 4 , the electronic device 400 includes at least one memory 401 and at least one processor 402. Instructions are stored in the at least one memory 401, and when the instructions are executed by the at least one processor 402, a model forgetting method according to an exemplary embodiment of the present disclosure is executed.
[0120] As an example, the electronic device 400 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instructions. Here, the electronic device 400 does not have to be a single electronic device, and can also be any assembly of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 400 can also be a part of an integrated control system or a system manager, or can be configured to be interconnected with a local or remote (e.g., via wireless transmission) interface as a portable electronic device.
[0121] In the electronic device 400, the processor 402 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. As an example and not a limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0122] The processor 402 can execute instructions or code stored in the memory 401, where the memory 401 can also store data. The instructions and data can also be sent and received via the network interface device over the network, where the network interface device can employ any known transmission protocol.
[0123] The memory 401 can be integrated with the processor 402. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the memory 401 can include separate devices, such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory 401 and the processor 402 can be operatively coupled or can communicate with each other, for example, via I / O ports, network connections, etc., such that the processor 402 can read files stored in the memory.
[0124] In addition, the electronic device 400 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the electronic device 400 can be connected to each other via a bus and / or network.
[0125] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the above model forgetting method. Examples of the computer-readable storage medium here include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0126] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program, which when executed by a processor implements the model forgetting method according to the present disclosure.
[0127] According to the model forgetting method, device, electronic device, storage medium, and computer program product of the present disclosure, by projecting the gradients of the forgetting loss and the retention loss in the orthogonal directions of each other respectively, it can be ensured that the conflicting parts between the two are removed. Moreover, by adding the two obtained projections, an update direction that takes into account both optimization objectives and does not introduce conflicts can be obtained. By performing gradient descent on the final loss of the final model according to this update direction to adjust the trainable parameters of the final model, it can be ensured that the optimization direction of the forgetting task is consistent with the optimization direction of the retention task, thereby avoiding unreasonable parameter changes. That is, in the present disclosure, by constraining the gradient direction in the optimization process, a better balance can be achieved between the two optimization objectives of machine forgetting, that is, while effectively eliminating the influence of forgotten data on the model, the characteristics of the retained data can be protected to the greatest extent. Further, by performing gradient projection, the stability and adaptability of the model can also be improved, so that the model can obtain accurate and consistent performance in different forgetting scenarios.
[0128] According to an exemplary embodiment of the present disclosure, by only fine-tuning the parameters of several layers close to the output layer of the original model, the computational complexity can be reduced and the processing speed can be accelerated.
[0129] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0130] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A model forgetting method, characterized in that, Including: Performing a first training on the original model based on the forgotten data to obtain a forgotten model, and performing a second training on the original model based on the retained data to obtain a retained model, where the forgotten data and the retained data are data included in the original training dataset, and the original training dataset includes multiple image samples and is used to train the original model for classifying images; Constructing a final model based on the original model, the forgotten model, and the retained model; Inputting the forgotten data and the retained data into the final model respectively to obtain a forgotten output and a retained output; Calculating a forgotten loss based on the forgotten output and the forgotten labels of the forgotten data; Calculating a retained loss based on the retained output and the retained labels of the retained data; Calculating a final loss of the final model based on the forgotten loss and the retained loss; Taking partial derivatives of the trainable parameters of the final model with respect to the final loss to obtain gradients of the forgotten loss and the retained loss; Projecting the gradients of the forgotten loss and the gradients of the retained loss in corresponding orthogonal directions respectively to obtain an orthogonal forgotten gradient and an orthogonal retained gradient; Calculating the sum of the orthogonal forgotten gradient and the orthogonal retained gradient to obtain an optimized gradient direction; Performing gradient descent on the final loss in accordance with the optimized gradient direction to adjust the trainable parameters, thereby training the final model.
2. The model forgetting method according to claim 1, wherein Calculating the orthogonal forgotten gradient through the following formula: Calculating the orthogonal retained gradient through the following formula: ; wherein, is the orthogonal forgetting gradient, is the orthogonal retention gradient, is the gradient of the forgetting loss, is the gradient of the retention loss, is the norm of the vector.
3. The model forgetting method according to claim 1, wherein The constructing a final model based on the original model, the forgotten model, and the retained model includes: Constructing the final model through the following formula: Among them, is the said final model, is the said original model, is the forgetting vector, is the retention vector, is the said forgetting model, is the said retention model, are the trainable parameters of the said final model.
4. The model forgetting method according to claim 1, wherein The calculating a final loss of the final model based on the forgotten loss and the retained loss includes: Calculating the final loss of the final model through the following formula: Among them, is the final loss of the final model, is a hyperparameter, is the forgetting loss, is the retention loss.
5. The model forgetting method according to claim 4, wherein The forgotten loss is represented by the following formula: Among them, is the forgotten data, is the forgetting label of the forgotten data, is the final model; The retained loss is represented by the following formula: Among them, is the retained data, is the retention label of the retained data, is the final model; Among them, is the input, is the output, is the number of categories, is the label corresponding to the input , is the score predicted by the model for the input under the parameter for belonging to the i-th category, is the score predicted by the model for the input under the parameter for belonging to the category.
6. The model forgetting method according to claim 1, wherein The performing a first training on the original model based on the forgotten data to obtain a forgotten model includes: Adjusting the parameters of a preset number of layers close to the output layer of the original model based on the forgotten data to perform the first training on the original model to obtain the forgotten model.
7. A model forgetting device, characterized in that, Including: A fine-tuning module configured to perform a first training on the original model based on the forgotten data to obtain a forgotten model, and perform a second training on the original model based on the retained data to obtain a retained model, where the forgotten data and the retained data are data included in the original training dataset, and the original training dataset includes multiple image samples and is used to train the original model for classifying images; A construction module configured to construct a final model based on the original model, the forgotten model, and the retained model; A data input module configured to input the forgotten data and the retained data into the final model respectively to obtain a forgotten output and a retained output; A forgotten loss calculation module configured to calculate a forgotten loss based on the forgotten output and the forgotten labels of the forgotten data; A retention loss calculation module, configured to calculate a retention loss based on the retention output and the retention label of the retention data; A final loss calculation module, configured to calculate a final loss of the final model based on the forgetting loss and the retention loss; A gradient acquisition module, configured to obtain gradients of the forgetting loss and the retention loss by taking partial derivatives of the trainable parameters of the final model using the final loss; A projection module, configured to project the gradients of the forgetting loss and the retention loss in corresponding orthogonal directions respectively to obtain an orthogonal forgetting gradient and an orthogonal retention gradient; An optimized gradient direction calculation module, configured to calculate the sum of the orthogonal forgetting gradient and the orthogonal retention gradient to obtain an optimized gradient direction; A gradient descent module, configured to perform gradient descent on the final loss in accordance with the optimized gradient direction to adjust the trainable parameters, thereby training the final model.
8. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the model forgetting method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the model forgetting method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the model forgetting method according to any one of claims 1 to 6.
Citation Information
Cited By
Image risk concept continuous erasing method and system based on feature orthogonality
CN120580513A
Video content forgetting method and system based on double-layer optimization, and medium
CN121174016A