Communication efficient federal forgetting method based on multi-objective gradient optimization

Through the federated forgetting method of multi-objective gradient optimization, client local training is combined with the Pareto optimal update of the central server, which solves the high cost and high participation threshold problems of the existing federated forgetting method, achieves efficient data deletion and model performance preservation, and is suitable for resource-constrained and privacy-sensitive application scenarios.

CN120688582APending Publication Date: 2025-09-23HARBIN INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510670321.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing federated forgetting methods have high computational costs, high participation thresholds, heavy storage burdens, and difficulty balancing the conflict between forgetting and retaining knowledge. They are unable to adapt to resource-constrained and privacy-sensitive application scenarios.

Method used

A communication-efficient federated forgetting method with multi-objective gradient optimization is adopted. Through local forgetting training on the client and Pareto-optimal gradient update on the central server, it achieves the goal of eliminating the need for global retraining and individual client operations. It combines knowledge distillation and multi-objective optimization strategies to balance forgetting and retention performance.

Benefits of technology

It reduces computing and communication costs, is suitable for resource-constrained and privacy-sensitive application scenarios, achieves efficient data deletion and model performance preservation, and is suitable for fields such as medical care, finance, and edge computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688582A_ABST
    Figure CN120688582A_ABST
Patent Text Reader

Abstract

According to the communication efficient federal forgetting method based on multi-objective gradient optimization, accurate data deletion can be realized only by independently executing local distillation training by a forgetting party client, other clients do not need to participate again or global model retraining is not needed, and the problems that an existing method is high in computing resource and high in participation threshold are effectively solved. According to the method, a Pareto optimal direction guided multi-objective optimization strategy is further designed, the influence of forgotten data can be minimized at the same time, effective knowledge of other clients can be reserved, and better stability and compatibility are shown in the real environment of heterogeneous data distribution or inconsistent client states.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a communication efficient federated forgetting method based on multi-objective gradient optimization. Background Art

[0002] With the widespread deployment of artificial intelligence in tasks such as image recognition, text processing, and medical decision-making, the need for data privacy protection and compliance is growing. This is particularly true in federated learning systems, where multiple data holders must collaborate on model training without sharing data. With the implementation of the "right to be forgotten" in GDPR, federated forgetting has become a new key capability for AI system compliance.

[0003] The existing federated forgetting methods have the following major problems:

[0004] 1. Reliance on global retraining: Some methods require retraining the entire federated model from scratch, which is computationally expensive and cannot be adapted to resource-constrained scenarios.

[0005] 2. Strong reliance on the participation of all clients: Some methods require all clients to participate in the forgetting process, which is almost impossible in the reality where clients are unavailable or participate asynchronously.

[0006] 3. Need to save historical model updates: For example, some gradient inversion methods rely on storing historical update records of all clients, which significantly increases the storage burden and the risk of privacy leakage.

[0007] 4. Unable to effectively balance the conflict between forgetting and retaining knowledge: When it is necessary to delete erroneous data or satisfy individual withdrawal requests, it is easy to destroy the overall performance of the global model and affect the accuracy of the retained data.

[0008] Therefore, there is an urgent need for an efficient federated forgetting method that does not require global retraining, supports single-client operation, and preserves performance. Summary of the Invention

[0009] The purpose of the present invention is to solve the problems in the prior art and propose a communication efficient federated forgetting method based on multi-objective gradient optimization.

[0010] The present invention is implemented through the following technical solution. The present invention proposes a communication efficient federated forgetting method based on multi-objective gradient optimization, which includes the following steps:

[0011] Step 1: Initialization and normal federated training. The central server initializes the federated global model parameters. In the standard federated learning phase, each client trains a local model based on local private data and periodically uploads local model updates to the central server. The central server aggregates the model updates from each client to form a new global model.

[0012] Step 2: A client user initiates a forget request, requesting to "forget" some data, that is, to eliminate the impact of the data on the global model;

[0013] Step 3: The forget client performs forget training locally, using gradient ascent on the forgotten data to eliminate the data's expressiveness in the model; using labels to calculate cross-entropy loss on the retained data, and aligning the global model and the training model through knowledge distillation; and uploading the trained model update to the central server;

[0014] Step 4: The central server calculates the optimal Pareto update gradient. The central server uses the last round of gradient updates received from other clients to construct a multi-objective optimization problem, looking for the gradient direction that balances forgetting performance and retention performance. The Pareto optimal gradient direction is calculated and used to guide the update process.

[0015] Step 5: Combine the forget gradient and Pareto gradient to update the global model;

[0016] Step 6: Perform steps 3-5 for T rounds until the model reaches the forgetting target and completes the forgetting process.

[0017] Furthermore, in step 1, the central server initializes the global model parameters w o , and distributed to each client; in each round of federated training, each client k performs several rounds of gradient descent on the local dataset to update the local model parameters w k , and upload the model update Δw to the central server k The central server uses weighted averaging to aggregate the model updates uploaded by the client to obtain a new global model. During the training phase, the standard federated averaging algorithm is used to perform several rounds of communication training.

[0018] Furthermore, in step 2, suppose a client u initiates a forget request to delete part of the local training data set D f Impact on the global model.

[0019] Furthermore, in step 3, client u performs forgetting training locally, with the goal of retaining as much knowledge of the undeleted data as possible while weakening the influence of the data to be forgotten. This process does not require communication with other clients.

[0020] The client defines the following local forgetting loss function:

[0021]

[0022] Where: L r To retain the loss, cross entropy loss plus KL divergence or decoupled distillation DKD can be used to maintain the prediction consistency on the retained data; Lf is the forgetting loss, which weakens the model's memory of forgotten data through reverse distillation or gradient ascent strategy; λ1 and λ2 are weighting coefficients;

[0023] After local training optimization, the client generates a local update Δw u , and upload it to the central server.

[0024] Furthermore, in step 4, the server does not need to restart the federated training process. It only uses the gradient set obtained from the last upload by other clients to construct a multi-objective optimization problem:

[0025] w t+1 =w t -ηd t ,

[0026]

[0027] In this way, the performance degradation caused by forgetting client updates is minimized, and the consistency of existing gradient information of other clients is maximized. By solving this multi-objective optimization problem, the central server obtains a Pareto optimal direction d t .

[0028] Furthermore, in step 5, the central server forgets the client's uploaded Δw u With Pareto direction d t , perform a global model update:

[0029] ω t+1 =ω t +[αΔω u -(1-α)||Δω u ||d t ]

[0030] Where: w t+1 is the current global model parameter; α∈(0,1) controls the balance between retention and forgetting; ||Δw u || is used to normalize the direction strength; d t The updated direction is output by the multi-objective optimal direction generation module.

[0031] The present invention has the following beneficial effects:

[0032] (1) The client-decoupled federated forgetting method proposed in the present invention only requires the forgetting client to perform local distillation training alone to achieve accurate data deletion, without the need for other clients to re-participate or global model retraining, effectively solving the problems of high computing resources and high participation threshold of existing methods.

[0033] (2) By introducing the final round gradient reuse mechanism, the present invention can reconstruct the knowledge contributions of multiple parties using only the gradient information in one communication, avoiding the problems of high storage and high communication overhead in traditional solutions and greatly reducing the system deployment cost.

[0034] (3) The present invention designs a multi-objective optimization strategy guided by the Pareto optimal direction, which can simultaneously minimize the impact of forgotten data and retain the effective knowledge of other clients, and show better stability and compatibility in real environments with heterogeneous data distribution or inconsistent client status.

[0035] (4) The present invention has extremely high communication and computing efficiency. Its communication cost can be reduced by more than 90% compared with traditional federated retraining. It is particularly suitable for resource-constrained or privacy-sensitive application scenarios such as medical care, finance, and edge computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a flow chart of a communication-efficient federated forgetting method based on multi-objective gradient optimization described in the present invention. DETAILED DESCRIPTION

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0038] See Figure 1 The present invention proposes a communication efficient federated forgetting method based on multi-objective gradient optimization, the method comprising the following steps:

[0039] Step 1: Initialization and normal federated training. The central server initializes the federated global model parameters. In the standard federated learning phase, each client trains a local model based on local private data and periodically uploads local model updates to the central server. The central server aggregates the model updates from each client to form a new global model.

[0040] In step 1, the central server initializes the global model parameters w o , and distributed to each client; in each round of federated training, each client k performs several rounds of gradient descent on the local dataset to update the local model parameters w k , and upload the model update Δw to the central server k The central server uses weighted averaging to aggregate the model updates uploaded by the client to obtain a new global model. During the training phase, the standard federated averaging algorithm is used to perform several rounds of communication training.

[0041] Step 2: A client user initiates a forget request, requesting to "forget" some data, that is, to eliminate the impact of the data on the global model;

[0042] In step 2, suppose a client u initiates a forget request to delete part of the local training dataset D f Impact on the global model.

[0043] Step 3: The forget client performs forget training locally, using gradient ascent on the forgotten data to eliminate the data's expressiveness in the model; using labels to calculate cross-entropy loss on the retained data, and aligning the global model and the training model through knowledge distillation; and uploading the trained model update to the central server;

[0044] In step 3, client u performs forgetting training locally, aiming to retain as much knowledge of the undeleted data as possible while weakening the influence of the data to be forgotten. This process does not require communication with other clients.

[0045] The client defines the following local forgetting loss function:

[0046]

[0047] Where: L r To retain the loss, cross entropy loss plus KL divergence or decoupled distillation DKD can be used to maintain the prediction consistency on the retained data; L f is the forgetting loss, which weakens the model's memory of forgotten data through reverse distillation or gradient ascent strategy; λ1 and λ2 are weighting coefficients;

[0048] After local training optimization, the client generates a local update Δw u , and upload it to the central server.

[0049] Step 4: The central server calculates the optimal Pareto update gradient. The central server uses the last round of gradient updates received from other clients to construct a multi-objective optimization problem, looking for the gradient direction that balances forgetting performance and retention performance. The Pareto optimal gradient direction is calculated and used to guide the update process.

[0050] In step 4, the server does not need to restart the federated training process. It only uses the gradient set obtained from the last upload by other clients to construct a multi-objective optimization problem:

[0051] ω t+1 =ω t -ηd t ,

[0052]

[0053] In this way, the performance degradation caused by forgetting client updates is minimized, and the consistency of existing gradient information of other clients is maximized. By solving this multi-objective optimization problem, the central server obtains a Pareto optimal direction d t .

[0054] Step 5: Combine the forget gradient and Pareto gradient to update the global model;

[0055] In step 5, the central server forgets the client's uploaded Δw u With Pareto direction d t , perform a global model update:

[0056] ω t+1 =ω t +[αΔω u -(1-α)||Δω u ||d t ]

[0057] Where: w t+1 is the current global model parameter; α∈(0,1) controls the balance between retention and forgetting; ||Δw u || is used to normalize the direction strength; d t The updated direction is output by the multi-objective optimal direction generation module.

[0058] Step 6: Perform steps 3-5 for T rounds until the model reaches the forgetting target and completes the forgetting process.

[0059] In step 6, steps 3 to 5 are performed for T rounds to obtain a forgetting model trained only through local client communication, thus achieving the forgetting goal.

[0060] The above description only expresses the preferred embodiments of the present invention and does not limit the present invention in any other form. Any technician familiar with the present invention may use the above disclosure to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A communication-efficient federated forgetting method based on multi-objective gradient optimization, characterized by: The method comprises the following steps: Step 1: Initialization and normal federated training. The central server initializes the federated global model parameters. In the standard federated learning phase, each client trains a local model based on local private data and periodically uploads local model updates to the central server. The central server aggregates the model updates from each client to form a new global model. Step 2: A client user initiates a forget request, requesting to "forget" some data, that is, to eliminate the impact of the data on the global model; Step 3: The forget client performs forget training locally, using gradient ascent on the forgotten data to eliminate the data's expressiveness in the model; using labels to calculate cross-entropy loss on the retained data, and aligning the global model and the training model through knowledge distillation; and uploading the trained model update to the central server; Step 4: The central server calculates the optimal Pareto update gradient. The central server uses the last round of gradient updates received from other clients to construct a multi-objective optimization problem, looking for the gradient direction that balances forgetting performance and retention performance. The Pareto optimal gradient direction is calculated and used to guide the update process. Step 5: Combine the forget gradient and Pareto gradient to update the global model; Step 6: Perform steps 3-5 for T rounds until the model reaches the forgetting target and completes the forgetting process.

2. The method according to claim 1, characterized in that In step 1, the central server initializes the global model parameters w o , and distributed to each client; in each round of federated training, each client k performs several rounds of gradient descent on the local dataset to update the local model parameters w k , and upload the model update Δw to the central server k The central server uses weighted average to aggregate the model updates uploaded by the clients to obtain a new global model. During the training phase, the standard federated averaging algorithm is used to perform several rounds of communication training.

3. The method according to claim 2, characterized in that In step 2, suppose a client u initiates a forget request to delete part of the local training dataset D f Impact on the global model.

4. The method according to claim 3, characterized in that In step 3, client u performs forgetting training locally, aiming to retain as much knowledge of the undeleted data as possible while weakening the influence of the data to be forgotten. This process does not require communication with other clients. The client defines the following local forgetting loss function: Where: L r To retain the loss, cross entropy loss plus KL divergence or decoupled distillation DKD can be used to maintain the prediction consistency on the retained data; L f is the forgetting loss, which weakens the model's memory of forgotten data through reverse distillation or gradient ascent strategy; λ1 and λ2 are weighting coefficients; After local training optimization, the client generates a local update Δw u , and upload it to the central server.

5. The method according to claim 4, characterized in that In step 4, the server does not need to restart the federated training process. It only uses the gradient set obtained from the last upload by other clients to construct a multi-objective optimization problem: This minimizes the performance degradation caused by forgotten client updates and maximizes the consistency of existing gradient information of other clients. By solving the multi-objective optimization problem, the central server obtains a Pareto optimal direction d t .

6. The method according to claim 5, characterized in that In step 5, the central server forgets the client's uploaded Δw u With Pareto direction d t , perform a global model update: w t+1 =w t +[αΔw u -(1-a)||Δw u ||d t ] Where: w t+1 is the current global model parameter; α∈(0,1) controls the balance between retention and forgetting; ||Δw u || is used to normalize the direction strength; d t The updated direction is output by the multi-objective optimal direction generation module.

Citation Information

Cited By

  • Longitudinal federal model forgetting method and device for lightweight adaptive optimizer scheduling

    CN121936627A