No-forgetting and no-data distillation method for heterogeneous federal learning
By using generators to synthesize virtual samples and historical difference generators in federated learning, combined with heterogeneity fine-tuning, the problems of data heterogeneity and catastrophic forgetting in federated learning are solved, achieving better model performance and robustness.
Patent Information
- Application Number
- CN202510371648.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
AI Technical Summary
There are problems of data heterogeneity and catastrophic forgetting in federated learning, resulting in bias and performance degradation of global models during aggregation.
Using the method of forget-free and data-free distillation, virtual samples are synthesized through the generator, fidelity, diversity and transferability loss functions are jointly optimized, elastic weight consolidation penalty terms are introduced, historical difference generator parameters are updated, and global models are optimized through heterogeneity fine-tuning.
It effectively solves the problems of data heterogeneity and catastrophic forgetting, improves the robustness and generalization performance of the model, eliminates the need for additional public data sets, and reduces the cost of data acquisition.
Smart Images

Figure CN120218283A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a method for no-forgetting and no-data-distillation for heterogeneous federated learning. Background Art
[0002] With the popularization of mobile devices, Internet of Things devices, and edge computing, data shows a trend of distributed storage and processing. Traditional centralized machine learning methods need to centralize data to a central server for processing, which not only brings huge communication overhead but also faces the risk of data privacy leakage. Federated learning, as an emerging distributed machine learning paradigm, allows clients to train models locally and only upload model parameters to the server for aggregation, thus effectively protecting data privacy. However, federated learning faces many challenges in practical applications, among which the most prominent ones are data heterogeneity and catastrophic forgetting problems.
[0003] Data heterogeneity refers to the significant differences in data distributions among different clients, that is, the client data distributions have the characteristics of non-independent and identically distributed. For example, in the medical field, the medical record data of different hospitals may cover different disease types and patient groups; in the financial field, the customer data of different banks may have different consumption habits and risk preferences. This data heterogeneity will cause biases in the global model during the aggregation process, making it difficult to effectively learn the knowledge of all clients, thereby reducing the generalization performance of the model.
[0004] Catastrophic forgetting refers to the phenomenon that a machine learning model forgets the knowledge and skills learned previously when learning new tasks or new data. The global model encounters different data distributions in each round of training, which causes drastic changes in model parameters and forgets the knowledge learned in previous rounds. This phenomenon will seriously reduce the stability and performance of the model.
[0005] To address the above challenges, researchers have proposed various solutions, among which the knowledge distillation technology shows great potential. Knowledge distillation can effectively improve the performance of the student model by transferring the knowledge of the teacher model to the student model. In federated learning, the local model can be regarded as the teacher model and the global model as the student model, and the knowledge of the local model can be integrated into the global model through knowledge distillation.
[0006] However, traditional knowledge distillation methods usually require an additional public dataset to align the output distributions of the teacher model and the student model. In practical applications, it is often very difficult or even impossible to obtain such a public dataset. In addition, when the public dataset has a large difference from the client data distribution, the effect of knowledge distillation will be greatly reduced. Summary of the Invention
[0007] To solve the above problems, the object of the present invention is to provide a method for heterogeneous federated learning without forgetting and without data distillation.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A method for heterogeneous federated learning without forgetting and without data distillation, comprising the following steps:
[0010] 1) Initialize the generator on the server side, and train the generator by jointly optimizing the fidelity loss, diversity loss, and transferability loss to ensure that the generated samples are aligned with the client data distribution and are diverse. At the same time, introduce an Elastic Weight Consolidation (EWC) penalty term to prevent the output distribution of the generator from changing drastically between communication rounds;
[0011] 2) In each round of communication, update the parameters of the historical difference generator by accumulating the differences of the historical generators, so that it can synthesize samples containing the knowledge of previous rounds, thereby retaining the knowledge learned by the global model in previous rounds and alleviating the catastrophic forgetting problem;
[0012] 3) Use the samples synthesized by the current round generator and the historical difference generator for knowledge distillation. By minimizing the output difference (measured using KL divergence) between the global model and the local model on the generated samples, ensure that the global model can learn both the local knowledge of the current round and retain the knowledge of previous rounds;
[0013] 4) After the knowledge distillation is completed, perform heterogeneous fine-tuning on the global model by introducing the differences between the historical global model and the local model, so that the global model can better adapt to the data heterogeneity between clients, thereby improving the robustness and generalization performance of the model.
[0014] For further optimization of this technical solution, in step 1), the training method of the generator includes:
[0015] 1.1) Initialize the generator on the server side. The goal of the generator is to synthesize virtual samples similar to the client data distribution to replace the real common data set;
[0016] 1.2) Train the generator by jointly optimizing three loss functions:
[0017] The fidelity loss function is used to ensure that the generated samples are aligned with the client data distribution,
[0018] The diversity loss function is used to encourage the generator to generate diverse samples,
[0019] The transferability loss function is used to ensure that the generated samples can effectively transfer the knowledge of the local model;
[0020] 1.3) Introduce an elastic weight consolidation penalty term to prevent the output distribution of the generator from changing drastically between different communication rounds, thereby stabilizing the output distribution of the generator.
[0021] For further optimization of this technical solution, the three loss functions are respectively:
[0022] Fidelity loss function: Ensure that the generated samples are aligned with the client data distribution, and use the cross-entropy loss function to measure the similarity between the generated samples and the true labels. Each client k (1 ≤ k ≤ K) has a private dataset and maintains a local model f k (w k ) with parameters w k ;
[0023] Diversity loss function: Encourage the generator to generate diverse samples and avoid generating overly similar samples. Optimize using the distance metric between samples. Here, |B| is the batch size, ‖·‖2 represents the distance function. In addition, z i and z j are the generator noise inputs drawn from the normal distribution for synthesizing samples and
[0024] Transferability loss function: By minimizing the output difference between the global model and the local model on the generated samples, ensure that the generated samples can effectively transfer the knowledge of the local model;
[0025] The elastic weight consolidation penalty term: and are the generators of the current round and the previous round respectively. F i represents the Fisher information matrix, where represents the samples synthesized by the generator in the previous round, is the overall objective function for training the generator in the previous round; The diagonal elements of the FIM are used as weights in the penalty term to penalize excessive modification of the important parameters of the previous-round generator during the training of the current generator;
[0026] The overall objective function for training the generator during communication round t (t > 0) is where λ is an adjustable hyperparameter.
[0027] For further optimization of this technical solution, in step 2), the update method of the historical difference generator includes:
[0028] 2.1) In each round of communication, the parameters of the generator are updated. To retain the knowledge of previous rounds, a historical difference generator is introduced;
[0029] 2.2) The parameters of the historical difference generator are updated by accumulating the differences of the historical generator, enabling it to synthesize samples containing the knowledge of previous rounds;
[0030] 2.3) Through the historical difference generator, the global model can retain the knowledge of previous rounds and alleviate the catastrophic forgetting problem.
[0031] For further optimization of this technical solution, the parameters of the historical difference generator are updated by the following formula: where, are the parameters of the historical difference generator, are the parameters of the generator in the current round.
[0032] For further optimization of this technical solution, in step 3), the method of knowledge distillation without forgetting includes:
[0033] 3.1) On the server side, knowledge distillation is performed using the samples synthesized by the generator. The local model is regarded as the teacher model, and the global model is regarded as the student model;
[0034] 3.2) By minimizing the output difference between the global model and the local model on the generated samples, using the KL divergence metric, the knowledge of the local model is transferred to the global model;
[0035] 3.3) To alleviate catastrophic forgetting, knowledge distillation is simultaneously performed using the samples synthesized by the current-round generator and the historical difference generator. The samples of the current-round generator are used to transfer the local knowledge of the current round, while the samples of the historical difference generator are used to retain the knowledge of previous rounds;
[0036] 3.4) The loss function of knowledge distillation consists of two parts: the KL divergence loss of the samples of the current-round generator and the KL divergence loss of the samples of the historical difference generator. The contributions of the two are balanced by adjusting the weight parameter.
[0037] For further optimization of this technical solution, the knowledge distillation loss function: The loss function consists of two parts: where, L kl is the KL divergence loss of the samples of the current-round generator, L his is the KL divergence loss of the samples of the historical difference generator, and ε is the weight parameter.
[0038] For further optimization of this technical solution, in step 4), the method of global model heterogeneity fine-tuning includes:
[0039] 4.1) After knowledge distillation is completed, the global model is further optimized through a heterogeneity fine-tuning process;
[0040] 4.2) The parameter update of the global model is carried out by accumulating the differences between the historical global model and the local models, ensuring that the global model can capture the heterogeneous knowledge among the clients.
[0041] In each round of communication, the server updates the global model parameters by means of aggregated averaging, ensuring that the global model can capture the heterogeneous knowledge among the clients.
[0042] In each round of communication, the server distributes the updated global model parameters to each client for the next round of local training;
[0043] 4.5) Through heterogeneity fine-tuning, the global model can better adapt to the data heterogeneity among the clients, improving the robustness and generalization performance of the model.
[0044] For a further optimization of this technical solution, step 4.1) specifically includes that in the basic federated learning setting, each client k (1 ≤ k ≤ K) has a private dataset and maintains a local model f k (w k ), with parameters w k ; before communication round t, client k downloads the global model parameters w t-1 from the server and updates its local model parameters k based on the local dataset X as shown in the formula: where β is the learning rate of local training. During communication round t, the server receives the local model parameters and aggregates them to obtain new global model parameters w t , as shown in the formula: Then, the server distributes w t to the clients for subsequent local training.
[0045] For a further optimization of this technical solution, the parameters of the global model are updated by the following formula: where, are the fine-tuned global model parameters, w t are the global model parameters after knowledge distillation, are the local model parameters of the k-th client.
[0046] Different from the prior art, the above technical solution has the following beneficial effects:
[0047] 1) No additional public dataset: The present invention synthesizes virtual samples through a generator for knowledge distillation, without relying on an additional public dataset, reducing the data acquisition cost and avoiding the problem of degradation of knowledge distillation effect caused by the inconsistency between public data and client data distribution.
[0048] 2) Effectively cope with data heterogeneity: By introducing the Elastic Weight Consolidation (EWC) mechanism and the historical difference generator, the present invention can stabilize the output distribution of the generator, enabling the global model to better adapt to the data heterogeneity between clients, thereby improving the generalization performance of the model.
[0049] 3) Alleviate the catastrophic forgetting problem: The present invention retains the knowledge of historical rounds through the historical difference generator and combines the dual-sample distillation mechanism to ensure that the global model does not forget the previously learned knowledge while learning new knowledge, significantly alleviating the catastrophic forgetting problem.
[0050] 4) Improve the robustness and performance of the model: Through the heterogeneity fine-tuning process, the global model can better capture the heterogeneous knowledge between clients, generating a more expressive and robust model, thus showing better performance in the non-independent and identically distributed scenario. Description of the Drawings
[0051] Figure 1 It is a framework diagram of a no-forgetting and no-data distillation method for heterogeneous federated learning. Detailed Embodiments
[0052] To elaborate on the technical content, structural features, achieved objectives, and effects of the technical solution in detail, the following is described in detail in combination with specific implementation examples and with reference to the accompanying drawings.
[0053] As Figure 1 shown, it is a framework diagram of a no-forgetting and no-data distillation method based on heterogeneous federated learning. This technology includes the following steps carried out in sequence:
[0054] 1) Generator training and high-quality sample generation
[0055] Initialize the generator on the server side and train the generator by jointly optimizing the fidelity loss, diversity loss, and transferability loss to ensure that the generated samples are aligned with the client data distribution and are diverse. At the same time, introduce the Elastic Weight Consolidation (EWC) penalty term to prevent the generator output distribution from changing drastically between communication rounds.
[0056] 1.1) Initialize the generator on the server side. The goal of the generator is to synthesize virtual samples similar to the client data distribution to replace the real public dataset.
[0057] 1.2) Train the generator by jointly optimizing three loss functions:
[0058] Fidelity loss: Ensure that the generated samples are aligned with the client data distribution, and use the cross - entropy loss function to measure the similarity between the generated samples and the true labels.
[0059] Fidelity loss: Ensure that the generated samples are aligned with the client data distribution, and use the cross - entropy loss function to measure the similarity between the generated samples and the true labels. Each client k (1 ≤ k ≤ K) has a private dataset and maintains a local model f k (w k ) with parameters w k .
[0060] Diversity loss: Encourage the generator to generate diverse samples, avoid generating overly similar samples, and optimize using the distance metric between samples.
[0061] Diversity loss function: Encourage the generator to generate diverse samples, avoid generating overly similar samples, and optimize using the distance metric between samples. Where |B| is the batch size and ‖·‖2 represents the distance function. Additionally, z i and z j are the generator noise inputs drawn from the normal distribution for synthesizing samples and
[0062] Transferability loss function: Ensure that the generated samples can effectively transfer the knowledge of the local model by minimizing the output difference between the global model and the local model on the generated samples.
[0063] Transferability loss function: Ensure that the generated samples can effectively transfer the knowledge of the local model by minimizing the output difference between the global model and the local model on the generated samples.
[0064] 1.3) Introduce the Elastic Weight Consolidation (EWC) penalty term to prevent the output distribution of the generator from changing drastically between different communication rounds, thus stabilizing the output distribution of the generator.
[0065] Elastic Weight Consolidation (EWC) penalty term: and are the generators of the current round and the previous round respectively, and F i represents the Fisher information matrix, where represents the samples synthesized by the generator in the previous round, It is the overall objective function for the training of the generator in the previous round. The diagonal elements of FIM are used as weights in the penalty term to penalize excessive modification of the important parameters of the previous-round generator during the current generator training. To prevent drastic changes in the output distribution of the generator between different communication rounds, the EWC penalty term is introduced to protect the important weights of the generator in the previous round from being overly modified, thus stabilizing the output distribution of the generator.
[0066] The overall objective function for the generator training during communication round t (t > 0) is where λ is an adjustable hyperparameter.
[0067] 1.4) The generator generates high-quality virtual samples by optimizing the above loss function for the subsequent knowledge distillation process.
[0068] 2) Historical difference generator update
[0069] In each communication round, the parameters of the historical difference generator are updated by accumulating the differences of the historical generators, enabling it to synthesize samples containing the knowledge of historical rounds, thereby retaining the knowledge learned by the global model in previous rounds and alleviating the catastrophic forgetting problem.
[0070] 2.1) In each communication round, the parameters of the generator are updated. To retain the knowledge of historical rounds, a historical difference generator is introduced, and its parameters are updated by accumulating the differences of the historical generators.
[0071] 2.2) The parameters of the historical difference generator are updated by accumulating the differences of the historical generators, enabling it to synthesize samples containing the knowledge of historical rounds.
[0072] Update of the parameters of the historical difference generator: The parameters of the historical difference generator are updated by the following formula: where, are the parameters of the historical difference generator, are the parameters of the generator in the current round. In this way, the historical difference generator can synthesize samples containing the knowledge of historical rounds.
[0073] 2.3) Through the historical difference generator, the global model can retain the knowledge of historical rounds and alleviate the catastrophic forgetting problem.
[0074] 3) Knowledge distillation without forgetting
[0075] Knowledge distillation is performed using the samples synthesized by the current-round generator and the historical difference generator. By minimizing the output difference (measured using KL divergence) between the global model and the local model on the generated samples, it is ensured that the global model can learn both the local knowledge of the current round and retain the knowledge of historical rounds.
[0076] 3.1) On the server side, knowledge distillation is performed using samples synthesized by the generator. The local model is regarded as the teacher model, and the global model is regarded as the student model.
[0077] 3.2) By minimizing the output difference (using KL divergence metric) between the global model and the local model on the generated samples, the knowledge of the local model is transferred to the global model.
[0078] 3.3) To alleviate catastrophic forgetting, knowledge distillation is performed using samples synthesized by both the current-round generator and the historical-difference generator. Samples from the current-round generator are used to transfer the local knowledge of the current round, while samples from the historical-difference generator are used to retain the knowledge of historical rounds;
[0079] 3.4) The knowledge distillation loss function consists of two parts: the KL divergence loss of the current-round generator samples and the KL divergence loss of the historical-difference generator samples. The contribution of the two is balanced by adjusting the weight parameter. The KL divergence loss of the current-round generator samples is used to ensure that the global model learns the local knowledge of the current round; the KL divergence loss of the historical-difference generator samples is used to ensure that the global model retains the knowledge of historical rounds.
[0080] Knowledge distillation loss function: The loss function consists of two parts: where, L kl is the KL divergence loss of the current-round generator samples, L his is the KL divergence loss of the historical-difference generator samples, and ε is the weight parameter.
[0081] 4) Heterogeneous fine-tuning of the global model
[0082] After knowledge distillation, the global model is heterogeneously fine-tuned by introducing the differences between the historical global model and the local model, enabling the global model to better adapt to the data heterogeneity between clients, thereby enhancing the robustness and generalization performance of the model.
[0083] 4.1) After knowledge distillation, the global model is further optimized through the heterogeneous fine-tuning process.
[0084] 4.2) The global model parameter update is performed by accumulating the differences between the historical global model and the local model, ensuring that the global model can capture the heterogeneous knowledge between clients.
[0085] In the basic federated learning setting, each client k (1 ≤ k ≤ K) has a private dataset and maintains a local model f k (w k ), with parameters w k; Before communication round t, client k downloads the global model parameters w from the server t-1 and updates its local model parameters based on the local dataset X k as follows: where β is the learning rate for local training. During communication round t, the server receives the local model parameters and aggregates them to obtain new global model parameters w as shown in the formula: t Then, the server distributes w to the clients for subsequent local training. t
[0086] 4.3) Through heterogeneous fine-tuning, the global model can generate a more expressive and robust model, improving its performance in non-i.i.d. scenarios.
[0087] Heterogeneous fine-tuning: In addition to obtaining the global model through aggregation in basic federated learning, after knowledge distillation, the global model is further optimized through a heterogeneous fine-tuning process. By introducing the differences between historical global models and local models, the global model can better adapt to data heterogeneity among clients.
[0088] Global model parameter update: The parameters of the global model are updated by the following formula: where are the fine-tuned global model parameters, w t are the global model parameters after knowledge distillation, and are the local model parameters of the k-th client. In this way, the global model can better capture heterogeneous knowledge among clients, improving the robustness and generalization performance of the model.
[0089] 5) Model testing and application:
[0090] 5.1) Use the above-trained global model to test metrics such as accuracy to ensure the performance of the model under different client data distributions;
[0091] 5.2) Distribute the trained global model to each client for inference and prediction of local tasks, ensuring that each client can obtain a high-performance personalized model while protecting data privacy.
[0092] Using three publicly available benchmark datasets, CIFAR-10, CIFAR-100, and SVHN (for details, see https: / / www.cs.toronto.edu / ~kriz / cifar.html and http: / / ufldl.stanford.edu / housenumbers), the datasets are distributed to 10 clients according to different heterogeneity distributions (simulated by the Dirichlet distribution with parameters α being 1, 0.1, and 0.01 respectively) for model performance testing (where the number of joint training rounds is 200, the number of local training rounds for each client in each round of joint training is 5, and the learning rate is 0.1). Finally, the average accuracy of the global model under a framework of a method for heterogeneous federated learning without forgetting and without data distillation is significantly better than other federated learning methods. The specific results are shown in the following table:
[0093]
[0094]
[0095] Although the above embodiments have been described, once those skilled in the art learn the basic creative concept, additional changes and modifications can be made to these embodiments. Therefore, the above are only the embodiments of the present invention, and do not limit the patent protection scope of the present invention. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A forget-free and data-free distillation method for heterogeneous federated learning, characterized in that: The steps include: 1) Initialize the generator on the server side and train the generator by jointly optimizing the fidelity loss, diversity loss, and transferability loss to ensure that the generated samples are aligned with the client data distribution and diversified. At the same time, introduce elastic weight consolidation penalty terms to prevent the generator output distribution from changing drastically between communication rounds. 2) In each round of communication, the parameters of the history difference generator are updated by accumulating the differences of the history generator, so that it can synthesize samples containing historical round knowledge, thereby retaining the knowledge learned by the global model in previous rounds and alleviating the catastrophic forgetting problem; 3) Use the samples synthesized by the current round generator and the historical difference generator to perform knowledge distillation, and minimize the output difference between the global model and the local model on the generated samples to ensure that the global model can learn the local knowledge of the current round while retaining the knowledge of the historical rounds; 4) After knowledge distillation is completed, the global model is heterogeneously fine-tuned by introducing the differences between the historical global model and the local model, so that the global model can better adapt to the data heterogeneity between clients, thereby improving the robustness and generalization performance of the model.
2. The forgetting-free and data-free distillation method for heterogeneous federated learning as claimed in claim 1, It is characterized in that In the step 1), the training method of the generator includes: 1.1) Initialize the generator on the server side. The goal of the generator is to synthesize virtual samples that are similar to the client data distribution to replace the real public dataset; 1.2) Train the generator by jointly optimizing three loss functions: The truth loss function is used to ensure that the generated samples are aligned with the client data distribution. Diversity loss function, used to encourage the generator to generate diverse samples, Transferability loss function, used to ensure that the generated samples can effectively transfer the knowledge of the local model; 1.3) Introduce an elastic weight consolidation penalty term to prevent the output distribution of the generator from changing drastically between different communication rounds, thereby stabilizing the output distribution of the generator.
3. The forgetfulness-free and data-free distillation method for heterogeneous federated learning according to claim 2, characterized in that: The three loss functions are: Fidelity loss function: Ensure that the generated samples are aligned with the client data distribution and use the cross entropy loss function CE to measure the generated samples x s Similarity with the true label y, where each client k (1≤k≤K) has a private dataset And maintain a local model f k (w k ), parameter is w k ; Diversity loss function: Encourage the generator to generate diverse samples and avoid generating samples that are too similar. Use the distance metric between samples to optimize, where |B| is the batch size, ‖·‖2 represents the distance function, and z i and z j is the generator noise input drawn from a normal distribution and used to synthesize samples and Transferability loss function: By minimizing the global model f(w) and the local model f(w k ) in generating sample x s The output difference on the local model ensures that the generated samples can effectively transfer the knowledge of the local model. For the two probability distributions P and Q, the KL divergence is The elastic weights consolidate the penalty term: and are the generators of the current round and the previous round, respectively, and F i represents the Fisher information matrix, in represents the sample synthesized by the generator in the previous round, is the overall objective function of the previous round of generator training; the diagonal elements of FIM are used as weights in the penalty term to penalize excessive modifications of important parameters of the previous round of generator during the current generator training; The overall objective function of the generator training during communication round t (t>0) is Where λ is an adjustable hyperparameter and i ranges from 1 to t.
4. The forget-free and data-free distillation method for heterogeneous federated learning as claimed in claim 1, It is characterized in that In step 2), the updating method of the historical difference generator includes: 2.1) In each round of communication, the parameters of the generator are updated. In order to retain the knowledge of historical rounds, a historical difference generator is introduced; 2.2) The parameters of the history difference generator are updated by accumulating the differences of the history generator, so that it can synthesize samples containing historical round knowledge; 2.3) Through the historical difference generator, the global model can retain the knowledge of historical rounds and alleviate the catastrophic forgetting problem.
5. The forgetfulness-free and data-free distillation method for heterogeneous federated learning according to claim 4, characterized in that: The parameters of the historical difference generator are updated by the following formula: in, are the parameters of the history difference generator, are the parameters of the current round generator.
6. The forget-free and data-free distillation method for heterogeneous federated learning as claimed in claim 1, It is characterized in that In step 3), the method for distilling knowledge without forgetting includes: 3.1) On the server side, the samples synthesized by the generator are used for knowledge distillation, the local model is regarded as the teacher model, and the global model is regarded as the student model; 3.2) By minimizing the difference between the output of the global model and the local model on the generated samples, the knowledge of the local model is transferred to the global model using the KL divergence metric; 3.3) In order to alleviate catastrophic forgetting, samples synthesized by the current round generator and the historical difference generator are used for knowledge distillation. The samples of the current round generator are used to transfer the local knowledge of the current round, while the samples of the historical difference generator are used to retain the knowledge of the historical rounds. 3.4) The loss function of knowledge distillation consists of two parts: the KL divergence loss of the current round generator sample and the KL divergence loss of the historical difference generator sample. The contribution of the two is balanced by adjusting the weight parameters.
7. The forgetfulness-free and data-free distillation method for heterogeneous federated learning according to claim 6, characterized in that: The knowledge distillation loss function: The loss function consists of two parts: Among them, L kl is the KL divergence loss of the generator sample in the current round, L his is the KL divergence loss of the historical difference generator sample, ε is the weight parameter, and x his represents a sample synthesized by the history difference generator.
8. The forgetfulness-free and data-free distillation method for heterogeneous federated learning according to claim 1, characterized in that: In step 4), the method for fine-tuning the global model heterogeneity includes: 4.1) After knowledge distillation is completed, the global model is further optimized through a heterogeneous fine-tuning process; 4.2) The parameters of the global model are updated by accumulating the differences between the historical global model and the local model, ensuring that the global model can capture the heterogeneous knowledge between clients. In each round of communication, the server updates the global model parameters by means of aggregate averaging to ensure that the global model can capture the heterogeneous knowledge among clients. In each round of communication, the server distributes the updated global model parameters to each client for the next round of local training; 4.5) Through heterogeneous fine-tuning, the global model can better adapt to the data heterogeneity between clients and improve the robustness and generalization performance of the model.
9. The forgetfulness-free and data-free distillation method for heterogeneous federated learning according to claim 8, characterized in that: The step 4.1) is specifically included in the basic federated learning setting, where each client k (1≤k≤K) has a private dataset And maintain a local model f k (w k ), parameter is w k ; Before communication round t, client k downloads global model parameters w from the server t-1 , and based on the local dataset X k Update its local model parameters Such as the formula: Where β is the learning rate for local training, is the loss function l on the client local data X k and parameter w t-1 Gradients on; During communication round t, the server receives local model parameters and aggregates them to obtain new global model parameters w t , as shown in the formula: The server will then t Distribute to clients for subsequent local training.
10. The forgetfulness-free and data-free distillation method for heterogeneous federated learning according to claim 8, characterized in that: The parameters of the global model are updated by the following formula: in, are the fine-tuned global model parameters, is the global model parameter after the last round of fine-tuning, w t is the global model parameter after knowledge distillation, are the local model parameters of the kth client.
Citation Information
Cited By
Federal damping forgetting method and device, electronic equipment and storage medium
CN121638387A
Artificial intelligence edge computing terminal for fault diagnosis of industrial equipment
CN122241361A