Federal learning model repairing method based on data selection and related device
By obtaining the features of new data samples added to the federated learning model, and using the data selection method for annotation and retraining, the federated learning model's performance degradation in data distribution offset is solved, and the model repair efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510396207.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
The federated learning model's performance deteriorates due to data distribution offset when facing new data. The existing retraining methods are computationally expensive and the tag collection of new data is heavy, which limits the efficiency of model repair.
By obtaining the data characteristics of the newly added data samples, a preset data selection method is used to select a preset number of new data samples for annotation and retraining, and repair it in combination with the global model and the local model to optimize the model performance.
It effectively alleviates the contradiction between the huge number of new data samples and limited testing resources, significantly improves the repair efficiency and accuracy of the federated learning model, and is suitable for different data sets and models.
Smart Images

Figure CN120278296A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of artificial intelligence and machine learning, and relates to a method for repairing a federated learning model based on data selection and related devices. Background Art
[0002] Federated learning has become a powerful technology in recent years and has been widely applied in different fields such as digital health, wireless communication, Internet of Things, and mobile edge computing. The core of the framework design of federated learning is to train multiple deep neural networks using private data at the client side and aggregate these models together at the central server as a global model to solve specific tasks. The main purpose of federated learning is to protect the private information that each customer does not want to share while using this information to solve the same problem.
[0003] Although federated learning, as a machine learning technology, has been widely applied in different fields, it also has problems that are often encountered in machine learning: data distribution shift, that is, there is a difference between the test distribution after model deployment and the training distribution during training, resulting in a significant performance degradation of the trained model on new test data. This is a fundamental challenge in machine learning and data-driven frameworks. As long as the model is trained on limited training data, it is vulnerable to distribution shift, and this problem is also common in federated learning. Many studies have shown that federated learning has a data distribution shift problem. For example, an image processing model based on federated learning often has low accuracy of processing results due to data distribution shift when processing image samples with a large distribution difference from the training image samples.
[0004] Therefore, it is necessary to regularly repair the model to handle data distribution shift. Model repair, that is, retraining the model with new data, is a direct method to handle data distribution shift. However, labeling the newly collected data is a heavy task, and the computational cost of retraining all the data is very high. These problems limit the direct use of retraining-based repair techniques in federated learning models. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a method for repairing a federated learning model based on data selection and related devices.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] In the first aspect of the present invention, a method for repairing a federated learning model based on data selection is provided, including: obtaining new data samples of the federated learning model; wherein, the distribution difference between the new data samples and the training data samples of the federated learning model is greater than a preset first threshold; using the global model or local model of the federated learning model to obtain the data features of the new data samples; according to the data features of the new data samples, using a preset data selection method to select a preset number of new data samples as supplementary data samples and label them; retraining the local model with the labeled supplementary data samples and the training data samples, uploading the retrained local model to the central cloud server, and receiving the global model based on the retrained local model sent by the central cloud server to obtain a repaired global model.
[0008] Optionally, the obtaining the data features of the new data samples based on the global model or local model of the federated learning model includes: when the distribution difference between the new data samples and the training data samples of the federated learning model is greater than a preset second threshold, obtaining the data features of the new data samples through the local model of the federated learning model; when the distribution difference between the new data samples and the training data samples of the federated learning model is not greater than the preset second threshold, randomly selecting a model from the global model or local model of the federated learning model to obtain the data features of the new data samples.
[0009] Optionally, the obtaining the data features of the new data samples based on the global model or local model of the federated learning model includes: inputting the new data samples into the global model or local model of the federated learning model, and taking the output of the logical value layer of the global model or local model of the federated learning model as the data features of the new data samples.
[0010] Optionally, the using a preset data selection method to select a preset number of new data samples includes: sorting the new data samples according to the data features of the new data samples by using a data selection method based on uncertainty or a data selection method based on diversity; selecting the first preset number of new data samples in the sorted new data samples.
[0011] Optionally, it further includes: denoting the current repaired global model as the first model, and obtaining a newly obtained repaired global model denoted as the second model by replacing the model used when obtaining the data features of the new data samples; selecting the model with better performance from the first model and the second model as the final repaired global model.
[0012] Optionally, it further includes: obtaining the repaired global models under each data selection method by replacing the preset data selection method; selecting the repaired global model with the optimal model performance among the repaired global models under each data selection method as the final repaired global model.
[0013] Optionally, the federated learning model is an image processing model, and both the newly added data samples and the training data samples are image samples; or the federated learning model is an audio processing model, and both the newly added data samples and the training data samples are audio samples; or the federated learning model is a text processing model, and both the newly added data samples and the training data samples are text samples.
[0014] In a second aspect of the present invention, there is provided a federated learning model repair system based on data selection, including: a data acquisition module for acquiring newly added data samples of the federated learning model; wherein the distribution difference between the newly added data samples and the training data samples of the federated learning model is greater than a preset first threshold; a feature extraction module for obtaining data features of the newly added data samples by using the global model or the local model of the federated learning model; a data selection module for selecting a preset number of newly added data samples as supplementary data samples and annotating them according to the data features of the newly added data samples by using a preset data selection method; a model repair module for retraining the local model by using the annotated supplementary data samples and the training data samples, uploading the retrained local model to the central cloud server, and receiving the global model based on the retrained local model issued by the central cloud server to obtain a repaired global model.
[0015] In a third aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned federated learning model repair method based on data selection are implemented.
[0016] In a fourth aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned federated learning model repair method based on data selection are implemented.
[0017] Compared with the prior art, the present invention has the following beneficial effects:
[0018] The method for repairing a federated learning model based on data selection in the present invention first obtains new data samples, then uses the global model or local model of the federated learning model to obtain the data features of the new data samples, and selects a preset number of new data samples based on the data features by using a preset data selection method, so as to find the key new data samples beneficial to the repair of the federated learning model, effectively alleviating the contradiction between the large number of new data samples and the limited test resources, avoiding the problem of poor model repair effect caused thereby, and significantly improving the repair efficiency of the federated learning model. The model repair method has good versatility and is applicable to repair using different data selection methods on various data sets, not limited to specific data sets or models. It effectively solves the problem of how to efficiently repair the federated learning model when facing a large number of new data with unknown data distributions, that is, when the distribution of the new data samples is significantly different from that of the training data samples of the federated learning model. Description of the Drawings
[0019] Figure 1 It is a flowchart of the method for repairing a federated learning model based on data selection according to an embodiment of the present invention.
[0020] Figure 2 It is a schematic diagram of the principle of repairing a federated learning model based on data selection according to an embodiment of the present invention.
[0021] Figure 3 It is a performance comparison diagram of the repaired global model under different feature extraction models according to an embodiment of the present invention.
[0022] Figure 4 It is a performance comparison diagram of the repaired global model under different data selection methods according to an embodiment of the present invention.
[0023] Figure 5 It is a block diagram of the system structure for repairing a federated learning model based on data selection according to an embodiment of the present invention. Detailed Embodiments
[0024] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0025] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] The present invention will be further described in detail below with reference to the accompanying drawings:
[0027] See Figure 1 and 2 In an embodiment of the present invention, a method for repairing a federated learning model based on data selection is provided, which can effectively improve the quality and efficiency of repairing the federated learning model.
[0028] Specifically, the method for repairing a federated learning model based on data selection of the present invention includes the following steps:
[0029] S1: Obtain new data samples of the federated learning model; wherein, the distribution difference between the new data samples and the training data samples of the federated learning model is greater than a preset first threshold.
[0030] S2: Use the global model or local model of the federated learning model to obtain the data characteristics of the new data samples.
[0031] S3: According to the data characteristics of the new data samples, use a preset data selection method to select a preset number of new data samples as supplementary data samples and label them.
[0032] S4: Retrain the local model using the labeled supplementary data samples and training data samples, upload the retrained local model to the central cloud server, and receive the global model based on the retrained local model sent by the central cloud server to obtain a repaired global model.
[0033] Explanatory, federated learning is a deep learning technique that allows collaborative model training among multiple users using their private data. The user side trains local models and only sends model updates to the central cloud server, which then aggregates these updates into a global model, allowing the user side to protect their privacy while still contributing to the collective model. Federated learning enables collaborative model training among multiple user sides without sharing data. Generally, federated learning consists of a central cloud server and multiple user sides. In each round of federated learning training, the central cloud server first dispatches a global model to the user sides for local training. Then, each user side uses its private data to train the received global model, obtains the trained local model, and uploads it to the central cloud server. Finally, the central cloud server aggregates all the local models collected from the user sides to update the global model and sends it down to each user side.
[0034] Explanatory, due to data privacy considerations, the testing of federated learning models is deployed on the user side. Like other deep learning models, the deployed federated learning models also face the challenge of encountering new data with unknown distributions, which may lead to performance degradation over time. How to effectively maintain the performance of federated learning models by processing new data with minimal cost is an important issue.
[0035] Explanatory, the essence of the federated learning model repair mechanism based on data selection is to select some representative data from the entire data pool for annotation, and then retrain the federated learning model together with the original training data, so that the model repair work (label budget and computing resources) is controllable. Different from the scenario of single model repair, federated learning involves two types of models, namely the global model and the local model. Considering the settings of federated learning, it is extremely important to effectively use the data selection-based model repair mechanism on the federated learning model to achieve effective repair of the federated learning model.
[0036] The method for repairing a federated learning model based on data selection of the present invention first obtains new data samples, then uses the global model or local model of the federated learning model to obtain the data features of the new data samples, and selects a preset number of new data samples based on the data features by using a preset data selection method, so as to find the key new data samples beneficial to the repair of the federated learning model, effectively alleviating the contradiction between the large number of new data samples and the limited test resources, avoiding the problem of poor model repair effect caused thereby, and at the same time significantly improving the repair efficiency of the federated learning model. The model repair method has good versatility and is applicable to repair using different data selection methods on various data sets, not limited to specific data sets or models. It effectively solves the problem of how to efficiently repair the federated learning model when facing a large number of new data with unknown data distributions, that is, when the distribution difference between the new data samples and the training data samples of the federated learning model is relatively large.
[0037] In a possible implementation manner, obtaining the data features of the new data samples based on the global model or local model of the federated learning model includes: when the distribution difference between the new data samples and the training data samples of the federated learning model is greater than a preset second threshold, obtaining the data features of the new data samples through the local model of the federated learning model; when the distribution difference between the new data samples and the training data samples of the federated learning model is not greater than the preset second threshold, randomly selecting a model from the global model or local model of the federated learning model to obtain the data features of the new data samples.
[0038] Explanatorily, due to the characteristics of the federated learning framework, each client has two models that can be used to extract data features: the local model trained by the client and the global model aggregated by the central server. According to the user's needs, a model can be freely selected to predict the unlabeled data.
[0039] Exemplarily, it is recommended to use the local model when the distribution difference between the new data samples and the training data samples of the federated learning model is relatively large. For example, a reference value, that is, the second threshold, can be preset in advance. When the distribution difference between the new data samples and the training data samples of the federated learning model is greater than the preset second threshold, the local model of the federated learning model is selected as the extraction model for the data features of the new data samples. Based on this selection, the characteristics of the local data can be better adapted, and the pertinence and accuracy of data processing can be improved.
[0040] In a possible implementation manner, obtaining the data features of the new data samples based on the global model or local model of the federated learning model includes: inputting the new data samples into the global model or local model of the federated learning model, and taking the output of the logical value layer of the global model or local model of the federated learning model as the data features of the new data samples.
[0041] Explanatory, the Logits layer is the raw and unprocessed scores or scores of the model output layer, especially in classification problems. It is the output result of the neural network model in the last layer (fully connected layer or output layer), providing score information for each category. These raw scores can be regarded as a measure of the model's confidence or probability for each category, but they do not directly represent probabilities. Specifically, the output of the Logits layer is the linear output of the model for each category and has not been processed by the activation function.
[0042] In a possible implementation manner, the step of selecting a preset number of new data samples by using the preset data selection method includes: sorting the new data samples according to the data characteristics of the new data samples by using a data selection method based on uncertainty or a data selection method based on diversity; and selecting the first preset number of new data samples in the sorted new data samples.
[0043] Explanatory, the data selection method is used to select new data samples by using the extracted data characteristics to solve the problem of data distribution shift caused by new data. Usually, all new data samples are sorted according to the data selection method, and then the TOP-N new data samples are selected according to the quantity requirement N.
[0044] The data selection method based on uncertainty generally includes sensitivity analysis method, probability analysis method, Monte Carlo simulation method, scenario analysis method, Bayesian analysis method, and DeepGini method, etc. The data selection method based on diversity generally includes clustering method, K-medoids method, maximum minimum distance method, and method based on feature similarity, etc. Among them, the DeepGini method simplifies the problem of measuring the misclassification probability to the problem of measuring the impurity of the test set. It is based on an intuitive idea: if a model outputs the same probability for each category, then the test data is easily misclassified by the model. Therefore, the DeepGini method prioritizes the test data by calculating the possibility of each test data being misclassified.
[0045] In a possible implementation manner, the method for repairing a federated learning model based on data selection further includes: denoting the currently repaired global model as the first model, and obtaining a re-obtained repaired global model by replacing the model used when obtaining the data characteristics of the new data samples, and denoting it as the second model; and selecting the model with better model performance between the first model and the second model as the final repaired global model.
[0046] Explanatory. Based on the above design, it is possible to flexibly select a better model as the final repaired global model, making full use of the advantages of different data feature extraction models. By actual comparison, it is ensured that the finally selected model has better performance, thereby improving the accuracy and reliability of the federated learning model repair, and enhancing the adaptability and generalization ability of the repaired global model to different data distributions.
[0047] In a possible implementation, the federated learning model repair method based on data selection further includes: obtaining the repaired global models under each data selection method by replacing the preset data selection method; selecting the repaired global model with the optimal model performance among the repaired global models under each data selection method as the final repaired global model.
[0048] Explanatory. By trying different data selection methods and obtaining the corresponding repaired global models, it is possible to comprehensively evaluate the effects of various data selection strategies, with stronger robustness and generalization ability, enhancing the adaptability of the repaired model to different data features and distributions, and thus achieving better results in practical applications.
[0049] In a possible implementation, the federated learning model is an image processing model, and both the newly added data samples and the training data samples are image samples; or the federated learning model is an audio processing model, and both the newly added data samples and the training data samples are audio samples; or the federated learning model is a text processing model, and both the newly added data samples and the training data samples are text samples.
[0050] Explanatory. The federated learning model repair method based on data selection can be applied to the fields of image processing, audio processing, and text processing. For image processing models, such as image classification models and object detection models, it has good application effects and effectively maintains the performance of image processing models under different data distributions.
[0051] When the federated learning model is an image processing model, the present invention provides an image processing model repair method based on data selection. By repairing the image processing model according to the newly added data samples, that is, the newly added image samples, it is finally ensured that when the image processing model processes newly added image samples with a large distribution difference from the training data samples, that is, the training image samples, a certain processing accuracy can be guaranteed.
[0052] In a possible implementation, taking an image classification model as an example, the federated learning model repair method based on data selection of the present invention is described.
[0053] For three federated learning-based image classification models trained on the Fashion-MNIST dataset using three training sets with different data distributions, the federated learning model repair method based on data selection of the present invention is implemented to test the model repair effects of data selection using the global model and the local model. The specific process is as follows.
[0054] The global model or the local model of the image classification model at the client is used to extract features from the newly added image samples. Depending on the data selection method, the model layer for feature extraction may vary. In this embodiment, the output of the Logits layer of the model is selected as the data feature of the newly added image samples.
[0055] Based on the data features of the newly added image samples, the newly added image samples are selected. In this embodiment, the DeepGini method is used to select the newly added image samples according to the uncertainty of the data features of the newly added image samples. The calculation method of the uncertainty is as follows:
[0056]
[0057] where x is the newly added image sample, p(·) is the feature extraction function, K is the total number of features, and j is a specified feature.
[0058] Finally, the selected newly added image samples are labeled as supplementary image samples, the supplementary image samples are merged with the original training image samples, and the local model is retrained and then sent back to the central server for aggregation to obtain the repaired global model. Among them, based on the distinction of the model used when obtaining the data features of the newly added data samples, two repaired global models can be obtained, namely the repaired global model based on the global model and the repaired global model based on the local model.
[0059] Evaluate the repair effects of the two repaired global models. The results are as Figure 3 shown. The experimental results show that the repaired global model based on the global model has a better model repair effect. The method proposed by the present invention can effectively evaluate the advantages and disadvantages of using the global model or the local model for repair data selection.
[0060] In a possible embodiment, the federated learning model repair method based on data selection of the present invention is implemented for different data selection methods to evaluate the repair effects of different data selection methods. The specific process is as follows.
[0061] In this embodiment, the data selection methods adopted include Random, LoGo, LC-g, MCP-g, Gini-g, LC-l, MCP-l, and Gini-l. Among them, LoGo is the method proposed in Document 1 (Document 1: Kim, SangMook, et al. "Re-thinking federated active learning based on inter-class diversity." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023.), LC (Least Confidence) is the least confidence strategy, MCP (Multiple-boundary Clustering and Prioritization) is the multi-boundary clustering and prioritization strategy, Gini (DeepGini) is the deep Gini coefficient strategy, and g and l respectively represent extracting data features of newly added data samples using the global model or the local model.
[0062] Finally, the repaired global models under different data selection methods are obtained, and the performance of the repaired global models under different data selection methods is compared to evaluate the repair effects of different data selection methods. The results are as Figure 4 shown. The experimental results show that the data selected by LoGo and LC-g can better repair the model. The method proposed in the present invention can effectively evaluate the advantages and disadvantages of using different data selection methods for model repair.
[0063] Generally speaking, the federated learning model repair method based on data selection in the present invention solves the challenge of maintaining the federated learning model in the face of new unknown data distributions. Especially when the number of newly received unlabeled data is large, using this method can efficiently complete the repair data selection, find the key data beneficial to model repair, and avoid the poor repair effect caused by the conflict between the large amount of new data and limited test resources. The repair and evaluation framework of the present invention has good versatility, is applicable to repairing and evaluating different data selection methods on various data sets, is not limited to specific data sets or models, and provides richer information for the interpretability of the federated learning model behavior and the repair of the model.
[0064] The following is the device embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For the details not disclosed in the device embodiment, please refer to the method embodiment of the present invention.
[0065] See Figure 5, in another embodiment of the present invention, a federated learning model repair system based on data selection is provided, which can be used to implement the above-mentioned federated learning model repair method based on data selection. Specifically, the federated learning model repair system based on data selection includes a data acquisition module, a feature extraction module, a data selection module, and a model repair module.
[0066] Among them, the data acquisition module is used to acquire new data samples of the federated learning model; among them, the distribution difference between the new data samples and the training data samples of the federated learning model is greater than a preset first threshold; the feature extraction module is used to use the global model or local model of the federated learning model to acquire the data features of the new data samples; the data selection module is used to select a preset number of new data samples as supplementary data samples and label them according to the data features of the new data samples by using a preset data selection method; the model repair module is used to retrain the local model with the labeled supplementary data samples and training data samples, upload the retrained local model to the central cloud server, and receive the global model based on the retrained local model sent by the central cloud server to obtain a repaired global model.
[0067] All relevant contents of each step involved in the embodiment of the above-mentioned federated learning model repair method based on data selection can be cited in the function description of the corresponding functional modules of the federated learning model repair system in the embodiment of the present invention, and will not be elaborated here.
[0068] The division of modules in the embodiments of the present invention is illustrative, only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present invention, the functional modules can be integrated in one processor, or exist separately physically, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0069] In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the method for repairing a federated learning model based on data selection.
[0070] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the method for repairing a federated learning model based on data selection in the above embodiments.
[0071] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0072] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0073] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for repairing a federated learning model based on data selection, characterized in that Including: Obtain new data samples of the federated learning model; wherein, the distribution difference between the new data samples and the training data samples of the federated learning model is greater than a preset first threshold; Use the global model or local model of the federated learning model to obtain the data features of the new data samples; According to the data features of the new data samples, select a preset number of new data samples using a preset data selection method as supplementary data samples and label them; Retrain the local model using the labeled supplementary data samples and training data samples, upload the retrained local model to the central cloud server, and receive the global model based on the retrained local model sent by the central cloud server to obtain a repaired global model.
2. The method for repairing a federated learning model based on data selection according to claim 1, wherein The obtaining of the data features of the new data samples using the global model or local model of the federated learning model includes: When the distribution difference between the new data samples and the training data samples of the federated learning model is greater than a preset second threshold, obtain the data features of the new data samples through the local model of the federated learning model; When the distribution difference between the new data samples and the training data samples of the federated learning model is not greater than a preset second threshold, randomly select a model from the global model or local model of the federated learning model to obtain the data features of the new data samples.
3. The method for repairing a federated learning model based on data selection according to claim 1, wherein The obtaining of the data features of the new data samples using the global model or local model of the federated learning model includes: Input the new data samples into the global model or local model of the federated learning model, and use the output of the logical value layer of the global model or local model of the federated learning model as the data features of the new data samples.
4. The method for repairing a federated learning model based on data selection according to claim 1, wherein The selecting of a preset number of new data samples using a preset data selection method includes: According to the data features of the new data samples, sort the new data samples using a data selection method based on uncertainty or a data selection method based on diversity; Select the first preset number of new data samples in the sorted new data samples.
5. The method for repairing a federated learning model based on data selection according to claim 1, wherein Also included: Record the current repaired global model as the first model, and obtain a newly obtained repaired global model by replacing the model used to obtain the data features of the new data samples, and record it as the second model; Select the model with better model performance between the first model and the second model as the final repaired global model.
6. The method for repairing a federated learning model based on data selection according to claim 1, wherein Also included: Obtain the repaired global model under each data selection method by replacing the preset data selection method; Select the repaired global model with the optimal model performance among the repaired global models under each data selection method as the final repaired global model.
7. The method for repairing a federated learning model based on data selection according to claim 1, wherein The federated learning model is an image processing model, and both the new data samples and the training data samples are image samples; or the federated learning model is an audio processing model, and both the new data samples and the training data samples are audio samples; or the federated learning model is a text processing model, and both the new data samples and the training data samples are text samples.
8. A federated learning model repair system based on data selection, characterized in that, Including: A data acquisition module for obtaining new data samples of the federated learning model; wherein, the distribution difference between the new data samples and the training data samples of the federated learning model is greater than a preset first threshold; A feature extraction module, which is used to obtain the data features of newly added data samples by using the global model or local model of the federated learning model; A data selection module, which is used to select a preset number of newly added data samples as supplementary data samples and label them according to the data features of the newly added data samples by using a preset data selection method; A model repair module, which is used to retrain the local model by using the labeled supplementary data samples and training data samples, upload the retrained local model to the central cloud server, and receive the global model based on the retrained local model sent by the central cloud server to obtain a repaired global model.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the federated learning model repair method based on data selection according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the federated learning model repair method based on data selection according to any one of claims 1 to 7 are implemented.