Federal learning client selection method based on multi-factor dynamic evaluation
By introducing multi-factor dynamic evaluation and backpack problem selection algorithms into the federated learning system, combining blockchain and mutual trust calculation, the problem of single-selected client selection and insufficient identification of malicious nodes in the existing technology is solved, significantly improving system efficiency and security.
Patent Information
- Application Number
- CN202510602120.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing federated learning system has a single standard in client selection, and it is difficult to dynamically integrate multi-dimensional features, resulting in high-potential clients being misscreened, neglecting client heterogeneity, lacking identification mechanisms for malicious nodes and lazy clients, affecting system performance and security.
A federated learning client selection method based on multi-factor dynamic evaluation is proposed. By constructing a three-dimensional evaluation system of ‘effort value-group reputation-data utility’, combined with the participant selection algorithm of backpacking problems, the client is selected to participate in federated learning, maximize the comprehensive value, and realize malicious node filtering and lazy client incentives through blockchain initialization and mutual trust calculation.
It significantly improves the overall efficiency and security of the federated learning system, accurately identify high-value clients through multi-factor evaluation, effectively filter malicious nodes, and realize efficient resource utilization and system reliability.
Smart Images

Figure CN120106183A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of federated learning and relates to a method for selecting a federated learning client based on multi-factor dynamic evaluation. Background Art
[0002] Federated learning technology uses a distributed training mechanism to achieve multi-party data value mining while protecting privacy, becoming a key technology to solve the problem of data islands. However, in actual deployment, the federated learning system faces the core challenge of uneven client participation quality: malicious nodes may upload low-quality models to interfere with global training, and resource-constrained clients have limited contributions due to insufficient data volume or computing power. The existing client selection mechanism mostly relies on single-dimensional indicators (such as data volume or historical contribution), making it difficult to accurately screen high-value participants, which seriously restricts the effectiveness of subsequent incentive mechanisms and affects the robustness and convergence efficiency of the global model.
[0003] Existing research generally focuses on the design of incentive mechanisms, but lacks in-depth optimization of the client screening process, resulting in the following problems: First, the client selection criteria are simplistic and only rely on static reputation or short-term data volume, and fail to dynamically integrate multi-dimensional features such as model update quality (such as accuracy increment), group trust, and data utility, resulting in high-potential clients being misscreened; second, the impact of client heterogeneity on the screening results is ignored. For example, clients with small data volume but high accuracy increment may be excluded due to imbalanced indicator weights; third, there is a lack of dynamic identification mechanism for "free-riding" behavior and malicious nodes. The infiltration of low-quality participants will aggravate resource waste and threaten system security.
[0004] For example, patent publication number CN114841364A discloses a federated learning method that meets the needs of personalized local differential privacy. It does not consider the free-riding and malicious users between users in federated learning, and only regards federated learning users as mutually trusted individuals, failing to reasonably screen the participation of clients. 2) The client selection evaluation criteria are simplistic, considering only one factor, the privacy budget, and ignoring the long-term or diverse behavior of clients.
[0005] Existing federated learning solutions generally do not consider the cooperative game phenomenon between federated learning clients, but only consider their individual interactions with model owners. In practice, clients may reward each other for greater rewards. For example, some clients may not be able to participate in federated learning due to small amounts of data or poor performance, or the rewards they receive are far less than the cost of their training. Federated learning mostly pursues high-precision simulations, so its rewards will increase significantly as the accuracy increases. Therefore, multiple clients can be aggregated into a client group, participate in federated learning through the client group, and then obtain the possibility of participating in federated learning and higher rewards. Summary of the invention
[0006] In view of the above problems, the present invention proposes a method for selecting federated learning clients based on multi-factor dynamic evaluation. By constructing a three-dimensional evaluation system of "effort value-group reputation-data utility", the ratio of the model accuracy increment to the data volume in a single round of training of the quantified client is used as the effort value, combined with long-term credibility as the group reputation, and the participant selection algorithm based on the knapsack problem is used to select clients to participate in federated learning, maximizing the comprehensive value of participating clients under the premise of meeting system resource constraints. It can provide a high-quality participant pool for subsequent incentive mechanisms, realize the coordinated optimization of malicious node filtering, lazy client incentives and class imbalance mitigation from the source, significantly improve the overall performance and security of the federated learning system, and provide reliable technical support for its large-scale application.
[0007] The present invention is implemented by the following technical solution: A method for selecting a federated learning client based on multi-factor dynamic evaluation, comprising the following steps: Step 1: The blockchain initializes all model owners and clients and defines the parameters of the model owners and clients; Step 2: The model owner publishes the federated learning task, including the initial global model, model requirements, and the amount of data involved in training; Step 3: The model owner calculates the active factors of the client in the current state, and then uses the participant selection algorithm based on the knapsack problem to select the client to participate in federated learning; Step 4: Each client downloads the global model, uses the local data set to train the distributed local model, and sends the trained local model to the blockchain; Step 5: The client calculates the cosine similarity between itself and other clients’ local models, and uses the cosine similarity to calculate the mutual trust between clients; the model owner reads the mutual trust between each client and calculates the group reputation of each client; Step 6: The model owner evaluates the client's effort value in this round of training based on the client's local model, which serves as a parameter for calculating the positive factor in the next round; Step 7: The model owner aggregates the local model parameters to update the global model. Step 8: Repeat steps 3 to 7 until the global model reaches the accuracy required for the federated learning task.
[0008] Further preferably, the active factor of the i-th client in the t-th round is calculated as follows: ; in, and are the group reputation and effort value of the i-th client in round t-1 respectively.
[0009] Further optimization, the participant selection algorithm process based on the knapsack problem is as follows: By traversing the current state of dp[i][w], the client with the maximum positive factor that can be selected under the data limit is gradually saved to select the optimal solution, where dp[i][w] represents the maximum positive factor of the previous i clients when the data limit is w. The update method of dp[i][w] is: dp[i][w]=max(dp[i-1][w],dp[i-1][wd[i-1]]+Post[i-1]); Among them, d[i-1] represents the local data volume of the i-1th client, and Post[i-1] represents the active factor of the i-1th client.
[0010] Further preferably, the mutual trust is calculated as follows: ; in, is the mutual trust between the i-th client in the t-th round and the j-th client in the t-th round, The value range of is [0,1], is the mutual trust between the i-th client and the j-th client in the t-1th round, is the freshness index, and its value range is [0,1].
[0011] Further preferably, the group reputation of each client is defined as the sum of the mutual trust of the client in the client group, expressed as follows: ; in, is the group reputation of the i-th client in the t-th round, with a value range of [0,n], where n is the number of clients.
[0012] Further preferably, the effort value is set as the relationship between the model increment and the data volume of the client local model, which is expressed as follows: ; in, is the effort value of the i-th client in the t-th round, is the local model accuracy of the tth round of the i-th client, is the accuracy of the global model in round t-1, Represents the data volume of the i-th client.
[0013] Technical effects of the present invention: (1) In view of the shortcomings of the existing solutions in the prior art that most of them adopt static reputation mechanisms and are difficult to identify malicious nodes that are periodically disguised, the present invention designs a client credibility calculation method based on group reputation, calculates the credibility of the client through the mutual trust between clients, and thus highlights the social attributes of federated learning in practical applications. By introducing time decay factors and behavior consistency detection, malicious clients that are disguised for a long time can be effectively identified and filtered.
[0014] (2) Traditional methods only consider a single indicator such as data volume or number of participations. We designed a positive factor to measure the client's enthusiasm, which comprehensively considers the ratio of the model accuracy increment and the client's training data volume, and can more accurately identify clients with strong training capabilities. This positive factor also takes into account the client's long-term credibility and fairly evaluates the client's past efforts, quantifying the client's improvement in model accuracy per unit data volume, thereby more fairly evaluating the client's enthusiasm.
[0015] (3) A multi-factor participant selection mechanism is proposed. This multi-factor includes the credibility and effort value in the positive factors, as well as the amount of local data of the client when the algorithm is performed. Based on this algorithm, the client selection scheme with the highest sum of positive factors can be screened out from the client group. Existing schemes mostly use methods with a single optimization objective, such as greedy algorithms. The present invention innovatively proposes a selection mechanism based on multiple factors, which jointly optimizes the group reputation, positive factors and data volume, and achieves the optimization of participant selection under the premise of satisfying resource constraints. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 The figure is a flow chart of the method of the present invention.
[0017] Figure 2 Model accuracy curves trained on the MNIST dataset for three different selection methods.
[0018] Figure 3 Model accuracy curves trained on the CIFAR dataset for three different selection methods. DETAILED DESCRIPTION
[0019] The present invention is further explained in detail below in conjunction with embodiments.
[0020] Reference Figure 1 , a federated learning client selection method based on multi-factor dynamic evaluation includes steps 1 to 8.
[0021] Step 1: The blockchain initializes all model owners and clients and defines the parameters of the model owners and clients.
[0022] 1 model owner m and n clients The federated learning system consists of is the i-th client, i={1, 2,…, n}. Under a fixed budget, after T rounds of synchronous federated learning training, the model accuracy reaches the accuracy required by the federated learning task. The accuracy of the trained local model is At the same time, the global model of round t is defined as The accuracy of the global model aggregated from all local models is .
[0023] Step 2: The model owner publishes the federated learning task, including the initial global model, model requirements, and the amount of data involved in training.
[0024] The model owner will give different initial global models according to different federated learning tasks, such as convolutional neural networks used in biomedical environments, recurrent neural networks for natural language processing, and so on.
[0025] Federated learning tasks will publish model requirements. The final global model of federated learning is the result of collaborative training of multiple clients. Its design and implementation need to meet multiple requirements, including performance, privacy protection, communication efficiency, computing efficiency, fairness, security, scalability, and interpretability. These requirements jointly determine the practicality and reliability of the federated learning model. In practical applications, it is necessary to weigh these requirements according to specific scenarios and needs to design the optimal global model.
[0026] The amount of training data is not a necessary condition for federated learning, but in most federated learning, the limited amount of training data cannot be covered by the limited amount of federated learning rewards. Therefore, in order to maximize the utility of the federated learning system, federated learning will have a corresponding incentive mechanism, which will reasonably plan the reward method for participating in federated learning and the amount of data involved in training. Assuming that there is a limit on the amount of training data involved in the federated learning task, the amount of data involved in the tth round of training is limited to .
[0027] Step 3: The model owner calculates the active factor of the client in the current state, and then uses the participant selection algorithm based on the knapsack problem to select the client to participate in federated learning.
[0028] First, define the client's positive factor. The positive factor represents the client's positive situation in federated learning. This positive situation will take into account two factors: 1. The reliability of the client, 2. The relative contribution of the client in federated learning.
[0029] The reliability of the client is defined as the client's group reputation, and the relative effort in federated learning is defined as the client's effort value. The positive factor will consider both the group reputation and the effort value to make a fair assessment of the client. The positive factor of the i-th client in the t-th round is calculated as follows: ; in and It represents the group reputation and effort value of the clients in the t-1th round.
[0030] The effort value is calculated by the model owner, and the group reputation is calculated between clients, which also avoids the client's "free-riding" behavior and "Byzantine" attack. "Free-riding" behavior refers to the situation in which some clients in the federated learning system save their own resources in order to minimize local computing input (such as reducing training rounds, reducing data quality, or directly forwarding the global model) while still obtaining the aggregated global model benefits. This phenomenon stems from the one-sidedness of the traditional incentive mechanism in the evaluation of contribution, which only relies on the surface contribution of model parameter updates and ignores the actual resource investment. "Free-riding" behavior not only leads to an imbalance between the effort value and the reward of the participants, but also may cause the problem of category imbalance in model training - when most clients choose to prioritize the categories with less training data, the recognition accuracy of the global model in a few categories is significantly reduced. "Byzantine" attack is an attack method in which malicious nodes in a distributed system destroy the security of the federated learning system by sending tampered model parameters (such as gradient reversal, weight noise injection, or label pollution).
[0031] Based on the above positive factors, a participant selection algorithm based on the knapsack problem is designed to screen out the optimal client selection scheme. Specifically, this algorithm can screen out the clients with the highest positive factors from the federated learning clients when the amount of data participating in the training is determined. This will be executed in each round of federated learning to ensure that each round is the optimal client selection scheme. Among them, the positive factor of each client is regarded as the value of the item, the amount of local data of the client is regarded as the knapsack capacity, and the data amount limit of the federated learning task is Considered as backpack capacity: The participant selection algorithm based on the knapsack problem proceeds as follows: By traversing the current state of dp[i][w], the optimal solution is selected by gradually saving the client with the maximum positive factor that can be selected under the data limit, where dp[i][w] represents the maximum positive factor of the previous i clients when the data limit is w. The update method of dp[i][w] is: dp[i][w]=max(dp[i-1][w],dp[i-1][wd[i-1]]+Post[i-1]); Among them, d[i-1] represents the local data volume of the i-1th client, and Post[i-1] represents the active factor of the i-1th client.
[0032] According to the participant selection algorithm based on the knapsack problem, the model owner can select the client participation plan with the highest positive factor under the restriction of the amount of participating data. The higher the positive factor, the greater the probability of taking positive actions when participating in federated learning.
[0033] Step 4: Each client downloads the global model, uses the local data set to perform distributed local model training, and sends the trained local model to the blockchain.
[0034] Each client uses the local dataset for training, and the loss function of the i-th client training is Can be defined as: ; in, is the local dataset of the ith client, represents the data volume of the i-th client, is the kth data sample, is the label of the kth data sample, k is the data sample number, is the local model trained in round t of the i-th client, Represents the data sample used by the i-th client In local model The loss function of .
[0035] In different federated learning, the loss function is defined in different forms. For example, in federated learning using linear regression, the loss function is Can be defined as: ; in, represents the transpose of the kth data sample.
[0036] At the same time, the goal of federated learning is to optimize a model parameter by minimizing the loss function of each client. : ; in, Represents the global loss function of the federated learning model, which is obtained by weighting the loss functions of each client. Indicates the total amount of data involved in federated learning.
[0037] Step 5: The client calculates the local model similarity between itself and other clients, and uses the local model similarity to calculate the mutual trust between clients; the model owner reads the mutual trust between each client and calculates the group reputation of each client.
[0038] When the client sends the local model to the blockchain, due to the immutability of the blockchain, each client can read each other's local model and then evaluate each other's credibility. The specific evaluation process is as follows: First, the client calculates the cosine similarity between the two local models: ; in, is the cosine similarity between the i-th client and the j-th client, is the local model trained in round t of the i-th client, is the local model trained in round t for the j-th client, is the model of the local model trained in the tth round on the i-th client, is the model of the local model trained in the tth round on the jth client.
[0039] The mutual trust between clients is set as a dynamic value based on time changes to evaluate the long-term credibility of clients. The freshness index of mutual trust is introduced , the size of the freshness index represents the importance of recent cosine similarity in the model: ; in, is the mutual trust between the i-th client in the t-th round and the j-th client in the t-th round, The value range of is [0,1], is the mutual trust between the i-th client and the j-th client in the t-1th round, is the freshness index, ranging from [0,1]; round 0 (i.e. initial value) =1, , , , represents the mutual trust between the j-th client and the i-th client in the t-th round, Represents the mutual trust of the i-th client in the t-th round; The group reputation of each client is defined as the sum of the mutual trust of the client in the client group, expressed as follows: ; in, is the group reputation of the i-th client in the t-th round, with a value range of [0,n], where n is the number of clients. By calculating the group reputation, when most clients take positive actions, the clients with positive actions will be rewarded with reputation, while the clients with negative actions will be punished with reputation, thus reducing the risk of Byzantine attacks.
[0040] Step 6: The model owner evaluates the client's effort value in this round of training based on the client's local model, which serves as a parameter for calculating the positive factor in the next round.
[0041] The model owner will evaluate the efforts of each client in federated learning. For fairness, the client's effort value is defined. The effort value represents the contribution made by the client per unit of data. This indicator can avoid the deviation caused by relying solely on the amount of data or training effect. Clients that use data efficiently can get higher contribution evaluations even if the amount of data is small. Even if a client has a small amount of data, if its accuracy increment is significantly higher than other clients, it means that it has made greater contributions to the global model through efficient training strategies or stronger computing power under limited resources.
[0042] The effort value is set as the relationship between the model increment and the data volume of the client local model, which is expressed as follows: ; in, is the effort value of the i-th client in the t-th round, is the local model accuracy of the tth round of the i-th client, is the accuracy of the global model in round t-1, Represents the data volume of the i-th client.
[0043] Step 7: Model owners aggregate local model parameters to update the global model.
[0044] After receiving the local models of each client and calculating their effort values, the model owner aggregates the local models of all clients participating in the training. The aggregation method adopted in this embodiment is the federated average algorithm. The Federated Average Algorithm (FedAvg) is a distributed model training method for federated learning, which aims to protect data privacy while achieving multi-party collaborative modeling. Specifically, it is a global model parameter update method based on weighted average, and the local model parameters are uploaded to the central server for aggregation, thereby updating the global model.
[0045] The global model will be aggregated based on the size of local data of each client. The specific aggregation method is: ; in, is the global model of round t, is the local model of the j-th client in the t-th round, , J is the client set participating in this round of training, is the data limit for the tth round of training, is the local dataset of the jth client, is the data volume of the jth client.
[0046] Step 8: Repeat steps 3 to 7 until the global model reaches the accuracy required for the federated learning task.
[0047] The Pytorch 1.2.0 framework is used to build a federated learning experimental environment. Two benchmark datasets, MNIST and CIFAR-10, are selected for algorithm verification. These two datasets are widely used to evaluate and verify the performance of the proposed federated learning participant selection mechanism. MNIST consists of handwritten digit images, including 60,000 28×28 pixel grayscale training images and 10,000 test images, covering 0-9 handwritten digit classification tasks. The images have been grayscale processed as a basic test benchmark for computer vision. The CIFAR-10 dataset contains 50,000 32×32 pixel RGB training images and 10,000 test images, divided into 10 categories of object recognition tasks. The data is divided into multi-client simulation scenarios, and each client holds a non-overlapping data subset.
[0048] In order to simulate the non-independent and identically distributed (Non-IID) characteristics of data in real scenarios, the experiment distributed the training data unevenly so that each client could obtain local data sets of different sizes and data distributions. The model training adopted the stochastic gradient descent (SGD) optimization algorithm, and the parameter update followed the FedAvg framework. Specifically, each participating client performed SGD training based on local data, and the server aggregated the model parameters submitted by the client through weighted averaging. This experimental setting can effectively evaluate the robustness of the proposed method under non-ideal data distribution conditions and provide a reliable benchmark environment for subsequent performance comparisons.
[0049] In order to verify the effectiveness of the method of the present invention, the proportion of clients that carry out malicious tampering attacks and free-riding attacks is set to 0.1.
[0050] The system parameters are shown in Table 1: Table 1 Experimental setup system parameters
[0051] The model performance under three client selection methods is compared and analyzed, including: (1) a federated learning client selection method based on multi-factor dynamic evaluation; (2) a federated learning client selection method based on a greedy algorithm, which gives priority to clients with the highest positive factors; (3) a federated learning client selection method based on random selection, which randomly selects clients. Figure 2 and Figure 3 The changing trends of model accuracy with global training rounds under different methods are demonstrated. The experimental results show that with the increase of global training rounds, the model accuracy of the three methods improves with global iterations, but the method proposed in the present invention shows significant advantages in final model performance and convergence speed.
[0052] The performance advantages are mainly reflected in the following aspects: first, compared with the greedy algorithm that only considers a single indicator of positive factors, the method of the present invention achieves a better client combination selection by comprehensively evaluating multiple factors such as positive factors and data volume restrictions; second, compared with the random selection scheme, the method of the present invention can more effectively utilize the computing resources of high-quality clients. This improvement stems from the collaborative processing of two key challenges in federated learning: ensuring the reliability of participating clients and the adequacy of training data, thereby achieving a better balance between model accuracy and training efficiency.
[0053] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0054] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. The federated learning client selection method based on multi-factor dynamic evaluation is characterized by: The following steps are involved: Step 1: The blockchain initializes all model owners and clients and defines the parameters of the model owners and clients; Step 2: The model owner publishes the federated learning task, including the initial global model, model requirements, and the amount of data involved in training; Step 3: The model owner calculates the active factors of the client in the current state, and then uses the participant selection algorithm based on the knapsack problem to select the client to participate in federated learning; Step 4: Each client downloads the global model, uses the local data set to train the distributed local model, and sends the trained local model to the blockchain; Step 5: The client calculates the cosine similarity between itself and other clients’ local models, and uses the cosine similarity to calculate the mutual trust between clients; the model owner reads the mutual trust between each client and calculates the group reputation of each client; Step 6: The model owner evaluates the client's effort value in this round of training based on the client's local model, which serves as a parameter for calculating the positive factor in the next round; Step 7: The model owner aggregates the local model parameters to update the global model. Step 8: Repeat steps 3 to 7 until the global model reaches the accuracy required for the federated learning task.
2. According to the method for selecting a federated learning client based on multi-factor dynamic evaluation in claim 1, it is characterized in that The positive factor of the i-th client in round t is calculated as follows: ; in, and are the group reputation and effort value of the i-th client in round t-1 respectively.
3. The method for selecting a federated learning client based on multi-factor dynamic evaluation according to claim 1 is characterized in that: The participant selection algorithm based on the knapsack problem proceeds as follows: By traversing the current state of dp[i][w], the client with the maximum positive factor that can be selected under the data limit is gradually saved to select the optimal solution, where dp[i][w] represents the maximum positive factor of the previous i clients when the data limit is w. The update method of dp[i][w] is: dp[i][w]=max(dp[i-1][w],dp[i-1][wd[i-1]]+Post[i-1]); Among them, d[i-1] represents the local data volume of the i-1th client, and Post[i-1] represents the active factor of the i-1th client.
4. The method for selecting a federated learning client based on multi-factor dynamic evaluation according to claim 1 is characterized in that: The mutual trust is calculated as follows: ; in, is the mutual trust between the i-th client in the t-th round and the j-th client in the t-th round, The value range of is [0,1], is the mutual trust between the i-th client and the j-th client in the t-1th round, is the freshness index, and its value range is [0,1].
5. The method for selecting a federated learning client based on multi-factor dynamic evaluation according to claim 4 is characterized in that: The group reputation of each client is defined as the sum of the mutual trust of the client in the client group, expressed as follows: ; in, is the group reputation of the i-th client in the t-th round, with a value range of [0,n], where n is the number of clients.
6. The method for selecting a federated learning client based on multi-factor dynamic evaluation according to claim 1 is characterized in that: The effort value is set as the relationship between the model increment and the data volume of the client local model, which is expressed as follows: ; in, is the effort value of the i-th client in the t-th round, is the local model accuracy of the tth round of the i-th client, is the accuracy of the global model in round t-1, Represents the data volume of the i-th client.
Citation Information
Patent Citations
Federal learning method meeting personalized local differential privacy requirements
CN114841364A
Robustness federated learning model aggregation method based on truth value discovery
CN114186237A
Wireless federated learning asynchronous training method based on optimization direction guidance
CN115618963A
Federal learning incentive method and system based on reputation reverse auction, and storage medium
CN116720593A
Personalized federated learning method and system based on similar features and collaboration
WO2024169374A1