Fair federal learning method for non-independent identically distributed data

By calculating client influence in federated learning and adjusting its aggregate weights, the client fairness problem under non-independent and homogeneous data is solved, and the accuracy and fairness of the global model are improved.

CN120012955APending Publication Date: 2025-05-16CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510076287.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In federated learning, the data between clients is often non-independent and homogeneously distributed, which makes it difficult for conventional federated learning algorithms to achieve the expected goals and a fairness problem between participating clients.

Method used

A fair federated learning method for non-independent and homogeneous data is proposed. Through the system initialization module, local module, influence calculation module, aggregation weight adjustment module and client selection module, the influence of each client is calculated and its aggregation weight is adjusted to ensure that clients with higher influence participate in more training.

Benefits of technology

By improving the accuracy of the global model on poorly performed clients and improving the fairness of federated learning, the experimental results show that the proposed method has improved both Jain’s Fairness Index and variance on multiple data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012955A_ABST
    Figure CN120012955A_ABST
Patent Text Reader

Abstract

The invention discloses a fair federated learning method for non-independent identically distributed data, and belongs to the technical field of federated learning. Although many fairness schemes for federated learning exist at present, most of the schemes need to increase client resource overhead. In the implementation process, an auxiliary data set is introduced, cross entropy and marginal loss are combined, and the performance of a global model when a client does not participate in aggregation is compared with the performance of the global model when the client participates in aggregation, so that the influence of the client is measured; establishing an aggregation weight adjustment mechanism by utilizing the solved influence, and ensuring that a client with relatively large influence on the global model obtains a higher aggregation weight; and according to the influence of the participating clients in each round, the number of times that the clients with relatively high influence participate in federal training is increased. Meanwhile, by the adoption of the method, client computing resource consumption and any additional communication do not need to be increased, the virtuous circle in the federal learning process is promoted, and the enthusiasm of participants is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning, and in particular to a fair collaboration method between clients in federated learning. Background Art

[0002] As a new type of production factor, data resources have an increasingly strong demand for openness, sharing and circulation. At the same time, as the scale of data continues to expand and data security awareness continues to increase, traditional centralized data collection and processing methods have risks such as high communication costs, high device power consumption, and infringement of user privacy and security. In order to maximize the value of data while ensuring data security as much as possible, the method of federated learning came into being.

[0003] Federated learning is a machine learning technology that trains models in a distributed manner. Through local training and central aggregation, each node can jointly train a global model without leaving the local data, which effectively solves the problems of data silos and data privacy. However, in practical applications, federated learning faces a key problem, that is, the data between clients are often non-independent and identically distributed (Non-IID, non-independent and identically distributed), with large differences. Therefore, conventional federated learning algorithms may find it difficult to achieve the desired goals, and in such cases, fairness issues will also arise among participating clients.

[0004] Since federated learning is a collaborative learning involving multiple users, the impact of different clients on the learning process may vary greatly, and this impact is affected by the combined effect of multiple factors. In particular, when the data between clients is not independent and identically distributed, the local model will have different update directions, causing the local model update obtained by client optimization to deviate from the global model update obtained by server aggregation, which in turn causes significant differences between the models of different clients. Each client cannot be treated fairly, which will seriously affect the enthusiasm of users to participate in federated learning and is not conducive to long-term development in practical applications.

[0005] Existing technologies have certain limitations in solving the fairness problem of federated learning, and the industry has conducted in-depth research from multiple aspects. Existing research includes improving data distribution, adopting personalized strategies, limiting local updates, optimizing client selection strategies, improving aggregation algorithms, and other directions. Among them, there are some relatively well-known solutions, but they also have shortcomings. For example, the FedProx algorithm allows the client to determine the degree of local training based on computing power, and prevents the local model from deviating too much from the global model through the proximal term. However, the proximal term constraint is isotropic, which limits the algorithm's ability to deal with non-independent and identically distributed data; the FedMGDA+ algorithm focuses on the results of the current model on all clients to ensure that the model reaches Pareto optimality, but Pareto optimality cannot guarantee that the client accuracy distribution is uniform enough, and there may still be unfairness.

[0006] In addition, since the client's computing resources and communication bandwidth are limited, and federated learning has strict requirements for privacy protection, it is a challenging task to improve the fairness of federated learning without increasing the client's computing burden and communication cost and directly sharing the original data. Summary of the invention

[0007] In order to overcome the above existing problems, a fair federated learning method for non-independent and identically distributed data is proposed. The technical solution of the present invention is as follows:

[0008] A fair federated learning method for non-independent and identically distributed data, characterized by comprising five parts: a system initialization module, a local module, an influence calculation module, an aggregation weight adjustment module, and a client selection module; wherein the system initialization module is used to complete the initialization of the federated learning system;

[0009] The local module is used to receive the global model and upload the local model on the client, and complete the local model training and other contents;

[0010] The influence calculation module is used to measure the influence of each client;

[0011] The aggregation weight adjustment module is used to adjust the aggregation weight of the client when participating in the aggregation according to the calculated client influence;

[0012] The client selection module is used to increase the opportunity for clients with higher influence to participate in federated learning according to the influence of the participating clients in each round.

[0013] Furthermore, the system initialization module is used to complete the initialization of the federated learning system, specifically including:

[0014] The server needs to hold the corresponding auxiliary dataset, which is used for influence calculation and can be smaller in size, but is not included in the client's data distribution and does not participate in the training process. This auxiliary dataset can be collected and held by the server normally, or customized according to specific scenarios.

[0015] In addition, global model initialization, parameter configuration, client registration, etc. remain consistent with the traditional federated averaging algorithm.

[0016] Furthermore, the local module is used to perform related tasks on the client, specifically including:

[0017] Receiving the global model from the server, training the local model, uploading the trained local model to the server, etc. are consistent with the traditional federated averaging algorithm and do not add additional computing or communication overhead.

[0018] Furthermore, the influence calculation module is used to measure the influence of each client, specifically including:

[0019] In a communication round, the server first calculates for each participating client. If the client l∈S t If a client is excluded from the global model aggregation process, the model performance obtained will be different from that obtained when all selected clients participate in the aggregation. P(w) is the performance evaluation of model w, which is used to measure the model's predictive ability for unknown samples: P(w) = -∑ x∈X ∑ i y i ·logq(x i ,w). At this time, all selected clients participate in the aggregation model The update is After excluding client l, the remaining selected clients participate in the aggregation model Updated to Then, the model needs to be calculated With model The difference in performance is taken as the influence of the excluded client l, denoted as The influence of the lth client in the tth communication round can be obtained as Furthermore, the aggregation weight adjustment module is used to adjust the aggregation weight of the client when participating in the aggregation according to the calculated client influence, specifically including:

[0020] Observe the changes in client influence. When client l has more unique data categories or larger data volumes than other clients, it usually has a greater impact on the update of the global model. Excluding client l will lead to significant changes in the performance of the global model.

[0021] like Right now This shows that the model The difference between the predicted results and the true value on the dataset X compared to the model is larger, then when client l does not participate in the model aggregation in round t, the resulting global model performs worse in terms of prediction accuracy. That is, in the global model aggregation process in round t, the participation of client l has a positive impact on improving the performance of the global model. And, The larger the value of , the greater the positive impact, thus allocating more aggregation weight to client l. On the contrary, if The smaller the value of , the more positive the impact of client l on the global model performance is, and may even have a negative impact. Therefore, less aggregation weight should be assigned to client l.

[0022] because The value of may be negative, which is not conducive to the subsequent weight calculation of client l when participating in the aggregation. Therefore, linear normalization is used to map it to the range of [c, 1+c], that is, The parameter c is used to prevent the training process from being dominated by individual clients, and to avoid the situation where the client with the least influence cannot participate in this round of aggregation because its influence becomes 0 after normalization. The value of is normalized to Furthermore, the client selection module is used to increase the chances of clients with higher influence to participate in federated learning according to the influence of the participating clients in each round, specifically including:

[0023] In each communication round, S t The clients in are sorted by their influence, and the α clients with the highest influence are selected to continue to participate in the next round of training. The α selected clients need to have a positive influence that is greater than the average influence of all the clients participating in the training in the current cycle. Clients with higher positive influence can participate in the next round of training, thereby promoting the improvement of global model fairness.

[0024] At the same time, in order to ensure that clients with lower influence do not lose the opportunity to participate in training, an upper limit needs to be set to limit the number of clients that continue to participate in the next round of training, that is, α needs to satisfy: γ∈[0,1), which is used to limit the maximum number of clients that can be selected to continue to participate in the next round of training.

[0025] A fair federated learning method for non-independent and identically distributed data includes five parts: system initialization module, local module, influence calculation module, aggregation weight adjustment module, and client selection module; among them,

[0026] The system initialization module is used to complete the initialization of the federated learning system and build an auxiliary data set;

[0027] The local module is used to receive the global model and upload the local model on the client, and complete local model training, etc.

[0028] The influence calculation module is used to compare the performance of the global model when the client does not participate in the aggregation in each communication round with the performance of the global model when it participates in the aggregation, so as to measure the influence of each client;

[0029] The aggregation weight adjustment module is used to adjust the aggregation weight of the client when participating in the aggregation according to the calculated client influence; the client selection module is used to increase the chances of clients with higher influence to participate in federated learning according to the influence of the participating clients in each round.

[0030] The advantages and beneficial effects of the present invention are as follows:

[0031] Compared with the prior art, the invention has the following three advantages:

[0032] (1) Design a client influence calculation mechanism based on marginal loss and cross entropy. Through this influence, an aggregation weight adjustment mechanism is established to ensure that clients with greater influence on the global model receive higher aggregation weights, thereby improving the accuracy of the global model on clients with poor performance.

[0033] (2) An influence-driven client selection strategy is proposed to increase the number of times clients with higher influence participate in federated training, ensuring that all types of data are fully learned and further improving the fairness of federated learning.

[0034] (3) Jain's Fairness Index is introduced to measure client fairness, and the proposed method is evaluated on multiple datasets. Experimental results show that the accuracy, variance and Jain's Fairness Index of the proposed method on clients with poor performance are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a system model diagram of the present invention;

[0036] Figure 2 The single client influence calculation process of the present invention;

[0037] Figure 3 It is a specific flow chart of a specific implementation mode of the present invention. DETAILED DESCRIPTION

[0038] The technical solution of the present invention will be described clearly and in detail below in combination with the system model diagram and specific embodiments of the present invention.

[0039] The technical solution of the present invention to solve the above technical problems is:

[0040] Reference Figure 1-Figure 3 This embodiment uses the CIFAR-10 dataset, which is a small dataset for identifying universal objects. It has 10 categories, and each image is a 32×32 3-channel color RGB image. It contains 50,000 training samples and 10,000 test samples. The specific implementation is as follows:

[0041] This embodiment is based on Python 3.7 programming language, combined with PyTorch 2.0.0 deep learning framework, and Anaconda as a package management and environment management tool, and implemented using PyCharm as an integrated development environment.

[0042] 1. In the initial case, the federated learning system is initialized. First, 1 / 5 of the data from each category in the test set of the CIFAR-10 dataset is randomly selected to form the auxiliary dataset X, and the remaining 4 / 5 of the data is retained as the test set. Subsequently, 100 clients are set, and 10 clients are selected in each round to participate in the training. Each client has only 3 categories of data sets, and the amount of data is randomly allocated. In addition, the necessary parameter settings are completed, and the global model and other configurations are initialized.

[0043] 2. Perform related tasks on the client. The client needs to download the global model sent from the server and use the local data set to train the local model. Then, upload the trained local model to the server. This process is consistent with the traditional federated average algorithm and does not increase the client's additional computing or communication overhead.

[0044] 3. Calculate the influence of the client based on marginal loss and cross entropy. In a communication round, the server first calculates for each participating client. If the client l∈S t If a client is excluded from the global model aggregation process, the model performance obtained will be different from that obtained when all selected clients participate in the aggregation, namely: P(w) is the performance evaluation of model w, which is used to measure the model's predictive ability for unknown samples, that is, P(w) = -∑ x∈X ∑ i y i ·logq(x i ,w). At this time, all selected clients participate in the aggregation model The update is After excluding client l, the remaining selected clients participate in the aggregation model Updated to

[0045] Then, the model needs to be calculated With model The difference in performance is taken as the influence of the excluded client l, denoted as The influence of the lth client in the tth communication round can be obtained as

[0046] 4. Observe the changes in client influence. When client l has more unique data categories or a larger amount of data than other clients, it usually has a greater impact on the update of the global model. Excluding client l will lead to a significant change in the performance of the global model.

[0047] like Right now This shows that the model The difference between the predicted results and the true value on the dataset X compared to the model is larger, then when client l does not participate in the model aggregation in round t, the resulting global model performs worse in terms of prediction accuracy. That is, in the global model aggregation process in round t, the participation of client l has a positive impact on improving the performance of the global model. And, The larger the value of , the greater the positive impact, thus allocating more aggregation weight to client l. On the contrary, if The smaller the value of , the more positive the impact of client l on the global model performance is, and may even have a negative impact. Therefore, less aggregation weight should be assigned to client l.

[0048] because The value of may be negative, which is not conducive to the subsequent weight calculation of client l when participating in the aggregation. Therefore, linear normalization is used to map it to the range of [c, 1+c], that is, The parameter c is set to 0.5 in this embodiment to prevent the training process from being dominated by individual clients and to avoid the situation where the client with the least influence cannot participate in this round of aggregation because its influence becomes 0 after normalization. The value of is normalized to

[0049] 5. In each communication round, S t The clients in are sorted by their influence, and the α clients with the highest influence are selected to continue to participate in the next round of training. The α selected clients need to have a positive influence that is greater than the average influence of all the clients participating in the training in the current cycle. Clients with higher positive influence can participate in the next round of training, thereby promoting the improvement of global model fairness.

[0050] At the same time, in order to ensure that clients with lower influence do not lose the opportunity to participate in training, an upper limit needs to be set to limit the number of clients that continue to participate in the next round of training, that is, α needs to satisfy: In the first round, a predefined number of clients are directly randomly selected to participate in the training, and the client selection strategy is enabled in subsequent rounds. In this embodiment, γ=0.5, which means that a maximum of 5 clients can be selected in each round to continue to participate in the next round of training.

[0051] 6. In order to measure fairness, Jain's Fairness Index is introduced, that is, It satisfies the independence of the number of users, that is, the index is applicable to any number of users; it satisfies the independence of scale and measurement, that is, the index does not depend on the scale of the allocation; it satisfies the boundedness, so that fairness can be expressed as a percentage; it satisfies the continuity, that is, the index should be continuous. The range of J is The larger the value of J, the fairer the resource allocation. This embodiment combines the variance and Jain's Fairness Index to measure fairness.

[0052] In the implementation example, the Jain's Fairness Index value of the present invention is 0.9785, and the variance is 63.73; while the Jain's Fairness Index value of the traditional federated average algorithm is 0.9278, and the variance is 202.53. In comparison, the improvement of the present invention is more obvious

[0053] The above embodiments should be understood as being only used to illustrate the technical solution of the present invention and not to limit the data set and parameter range to which the present invention is applicable. After reading the contents of the present invention, technicians can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A fair federated learning method for non-independent and identically distributed data, characterized in that: include: The system consists of five parts: system initialization module, local module, influence calculation module, aggregation weight adjustment module, and client selection module. The system initialization module is used to complete the initialization of the federated learning system. The local module is used to receive the global model and upload the local model on the client, and complete local model training, etc. The influence calculation module is used to measure the influence of each client. The aggregation weight adjustment module is used to adjust the aggregation weight of the client when participating in aggregation according to the calculated client influence. The client selection module is used to increase the opportunity for clients with higher influence to participate in federated learning according to the influence of the participating clients in each round.

2. A fair federated learning method for non-independent and identically distributed data according to claim 1, characterized in that: In the system initialization module, the server side needs to hold a corresponding auxiliary data set, which is used for influence calculation.

3. The auxiliary data set according to claim 2, characterized in that: The auxiliary data set can be normally collected and held by the server, or customized according to a specific scenario, and has a smaller scale, but is not included in the data distribution of the client and does not participate in the training process.

4. A fair federated learning method for non-independent and identically distributed data according to claim 1, characterized in that: The influence calculation module is used to measure the influence of each client, specifically including: Combining marginal loss and cross entropy, the performance of the global model of a single client in each communication round when it does not participate in aggregation is compared with the performance of the global model when it participates in aggregation to measure the influence of the client. P(w) is the performance evaluation of model w, which is used to measure the model's predictive ability for unknown samples. P(w) = -∑ x∈x ∑ i y i ·logq(x i ,w). For all selected clients participating in the aggregation model And after excluding client l, the model obtained by the remaining selected clients participating in the aggregation The influence of the lth client in the tth communication round can be obtained as 5. A fair federated learning method for non-independent and identically distributed data according to claim 1, characterized in that: The aggregation weight adjustment module is used to adjust the aggregation weight of the client when participating in the aggregation according to the calculated client influence, specifically including: like The larger the value of , the more aggregation weight is assigned to client l. On the contrary, if The smaller the value of is, the less aggregation weight is assigned to client l. Then, linear normalization is used to map it to the range of [c, 1+c], that is, Finally, The value of is normalized to 6. A fair federated learning method for non-independent and identically distributed data according to claim 1, characterized in that: The client selection module is used to increase the chances of clients with higher influence to participate in federated learning according to the influence of the participating clients in each round, specifically including: In each communication round, S t The clients in are sorted according to their influence, and the α clients with the highest influence are selected to continue to participate in the next round of training. α satisfies At the same time, α also needs to satisfy γ∈[0,1), which is used to limit the maximum number of clients that can be selected to continue to participate in the next round of training.

7. A fair federated learning method for non-independent and identically distributed data according to claim 1, characterized in that: include: There are five parts in total: system initialization module, local module, influence calculation module, aggregation weight adjustment module, and client selection module; among them, The system initialization module is used to complete the initialization of the federated learning system and build an auxiliary data set; The local module is used to receive the global model and upload the local model on the client, and complete local model training, etc. The influence calculation module is used to compare the performance of the global model when the client does not participate in the aggregation in each communication round with the performance of the global model when it participates in the aggregation, so as to measure the influence of each client; The aggregation weight adjustment module is used to adjust the aggregation weight of the client when participating in the aggregation according to the calculated client influence; The client selection module is used to increase the chances of clients with higher influence to participate in federated learning based on the influence of the participating clients in each round.