Privacy protection method, device and storage medium for federated learning with heterogeneous devices
By setting training delay and data similarity constraints in federated learning, and using multi-arm robbers and Pareto algorithms for server clustering and model aggregation, the problem of excessive model training delay caused by device heterogeneity is solved, and the training efficiency is improved.
Patent Information
- Application Number
- CN202510163949.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The existing federated learning privacy protection method has a longer training delay and low training efficiency when the device is heterogeneous.
By setting training delay constraints and data similarity constraints, the multi-arm robber algorithm is used to predict the computing delay of the server, cluster the server based on the Pareto algorithm, and select cluster head servers for model aggregation and communication.
The communication consumption of federated learning is reduced, the training efficiency of local models and the communication efficiency of model parameter exchange is improved, thereby improving the training efficiency of global models.
Smart Images

Figure CN119646884B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of information security technology, and in particular, relates to a method, device and storage medium for protecting the privacy of federated learning for heterogeneous devices. Background Art
[0002] With the rapid development of artificial intelligence technology, data privacy and security have become an important issue. In an environment where data silos conflict with the need for data fusion, federated learning has emerged as a new type of distributed machine learning technology, which effectively avoids direct data transmission and greatly reduces the risk of data leakage.
[0003] In the existing federated learning privacy protection method, each participant uses local data to train a local model, obtains different model parameters, and uploads the obtained model parameters to the cloud at the same time. The cloud completes the aggregation and update of the model parameters, and returns the updated parameters to the terminals of each participant. The terminals of each participant start the next iteration and repeat the above steps until the entire training process converges. However, this method can only be implemented in the ideal situation where the computing performance and communication environment of all participants are the same. It does not take into account the situation of heterogeneous devices. When the computing power and communication capabilities of each participant vary greatly, it will lead to longer model training delays and lower training efficiency.
[0004] Therefore, how to improve the efficiency of model training in federated learning has become an urgent problem to be solved. Summary of the invention
[0005] The embodiments of the present application provide a federated learning privacy protection method, device, and storage medium for device heterogeneity, aiming to improve the efficiency of model training in federated learning.
[0006] In a first aspect, an embodiment of the present application provides a method for privacy protection of federated learning for device heterogeneity, the method comprising: setting training delay constraints and data similarity constraints; based on a multi-armed bandit algorithm, using the historical training time of each server for the local model to predict the current training time, and obtaining the calculation delay of each server; based on a Pareto algorithm, using the training delay constraints, the data similarity constraints and the calculation delay of each server, clustering each server to obtain at least one cluster, each cluster including at least two servers; determining a cluster head server in each cluster based on the training delay of each server, so as to aggregate the local models of the servers in each cluster through each cluster head server, and communicate with a central server, the training delay including the calculation delay and the communication delay.
[0007] In a possible implementation, the Pareto algorithm is based on the training delay constraint, the data similarity constraint and the calculation delay of each server to cluster each server to obtain at least one cluster, each cluster including at least two servers, including: calculating the data similarity between each server to obtain a data similarity calculation result; based on the calculation delay of each server, calculating the training delay of each server to obtain a training delay calculation result; based on the Pareto algorithm, the training delay constraint, the data similarity constraint, the data similarity calculation result and the training delay calculation result are used to cluster each server.
[0008] In a possible implementation, the calculating the data similarity between the servers to obtain the data similarity calculation result includes: calculating the gradient of the non-sensitive data of each server in the pre-training; calculating the cosine similarity between the gradients of the non-sensitive data of each server; and determining the data similarity calculation result based on the cosine similarity between the gradients of the non-sensitive data of each server.
[0009] In a possible implementation, the calculating of the training delay of each server based on the calculation delay of each server to obtain the training delay calculation result includes: calculating the communication delay of each server using a first calculation formula based on Shannon's theorem; calculating the training delay of each server based on the calculation delay and communication delay of each server; determining the training delay calculation result based on the training delay of each server; the first calculation formula is:
[0010] ,
[0011] Among them, T x,r is the communication delay of the target server, Z x is the network bandwidth of the target server, p x is the channel transmission power, B x is the channel gain between the target server and the central server, v is the white noise power, and the target server is any of the servers.
[0012] In a possible implementation, the clustering of each of the servers is performed based on the Pareto algorithm using the training delay constraint, the data similarity constraint, the data similarity calculation result, and the training delay calculation result, including: constructing an objective function based on the Pareto algorithm using the data similarity calculation result and the training delay calculation result; solving the objective function based on the Karush-Kuhn-Tucker condition using the training delay constraint and the data similarity constraint to achieve clustering of each of the servers; the objective function is:
[0013] ,
[0014] Wherein, T is the training delay calculation result, is the data similarity calculation result.
[0015] In a possible implementation, the method of predicting the current training time based on the multi-armed bandit algorithm by using the historical training time of each server for the local model to obtain the calculation delay of each server includes: predicting the current training time based on the multi-armed bandit algorithm by using the historical training time of each server for the local model and using a second calculation formula to obtain the calculation delay of each server; the second calculation formula is:
[0016] ,
[0017] in, Indicates the current status of the target server. represents the current computing performance of the target server, Indicates whether the target server has participated in the previous training task, which is 1 if it is and 0 if it is not; represents the historical training information of the target server, represents the historical training time of the local model by the target server, Indicates the preparation time for the target server to participate in this training. Indicates the computing delay of the target server, where the target server is any of the servers.
[0018] In a possible implementation, the training delay constraint includes that the total training delay of all servers in each of the clusters is equal to the minimum training delay among servers outside the cluster; and the data similarity constraint includes maximizing the data similarity between the servers in each of the clusters.
[0019] In a possible implementation manner, determining the cluster head server in each of the clusters based on the training delay of each of the servers includes: determining a server with the smallest training delay in each of the clusters as the cluster head server in each of the clusters.
[0020] In a second aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method described in the first aspect or any one of the implementation methods thereof is implemented.
[0021] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect or any one of the implementation methods thereof is implemented.
[0022] In a fourth aspect, an embodiment of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect or any one of the implementation methods thereof.
[0023] Compared with the prior art, the embodiments of the present application have the following beneficial effects: based on the multi-armed bandit algorithm, the historical training time of the local model of each server is used to predict the current training time to obtain the computing delay of each server; based on the Pareto algorithm, each server is clustered using the set training delay constraint and data similarity constraint and the predicted computing delay of each server; based on the training delay of each server including the computing delay and the communication delay, the cluster head server in each cluster is determined, so that the local models of the servers in each cluster are aggregated through each cluster head server, and the local models are communicated with the central server, the computing performance, communication environment and data similarity of each server are considered, and the idea of clustering is used to combine servers with weaker performance and longer communication delay to perform local training on the model, and only the cluster head server exchanges model parameters with the central server, which avoids each server from communicating directly with the central server, reduces the communication consumption of federated learning, improves the training efficiency of the local model and the communication efficiency of the model parameter exchange, and thus improves the training efficiency of the global model in federated learning.
[0024] It can be understood that the electronic device, computer-readable storage medium, and computer program product provided in the embodiments of the present application have the same beneficial effects as the above-mentioned device heterogeneous federated learning privacy protection method, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0026] Figure 1 A flowchart of a method for protecting privacy in federated learning for heterogeneous devices provided in one embodiment of the present application;
[0027] Figure 2 A schematic diagram of the architecture of a federated learning privacy protection system for heterogeneous devices provided in one embodiment of the present application;
[0028] Figure 3 A flowchart of another method for protecting privacy in federated learning for heterogeneous devices provided in one embodiment of the present application;
[0029] Figure 4 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0030] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0031] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0032] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0033] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0034] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0035] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0036] To facilitate understanding, some concepts involved in the embodiments of the present application are first explained.
[0037] (1) Federated Learning
[0038] Federated learning is a distributed learning framework involving multiple parties, which can effectively help multiple institutions or enterprises use data and conduct machine learning modeling while meeting the requirements of user sensitive data protection and privacy security. Federated learning requires all participants to train a global model without exchanging data samples, that is, each edge device updates the local model through local data and sends the relevant information and parameters of the generated model back to the global server; then, the global server updates the shared model by aggregating the received weights.
[0039] (2) MAB technology
[0040] The multi-armed bandit (MAB) technique is a sequential decision-making technique where the decision maker has multiple arms to choose from, each with an unknown reward distribution. The goal of the decision maker is to choose the arm in each round to obtain the maximum expected reward. If each arm is represented by an observable feature vector, the problem is called a contextual bandit problem. The MAB problem can be represented as a tuple . Where K is the number of rounds of participation, is the set of all arms in round k, and are the context vector and the coefficient vector respectively. In machine learning, minimizing the loss function generally indicates model convergence. Therefore, the problem of maximizing the benefit can be transformed into the problem of minimizing the loss, that is, in each round of task k∈K, a set of feature vectors, i.e., a set of arms (usually called super arms), is selected to extract the loss from each selected arm. Then the target is to select S each time. k To minimize losses .
[0041] (3) Pareto optimality
[0042] Pareto optimality, also known as Pareto efficiency, is an ideal state of rationally allocating existing resources under two or more constraints. The Pareto model is often used to solve dual-objective or multi-objective optimization problems. In dual-objective optimization, assuming that there are two objective functions, for solution A, if no other solution can be found in the entire variable space that is better than solution A, then solution A is the Pareto optimal solution. In practice, there is generally not only one optimal solution, but multiple solutions. All Pareto optimal solutions constitute the Pareto optimal solution set, and these solutions constitute the Pareto optimal frontier or Pareto frontier surface of the problem through the mapping of the objective function. For dual-objective problems, the Pareto optimal frontier is usually a line. For multi-objective problems, the Pareto optimal frontier is usually a hypersurface. Generally speaking, decision makers make decisions in the Pareto optimal solution set according to different situations.
[0043] The technical solution of the present application will be described in detail below with reference to the accompanying drawings.
[0044] Figure 1 A flowchart of a device heterogeneous federated learning privacy protection method provided in an embodiment of the present application is provided. For the sake of convenience, only the part related to the present embodiment is shown, which specifically includes the following steps:
[0045] S110, setting training delay constraints and data similarity constraints.
[0046] In a possible implementation, the training delay constraint includes that the total training delay of all servers in each cluster is equal to the minimum training delay of the servers outside the cluster; the data similarity constraint includes maximizing the data similarity between the servers in each cluster.
[0047] In the specific implementation, since the performance of the server and the communication environment will affect the model generation time, when each server performs a separate model training, the delay of the model generation with different performance is also different. The weaker the computing performance, the longer the delay. At the same time, due to the large amount of data in the model, its transmission delay is also not negligible. Too long communication delay will also lead to a longer time for the central server to aggregate the model. Therefore, by using the clustering method, the servers with poor performance and poor communication environment are clustered together, which can generate models with a smaller delay than the server with the strongest computing performance. The cluster head server transmits the model to the servers in the cluster, which reduces the delay of the central server transmitting the model to each server, thereby reducing the time for model generation and aggregation.
[0048] As an example, when clustering, the training delay constraint is that the number of servers in each cluster is d = {d 1 ,d 2 ,...,d xThe total training delay of} is equal to the smallest training delay among the servers outside the cluster, and the training delay includes the computation delay and the communication delay.
[0049] Exemplarily, the expression of the training delay constraint is:
[0050] ,
[0051] Among them, T x,b represents the computation delay of each server in the cluster, T x,r represents the communication delay of each server in the cluster, Indicates the computational delay of the out-of-cluster server with the smallest training delay. Indicates the communication delay corresponding to the out-of-cluster server with the smallest training delay.
[0052] In the specific implementation, since the data similarity between the servers in the cluster will affect the quality of the aggregate model, the higher the data similarity between the servers in the cluster, the better the aggregate model. Therefore, data similarity constraints are proposed to ensure the quality of the aggregate model in the cluster.
[0053] As an example, calculate the d of each server in each cluster in the pre-training x The non-sensitive data gradients of the servers are obtained and the cosine similarity between the gradients of the non-sensitive data of each server is calculated and used as the data similarity of the servers in each cluster.
[0054] Exemplarily, the calculation formula of cosine similarity is as follows:
[0055]
[0056] in, represents the cosine similarity between the gradients of non-sensitive data of any two servers in the cluster, Represents the gradient of non-sensitive data.
[0057] S120, based on the multi-armed bandit algorithm, the historical training time of each server for the local model is used to predict the current training time, and the calculation delay of each server is obtained.
[0058] In a possible implementation, based on the multi-armed bandit algorithm, the historical training time of each server for the local model is used to predict the current training time using the second calculation formula to obtain the calculation delay of each server; wherein the second calculation formula is:
[0059] ,
[0060] in, Indicates the current status of the target server. Indicates the current computing performance of the target server. Indicates whether the target server has participated in the previous training task, which is 1 if it is and 0 if it is not; Represents the historical training information of the target server, Indicates the historical training time of the local model by the target server. Indicates the preparation time for the target server to participate in this training. Indicates the computing latency of the target server, which is any server participating in federated learning.
[0061] In the specific implementation, due to the computing performance constraints of each server, the training time of the local model is different. At the same time, it is affected by factors such as the communication environment and system stability. The training time cannot be accurately calculated before model training. Therefore, based on MAB technology, the historical training time is used to predict the current training time, and the prediction result is used as the computing delay of each server.
[0062] As an example, suppose server d x In one round of training, the predicted time to complete task k can be expressed as the second calculation formula. In the second calculation formula, is a fixed value that is calculated based on the server's historical information. Indicates server d x The historical performance before computing task k (i.e., the previous context information and the computing time of model training) is expressed as follows:
[0063]
[0064] in, Indicates server d x The context information of the local model training before computing task k, Indicates server d x The computation time when participating in local model training for the pth time is replace You can get:
[0065] ,
[0066] make , , You can get:
[0067]
[0068] Introducing exploration parameters , the confidence interval upper bound algorithm (Upper Confidence Bound, UCB) is used to solve the above formula. The calculation formula is as follows:
[0069] .
[0070] Exemplarily, the server d can be estimated through the above steps. x The training time of the local model After calculating the training time of all servers for the local model After that, Sort from high to low, take the server with the smallest computing delay as the benchmark, form a cluster with servers with higher computing delay, use MAB technology to predict the computing delay of the cluster, and perform continuous combination calculations until it meets the training delay constraint and data similarity constraint, until it is divided into n clusters. Finally, add cluster labels to the servers of the federated learning task, and the central server distinguishes servers of different clusters through cluster labels. The number of clusters n can be customized according to actual conditions, and this application does not limit this.
[0071] Furthermore, in order to obtain the optimal clustering combination under the constraints of training delay and data similarity, the Pareto algorithm is used to model and calculate the clustering.
[0072] S130, based on the Pareto algorithm, using the training delay constraint, the data similarity constraint and the computing delay of each server, clustering each server to obtain at least one cluster, each cluster including at least two servers.
[0073] In a possible implementation, the data similarity between the servers is calculated to obtain the data similarity calculation result; based on the calculation delay of each server, the training delay of each server is calculated to obtain the training delay calculation result; based on the Pareto algorithm, the servers are clustered using the training delay constraint, the data similarity constraint, the data similarity calculation result and the training delay calculation result.
[0074] As an example, the gradient of non-sensitive data of each server in pre-training is calculated; the cosine similarity between the gradients of non-sensitive data of each server is calculated; and the data similarity calculation result is determined based on the cosine similarity between the gradients of non-sensitive data of each server.
[0075] Exemplarily, the above-mentioned cosine similarity calculation formula is used to calculate the cosine similarity between the gradients of the non-sensitive data of each server, and the cosine similarity between the gradients of the non-sensitive data of each server is used as the data similarity between the servers to obtain the data similarity calculation result.
[0076] As an example, based on Shannon's theorem, the communication delay of each server is calculated using the first calculation formula; based on the calculation delay and communication delay of each server, the training delay of each server is calculated; based on the training delay of each server, the training delay calculation result is determined; wherein, the first calculation formula is:
[0077] ,
[0078] Among them, T x,r is the communication delay of the target server, Z x is the network bandwidth of the target server, p x is the channel transmission power, B x is the channel gain between the target server and the central server, v is the white noise power, and the target server is any server participating in federated learning.
[0079] Exemplarily, as shown in the training delay calculation formula, the training delay T is calculated by the calculation delay T x,b With the communication delay T x,r The delay T is calculated x,b In step S120, the communication delay T is calculated. x,r Based on Shannon's theorem, the first calculation formula is used to calculate the training delay. The formula for calculating the training delay is as follows:
[0080] .
[0081] As an example, based on the Pareto algorithm, the objective function is constructed using the data similarity calculation results and the training delay calculation results; based on the Karush-Kuhn-Tucker condition, the objective function is solved using the training delay constraint and the data similarity constraint to achieve clustering of each server; wherein the objective function is:
[0082] ,
[0083] Among them, T is the calculation result of training delay, The result of data similarity calculation.
[0084] For example, in the clustering process, it is hoped that the training delay is minimized and the data similarity is maximized, that is, the cosine similarity is maximized, so an objective function can be established, namely:
[0085] ,
[0086] Then, based on the Karush-Kuhn-Tucher (KKT) condition to minimize the accumulated loss function, the optimization problem is transformed and solved to achieve clustering of each server.
[0087] S140, determining the cluster head server in each cluster based on the training delay of each server, so as to aggregate the local models of the servers in each cluster through each cluster head server and communicate with the central server, wherein the training delay includes calculation delay and communication delay.
[0088] In a possible implementation manner, a server with the smallest training delay in each cluster is determined as a cluster head server in each cluster.
[0089] As an example, the selection of cluster head servers is to select the cluster head servers of each cluster during the pre-training process to undertake the aggregation of local models within the cluster. Therefore, each cluster head server needs to have a certain computing power. Usually, the server with the strongest performance in each cluster is used as the cluster head server. For example, the server with the smallest training delay in each cluster is used as the cluster head server. Since the selection of servers within the cluster is based on the data characteristics of the cluster head servers, the cluster head servers will also have certain rewards when aggregating local models within the cluster.
[0090] Exemplarily, the expression of the cluster head server is as follows:
[0091] ,
[0092] Where E represents the performance of the server.
[0093] As an example, after completing clustering and selecting cluster head servers in each cluster, the federated learning privacy protection system for heterogeneous devices is as follows: Figure 2 As shown, it includes a central server, three high-performance servers and multiple clusters, each of which includes a cluster head server and multiple low-performance servers; among them, the three high-performance servers each train a local model through local data, and send the relevant information and parameters of the generated model to the central server; multiple servers in each cluster jointly train the local model, the cluster head server aggregates the local models trained by the servers in the cluster, and sends the relevant information and parameters of the models trained in the cluster to the central server, and the central server trains the global model based on the received information and parameters.
[0094] The technical solution provided in this embodiment is based on the multi-armed bandit algorithm, and uses the historical training time of each server for the local model to predict the current training time to obtain the calculation delay of each server; based on the Pareto algorithm, the servers are clustered by using the set training delay constraint and data similarity constraint and the predicted calculation delay of each server; the cluster head server in each cluster is determined based on the training delay of each server including the calculation delay and the communication delay, so as to aggregate the local models of the servers in each cluster through each cluster head server and communicate with the central server, taking into account the computing performance, communication environment and data similarity of each server, and using the idea of clustering, the servers with weaker performance and longer communication delay are combined to perform local training on the model, and only the cluster head server exchanges model parameters with the central server, which avoids each server from communicating directly with the central server, reduces the communication consumption of federated learning, improves the training efficiency of the local model and the communication efficiency of the model parameter exchange, and thus improves the training efficiency of the global model in federated learning.
[0095] Figure 3 A flowchart of another method for protecting privacy in heterogeneous federated learning provided in an embodiment of the present application is provided. For ease of explanation, only the part related to the present embodiment is shown. The method provided in the present embodiment specifically includes the following steps:
[0096] Step 1: Definition of training delay constraints and data similarity constraints
[0097] (1) Definition of training delay constraint
[0098] By clustering servers with poor performance and poor communication environment, the time difference between the model generation delay and the server with the strongest computing performance can be reduced. The cluster head server transmits the model to the servers in the cluster, which reduces the time delay for the central server to transmit the model to each server, thereby reducing the time for model generation and aggregation. When clustering, the training delay constraint is defined as d = {d 1 ,d 2 ,...,d x The total training delay of the cluster is equal to the minimum training delay of the servers outside the cluster. The expression of the training delay constraint is:
[0099] ,
[0100] Among them, T x,b represents the computation delay of each server in the cluster, T x,r represents the communication delay of each server in the cluster, Indicates the computational delay of the out-of-cluster server with the smallest training delay. Indicates the communication delay corresponding to the out-of-cluster server with the smallest training delay.
[0101] (2) Definition of data similarity constraints
[0102] Since the data similarity in the server will affect the quality of the aggregate model, the higher the data similarity of each server, the better the aggregate model. Therefore, the data similarity constraint is defined to maximize the data similarity between the servers in each cluster to ensure the quality of the aggregate model within the cluster.
[0103] Specifically, calculate the d of each server in each cluster in the pre-training x The non-sensitive data gradients are calculated, and the cosine similarity between the non-sensitive data gradients of each server is calculated, which is used as the data similarity of the servers in each cluster. The calculation formula of cosine similarity is as follows:
[0104] ,
[0105] in, represents the cosine similarity between the gradients of non-sensitive data of any two servers in the cluster, Represents the gradient of non-sensitive data.
[0106] Step 2: Select cluster head server
[0107] The selection of cluster head servers is to select the cluster head servers of each cluster during the pre-training process, which is responsible for the aggregation of local models in the cluster. Therefore, each cluster head server needs to have a certain computing power. Usually, the server with the strongest performance in each cluster is used as the cluster head server. For example, the server with the smallest training delay in each cluster is used as the cluster head server. Since the selection of servers in the cluster is based on the data characteristics of the cluster head server, the cluster head server will also have a certain reward when aggregating local models in the cluster. Specifically, the expression of the cluster head server is as follows:
[0108] ,
[0109] Where E represents the performance of the server.
[0110] Step 3: Calculation delay prediction based on MAB technology
[0111] In order to select servers within a cluster under the constraints of training delay and data similarity, servers are clustered based on MAB. The MAB algorithm is used to predict the computing delay of each server, thereby realizing server clustering.
[0112] Since the computing performance of each server is different, the training time of the local model is different. At the same time, it is affected by factors such as the communication environment and system stability. Therefore, it is impossible to accurately calculate the training time before model training. Therefore, based on MAB technology, the historical training time is used to predict the current training time, and clustering is performed according to the computing delay.
[0113] Assume that server d x In one round of training, the predicted time to complete task k can be expressed as follows:
[0114] ,
[0115] in, Indicates the captured server d x The current status of Indicates server d x Current computing performance, Indicates server d x Whether the participant has participated in the previous training task, which is 1 if yes and 0 if no; Indicates server d x Historical training information, Indicates server d x The historical training time for local model training, Indicates server d x Preparation time for participating in this training, Indicates server d x The computation delay.
[0116] in, Is a fixed value based on the server d x This value is calculated using historical information from Indicates server d x The historical performance before computing task k (i.e., the previous context information and the computing time of model training) is expressed as follows:
[0117]
[0118] in, Indicates server d x The context information of the local model training before computing task k, Indicates server d x The computation time when participating in local model training for the pth time is replace You can get:
[0119] ,
[0120] make , , You can get:
[0121] ,
[0122] Introducing exploration parameters , the UCB algorithm is used to solve the above equation, and the calculation formula is as follows:
[0123] .
[0124] Through the above steps, we can estimate the server d x Prediction training time for local models . After calculating the prediction training time of the local models of all servers After that, Sort from high to low, take the server with the smallest computing delay as the benchmark, form a cluster with servers with higher computing delay, use MAB technology to predict the computing delay of the cluster, and perform continuous combination calculations until it meets the training delay constraint and data similarity constraint, until it is divided into n clusters. Finally, add cluster labels to the servers of the federated learning task, and the central server distinguishes servers in different clusters through cluster labels. In order to obtain the optimal clustering combination under the constraints of training delay and data similarity, the Pareto algorithm is further used to model and calculate the clustering.
[0125] Step 4: Run independent tests to ensure that system functionality and performance meet expectations
[0126] Step 5: Clustering based on Pareto algorithm
[0127] (1) Data similarity calculation
[0128] First, calculate each server d in pre-training x The non-sensitive data gradients are calculated, and the cosine similarity between the gradients is calculated, and the cosine similarity between the gradients is used as the data similarity between the servers. The calculation formula of cosine similarity is as follows:
[0129] ,
[0130] in, represents the cosine similarity between the gradients of non-sensitive data of any two servers, Represents the gradient of non-sensitive data.
[0131] (2) Training delay calculation
[0132] The training delay T is calculated by the calculation delay T x,b With the communication delay T x,rThe delay T is calculated x,b It has been calculated in step 3, so the communication delay T is calculated based on Shannon's theorem x,r , the calculation formula is as follows:
[0133] ,
[0134] Among them, Z x For server d x The network bandwidth, p x is the channel transmission power, B x For server d x The channel gain between the central server and v is the white noise power.
[0135] Exemplarily, the calculation formula for training delay is as follows:
[0136]
[0137] (3) Problem transformation
[0138] In the clustering process, we hope to minimize the training delay and maximize the data similarity, that is, maximize the cosine similarity. Therefore, we can establish an objective function, namely:
[0139] ,
[0140] Then, based on the KKT condition, the accumulated loss function is minimized to transform and solve the optimization problem.
[0141] In summary, the privacy protection method for heterogeneous federated learning proposed in this application mainly includes the following key innovations:
[0142] (1) To address the problem of high communication overhead caused by model transmission during federated learning, the servers participating in federated learning are clustered. Servers with poor performance and communication environment are grouped together to train the model together. The server with the lowest training latency is selected as the cluster head server in the cluster, which is responsible for aggregating the local models of other servers in the cluster and submitting them to the central server, thereby reducing the communication overhead between the central server and all servers, thereby reducing the communication consumption of federated learning without affecting the accuracy of the model.
[0143] (2) In view of the fact that the current federated learning privacy protection scheme does not take into account the long model training delay caused by device heterogeneity, the idea of clustering is used to improve the efficiency of model training while ensuring the accuracy of model training. The training delay constraint and data similarity constraint are modeled, and servers with long training delay and large data volume similarity are clustered. In the clustering process, the MAB technology and the historical status information of each participating server are used to calculate the computing delay of the current server; then the multi-server clustering problem is transformed into a joint optimization problem of training delay and data similarity, and the Pareto algorithm is used to solve it, so as to obtain the optimal clustering combination.
[0144] Based on the above key innovations, the technical solution provided by this application has the following beneficial effects:
[0145] (1) By clustering servers with poor communication environments or low performance and selecting the server with the lowest training latency as the cluster head server, each server is prevented from communicating directly with the central server. This centralized model aggregation significantly reduces the amount of data transmission and provides a simplified communication architecture that facilitates supervision and maintenance of resource loads on different servers.
[0146] (2) By grouping devices with longer training delays and devices with high data similarity into one cluster, the training delay problem caused by device heterogeneity is reduced. With the help of MAB technology and the historical status information of the server, the computing delay of the server is dynamically calculated, which improves the real-time and flexibility of model training.
[0147] Figure 4 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 4 As shown, the electronic device 4 of this embodiment includes: at least one processor 40 ( Figure 4 Only one is shown in the figure), a memory 41, and a computer program 42 stored in the memory 41 and executable on at least one processor 40, the processor 40 executes the computer program 42 to implement the above Figure 1 or Figure 3 Steps in a method embodiment.
[0148] The electronic device 4 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device 4 may include but is not limited to a processor 40 and a memory 41. Those skilled in the art will appreciate that Figure 4 It is only an example of the electronic device 4 and does not constitute a limitation on the electronic device 4. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0149] The processor 40 may be a Central Processing Unit (CPU), or the processor 40 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0150] In some embodiments, the memory 41 may be an internal storage unit of the electronic device 4, such as the hard disk or memory of the electronic device 4. In other embodiments, the memory 41 may also be an external storage device of the electronic device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 4. Further, the memory 41 may also include both the internal storage unit and the external storage device of the electronic device 4. The memory 41 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of a computer program. The memory 41 may also be used to temporarily store data that has been output or is to be output.
[0151] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above various method embodiments can be implemented.
[0152] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk.
[0153] A computer-readable storage medium provided in an embodiment of the present application has the same beneficial effects as the above-mentioned device heterogeneous federated learning privacy protection method.
[0154] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0155] A computer program product provided in an embodiment of the present application has the same beneficial effects as the above-mentioned device heterogeneous federated learning privacy protection method.
[0156] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0157] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0158] In the embodiments provided in the present application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are merely schematic, for example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0159] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0160] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A privacy protection method for federated learning with heterogeneous devices, characterized in that: The method comprises: Set training delay constraints and data similarity constraints; Based on the multi-armed bandit algorithm, the historical training time of each server for the local model is used to predict the current training time, and the calculation delay of each server is obtained; Based on the Pareto algorithm, the servers are clustered using the training delay constraint, the data similarity constraint and the computing delay of each server to obtain at least one cluster, each of which includes at least two servers; Determine a cluster head server in each of the clusters based on a training delay of each of the servers, so as to aggregate local models of servers in each of the clusters through each of the cluster head servers and communicate with a central server, wherein the training delay includes the calculation delay and the communication delay; The training delay constraint includes that the total training delay of all servers in each cluster is equal to the minimum training delay of servers outside the cluster; and the data similarity constraint includes maximizing the data similarity between the servers in each cluster.
2. The method according to claim 1, characterized in that The Pareto algorithm is based on the training delay constraint, the data similarity constraint and the calculation delay of each server, and each server is clustered to obtain at least one cluster, each of which includes at least two servers, including: Calculating the data similarity between the servers to obtain a data similarity calculation result; Based on the calculation delay of each of the servers, the training delay of each of the servers is calculated to obtain a training delay calculation result; Based on the Pareto algorithm, the servers are clustered using the training delay constraint, the data similarity constraint, the data similarity calculation result and the training delay calculation result.
3. The method according to claim 2, characterized in that The calculating the data similarity between the servers to obtain the data similarity calculation result includes: Calculate the gradient of the non-sensitive data of each server in the pre-training; Calculating the cosine similarity between the gradients of the non-sensitive data of each of the servers; The data similarity calculation result is determined based on the cosine similarity between the gradients of the non-sensitive data of each server.
4. The method according to claim 2, characterized in that: The step of calculating the training delay of each server based on the calculation delay of each server to obtain the training delay calculation result includes: Based on Shannon's theorem, the communication delay of each server is calculated using the first calculation formula; Calculating the training delay of each server based on the computing delay and the communication delay of each server; Determining the training delay calculation result based on the training delay of each of the servers; The first calculation formula is: , Among them, T x,r is the communication delay of the target server, Z x is the network bandwidth of the target server, p x is the channel transmission power, B x is the channel gain between the target server and the central server, v is the white noise power, and the target server is any of the servers.
5. The method according to claim 2, characterized in that: The clustering of the servers based on the Pareto algorithm using the training delay constraint, the data similarity constraint, the data similarity calculation result and the training delay calculation result includes: Based on the Pareto algorithm, construct an objective function using the data similarity calculation result and the training delay calculation result; Based on the Karush-Kuhn-Tucker condition, the objective function is solved by using the training delay constraint and the data similarity constraint to achieve clustering of the servers; The objective function is: , Wherein, T is the training delay calculation result, is the data similarity calculation result.
6. The method according to claim 1, characterized in that The multi-armed bandit algorithm is based on the historical training time of each server for the local model to predict the current training time, and obtain the calculation delay of each server, including: Based on the multi-armed bandit algorithm, the historical training time of each server for the local model is used to predict the current training time using the second calculation formula to obtain the calculation delay of each server; The second calculation formula is: , in, Indicates the current status of the target server. represents the current computing performance of the target server, Indicates whether the target server has participated in the previous training task, which is 1 if it is and 0 if it is not; represents the historical training information of the target server, represents the historical training time of the local model by the target server, Indicates the preparation time for the target server to participate in this training. Indicates the computing delay of the target server, where the target server is any of the servers.
7. The method according to any one of claims 1 to 6, characterized in that: The determining of the cluster head server in each of the clusters based on the training delay of each of the servers comprises: The server with the smallest training delay in each of the clusters is determined as the cluster head server in each of the clusters.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Federal learning training and prediction method and system based on heterogeneous resources
CN114219097A
Mobile user equipment clustering training method for wireless federated learning
CN114553661A