Clustered Data Sharing Method, Device, and Storage Medium Based on Federated Learning System

By clustering distributed devices in the federated learning system and sharing training data using cluster head devices, the problems of slow convergence speed and low model accuracy caused by data heterogeneity are solved, and faster training speed and lower communication overhead are achieved.

CN116233954BActive Publication Date: 2025-07-08BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211575350.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-07-08
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

In federated learning, the problems of slow convergence speed, large communication overhead and low model accuracy due to unbalanced device data distribution.

Method used

The distributed devices are divided into clusters based on a pre-set clustering algorithm, and the training data is shared with the cluster head device to the member devices in the cluster, and iterate the initial model in collaboration with the central server to slow down the degree of data heterogeneity.

Benefits of technology

This improves the convergence speed of federated learning, reduces communication overhead, and improves the accuracy of the final model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116233954B_ABST
    Figure CN116233954B_ABST
Patent Text Reader

Abstract

The present invention provides a clustered data sharing method, apparatus and storage medium based on a federated learning system. The federated learning system includes K distributed devices and a central server, where K is an integer greater than 1. The method includes: dividing the K distributed devices into M clusters based on a pre-set clustering algorithm; M is an integer less than K, and at least one of the M clusters includes a cluster head device and in-cluster member devices; controlling the cluster head devices in each cluster to share training data with the in-cluster member devices; and based on a pre-set federated learning algorithm, collaboratively iteratively training a pre-set initial model through the training data of each distributed device and the central server to obtain a target model after federated learning training. By clustering the distributed devices and having the cluster head devices share training data with the in-cluster member devices, the present invention reduces the degree of data heterogeneity, decreases the communication overhead of federated learning training, and improves the accuracy of the finally trained target model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and in particular, to a cluster data sharing method, device, and storage medium based on a federated learning system. Background Art

[0002] Federated Learning (FL) is a machine learning (ML) technology that can complete model training without revealing private data. Its core idea is to perform distributed model training among multiple devices with local data. Specifically, multiple distributed devices locally train the model using local data, and the server aggregates the model parameters after training by each distributed device to obtain the model after federated learning. Compared with traditional centralized ML technologies, FL does not require uploading the local data of the user side to the server, and can better protect the data privacy of users.

[0003] In actual application scenarios, due to the limited geographical environment and observation capabilities of each device, the characteristics of the data generated by distributed devices often have imbalances in distribution, that is, the data generated by distributed devices belongs to non-independent and identically distributed (Non-IID). It can also be understood that there is a problem of data heterogeneity in the training data sources of current federated learning.

[0004] The problem of data heterogeneity will reduce the convergence speed of the federated learning algorithm, increase the communication overhead of federated learning training. In addition, data heterogeneity will also cause Stochastic Gradient Decent (SGD), resulting in an offset of the model update direction compared to the target direction, and thus leading to a low accuracy of the finally trained model. Summary of the Invention

[0005] The present invention provides a cluster data sharing method, device, and storage medium based on a federated learning system to solve the problems in the prior art that the convergence speed of the federated learning algorithm is low, the communication overhead of federated learning training is high, and the accuracy of the finally trained model is low.

[0006] The present invention provides a cluster data sharing method based on a federated learning system. The federated learning system includes K distributed devices and a central server, where K is an integer greater than 1;

[0007] The method includes:

[0008] Based on a pre-set clustering algorithm, K distributed devices are divided into M clusters; where M is an integer less than K, and there is at least one cluster among the M clusters that includes a cluster head device and in-cluster member devices.

[0009] Control the cluster head devices in each of the clusters to share training data with the in-cluster member devices.

[0010] Based on a pre-set federated learning algorithm, the initial model pre-set is collaboratively and iteratively trained through the training data of each of the distributed devices and the central server to obtain a target model after federated learning training.

[0011] According to a clustering data sharing method based on a federated learning system provided by the present invention, the division of K distributed devices into M clusters based on a pre-set clustering algorithm includes:

[0012] Establish a privacy constraint graph for the K distributed devices; where the privacy constraint graph includes K nodes corresponding to the K distributed devices and edges for connecting each of the nodes.

[0013] Calculate the intimacy relationship value between each of the distributed devices and other distributed devices among the K distributed devices as the first attribute value of the edge between the node corresponding to each of the distributed devices and the node corresponding to the other distributed devices; where the intimacy relationship value is used to characterize the trust degree between each of the distributed devices.

[0014] Delete the edges corresponding to the first attribute values less than the privacy threshold in the privacy constraint graph to obtain a privacy communication constraint graph.

[0015] Calculate the Earth Mover's Distance (EMD) of the system corresponding to the K distributed devices as the attribute value of the K nodes; where the EMD in the system is used to characterize the difference between the distribution of the training data of each of the distributed devices and the distribution of the global data of the K distributed devices.

[0016] In the privacy communication constraint graph, select the corresponding nodes as cluster head nodes in the order of the attribute values of the nodes from large to small until there is an edge between each of the other nodes and at least one cluster head node, to obtain M cluster head nodes; where the other nodes are the nodes among the K nodes except the cluster head nodes.

[0017] Calculate the inter-device EMD between each of the distributed devices and other distributed devices among the K distributed devices as the second attribute value of the edge between the node corresponding to each of the distributed devices and the node corresponding to the other distributed devices; where the inter-device EMD is used to characterize the difference in the distribution of the training data between each of the distributed devices.

[0018] When there is an edge only between the other node and one cluster head node, divide the other node into the cluster where the cluster head node with which there is an edge between the other node is located;

[0019] When there are edges between the other node and at least two cluster head nodes, divide the other node into the cluster where the cluster head node with the largest second attribute value corresponding to the edge between the other node is located;

[0020] Take the distributed devices corresponding to the M cluster head nodes as the cluster head devices of the M clusters, and take the distributed devices corresponding to the other nodes within the clusters where the respective cluster head nodes are located as the in-cluster member devices of the M clusters.

[0021] According to a clustering data sharing method based on a federated learning system provided by the present invention, before deleting the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph to obtain a privacy communication constraint graph, the method further includes:

[0022] Calculate the data transmission rate between each of the distributed devices and the other distributed devices, and use it as the third attribute value of the edge between the node corresponding to each of the distributed devices and the node corresponding to the other distributed devices;

[0023] The step of deleting the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph to obtain a privacy communication constraint graph includes:

[0024] Delete the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph, and delete the edges corresponding to the third attribute values less than the communication rate threshold to obtain the privacy communication constraint graph.

[0025] According to a clustering data sharing method based on a federated learning system provided by the present invention, before controlling the cluster head devices in each of the clusters to share training data with the in-cluster member devices, the method further includes:

[0026] Calculate the sharing delay t of each cluster head device sharing training data with the in-cluster member devices s , and the training delay t of performing one round of model training based on the federated learning algorithm FL ;

[0027] Based on t s and t FL , use formula (1) to calculate the target shared data volume N of each cluster head device S and the target central processing unit (CPU) frequency f of each distributed device for training the model:

[0028]

[0029] Among them, Ω() represents the number of iterations for training the model of the federated learning system;

[0030] Controlling each cluster head device in each cluster to share training data with the member devices within the cluster includes:

[0031] Controlling each cluster head device to share training data with a data volume of N S with the member devices within the cluster;

[0032] Based on the pre-set federated learning algorithm, through the training data of each distributed device and the central server, jointly iteratively training the pre-set initial model to obtain the target model after federated learning training, includes:

[0033] Based on the federated learning algorithm, with f being the CPU frequency of the training model of each distributed device, through the training data of each distributed device and the central server, jointly iteratively training the pre-set initial model to obtain the target model after federated learning training.

[0034] According to a method for sharing clustered data based on a federated learning system provided by the present invention, calculating the sharing delay t of each cluster head device sharing training data with the member devices within the cluster s , and the training delay t for one round of model training based on the federated learning algorithm FL , includes:

[0035] Using formula (2) to calculate the sharing delay t s :

[0036]

[0037] Among them, M represents the set of cluster head devices, n represents the nth cluster head device in the set of cluster head devices, C m represents the set of member devices within the cluster where the nth cluster head device is located, C represents the Cth member device within the cluster in the set of member devices within the cluster, a represents the number of bits occupied by a sample of a training data, represents the data volume of the mth cluster head device sharing training data with the member devices within the cluster, v m,c represents the data transmission rate between the mth cluster head device and the cth member device within the cluster;

[0038] Using formula (3) to calculate the training delay t FL :

[0039]

[0040] Among them, represents the downlink delay, Characterize the update delay, Characterize the uplink delay.

[0041] According to a method for sharing clustered data based on a federated learning system provided by the present invention, controlling each cluster head device in each cluster to share training data with the member devices in the cluster includes:

[0042] Controlling each cluster head device in each cluster to share training data with the member devices in the cluster by means of device-to-device (D2D) multicast.

[0043] The present invention also provides a device for sharing clustered data based on a federated learning system. The federated learning system includes K distributed devices and a central server, where K is an integer greater than 1;

[0044] The device includes:

[0045] A clustering module, configured to divide the K distributed devices into M clusters based on a pre-set clustering algorithm; where M is an integer less than K, and there is at least one cluster among the M clusters that includes a cluster head device and member devices in the cluster;

[0046] A control module, configured to control each cluster head device in each cluster to share training data with the member devices in the cluster;

[0047] A federated learning training module, configured to collaboratively and iteratively train a pre-set initial model with the training data of each distributed device and the central server based on a pre-set federated learning algorithm to obtain a target model after federated learning training.

[0048] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for sharing clustered data based on a federated learning system as described in any one of the above.

[0049] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for sharing clustered data based on a federated learning system as described in any one of the above.

[0050] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for sharing clustered data based on a federated learning system as described in any one of the above.

[0051] The clustering data sharing method, device, and storage medium based on a federated learning system provided by the present invention first divide K distributed devices of the federated learning system into M clusters based on a pre-set clustering algorithm, so that the cluster head devices in each cluster share training data with the member devices in the cluster, which can slow down the data heterogeneity degree between the training data of each distributed device. Then, based on the federated learning algorithm, the initial model pre-set is collaboratively iteratively trained through the training data of each distributed device and the central server to obtain the target model after federated learning training. Compared with the model training process of federated learning in the related art, in the embodiments of the present invention, by clustering the distributed devices so that the cluster head devices share training data with the member devices in the cluster, the degree of data heterogeneity is slowed down, thereby improving the convergence speed of the subsequent federated learning algorithm, reducing the communication overhead of federated learning training, and effectively improving the accuracy of the finally trained target model. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 is one of the schematic flowcharts of the clustering data sharing method based on the federated learning system provided by the present invention;

[0054] Figure 2 is the schematic diagram of the privacy constraint graph in the clustering data sharing method based on the federated learning system provided by the present invention;

[0055] Figure 3 is the schematic diagram of the privacy communication constraint graph in the clustering data sharing method based on the federated learning system provided by the present invention;

[0056] Figure 4 is the second schematic flowchart of the clustering data sharing method based on the federated learning system provided by the present invention;

[0057] Figure 5 is the schematic structural diagram of the clustering data sharing device based on the federated learning system provided by the present invention;

[0058] Figure 6 is the schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] The following describes the cluster data sharing method, device, and storage medium based on a federated learning system of the present invention with reference to the accompanying drawings.

[0061] The federated learning system of the embodiments of the present invention includes K distributed devices and a central server, where K is an integer greater than 1, and the distributed devices are, for example, terminals.

[0062] Figure 1 is one of the flow diagrams of the cluster data sharing method based on a federated learning system provided by the present invention. As Figure 1 shown, the method includes steps 101 to 103; where:

[0063] Step 101: Divide the K distributed devices into M clusters based on a pre-set clustering algorithm; where M is an integer less than K, and at least one of the M clusters includes a cluster head device and in-cluster member devices.

[0064] Step 102: Control the cluster head devices in each of the clusters to share training data with the in-cluster member devices.

[0065] Step 103: Based on a pre-set federated learning algorithm, jointly and iteratively train a pre-set initial model with the training data of each of the distributed devices and the central server to obtain a target model after federated learning training.

[0066] Specifically, in the related art, due to the limited geographical environment and observation capabilities of each device, the characteristics of the data generated by the distributed devices often have an unbalanced distribution. This unbalanced data distribution can be understood as data heterogeneity, which will reduce the convergence speed of the federated learning algorithm, increase the communication overhead of federated learning training, and in addition, will also cause SGD, resulting in an offset of the model update direction compared with the target direction, and thus leading to a low accuracy of the finally trained model.

[0067] Currently, there is a large amount of work to address the challenge of data heterogeneity in federated learning. Among them, the solutions based on effective algorithm design will increase the computational overhead of local devices and cannot significantly improve the model accuracy of federated learning in the case of data heterogeneity. In addition, many studies have made the cross-device data distribution more consistent through data augmentation methods. Considering that there are a large number of devices and private data in the Internet of Things (IoT) scenario, the above data augmentation-based methods require building an available, public, and ideal dataset on the central server side, which may consume a large amount of costs and there are security risks in the construction process. Therefore, it cannot be well applied to the scenario of large-scale user federated learning.

[0068] Aiming at the defects in the existing technologies, the embodiments of the present invention provide a clustering data sharing method based on a federated learning system, which can slow down the degree of data heterogeneity without relying on an additional dataset on the central server, thereby reducing the training cost (communication overhead) of federated learning and improving the accuracy of the trained model.

[0069] In the embodiments of the present invention, first, based on a preset clustering algorithm, K distributed devices of the federated learning system are divided into M clusters. Among them, there is at least one cluster including a cluster head device and in-cluster member devices in the M clusters. It should be noted that there can be clusters including only cluster head devices in the M clusters; in a cluster, it can be set that there is only one cluster head device, and the other devices in the cluster are all in-cluster member devices.

[0070] After clustering, the cluster head device in each cluster shares training data with the in-cluster member devices, which can, to a certain extent, slow down the degree of data heterogeneity between the training data of each distributed device. Then, based on the federated learning algorithm, the initial model preset is iteratively trained in cooperation with the central server through the training data of each distributed device to obtain the target model after federated learning training.

[0071] It should be noted that the federated learning algorithm can select existing federated learning algorithms. The focus of the present invention is that within each cluster, all or part of the training data is shared by the cluster head device with the in-cluster member devices to slow down the degree of data heterogeneity between the devices in the cluster. Then, using the federated learning algorithm, each distributed device with the reduced data heterogeneity degree uses the training data to iteratively train the initial model to obtain the target model after federated learning training.

[0072] Optionally, the iterative training process of the model is as follows:

[0073] One round of training process of federated learning consists of four parts: model download, local update, model upload, and model aggregation:

[0074] 1) Model Download: The central server of the federated learning system first selects the set S of distributed devices participating in FL training in the μ-th round μ and broadcasts the initialized global model w g (i.e., the initial model) to the selected distributed devices;

[0075] 2) Local Update: After receiving the initialized global model w g from the central server, the distributed devices update the model w on their local datasets using the training data through SGD g to obtain the updated model w k , specifically using the following formula: where represents the gradient of the loss function, and η represents the learning rate;

[0076] 3) Model Upload: After completing the local model update, the distributed devices upload the local model to the Base Station (BS) via the uplink, and the uplink uses Orthogonal Frequency Division Multiple Access (OFDMA);

[0077] 4) Model Aggregation: When the steps of local model training are completed, the distributed devices send the local models to the BS for synchronous aggregation. The aggregation method of the global model is the weighted average of the uploaded models, specifically using formula (4):

[0078]

[0079] where n k represents the data volume of the training data of the k-th distributed device, and n represents the sum of the data volumes of the K distributed devices in the federated learning system;

[0080] After the global model aggregation is completed, it can be fed back to the users selected in the next round, and the above model training process is repeated for multiple rounds until convergence to obtain the target model after federated learning training.

[0081] The clustering data sharing method based on a federated learning system provided by an embodiment of the present invention first divides K distributed devices of the federated learning system into M clusters based on a preset clustering algorithm, so that the cluster head devices in each cluster share training data with the member devices in the cluster, which can slow down the data heterogeneity degree among the training data of the distributed devices. Then, based on the federated learning algorithm, the preset initial model is collaboratively iteratively trained with the training data of each distributed device and a central server to obtain a target model after federated learning training. Compared with the model training process of the federated learning in the related art, in the embodiment of the present invention, by clustering the distributed devices and enabling the cluster head devices to share training data with the member devices in the cluster, the degree of data heterogeneity is slowed down, thereby improving the convergence speed of the subsequent federated learning algorithm, reducing the communication overhead of the federated learning training, and effectively improving the accuracy of the finally trained target model.

[0082] Optionally, the implementation manner of dividing the K distributed devices into M clusters based on the preset clustering algorithm may include:

[0083] Establish a privacy constraint graph of the K distributed devices; wherein, the privacy constraint graph includes K nodes corresponding to the K distributed devices and edges for connecting the nodes;

[0084] Calculate the intimacy relationship value between each distributed device and other distributed devices among the K distributed devices as the first attribute value of the edge between the node corresponding to each distributed device and the node corresponding to the other distributed devices; wherein, the intimacy relationship value is used to characterize the trust degree between the distributed devices;

[0085] Delete the edges corresponding to the first attribute values less than the privacy threshold in the privacy constraint graph to obtain a privacy communication constraint graph;

[0086] Calculate the Earth Mover’s Distance (EMD) in the system corresponding to the K distributed devices as the attribute value of the K nodes; wherein, the EMD in the system is used to characterize the difference between the distribution of the training data of the distributed devices and the distribution of the global data of the K distributed devices;

[0087] In the privacy communication constraint graph, select the corresponding nodes as cluster head nodes in the order from large to small of the attribute values (EMD in the system) of the nodes until there are edges between all other nodes and at least one cluster head node, so as to obtain M cluster head nodes; wherein, the other nodes are the nodes except the cluster head nodes among the K nodes;

[0088] Calculate the inter-device EMD between each of the distributed devices and the other distributed devices among the K distributed devices, and use it as the second attribute value of the edge between the node corresponding to each distributed device and the node corresponding to the other distributed devices; wherein, the inter-device EMD is used to characterize the difference in the distribution of training data between each of the distributed devices.

[0089] In the case where there is an edge only between the other node and one cluster head node, divide the other node into the cluster where the cluster head node with which there is an edge between the other node is located.

[0090] In the case where there are edges between the other node and at least two cluster head nodes, divide the other node into the cluster where the cluster head node with the largest second attribute value (inter-device EMD) corresponding to the edge between the other node and the other nodes is located.

[0091] Take the distributed devices corresponding to the M cluster head nodes as the cluster head devices of the M clusters, and take the distributed devices corresponding to the other nodes within the clusters where each cluster head node is located as the intra-cluster member devices of the M clusters.

[0092] Specifically, first establish a privacy constraint graph of K distributed devices K represents a set of points, including K nodes corresponding to K distributed devices, and ε represents a set of edges, including the edges used to connect each node.

[0093] For example, Figure 2 is a schematic diagram of the privacy constraint graph in the cluster data sharing method based on the federated learning system provided by the present invention, as Figure 2 shown, K is 7, corresponding to nodes numbered 1-7.

[0094] First calculate the intimacy relationship value e between each distributed device and the other distributed devices among the K distributed devices, and use it as the first attribute value of the edge between the node corresponding to each distributed device and the node corresponding to the other distributed devices; wherein, the intimacy relationship value is used to characterize the degree of trust between each of the distributed devices.

[0095] Delete the edges corresponding to the first attribute values less than the privacy threshold in the privacy constraint graph to obtain a privacy communication constraint graph.

[0096] Such as Figure 2As shown, assuming that the intimacy relationship value between Node 1 and Node 2 is equal to 0.5, the first attribute value of the edge connecting Node 1 and Node 2 is set to 0.5, which is represented as (0.5) in the figure. Similarly, calculate and set the first attribute value of the edge between each pair of nodes. Assume that the first attribute value of the edge between Node 1 and Node 6 is 1, represented as (1); the first attribute value of the edge between Node 2 and Node 6 is 2, represented as (2); the first attribute value of the edge between Node 2 and Node 7 is 2, represented as (2); the first attribute value of the edge between Node 5 and Node 7 is 2, represented as (2); the first attribute value of the edge between Node 3 and Node 7 is 2, represented as (2); the first attribute value of the edge between Node 4 and Node 7 is 2, represented as (2). Figure 2 Only some edges and their corresponding first attribute values are shown in Figure 2 . For the edges not shown, it can be considered that their first attribute values are less than the privacy threshold, and the privacy threshold is, for example, 1.

[0097] Figure 3 It is a schematic diagram of the privacy communication constraint graph in the cluster data sharing method based on the federated learning system provided by the present invention. As Figure 3 shown, compared with Figure 2 , since the first attribute value of the edge between Node 1 and Node 2 is less than the privacy threshold, the edge between Node 1 and Node 2 is deleted.

[0098] Optionally, the intimacy relationship value can also be called social intimacy, which is used to characterize the trust degree between two distributed devices in the D2D network. In particular, the intimacy relationship value can be equal to 1, indicating the trust relationship with itself; the intimacy relationship value can also be equal to 0, indicating no trust relationship with other distributed devices; the higher the intimacy relationship value, the closer the relationship between the two distributed devices.

[0099] The intimacy relationship value e(k, j) between the kth distributed device and the jth distributed device can be calculated by the following formula:

[0100] where φ k,j represents the number of interactions between the kth distributed device and the jth distributed device, which can be obtained from the prior information in the environment. When a distributed device communicates with another distributed device frequently for a long time, the connection between the distributed devices can be considered reliable. When the communication distance d k,j between the kth distributed device and the jth distributed device exceeds the threshold d th , it can be considered that there is no D2D connection between the two distributed devices.

[0101] Delete the edges corresponding to the first attribute values less than the privacy threshold in the privacy constraint graph. For example, delete the edge between node 1 and node 2 to obtain the privacy communication constraint graph. It can be understood that the relationship between the two nodes connected by the remaining edges in the privacy communication constraint graph is closer. When selecting the cluster head devices corresponding to the in-cluster member devices later, the privacy communication constraint graph can be used for selection to ensure that the privacy data of the distributed devices is not leaked as much as possible when the cluster head devices share data;

[0102] After obtaining the privacy communication constraint graph, first calculate the EMD in the system corresponding to the K distributed devices as the attribute values of the K nodes. Among them, the EMD in the system is used to characterize the difference between the distribution of the training data of each distributed device and the distribution of the global data of the K distributed devices. As Figure 3 shown, assume that the EMDs in the systems corresponding to nodes 1 - 7 are 1 - 7 respectively, that is, the attribute values corresponding to nodes 1 - 7 are 1 - 7 respectively;

[0103] Then calculate the inter-device EMD between each distributed device and other distributed devices among the K distributed devices as the second attribute value of the edge between the node corresponding to each distributed device and the node corresponding to other distributed devices, that is, as the second attribute value of the edge between each node and other nodes; among them, the inter-device EMD is used to characterize the difference in the distribution of the training data between each distributed device;

[0104] It should be noted that only the second attribute values corresponding to the existing edges in the privacy communication constraint graph can be calculated. Assume that the second attribute value of the edge between node 1 and node 6 is calculated as 1, the second attribute value of the edge between node 2 and node 6 is calculated as 2, the second attribute value of the edge between node 2 and node 7 is calculated as 4, the second attribute value of the edge between node 5 and node 7 is calculated as 3, the second attribute value of the edge between node 3 and node 7 is calculated as 5, and the second attribute value of the edge between node 4 and node 7 is calculated as 6.

[0105] After calculating the EMD in the system and the inter-device EMD, the following begins to use the privacy communication constraint graph for clustering to select the cluster head devices and the in-cluster member devices:

[0106] In the privacy communication constraint graph, select the corresponding nodes as the cluster head nodes in the order of the EMD in the system from large to small until there is an edge between each of the other nodes except the cluster head nodes among the K nodes and at least one cluster head node, and obtain M cluster head nodes;

[0107] For example, as Figure 3As shown in the figure, first select node 7 with the largest EMD in the system as the cluster head node. At this time, there are edges between nodes 2, 3, 4, and 5 and node 7, but there are no edges between nodes 1, 6 and node 7. Therefore, continue to select node 6 with the second largest EMD in the system as the cluster head node. At this time, for other nodes that are not cluster head nodes, there are edges between them and at least one cluster head node. Then, determine nodes 6 and 7 as the final cluster head nodes;

[0108] Next, start to select the intra-cluster member nodes corresponding to the cluster head nodes 6 and 7 respectively. Specifically, it is divided into the following two cases:

[0109] 1) In the case where there is an edge between other nodes and only one cluster head node, divide other nodes into the cluster where the cluster head node with an edge between it and other nodes is located;

[0110] For example, as Figure 3 shown in the figure, there is an edge only between node 1 and cluster head node 6. Therefore, directly divide node 1 into the cluster where cluster head node 6 is located, and node 1 is used as an intra-cluster member node in the cluster where cluster head node 6 is located; Similarly, there are edges only between nodes 3, 4, and 5 and cluster head node 7. Therefore, directly divide nodes 3, 4, and 5 into the cluster where cluster head node 7 is located, and nodes 3, 4, and 5 are used as intra-cluster member nodes in the cluster where cluster head node 7 is located.

[0111] 2) In the case where there are edges between other nodes and at least two cluster head nodes, divide other nodes into the cluster where the cluster head node with the largest EMD between the devices corresponding to the edges between it and other nodes is located.

[0112] For example, as Figure 3 shown in the figure, there is an edge between node 2 and both cluster head node 6 and cluster head node 7. Then, compare the second EMD of the edge between node 2 and cluster head node 6 with the second EMD of the edge between node 2 and cluster head node 7. Since the second EMD of the edge between node 2 and cluster head node 6 is 2, while the second EMD of the edge between node 2 and cluster head node 7 is 4, select the cluster where cluster head node 7 with a larger second EMD is located as the cluster for node 2, that is, node 2 is used as an intra-cluster member node in the cluster where cluster head node 7 is located.

[0113] After selecting the cluster head nodes and their corresponding intra-cluster member nodes, take the distributed devices corresponding to the M cluster head nodes as the cluster head devices of the M clusters, and take the distributed devices corresponding to the other nodes in the clusters where each cluster head node is located as the intra-cluster member devices of the M clusters.

[0114] Optionally, before deleting the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph to obtain the privacy communication constraint graph, the data transmission rate between each of the distributed devices and the other distributed devices may be calculated as the third attribute value of the edges between the nodes corresponding to each of the distributed devices and the nodes corresponding to the other distributed devices;

[0115] The implementation manner of deleting the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph to obtain the privacy communication constraint graph may include:

[0116] Delete the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph, and delete the edges corresponding to the third attribute values less than the communication rate threshold, to obtain the privacy communication constraint graph.

[0117] Specifically, considering that the amount of shared data will also affect the optimization objective, a rate threshold v th may be set as the communication rate threshold. When selecting the intra-cluster member nodes of the cluster head node, the intra-cluster member nodes with a data transmission rate less than the communication rate threshold are not considered to avoid transmitting too little data. Specifically, the data transmission rate between each distributed device and other distributed devices may be calculated as the third attribute value of the edges between the nodes corresponding to each distributed device and the nodes corresponding to the other distributed devices. When generating the privacy communication constraint graph, not only the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph are deleted, but also the edges corresponding to the third attribute values less than the communication rate threshold are deleted, to obtain the privacy communication constraint graph, and then the cluster head device and the intra-cluster member devices may be selected based on the privacy communication constraint graph.

[0118] In the embodiments of the present invention, the data transmission rate constraint condition is considered to avoid transmitting too little data between distributed devices, further alleviating the problem of data heterogeneity in the federated learning system, further improving the convergence speed of the subsequent federated learning algorithm, reducing the communication overhead of federated learning training, and improving the accuracy of the final trained target model.

[0119] Optionally, before controlling each cluster head device in each cluster to share training data with the intra-cluster member devices, the sharing delay t s for each cluster head device to share training data with the intra-cluster member devices may be calculated, and FL the training delay t

[0120] for one round of model training based on the federated learning algorithm; s Based on t FL and t Sand the target central processing unit (CPU) frequency f of each of the distributed device training models:

[0121]

[0122] where Ω() represents the number of iterations of training the model of the federated learning system;

[0123] Controlling each cluster head device in each cluster to share training data with the member devices in the cluster includes:

[0124] Controlling each cluster head device to share training data with a data volume of N S to the member devices in the cluster;

[0125] The implementation manner of obtaining the target model after federated learning training by collaboratively iteratively training the preset initial model through the training data of each distributed device and the central server based on the preset federated learning algorithm may include:

[0126] Based on the federated learning algorithm, with f as the CPU frequency of each distributed device training model, the preset initial model is collaboratively iteratively trained through the training data of each distributed device and the central server to obtain the target model after federated learning training.

[0127] Specifically, the sharing delay t s for each cluster head device to share training data with the member devices in the cluster can be calculated, and the training delay t FL for one round of model training based on the federated learning algorithm, and then based on t s and t FL , the target shared data volume N S for each cluster head device and the target CPU frequency f of each distributed device training model are calculated. Here, the solved N S and f can be understood as the optimal combination of training the model of the federated learning system, comprehensively considering the delay of training the model and the accuracy of the trained model.

[0128] It should be noted that for the solution of formula (1), there are the following constraint conditions: θ≥θ th , γ k ≥γ th , k∈K and 0≤f≤f max .

[0129] The difficulty of this sub-problem lies in that the expression of Ω(N S ) of the communication rounds with respect to the shared data volume is unknown. The method of data fitting is used to determine the expression of the communication rounds. According to the analysis, it can be determined that Ω(N S) The basic function form is Through Ω(N S ) The basic form can be used to determine that the sub - problem is a convex problem, and it is easy to directly obtain the optimal solution of this problem with existing algorithms. For example, the gradient descent method, interior point method, KKT conditions, etc. can be used for solving to obtain N S and f.

[0130] Solve for N S and f, and then control each cluster - head device to share the amount of data N S of training data with the in - cluster member devices; and based on the federated learning algorithm, using f as the CPU frequency for training the model of each distributed device, the initial model preset is collaboratively and iteratively trained through the training data of each distributed device and the central server to obtain the target model after federated learning training, so as to achieve a better model training effect.

[0131] Optionally, the implementation of calculating the sharing delay t s for each cluster - head device to share training data with the in - cluster member devices, and the training delay t FL for one - round model training based on the federated learning algorithm may include:

[0132] Using formula (2), calculate the sharing delay t s :

[0133]

[0134] where M represents the set of cluster - head devices, m represents the m - th cluster - head device in the set of cluster - head devices, C m represents the set of in - cluster member devices in the cluster where the m - th cluster - head device is located, c represents the c - th in - cluster member device in the set of in - cluster member devices, a represents the number of bits occupied by a sample of a training data, represents the amount of training data that the m - th cluster - head device shares with the in - cluster member devices, v m,c represents the data transmission rate between the m - th cluster - head device and the c - th in - cluster member device;

[0135] Using formula (3), calculate the training delay t FL :

[0136]

[0137] where, represents the downlink delay, represents the update delay, represents the uplink delay.

[0138] Specifically, formula (2) can be used to calculate the sharing delay t s, and use formula (3) to calculate the training delay t FL ;

[0139] Next, the downlink delay update delay and uplink delay in formula (3) will be described.

[0140] 1) Downlink delay is the downlink data delay when the BS broadcasts the global model to the distributed devices. Specifically, the downlink delay can be calculated by formula (5)

[0141]

[0142] where D w represents the number of bits occupied by the global model. Since the global model and the local model in FL use the same architecture, D w is also the number of bits of the local model, B D represents the broadcast bandwidth of the BS, P B represents the transmission power of the base station, h k represents the channel gain between the BS and the k-th distributed device, and N0 represents the power spectral density of the noise.

[0143] 2) Update delay Generally, local update is to perform the SGD algorithm to minimize the loss function. Set the number of rounds (Epoch) of the update algorithm to E, and the corresponding calculation delay when the k-th distributed device performs one round of local update is formula (6):

[0144]

[0145] where L k represents the number of CPU cycles required to train each training sample on the k-th distributed device, f k represents the CPU frequency of the k-th distributed device;

[0146] The energy consumption of each user during local update can be calculated using (7)

[0147]

[0148] where ρ k represents the energy consumption coefficient, which depends on the hardware attributes of the k-th distributed device.

[0149] 3) Uplink delay The delay for the k-th distributed device to upload the local model to the BS can be specifically calculated by formula (8) for the uplink delay

[0150]

[0151] Among them, B U represents the bandwidth allocated to the participant for communication with the BS, and P k represents the transmission power of the k-th distributed device.

[0152] The energy consumption generated by the transmission model can be calculated using (9)

[0153]

[0154] Optionally, the implementation manner of controlling each cluster head device in each cluster to share training data with the in-cluster member devices may include:

[0155] Controlling each cluster head device in each cluster to share training data with the in-cluster member devices in a Device-to-Device (D2D) multicast manner.

[0156] Specifically, the cluster head device sharing training data with the in-cluster member devices in a D2D multicast manner improves the data security of the distributed devices as compared with sharing training data through the base station.

[0157] The following gives an example to illustrate the cluster data sharing method based on the federated learning system provided by the embodiments of the present invention.

[0158] 1. Assume that in a wireless edge computing cell system, there are K intelligent mobile devices as distributed devices, and a BS is equipped. Each user has one distributed device. The local samples collected by each distributed device are Among them, is the input feature vector of the sample, is the output feature vector of the sample, and n k is the number of samples of the user's distributed device. The user and the BS cooperate to complete a deep learning (DL) task under a client-server architecture. This edge environment not only includes common uplink and downlink transmissions, but also nearby users can use device-to-device D2D communication.

[0159] The local data generated by the user's distributed device may have significant differences in statistical distribution, which may lead to the generation of non-IID. For a clearer illustration, the mathematical definitions of IID and non-IID are given first.

[0160] IID is a common assumption in Distributed Learning, where the data distributions of all users follow the same global distribution P g (x,y).

[0161] However, due to geographical environment and observation ability limitations of users' distributed devices, local data no longer satisfies the IID assumption. That is, the data of each user's distributed device follows its own different distribution P k (x,y), which can also be written as P k (y)P k (x|y).

[0162] A common non-IID type in federated learning is data distribution skew. In this type of situation, the conditional distributions P k (x|y) of all users' distributed devices are the same, but the marginal distributions P k (y) are different. Using EMD, that is, using D EMD (k) to quantify the heterogeneity of data on distributed devices, as shown in formula (10):

[0163]

[0164] In addition, the weighted average sum of the EMD values of all users' distributed devices is defined as the overall EMD value of the system, as shown in formula (11):

[0165]

[0166] Through research and experiments, it can be found that the smaller the EMD value of the system, the lower the value of the loss function (Loss) of FL training, and the faster the convergence speed. This phenomenon indicates that reducing the EMD value of the system can accelerate the training of FL and improve the model accuracy.

[0167] To alleviate the data imbalance characteristic in FL, the embodiment of the present invention proposes a clustering data sharing framework that can reduce communication overhead. By exchanging a small amount of training data within clusters with efficient communication and privacy protection characteristics, the distribution of training data can be made more similar among distributed devices.

[0168] Figure 4 is the second schematic diagram of the process of the clustering data sharing method based on the federated learning system provided by the present invention, as Figure 4 shown, the framework includes two stages:

[0169] (I) Data processing stage

[0170] In this stage, clustering is performed according to factors such as the data distribution, channel state, and credibility of the user's distributed devices. Within each cluster, the cluster head (CH) device shares a portion of its training data with other cluster member devices (CMs) within the cluster through D2D multicast. The cluster head device is denoted as m ∈ M, and the corresponding set of cluster member devices is C m . Note that to avoid conflicts, a node user can join at most one cluster, that is After the data sharing is completed, the amount of local data on the k-th distributed device is Specifically, as shown in Equation (12):

[0171]

[0172] where represents the amount of data shared by the cluster head device m.

[0173] In particular, when node k is selected as the cluster head, P m (y) = P k (y), so the new data distribution on node k can be written as:

[0174] Intuitively, data sharing can change the original data distribution on the distributed devices and buffer the differences in data characteristics. The EMD value of the system after data sharing becomes Equation (13):

[0175]

[0176] (II) Federated Learning Training Stage

[0177] FL jointly trains a global model through cooperation among users and multiple users. The goal of the global model is to minimize the loss function, which can be expressed by Equation (14):

[0178]

[0179] where F k (·) represents the loss function of the k-th distributed device.

[0180] As is well known, data sharing inevitably brings privacy risks and communication costs. To mitigate these impacts, the present invention defines two metrics to further design the clustering algorithm.

[0181] 1. Social Closeness: Social closeness is an important indicator in the social network, which can reflect the intimacy between users' distributed devices. Simply put, users are willing to exchange data with other users with higher social closeness to ensure data privacy. In the framework of clustered data sharing, it is required that there is a certain social closeness between the member devices within a cluster and the cluster head device, that is Construct a graph (privacy constraint graph) to more clearly explain the social closeness relationship between users' distributed devices, where is the set of edges in the graph, representing the social closeness value between users' distributed devices.

[0182] 2. Transmission Delay: The cluster head device shares part of the training data set with the member devices within the cluster through D2D multicast communication. The data sharing transmission rate from the cluster head device m to each member device c∈c m is x m,c , and the specific calculation is as shown in formula (15):

[0183]

[0184] where represents the communication bandwidth of the cluster head device for broadcasting the model, h m,c represents the channel gain between the cluster head device m and the member device c within the cluster, represents the transmission power of the cluster head device m, I m represents the interference caused by the cluster head devices located in other service areas, and N0 represents the noise power spectral density.

[0185] In the multicast mode, the data sharing delay within a cluster is determined by the worst link, and at the same time, the total data sharing time cost t s depends on the maximum delay of the cluster, and the specific calculation is as shown in formula (2); including the sharing delay in the data preparation stage, the delay of the overall framework also includes the delay of model download (downlink delay), the delay of model update (update delay), and the delay of model upload (uplink delay), that is, the delay of one round of FL training is determined by the maximum value of the sum of the downlink, update, and uplink delays:

[0186] The energy consumption γ of one FL training round k consists of two parts: calculation and uplink, that is,

[0187] An appropriate clustered data sharing method is crucial for federated learning. First, the embodiments of the present invention propose an optimization problem to minimize the latency from data sharing and FL training and maximize the model accuracy as much as possible. Then, the impact of the clustered data sharing method on this problem is explored in detail and solved separately into two sub-problems.

[0188] The problem of minimizing the training latency can be formulated as Equation (16):

[0189]

[0190] The constraints include:

[0191] (1) θ ≥ θ th ;

[0192] (2)

[0193] (3)

[0194] (4),

[0195] (5) t s ≥ T th ;

[0196] (6) γ k ≥ γ th , k ∈ K;

[0197] (7) 0 ≤ f ≤ f max .

[0198] where C = {C1…, C m ,…} is the set composed of all in-cluster member devices, is the set composed of the data volume shared by each cluster head device, and Ω is the number of communication rounds required for the federated learning system to reach the target accuracy.

[0199] The above constraint (1) characterizes that the target accuracy of FL needs to be greater than the accuracy threshold; constraint (2) is the clustering constraint; constraint (3) is used to limit the upper bound of the shared data; constraint (4) is used to characterize the privacy requirements among users; the maximum limit of the transmission delay is set to T th , constraint (5) is used to maintain an appropriate sharing preparation time; constraint (6) is the energy consumption requirement for a user to participate in one round of FL training; constraint (7) is the requirement for the CPU operating frequency of distributed devices.

[0200] Since the optimization objective is a coupled problem, which involves both the shared data volume and the device computing power and depends on the clustering scheme. However, once the cluster head devices and in-cluster member devices are selected, the original problem can be easily solved.

[0201] The present invention decomposes the original problem into two sub-problems: 1) how to select cluster head devices and in-cluster member devices; 2) how to optimize the shared data volume and CPU frequency.

[0202] A. Determine the sub-problems of clustering:

[0203] Through experiments and analysis, it can be obtained that a clustering data sharing method that can reduce communication delay and improve model accuracy is equivalent to minimizing Therefore, through the joint design of data sharing and clustering strategies, while ensuring user privacy and low communication costs, the present invention designs an optimization framework for minimizing the system EMD:

[0204] Its constraint conditions are as follows:

[0205] (1)

[0206] (2)

[0207] (3) t s ≥T th .

[0208] There are the following difficulties in solving this optimization problem: First, the joint decision-making of clustering and shared data volume is coupled together, and the relationship between variables and objectives is unknown due to the lack of an explicit expression. Second, the selection of clustering and cluster head devices constructs an NP-hard problem, which is difficult to solve, and existing algorithms all rely on heuristic algorithms. In addition, the formation of clusters also includes restrictions on privacy protection and transmission efficiency, making the problem more complex. To solve these problems, the limitations of the constraint conditions can be temporarily ignored, and the impact of the clustering strategy on the optimization objective can be considered. The present invention designs the following three conditions to analyze the target benefits that can be obtained by the clustering algorithm from the perspectives of individuals, within-clusters, and between-clusters.

[0209] Condition 1 (from the individual perspective): After data sharing, the EMD value of any user is as small as possible, that is,

[0210] Without considering any constraint conditions, Condition 1 is equivalent to the original optimization problem. However, it is still difficult to meet this condition, so the objective can be transformed into the following two extended Conditions 2 and 3;

[0211] Condition 2 (from the within-cluster perspective): In a cluster, if the best data sharing effect is to be ensured, the EMDs of the cluster head device and the in-cluster member devices need to be as different as possible, that is,

[0212] The EMD values of the in-cluster member devices are known to the user. Therefore, to maximize the sharing effect, distributed devices with as small EMD as possible should be selected as cluster head devices. Simply put, the higher the data quality of a distributed device, the greater the likelihood of being selected as a cluster head device. Although Condition 2 ensures the sharing benefits within the cluster, the clustering still cannot be completed because of the uncertainty of the clustering boundary, or rather, it is not clear how many cluster head devices should be selected. Therefore, Condition 3 is needed to define the inter-cluster relationship.

[0213] Condition 3 (Inter-cluster angle): Given the clustering results M and C, an ideal clustering result should satisfy that for any member c ∈ C m the distribution distance from the current cluster head device m should be greater than the distribution distance from other cluster head devices m′:

[0214] where is the EMD between the in-cluster member device and the cluster head device. For critical nodes that can belong to multiple clusters, this condition can be used as a separation criterion. When the data distribution of a critical node is significantly different from the data distribution of its cluster head device, the critical node will be added to that cluster. Through the analysis of the two conditions, the following theorem holds:

[0215] Theorem 1: Without considering the constraint conditions, if M * and C * are the optimal solutions of the optimization problem, then M * and C * both satisfy Condition 2 and Condition 3.

[0216] However, due to the constraint conditions, there may be no connection between some nodes, making Theorem 1 unable to be directly used. At the same time, considering that the amount of shared data will also affect the optimization goal, a rate threshold v th is set to avoid transmitting too little data.

[0217] Taking into account the constraint conditions in total, a graph (privacy communication constraint graph) can be reconstructed, where the edge set

[0218] Applying Condition 2 and Condition 3 to the graph an adaptive clustering algorithm based on data distribution is proposed. The specific steps of the algorithm can be divided into two parts, the selection process of cluster head devices and the connection process of in-cluster member devices.

[0219] For the selection of cluster head devices, calculate D EMD (k) and sort it, then select cluster head nodes in descending order until all nodes are covered.

[0220] For the connection of cluster head devices, repeatedly select the node with the maximum edge until all nodes are selected. This algorithm can adaptively determine the number of clusters and has low complexity. In addition, the algorithm can be applied to all existing non-IID FL algorithms.

[0221] Input the privacy constraint graph and the constraint threshold e th ,v th ,T th ; Output M, C;

[0222] The specific process is as follows:

[0223] Step 1. For all distributed devices in K, calculate respectively:

[0224] (1) The transmission rate v between adjacent distributed devices k,j ;

[0225] (2) The EMD distance D of the data distribution between adjacent distributed devices EMD (k,j);

[0226] (3) The EMD distance D of the self-data distribution and the global data distribution EMD (k).

[0227] Step 2. Establish the privacy communication constraint graph Among them,

[0228] Step 3. Select the cluster head devices M in descending order of the D EMD (k) value, so that all nodes can be covered by the cluster head devices;

[0229] Step 4. Among the nodes that have not been assigned, select the node c with the largest and assign it to the cluster of the cluster head device m;

[0230] Step 5. Repeat Step 4 until all distributed devices are clustered.

[0231] B. Sub-problem of joint optimization of data volume and frequency

[0232] After determining the clustering result, the original problem becomes At this time, use the gradient descent method, interior point method, KKT conditions, etc. to solve, and obtain N S and f.

[0233] II. An embodiment of the present invention provides a clustered data sharing method based on a federated learning system. By quantifying the heterogeneity degree of data distribution in federated learning, a federated learning framework for clustered data sharing is proposed to estimate and eliminate the influence of distribution bias on the performance and model accuracy of federated learning;

[0234] Specifically, based on the federated learning framework for clustered data sharing, the first step is clustered data sharing. The users in the system are clustered according to the designed clustering algorithm, and the selected cluster head devices share part of the data with the intra-cluster member devices; the second step is that the base station broadcasts the model. Select the devices participating in FL training in this round and broadcast the global model to each distributed device; the third step is that the local devices update the model. After receiving the global model, use the dataset on the distributed devices to train the model; the fourth step is to upload the model. After the update is completed, the distributed devices upload the model to the base station through the wireless communication network; the fifth step is that the base station synchronously aggregates the model. The base station performs weighted aggregation on all the received local models and feeds them back to the distributed devices for the next round of training. Repeat the second step to the fifth step until the FL training is completed.

[0235] III. An embodiment of the present invention provides a clustering algorithm for minimizing the system data heterogeneity degree. An optimization problem of minimizing the distribution distance is formulated under privacy and communication constraints, and the optimization objectives are analyzed from three perspectives: individual, intra-cluster, and inter-cluster, so as to design a novel clustering algorithm.

[0236] The clustered data sharing device based on the federated learning system provided by the present invention will be described below. The clustered data sharing device based on the federated learning system described below can be correspondingly referred to the clustered data sharing method based on the federated learning system described above.

[0237] The federated learning system of the embodiment of the present invention includes K distributed devices and a central server, where K is an integer greater than 1, Figure 5 is a schematic structural diagram of the clustered data sharing device based on the federated learning system provided by the present invention. As Figure 5 shown, the clustered data sharing device 500 based on the federated learning system includes:

[0238] A clustering module 501, configured to divide the K distributed devices into M clusters based on a preset clustering algorithm; where M is an integer less than K, and at least one of the M clusters includes a cluster head device and intra-cluster member devices;

[0239] A control module 502, configured to control the cluster head devices in each of the clusters to share training data with the intra-cluster member devices;

[0240] The federated learning training module 503 is used to collaboratively and iteratively train a preset initial model with the training data of each of the distributed devices and the central server based on a preset federated learning algorithm, so as to obtain a target model after federated learning training.

[0241] For the clustered data sharing device based on the federated learning system provided by the embodiments of the present invention, first, the clustering module divides the K distributed devices of the federated learning system into M clusters based on a preset clustering algorithm, so that the control module controls the cluster head devices in each cluster to share training data with the member devices in the cluster, which can slow down the data heterogeneity degree among the training data of each distributed device. Then, the federated learning training module collaboratively and iteratively trains a preset initial model with the training data of each distributed device and the central server based on the federated learning algorithm, so as to obtain a target model after federated learning training. Compared with the model training process of federated learning in the related art, in the embodiments of the present invention, by clustering the distributed devices, the cluster head devices share training data with the member devices in the cluster, which slows down the degree of data heterogeneity, thereby improving the convergence speed of the subsequent federated learning algorithm, reducing the communication overhead of federated learning training, and effectively improving the accuracy of the finally trained target model.

[0242] Optionally, the clustering module 501 is specifically used for:

[0243] Establish a privacy constraint graph of the K distributed devices; wherein, the privacy constraint graph includes K nodes corresponding to the K distributed devices and edges for connecting the nodes;

[0244] Calculate the intimacy relationship value between each of the distributed devices and other distributed devices among the K distributed devices as the first attribute value of the edge between the node corresponding to each of the distributed devices and the node corresponding to the other distributed devices; wherein, the intimacy relationship value is used to characterize the trust degree between each of the distributed devices;

[0245] Delete the edges corresponding to the first attribute values less than the privacy threshold in the privacy constraint graph to obtain a privacy communication constraint graph;

[0246] Calculate the earth mover's distance EMD in the system corresponding to the K distributed devices as the attribute value of the K nodes; wherein, the EMD in the system is used to characterize the difference between the distribution of the training data of each of the distributed devices and the distribution of the global data of the K distributed devices.

[0247] In the privacy communication constraint graph, select the corresponding nodes as cluster head nodes in descending order of the attribute values of the nodes until there is an edge between each of the other nodes and at least one cluster head node, obtaining M cluster head nodes; where the other nodes are the nodes among the K nodes excluding the cluster head nodes.

[0248] Calculate the inter-device EMD between each distributed device and other distributed devices among the K distributed devices, as the second attribute value of the edge between the node corresponding to each distributed device and the node corresponding to the other distributed devices; where the inter-device EMD is used to characterize the difference in the distribution of training data between each distributed device.

[0249] In the case where there is an edge between the other nodes and only one cluster head node, divide the other nodes into the cluster where the cluster head node with an edge between it and the other nodes is located.

[0250] In the case where there are edges between the other nodes and at least two cluster head nodes, divide the other nodes into the cluster where the cluster head node with the largest second attribute value corresponding to the edge between it and the other nodes is located.

[0251] Take the distributed devices corresponding to the M cluster head nodes as the cluster head devices of the M clusters, and take the distributed devices corresponding to the other nodes within the clusters where each cluster head node is located as the in-cluster member devices of the M clusters.

[0252] Optionally, the clustering data sharing device 500 of the federated learning system further includes:

[0253] A processing module, configured to calculate the data transmission rate between each distributed device and the other distributed devices, as the third attribute value of the edge between the node corresponding to each distributed device and the node corresponding to the other distributed devices.

[0254] The clustering module 501 is further specifically configured to:

[0255] Delete the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph, and delete the edges corresponding to the third attribute values less than the communication rate threshold, to obtain the privacy communication constraint graph.

[0256] Optionally, the processing module is further configured to:

[0257] Calculate the sharing delay t of each cluster head device sharing training data with the in-cluster member devices s and the training delay t for one round of model training based on the federated learning algorithm FL ;

[0258] Based on t s and tFL Calculate the target shared data volume N of each of the cluster head devices using formula (1) S and the target central processing unit (CPU) frequency f of each of the distributed devices for training the model:

[0259]

[0260] where Ω() represents the number of iterations of training the model in the federated learning system;

[0261] The control module 502 is specifically configured to control each of the cluster head devices to share training data with a data volume of N S with the in-cluster member devices;

[0262] The federated learning training module 503 is specifically configured to: based on the federated learning algorithm, with f as the CPU frequency of each of the distributed devices for training the model, cooperate with the central server to iteratively train a pre-set initial model through the training data of each of the distributed devices, and obtain a target model after federated learning training.

[0263] Optionally, the processing module is further specifically configured to:

[0264] Calculate the sharing delay t using formula (2) s :

[0265]

[0266] where M represents the set of cluster head devices, m represents the m-th cluster head device in the set of cluster head devices, C m represents the set of in-cluster member devices in the cluster where the m-th cluster head device is located, c represents the c-th in-cluster member device in the set of in-cluster member devices, a represents the number of bits occupied by a sample of a training data, represents the data volume of the m-th cluster head device sharing training data with the in-cluster member devices, v m,c represents the data transmission rate between the m-th cluster head device and the c-th in-cluster member device;

[0267] Calculate the training delay t using formula (3) FL :

[0268]

[0269] where represents the downlink delay, represents the update delay, represents the uplink delay.

[0270] Optionally, the control module 502 is further specifically configured to: control the cluster head devices in each of the clusters to share training data with the member devices in the cluster by means of device-to-device (D2D) multicast.

[0271] Figure 6 is a schematic structural diagram of an electronic device provided by the present invention. As Figure 6 shown, the electronic device 600 may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call logic instructions in the memory 630 to execute a method for sharing clustered data based on a federated learning system, where the federated learning system includes K distributed devices and a central server, and K is an integer greater than 1;

[0272] The method includes: dividing the K distributed devices into M clusters based on a pre-set clustering algorithm; where M is an integer less than K, and at least one of the M clusters includes a cluster head device and member devices in the cluster;

[0273] Controlling the cluster head devices in each of the clusters to share training data with the member devices in the cluster;

[0274] Based on a pre-set federated learning algorithm, cooperatively iteratively training a pre-set initial model through the training data of each of the distributed devices and the central server to obtain a target model after federated learning training.

[0275] In addition, when the logic instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0276] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the clustering data sharing method based on the federated learning system provided by the above-mentioned various methods. The federated learning system includes K distributed devices and a central server, where K is an integer greater than 1;

[0277] The method includes: dividing the K distributed devices into M clusters based on a pre-set clustering algorithm; where M is an integer less than K, and at least one of the M clusters includes a cluster head device and in-cluster member devices;

[0278] Controlling the cluster head devices in each of the clusters to share training data with the in-cluster member devices;

[0279] Based on a pre-set federated learning algorithm, collaboratively and iteratively training a pre-set initial model through the training data of each of the distributed devices and the central server to obtain a target model after federated learning training.

[0280] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the clustering data sharing method based on the federated learning system provided by the above-mentioned various methods. The federated learning system includes K distributed devices and a central server, where K is an integer greater than 1;

[0281] The method includes: dividing the K distributed devices into M clusters based on a pre-set clustering algorithm; where M is an integer less than K, and at least one of the M clusters includes a cluster head device and in-cluster member devices;

[0282] Controlling the cluster head devices in each of the clusters to share training data with the in-cluster member devices;

[0283] Based on a pre-set federated learning algorithm, collaboratively and iteratively training a pre-set initial model through the training data of each of the distributed devices and the central server to obtain a target model after federated learning training.

[0284] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0285] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0286] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for sharing clustered data based on a federated learning system, characterized in that The federated learning system includes K distributed devices and a central server, where K is an integer greater than 1; The method includes: Dividing the K distributed devices into M clusters based on a pre-set clustering algorithm; where M is an integer less than K, and there is at least one cluster among the M clusters that includes a cluster head device and in-cluster member devices; Controlling the cluster head devices in each of the clusters to share training data with the in-cluster member devices; Based on a pre-set federated learning algorithm, collaboratively and iteratively training a pre-set initial model through the training data of each of the distributed devices and the central server to obtain a target model after federated learning training; The dividing the K distributed devices into M clusters based on a pre-set clustering algorithm includes: Building a privacy constraint graph of the K distributed devices; where the privacy constraint graph includes K nodes corresponding to the K distributed devices and edges for connecting each of the nodes; Calculate the intimacy relationship value between each of the distributed devices and other distributed devices among the K distributed devices, and use it as the first attribute value of the edge between the node corresponding to each distributed device and the node corresponding to the other distributed devices; wherein, the intimacy relationship value is used to characterize the trust degree between each of the distributed devices; wherein, the intimacy relationship value between the th distributed device and the th distributed device can be calculated by the following formula: ; Among them characterize the intimacy relationship value between the th distributed device and the th distributed device. characterize the number of interactions between the th distributed device and the th distributed device, which can be obtained from prior information in the environment; when a distributed device communicates frequently with another distributed device for a long time, the connection between the distributed devices can be considered reliable. When the th distributed device and the th distributed device have a communication distance exceeding the threshold, it can be considered that there is no device-to-device (D2D) connection between the two distributed devices; Deleting the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph to obtain a privacy communication constraint graph; Calculating the Earth Mover's Distance (EMD) of the system corresponding to the K distributed devices as the attribute values of the K nodes; where the EMD in the system is used to characterize the difference between the distribution of the training data of each of the distributed devices and the distribution of the global data of the K distributed devices; In the privacy communication constraint graph, selecting the corresponding nodes as cluster head nodes in the order of the attribute values of the nodes from largest to smallest until there are edges between all other nodes and at least one cluster head node to obtain M cluster head nodes; where the other nodes are the nodes among the K nodes except the cluster head nodes; Calculating the inter-device EMD between each of the distributed devices and other distributed devices among the K distributed devices as the second attribute values of the edges between the nodes corresponding to each of the distributed devices and the nodes corresponding to the other distributed devices; where the inter-device EMD is used to characterize the difference in the distribution of the training data between each of the distributed devices; In the case where there is only one edge between an other node and a cluster head node, dividing the other node into the cluster where the cluster head node with which there is an edge with the other node is located; In the case where there are edges between an other node and at least two cluster head nodes, dividing the other node into the cluster where the cluster head node corresponding to the largest second attribute value of the edge corresponding to the other node is located; Taking the distributed devices corresponding to the M cluster head nodes as the cluster head devices of the M clusters, and taking the distributed devices corresponding to the other nodes in the clusters where each of the cluster head nodes is located as the in-cluster member devices of the M clusters.

2. The method for sharing clustered data based on a federated learning system according to claim 1, wherein Before deleting the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph to obtain a privacy communication constraint graph, the method further includes: Calculating the data transmission rate between each of the distributed devices and other distributed devices as the third attribute values of the edges between the nodes corresponding to each of the distributed devices and the nodes corresponding to the other distributed devices; Deleting the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph to obtain a privacy communication constraint graph includes: Deleting the edges corresponding to the first attribute values less than the privacy degree threshold in the privacy constraint graph, and deleting the edges corresponding to the third attribute values less than the communication rate threshold to obtain the privacy communication constraint graph.

3. The method for sharing clustered data based on a federated learning system according to claim 1, wherein Before controlling the cluster head devices in each of the clusters to share training data with the intra-cluster member devices, the method further includes: Calculate the sharing delay of each of the cluster head devices sharing training data with the in-cluster member devices and the training delay for one round of model training based on the federated learning algorithm ; Based on and the target shared data volume of each of the cluster head devices is calculated using formula (1) and the target central processing unit (CPU) frequency for training the model by each of the distributed devices : (1) Among them, ( ) represents the number of iterations for training the model of the federated learning system; Controlling the cluster head devices in each of the clusters to share training data with the intra-cluster member devices includes: Control each of the cluster head devices to share the amount of data to the in-cluster member devices as training data; Based on a preset federated learning algorithm, collaboratively and iteratively training a preset initial model with the training data of each of the distributed devices and the central server to obtain a target model after federated learning training includes: Based on the federated learning algorithm, is the CPU frequency for training models for each of the distributed devices, and the initial model preset is collaboratively and iteratively trained with the training data of each of the distributed devices and the central server to obtain the target model after federated learning training.

4. The method for sharing clustered data based on a federated learning system according to claim 3, wherein Calculating the sharing delay for each of the cluster head devices to share training data with the in-cluster member devices and the training delay for one round of model training based on the federated learning algorithm including: Use formula (2) to calculate the shared delay : (2) Among them, represents the set of the cluster head devices, represents the th cluster head device in the set of the cluster head devices, represents the th set of the in-cluster member devices within the cluster where the th cluster head device is located, represents the th in-cluster member device in the set of the in-cluster member devices, represents the number of bits occupied by a sample of a training data, represents the data volume shared by the th cluster head device to the in-cluster member devices, represents the data transmission rate between the th cluster head device and the th in-cluster member device; Using formula (3), calculate the training time delay : (3) Among them, represents the downlink delay, represents the update delay, represents the uplink delay.

5. The method for sharing clustered data based on a federated learning system according to any one of claims 1 to 4, characterized in that, Controlling the cluster head devices in each of the clusters to share training data with the intra-cluster member devices includes: Controlling the cluster head devices in each of the clusters to share training data with the intra-cluster member devices in a device-to-device (D2D) multicast manner.

6. A clustered data sharing device based on a federated learning system, characterized in that, The federated learning system includes K distributed devices and a central server, where K is an integer greater than 1; The device includes: A clustering module, configured to divide the K distributed devices into M clusters based on a preset clustering algorithm; where M is an integer less than K, and at least one of the M clusters includes a cluster head device and intra-cluster member devices; A control module, configured to control the cluster head devices in each of the clusters to share training data with the intra-cluster member devices; A federated learning training module, configured to collaboratively and iteratively train a preset initial model with the training data of each of the distributed devices and the central server based on a preset federated learning algorithm to obtain a target model after federated learning training; The clustering module is specifically configured to establish a privacy constraint graph for the K distributed devices; wherein, the privacy constraint graph includes K nodes corresponding to the K distributed devices, and edges for connecting the nodes; calculate the intimacy relationship values between each of the distributed devices and other distributed devices among the K distributed devices, and use the values as the first attribute values of the edges between the nodes corresponding to each of the distributed devices and the nodes corresponding to the other distributed devices; wherein, the intimacy relationship values are used to characterize the trust levels between the distributed devices; wherein, the intimacy relationship value between the m-th distributed device and the n-th distributed device can be calculated through the following formula: ; where characterizes the intimacy relationship value between the th distributed device and the th distributed device. characterizes the number of interactions between the th distributed device and the th distributed device, which can be obtained from prior information in the environment; when a distributed device communicates frequently with another distributed device for a long time, the connection between the distributed devices can be considered reliable. When the communication distance between the th distributed device and the th distributed device exceeds a threshold , it can be considered that there is no device-to-device (D2D) connection between the two distributed devices; in the privacy constraint graph, delete the edges corresponding to the first attribute values less than the privacy degree threshold to obtain a privacy communication constraint graph; calculate the Earth Mover's Distance (EMD) of the system corresponding to the K distributed devices as the attribute value of the K nodes; where the EMD in the system is used to characterize the difference between the distribution of the training data of each distributed device and the distribution of the global data of the K distributed devices; in the privacy communication constraint graph, select the corresponding nodes as cluster head nodes in descending order of the attribute values of the nodes until there are edges between all other nodes and at least one cluster head node, obtaining M cluster head nodes; where the other nodes are the nodes among the K nodes except the cluster head nodes; calculate the inter-device EMD between each distributed device and the other distributed devices among the K distributed devices as the second attribute value of the edge between the node corresponding to each distributed device and the node corresponding to the other distributed devices; where the inter-device EMD is used to characterize the difference in the distribution of the training data between each distributed device; in the case where the other nodes have an edge with only one cluster head node, divide the other nodes into the cluster where the cluster head node with the edge with the other nodes is located; in the case where the other nodes have edges with at least two cluster head nodes, divide the other nodes into the cluster where the cluster head node with the largest second attribute value corresponding to the edge with the other nodes is located; take the distributed devices corresponding to the M cluster head nodes as the cluster head devices of the M clusters, and take the distributed devices corresponding to the other nodes in the clusters where each cluster head node is located as the intra-cluster member devices of the M clusters.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it implements the clustering data sharing method based on a federated learning system according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by a processor, it implements the clustering data sharing method based on a federated learning system according to any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the clustering data sharing method based on a federated learning system according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Reliable federal learning method, system, device and terminal under clustering network architecture

    CN115150918A

  • Clustering type federated learning driven mechanical fault diagnosis method and device and medium

    CN115438714A