An efficient hierarchical federated learning approach for heterogeneous resources and data

By using the K-means algorithm to group and sort mobile devices in the HFL framework, and using the ring-shaped sequential group pre-aggregation method, the problems of uneven computing efficiency and communication overhead of HFL in heterogeneous resources and data environments are solved, and efficient model training and accuracy improvement are achieved.

CN119728753BActive Publication Date: 2025-05-06HUNAN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510218358.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-06
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In a distributed storage computing environment, HFL faces problems with heterogeneous resources and data, resulting in uneven computing efficiency and large communication overhead, which affects the training accuracy and convergence speed of the model.

Method used

By building an HFL framework, using the K-means algorithm to group mobile devices and sort them according to computing capabilities, the pre-aggregation method within the ring-shaped sequence group is adopted to reduce the communication link with edge servers and cloud servers and improve communication efficiency.

Benefits of technology

It realizes that in heterogeneous resources and data environments, improves the efficiency and accuracy of model training, reduces training delay and communication overhead, and enhances the fusion of knowledge and the convergence of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119728753B_ABST
    Figure CN119728753B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of cloud-edge-end collaborative training, and discloses an efficient hierarchical federated learning method for heterogeneous resources and data, including: grouping the mobile devices MD on each edge server ES side according to computing power and data distribution, so that each group has the greatest data distribution similarity. At the same time, the devices with strong computing power are evenly distributed in different groups to reduce the differences in computing power between groups. At the same time, the device side of the present invention adopts an intra-group circular sequential training method to transmit the trained model to the next connected device, reducing the communication link between the mobile device MD and the edge server ES, and improving the efficiency of model training. The present invention uses communication between devices instead of communication between devices and servers, reduces the number of links for communication with the server, increases the communication bandwidth, and can achieve enhanced knowledge fusion while reducing communication overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cloud-edge-end collaborative training and relates to an efficient hierarchical federated learning method for heterogeneous resources and data. Background Art

[0002] The "cloud-edge-end" collaborative hierarchical federated learning (HFL) avoids the transmission of centralized data by training the model on the local device and sharing and aggregating the model on the server. HFL can complete distributed machine learning training among a large number of users without sharing local private data sets with external devices. It can use the local computing resources on the device side to train the model while protecting data privacy. The specific training process of HFL includes: the central server initializes the global model and sends it to each mobile device MD participating in the training. The mobile device MD trains the global model locally and uploads the updated model gradient to the edge server ES for partial aggregation. Finally, the aggregated model is uploaded to the cloud server CS for global aggregation and the global model is updated. The goal of collaborative learning is achieved by iterating the entire process until the global model converges or reaches the predetermined round.

[0003] However, when applying HFL to an actual distributed storage computing environment, there are many challenges. The first challenge is that there is a high degree of resource heterogeneity and data heterogeneity among large-scale participating devices. On the one hand, the HFL process involves a large number of mobile devices MD, and at the same time, each mobile device MD has significant differences in hardware, network environment, and computing power. This will cause some MDs to complete tasks faster during model training, while mobile device MDs with poor performance will take longer, extending the aggregation time of the global model. On the other hand, the data distribution of each mobile device MD may be different, so there is a problem of non-independent and identically distributed data of the mobile device MDs participating in the training. Non-independent and identically distributed training causes the trained global model to be more biased towards the data features on certain specific MDs, and is not applicable to other MDs.

[0004] The second challenge is that the HFL training process requires all MDs to periodically upload local model parameters or gradients to the ES, which involves a lot of communication overhead. At the same time, the frequent transmission of high-dimensional and large-scale data on limited communication bandwidth may also lead to communication congestion and delay. Many existing methods reduce the amount of data transmitted between MD and ES or CS through device sampling, gradient compression, model pruning, etc., but these methods all reduce communication overhead at the expense of reducing model accuracy, affecting the expressiveness and accuracy of the model.

[0005] The third challenge is that the existing HFL methods use synchronous or asynchronous mechanisms to aggregate model transmission. Synchronous aggregation requires all devices to upload model parameters at the same time, which increases bandwidth burden and prolongs waiting time. The model parameters obtained by the asynchronous aggregation central server during aggregation may come from different training stages, which may reduce the convergence of the global model. Summary of the invention

[0006] The purpose of the present invention is to solve the problem in the prior art that when HFL is applied to actual distributed storage calculations, the MD computing efficiency of mobile devices with different performances is different, and a large amount of communication overhead is occupied during the HFL training process, which increases the bandwidth burden, and to provide an efficient hierarchical federated learning method and system for heterogeneous resources and data.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] An efficient hierarchical federated learning method for heterogeneous resources and data, including:

[0009] Build the HFL framework, initialize the global model parameters inside the cloud server CS in the HFL framework, and send the global model parameters to the edge server ES; the edge server ES sends the initial global model to the mobile device MD participating in the training;

[0010] Based on the periodic data of the mobile device MD processing the sample data, the time for each mobile device MD to perform local model training is obtained;

[0011] Based on the transmission rate at which each mobile device MD uploads the local updated parameter model to the edge server ES, the transmission time when each mobile device MD transmits the parameter model to the edge server ES is obtained; and then the training time of the local model in each training iteration is obtained;

[0012] The edge server ES uses the K-means algorithm to group the mobile devices MD within the service range and sort them according to computing power;

[0013] Each mobile device MD passes the trained model to the next device it is connected to. The device continues to train the model using its local data set and aggregates the received model after the training is completed. When the training is completed on the last device in each group, the device will synchronously upload the locally trained and aggregated final model to the edge server ES for the first layer of aggregation.

[0014] Between the edge server ES and the cloud server CS, each edge server ES uploads part of the aggregated model to the buffer in the cloud, and determines whether the number of models in the buffer reaches the threshold. If so, a global model aggregation operation is performed.

[0015] A further improvement of the present invention is:

[0016] Furthermore, the HFL framework includes a cloud server CS, M edge servers and N mobile devices During operation, each edge server ES is located on a base station covering a local area, serving mobile devices MD within its coverage area, and each mobile device MD can only be connected to one edge server ES.

[0017] Furthermore, based on the periodic data of the mobile device MD processing the sample data, the time for each mobile device MD to perform local model training is obtained, specifically:

[0018] ;

[0019] in, The number of CPU cycles required to process one data sample for the mobile device MD, The number of CPU cycles required to perform one round of local training; The CPU frequency allocated by the mobile device MD for local training.

[0020] Further, based on the transmission rate at which each mobile device MD uploads the locally updated parameter model to the edge server ES, the transmission time when each mobile device MD transmits the parameter model to the edge server ES is obtained, specifically:

[0021] In the HFL system, the mobile device MD is connected to the edge server ES through a wireless network. The mobile device MD updates the local model Send to the edge server ES to which it is connected; the entire communication bandwidth is divided into several frequency bands based on frequency division multiplexing FDMA, and the mobile device MD always occupies this frequency band during the communication process, so that all mobile devices MD occupy different bandwidth resources for communication at the same time; therefore, the transmission rate at which each mobile device MD uploads the local updated parameter model to the edge server ES in the HFL system is:

[0022] ;

[0023] in, is the channel bandwidth, Indicates the number of packets of bandwidth, represents the gain of the channel, represents the transmission power between the mobile device MD and the edge server, represents noise; therefore, the transmission time when the mobile device MD transmits the parameter model to the edge server ES is:

[0024] ;

[0025] in, Indicates the parameter size of the local model, in bits; is the transmission rate;

[0026] The time for the local model to be trained in each training iteration is specifically:

[0027] ;

[0028] in, Calculate time locally;

[0029] In the HFL model training scenario, the communication between the mobile device MD and the edge server ES adopts the synchronous transmission method. In each training iteration, the completion time of the local model training is the longest among all the group training completion times, that is,

[0030] ;

[0031] in, is the number of groups.

[0032] Furthermore, the edge server ES uses the K-means algorithm to group the mobile devices MD within the service range and sort them according to computing power. Specifically, the communication between the mobile devices MD is used instead of the communication between the mobile devices MD and the edge server ES; the mobile devices MD are pre-grouped according to computing power, and devices with smaller differences in computing power are grouped into the same group. Then, the K-means algorithm is used to select devices with smaller data distribution similarity from each group, and the devices are regrouped into one group to obtain the final grouping result.

[0033] Furthermore, the mobile devices MD are pre-grouped according to their computing power, and the devices with smaller computing power differences are grouped into the same group. Then, the K-means algorithm is used to select devices with smaller data distribution similarity from each group and regroup them into one group to obtain the final grouping result. Specifically, in the initialization stage of group training, all devices under each edge server ES are regarded as a whole, and the mobile devices MD with different computing power and similar data distribution are grouped into one group. The number of groups is k, and the goal of grouping is to make each group have the smallest data distribution difference. The data distribution and computing power of each device are represented by a feature vector: Let The data distribution feature vector of each device is: , which can usually be represented by the category distribution or statistical characteristics of data samples; let The device settings The computing power of a device is ; Represent the characteristics of each mobile device MD as a vector ,in, is a weight factor used to balance the impact of data distribution characteristics and computing power characteristics;

[0034] The K-means goal is to minimize the sum of the squared distances from a device to the center of its cluster:

[0035] ;

[0036] in, is the index of the device, is the cluster index, Indicates A clustered set of devices, It is The center vector of each cluster, No. The center vectors of clusters are first randomly selected Device feature vectors are used as initial cluster centers ; For each device , calculate its distance to each cluster center, and assign the device to the cluster center closest to it: ;

[0037] in and are two different cluster center vectors, arrive The distance ratio arrive The distance is closer, for each cluster , calculate the new cluster center is the average value of all MD feature vectors in the cluster:

[0038] ;

[0039] Repeat the above steps until the cluster center no longer changes significantly or the maximum number of iterations is reached, and a group of devices with similar data distribution and balanced computing power is obtained.

[0040] Furthermore, after the mobile devices MD are grouped, the method further includes: the mobile device MD performs local training to update the gradient, and transmits the model gradient to the next connected mobile device MD for pre-aggregation in the ring sequence group until the last device in each group, specifically:

[0041] Each mobile device has its own specific user data set , whose size is ; Mobile devices participate in HFL collaborative training to optimize the global model; Among them, the optimization goal of local training is:

[0042] ;

[0043] in, Indicates the device In local dataset The loss function during the training process is: is the first samples; Represents the sum of loss functions of all devices.

[0044] In HFL, each edge server communicates with mobile devices within its coverage area, represented as , the size of each group is ; The data set of mobile devices contained in each edge server is represented as ; The MDs in each edge server group are grouped according to their computing power and data distribution characteristics, so that MDs with similar computing power and small feature distribution differences are evenly divided into different groups, expressed as , the size of each group is ; When the local model training reaches the preset number of iterations, all selected nodes within the service range of each edge server will upload the gradient update and aggregate all received gradient updates to obtain a new local gradient;

[0045] Obtain the optimal global model to minimize the global dataset The global loss function trained on ; where the global dataset is represented as , the size is , is the sum of the data sets of all devices within the service range of the edge server ES; when the edge server ES completes the aggregation, the gradient aggregated by the edge server ES is uploaded to the cloud server CS for aggregation, and finally the global gradient is obtained; the optimization target of the cloud server CS is expressed as:

[0046] ;

[0047] in, is the model gradient, is a single sample, is the data sample size for each group, is the data sample size of all groups; the parameter updating method is the gradient descent method, and the updating process of the model parameters is:

[0048] ;

[0049] in, represents the training round, is the learning rate when the local model parameters are updated;

[0050] When training the model in the HFL system, it is necessary to ensure that the training time does not exceed the required task deadline; the goal of this model training is to minimize the delay of model training and improve the training accuracy of the model without affecting the convergence performance of the model in the environment of model hierarchical aggregation; minimizing the loss and training delay in the model training process is expressed as:

[0051] ;

[0052] The training delay of the model cannot exceed the defined deadlines, specifically:

[0053] ;

[0054] in, Indicates the iteration round in the training process;

[0055] (1);

[0056] (2);

[0057] Formula 1 indicates that after all mobile devices MD under each edge server ES are regrouped, each mobile device MD can only be divided into one group; Formula 2 indicates that all mobile devices MD under each ES should participate in the group training.

[0058] Furthermore, each mobile device MD passes the trained model to the next device to which it is connected. The device continues to train the model using its local data set and aggregates the received model after the training is completed. When the training is completed on the last device in each group, the device synchronously uploads the locally trained and aggregated final model to the edge server ES for the first layer of aggregation. Specifically:

[0059] The training steps of each mobile device MD in each round include:

[0060] (1) The time required for the mobile device MD to train its own model locally The data are sent to the edge server ES, and then the edge server ES groups the devices according to time and sorts the mobile devices MD in each group according to the predetermined requirements;

[0061] (2) The mobile device MD receives the global model from the edge server ES and performs local training to update the model;

[0062] (3) The mobile device MD transmits the updated model to the connected mobile device MD based on the pre-set training sequence;

[0063] (4) After the current mobile device MD receives the local training model sent by its previous mobile device MD, it uses its local data set to continue local training on the received model; when the model is transmitted to a location without a successor device, the device uploads the locally updated model to the edge server ES for synchronization aggregation.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] The present invention groups the mobile devices MD on each edge server ES side according to computing power and data distribution, so that each group has the greatest data distribution similarity. At the same time, devices with strong computing power are evenly distributed in different groups to reduce the differences in computing power between groups. At the same time, the device side of the present invention adopts an intra-group circular sequential training method to transmit the trained model to the next connected device, reducing the communication link between the mobile device MD and the edge server ES, and improving the efficiency of model training. The present invention uses communication between devices instead of communication between devices and servers, reduces the number of links for communication with the server, increases the communication bandwidth, and can achieve enhanced knowledge fusion while reducing communication overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0067] Figure 1 A schematic diagram of a process flow of an efficient hierarchical federated learning method for heterogeneous resources and data of the present invention;

[0068] Figure 2 A schematic diagram of the SAEHFL method for efficient hierarchical federated learning of heterogeneous resources and data according to the present invention;

[0069] Figure 3 It is a schematic diagram of the structure of the efficient hierarchical federated learning system for heterogeneous resources and data of the present invention;

[0070] Figure 4 (a) is a diagram showing the relationship between training rounds and convergence under the MNIST dataset;

[0071] Figure 4(b) is a diagram showing the relationship between training rounds and convergence under the FEMNIST dataset;

[0072] Figure 4 (c) is a diagram showing the relationship between training rounds and convergence under the CIFAR10 dataset;

[0073] Figure 5 This is a comparison chart of the training time spent by different experimental methods on the MNIST dataset in the same training round;

[0074] Figure 6 This is a comparison chart of the training time spent by different experimental methods on the FEMNIST dataset in the same training round;

[0075] Figure 7 This is a comparison chart of the training time spent by different experimental methods on the CIFAR10 dataset in the same training round;

[0076] Figure 8 This is a comparison chart of the accuracy achieved by different experimental methods in the same training rounds under the MNIST dataset;

[0077] Fig. 9 This is a comparison chart of the accuracy achieved by different experimental methods in the same training rounds on the FEMNIST dataset;

[0078] Fig.10 This is a comparison chart of the accuracy achieved by different experimental methods in the same training rounds on the CIFAR10 dataset;

[0079] Fig.11 This is a diagram showing the relationship between the training time and the accuracy required to achieve the same target using different experimental methods on the MNIST dataset.

[0080] Fig.12 This is a diagram showing the relationship between the training time and the accuracy required to achieve the same target using different experimental methods on the FEMNIST dataset.

[0081] Fig.13 This is a diagram showing the relationship between the training time and the accuracy required to achieve the same target using different experimental methods on the CIFAR10 dataset.

[0082] Fig.14 This is a comparison of the training time taken by SAEHFL, SAEHFL-Randomization, SAEHFL-All Communications, and SAEHFL-Synchronous for the same number of epochs on the MNIST dataset;

[0083] Fig.15 This is a comparison of the training time taken by SAEHFL, SAEHFL-Randomization, SAEHFL-All Communications, and SAEHFL-Synchronous for the same number of epochs on the FEMNIST dataset;

[0084] Fig.16 This is a comparison of the training time taken by SAEHFL, SAEHFL-Randomization, SAEHFL-All Communications, and SAEHFL-Synchronous for the same number of epochs on the CIFAR10 dataset;

[0085] Fig.17 This is a comparison chart of the accuracy achieved by SAEHFL, SAEHFL-Randomization, SAEHFL-All Communications, and SAEHFL-Synchronous when trained for the same number of epochs on the MNIST dataset;

[0086] Fig.18 This is a comparison chart of the accuracy achieved by SAEHFL, SAEHFL-Randomization, SAEHFL-All Communications, and SAEHFL-Synchronous when trained for the same number of epochs on the FEMNIST dataset;

[0087] Fig.19 This is a comparison of the accuracy achieved by SAEHFL, SAEHFL-Randomization, SAEHFL-All Communications, and SAEHFL-Synchronous when trained for the same number of epochs on the CIFAR10 dataset. DETAILED DESCRIPTION

[0088] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0089] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0090] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0091] In the description of the embodiments of the present invention, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. indicate an orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed when in use, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0092] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", which does not mean that the structure must be completely horizontal, but can be slightly tilted.

[0093] In the description of the embodiments of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0094] The present invention is further described in detail below in conjunction with the accompanying drawings:

[0095] See also Figure 1 The present invention discloses an efficient hierarchical federated learning method for heterogeneous resources and data, comprising:

[0096] S101, constructing an HFL framework, initializing the global model parameters inside the cloud server CS in the HFL framework, and sending the global model parameters to the edge server ES; the edge server ES sends the initial global model to the mobile device MD participating in the training;

[0097] The HFL framework consists of a cloud server CS and M edge servers and N mobile devices During operation, each edge server ES is located on a base station covering a local area, serving mobile devices MD within its coverage area, and each mobile device MD can only be connected to one edge server ES.

[0098] S102, based on the periodic data of the mobile device MD processing the sample data, obtaining the time for each mobile device MD to perform local model training;

[0099] ;

[0100] in, The number of CPU cycles required to process one data sample for the mobile device MD, The number of CPU cycles required to perform one round of local training; The CPU frequency allocated by the mobile device MD for local training.

[0101] S103, based on the transmission rate at which each mobile device MD uploads the local updated parameter model to the edge server ES, obtain the transmission time when each mobile device MD transmits the parameter model to the edge server ES; and then obtain the training time of the local model in each training iteration process;

[0102] In the HFL system, the mobile device MD is connected to the edge server ES through a wireless network. The mobile device MD updates the local model Send to the edge server ES to which it is connected; the entire communication bandwidth is divided into several frequency bands based on frequency division multiplexing FDMA, and the mobile device MD always occupies this frequency band during the communication process, so that all mobile devices MD occupy different bandwidth resources for communication at the same time; therefore, the transmission rate at which each mobile device MD uploads the local updated parameter model to the edge server ES in the HFL system is:

[0103] ;

[0104] in, is the channel bandwidth, Indicates the number of packets of bandwidth, represents the gain of the channel, represents the transmission power between the mobile device MD and the edge server, represents noise; therefore, the transmission time when the mobile device MD transmits the parameter model to the edge server ES is:

[0105] ;

[0106] in, Indicates the parameter size of the local model, in bits; is the transmission rate;

[0107] The time for the local model to be trained in each training iteration is specifically:

[0108] ;

[0109] in, Calculate time locally;

[0110] In the HFL model training scenario, the communication between the mobile device MD and the edge server ES adopts the synchronous transmission method. In each training iteration, the completion time of the local model training is the longest among all the group training completion times, that is,

[0111] ;

[0112] in, is the number of groups.

[0113] S104, the edge server ES uses the K-means algorithm to group the mobile devices MD within the service range and sort them according to computing capabilities;

[0114] The communication between mobile devices MD is used instead of the communication between the mobile devices MD and the edge server ES; the mobile devices MD are pre-grouped according to their computing power, and the devices with smaller differences in computing power are grouped into the same group. Then, the K-means algorithm is used to select devices with smaller data distribution similarity from each group and regroup them into one group to obtain the final grouping result.

[0115] Mobile devices MD are pre-grouped according to computing power, and devices with smaller computing power differences are grouped into the same group. Then, the K-means algorithm is used to select devices with smaller data distribution similarity from each group and regroup them into one group to obtain the final grouping result. Specifically, in the initialization stage of group training, all devices under each edge server ES are regarded as a whole, and mobile devices MD with different computing power and similar data distribution are grouped into one group. The number of groups is k, and the goal of grouping is to make each group have the smallest data distribution difference; the data distribution and computing power of each device are represented by a feature vector: let The data distribution feature vector of each device is: , which can usually be represented by the category distribution or statistical characteristics of data samples; let The device settings Devices ; Represent the characteristics of each mobile device MD as a vector ,in, is a weight factor used to balance the impact of data distribution characteristics and computing power characteristics;

[0116] The K-means goal is to minimize the sum of the squared distances from a device to the center of its cluster:

[0117] ;

[0118] in, Indicates A clustered set of devices, It is The center vector of each cluster is randomly selected. Device feature vectors are used as initial cluster centers ; For each device , calculate its distance to each cluster center, and assign the device to the cluster center closest to it:

[0119] ;

[0120] in and are two different cluster center vectors, arrive The distance ratio arrive The distance is closer, for each , calculate the new cluster center is the average value of all MD feature vectors in the cluster:

[0121] ;

[0122] Repeat the above steps until the cluster center no longer changes significantly or the maximum number of iterations is reached, and a group of devices with similar data distribution and balanced computing power is obtained.

[0123] After the mobile devices MD are grouped, the method further includes: the mobile device MD performs local training to update the gradient, and transmits the model gradient to the next connected mobile device MD for pre-aggregation in the ring sequence group until the last device in each group, specifically:

[0124] Each mobile device has its own specific user data set , whose size is ; Mobile devices participate in HFL collaborative training to optimize the global model; Among them, the optimization goal of local training is:

[0125] ;

[0126] in, Indicates the device In local dataset The loss function during the training process is: is the first samples;

[0127] In HFL, each edge server communicates with mobile devices within its coverage area, represented as , the size of each group is ; The data set of mobile devices contained in each edge server is represented as ; The MDs in each edge server group are grouped according to their computing power and data distribution characteristics, so that MDs with similar computing power and small feature distribution differences are evenly divided into different groups, expressed as , the size of each group is ; When the local model training reaches the preset number of iterations, all selected nodes within the service range of each edge server will upload the gradient update and aggregate all received gradient updates to obtain a new local gradient;

[0128] Obtain the optimal global model to minimize the global dataset The global loss function trained on ; where the global dataset is represented as , the size is , is the sum of the data sets of all devices within the service range of the edge server ES; when the edge server ES completes the aggregation, the gradient aggregated by the edge server ES is uploaded to the cloud server CS for aggregation, and finally the global gradient is obtained; the optimization target of the cloud server CS is expressed as:

[0129] ;

[0130] in, is the model gradient, is a single sample, is the data sample size for each group, is the data sample size of all groups; the parameter updating method is the gradient descent method, and the updating process of the model parameters is:

[0131] ;

[0132] in, represents the training round, is the learning rate when the local model parameters are updated;

[0133] When training the model in the HFL system, it is necessary to ensure that the training time does not exceed the required task deadline; the goal of this model training is to minimize the delay of model training and improve the training accuracy of the model without affecting the convergence performance of the model in the environment of model hierarchical aggregation; minimizing the loss and training delay in the model training process is expressed as:

[0134] ;

[0135] The training delay of the model cannot exceed the defined deadlines, specifically:

[0136] ;

[0137] in, Indicates the iteration round in the training process;

[0138] ;

[0139] The above formulas respectively indicate that after all mobile devices MD under each edge server ES are regrouped, each mobile device MD can only be grouped into one group; and that all mobile devices MD under each edge server ES should participate in group training.

[0140] S105, each mobile device MD transmits the trained model to the next device it is connected to, and the device continues to train the model using its local data set and aggregates the received model after the training is completed; until the last device in each group is trained, the device will synchronously upload the final model trained locally and aggregated to the edge server ES for the first layer of aggregation;

[0141] At the beginning of each round of training, the cloud server CS sends the initialized global model to each edge server ES, and the edge server ES sends the initial global model to each mobile device MD. After each training iteration T times, the edge server ES partially aggregates all the received local update models; the training steps of each mobile device MD in each round include: (1) The time required for the mobile device MD to train its own model locally The data are sent to the edge server ES, and then the edge server ES groups the devices according to time and sorts the mobile devices MD in each group according to the predetermined requirements;

[0142] (2) The mobile device MD receives the global model from the edge server ES and performs local training to update the model;

[0143] (3) The mobile device MD transmits the updated model to the connected mobile device MD based on the pre-set training sequence;

[0144] (4) After the current mobile device MD receives the local training model sent by its previous mobile device MD, it uses its local data set to continue local training on the received model; when the model is transmitted to a location without a successor device, the device uploads the locally updated model to the edge server ES for synchronization aggregation.

[0145] S105 also includes: during the model transmission process, the bandwidth resource determines the speed of the model transmission. In order to ensure the reasonable allocation of resources and improve the efficiency of model transmission, a bandwidth allocation method based on code division multiplexing is proposed. Figure 2 , SAEHFL groups the devices within the edge server ES according to computing power during local training, ensuring that the average computing power difference between all groups under a single edge server ES is minimal, thereby ensuring that the training time difference of each group is minimal. Therefore, the idea of ​​code division multiplexing is used to evenly distribute the communication bandwidth under a single edge server ES to the communication nodes connected to the edge server ES to complete the model transmission. SAEHFL allocates the same communication bandwidth to each group. Specifically, when each edge server ES includes 16 mobile devices MD, divided into 4 groups, the number of mobile devices MD in each group is 4. At the same time, the communication bandwidth is 16 units, and the traditional communication method evenly distributes the communication bandwidth to all mobile devices MD, that is, each mobile device MD obtains 1 unit of communication bandwidth. However, SAEHFL adopts a group ring sequential communication aggregation method to reduce the communication link between the mobile device MD and the edge server ES, and increases the communication bandwidth between a single mobile device MD and the edge server ES through communication aggregation between devices. Each group contains 4 units of communication bandwidth. The mobile device MD that directly communicates with the edge server ES can directly use 4 units of communication bandwidth to upload the aggregated model in the group to the edge server ES, reducing the transmission time and improving the communication efficiency.

[0146] S106, between the edge server ES and the cloud server CS, each edge server ES uploads a partially aggregated model to the buffer in the cloud, and determines whether the number of models in the buffer reaches a threshold. If so, a global model aggregation operation is performed.

[0147] See also Figure 3 The present invention discloses an efficient hierarchical federated learning system for heterogeneous resources and data, including:

[0148] An initialization module, wherein the initialization module constructs an HFL framework, initializes global model parameters within the cloud server CS in the HFL framework, and sends the global model parameters to the edge server ES; the edge server ES sends the initial global model to the mobile device MD participating in the training;

[0149] A first acquisition module, wherein the first acquisition module acquires the time for each mobile device MD to perform local model training based on the periodic data of the mobile device MD processing the sample data;

[0150] A second acquisition module, wherein the second acquisition module acquires the transmission time when each mobile device MD transmits the parameter model to the edge server ES based on the transmission rate at which each mobile device MD uploads the local updated parameter model to the edge server ES; and further acquires the training time of the local model in each training iteration process;

[0151] A grouping module, wherein the edge server ES of the grouping module groups the mobile devices MD within the service range using a K-means algorithm and sorts them according to computing capabilities;

[0152] Aggregation module, in which each mobile device MD transmits the trained model to the next device to which it is connected. The device continues to train the model using its local data set and aggregates the received model after the training is completed. When the training is completed on the last device in each group, the device synchronously uploads the locally trained and aggregated final model to the edge server ES for the first layer of aggregation.

[0153] The judgment module is located between the edge server ES and the cloud server CS. Each edge server ES uploads a partially aggregated model to a buffer in the cloud to determine whether the number of models in the buffer reaches a threshold. If so, a global model aggregation operation is performed.

[0154] Example:

[0155] In this embodiment, the effectiveness of SAEHFL is verified from five aspects: the buffer size of SAEHFL under different data sets, the convergence of SAEHFL under different data sets, the comparison of training time of different methods, the training time to achieve the target accuracy, and the training epochs to achieve the target accuracy. The comparison methods include:

[0156] (1) SFedAvg: All devices participate in local training and perform synchronous aggregation HFL method. When the device with faster training completes the local model training, it needs to wait for other devices to complete local training before uploading the model to the server;

[0157] (2) SAFedAvg: All devices participate in local training and perform asynchronous aggregation in the HFL method. After each device completes local training, it is immediately uploaded to the server for aggregation and a new global aggregation model is obtained instantly;

[0158] (3) Random: In each round of training, the devices participating in the training are randomly selected and the synchronous aggregation HFL method is executed;

[0159] (4) FedProx: All devices participate in local training and execute the synchronous aggregation HFL method, but the local training rounds between devices are inconsistent. Devices with long training time have fewer training rounds, while devices with short training time have more training rounds.

[0160] (5) HFLCHFL: The model is trained by the local computing time, communication time and frequency of the device. In this embodiment, the learning rate is one of the factors that affect the speed of model convergence and the final accuracy. A learning rate that is too high may cause the model to skip the optimal solution when updating, fail to converge or converge to a suboptimal solution. Although a learning rate that is too low has a higher probability of obtaining the optimal solution, it will result in a long training time.

[0161] Semi-asynchronous communication allows the aggregation and update of the model to be triggered after the data of a part of the devices arrives, without waiting for all clients, which improves the communication and training efficiency of the system. The hierarchical federated learning system using semi-asynchronous communication contains a buffer in the cloud server to store the model or gradient updates sent by the edge server to the cloud server, and performs global aggregation operations when the updates stored in the buffer reach their threshold. Therefore, the buffer size in semi-asynchronous communication determines the number of received updates that need to be waited for to update the model on the server, and this choice directly affects the communication and training efficiency of the system. In order to balance the accuracy and training time of the model, Table 1 fixes the buffer size to 10%, 30%, 50%, 70%, and 90% (the ratio of the number of gradient updates received by the current cloud server to the total number of gradient updates), and records the training time and the accuracy of the final model when the training is fixed for 400 epochs. According to the experimental results, when the buffer is too small, the model update may be too dependent on the data of a few clients, resulting in deviations and low model accuracy. When the buffer is too large, although the stability and accuracy of the model update are improved, the delay and system waiting time are increased, and the update frequency and system robustness are reduced. In summary, in order to achieve a balance between model accuracy and training time, the buffer size on the MNIST, FEMNIST, and CIFAR10 datasets is set to 50%, and all the following experiments are completed under this setting.

[0162] Table 1 Comparison of buffer size on accuracy and training time under different datasets

[0163]

[0164] In this embodiment, convergence is an important indicator to measure the generalization performance of the model under different data distributions. In HFL, data from different MDs often have the characteristics of Non-IID, that is, there are significant differences in the data distribution between MDs. This data imbalance and heterogeneity may cause the model to perform poorly on some MDs, or even unstable convergence behavior. Therefore, the present invention uses non-independent and identically distributed data to verify the convergence of SAEHFL under different data sets. The results are shown in Figure 2. Figure 4 (a) Figure 4 (b) and Figure 4 (c) SAEHFL can adapt well to the three data sets of MNIST, FEMNIST, and CIFAR10, and its Loss value can converge to a lower range. Although SAEHFL uses a semi-asynchronous communication method between ES and CS, which will reduce the convergence speed. However, in the device training part, it groups devices with similar data distribution into one group for training aggregation. The update direction of the device model in the group is relatively consistent, which helps to produce smoother aggregation results, reduce the update differences of the model in the group, and reduce the global model fluctuations caused by data heterogeneity.

[0165] The present invention uses SAEHFL to train for 400 epochs on the datasets MNIST, FEMNIST, and CIFAR10, and records the training time required for the fixed epochs and the model accuracy achieved. Figure 5 , Figure 6 , Figure 7 , Figure 8 , Fig. 9 , Fig.10 The experimental results show that the training time required for SAEHFL to train 400 epochs is 1263.91s longer than that of SAFedAvg, but 3048.02s, 2794.08s, 3035.16s, and 2779.6s shorter than that of SFedAvg, Random, FedProx, and HFLCHFL. This is because SAFedAvg allows the server to aggregate and update immediately after receiving the update from the device without waiting for the completion of all devices, reducing the delay caused by synchronization blocking, allowing training to proceed more quickly, but its final accuracy is lower than all other methods. For example, when training for 200 epochs, the accuracy of SAEHFL is 0.966, while the accuracy of SFedAvg, SAFedAvg, Random, FedProx, and HFLCHFL is 0.979, 0.789, 0.95, 0.95, and 0.954. In summary, SAEHFL achieves a balance between training time and model accuracy. It uses semi-asynchronous communication to reduce waiting time during training. At the same time, it combines group training to improve the convergence speed of the model and ensure model accuracy.

[0166] The target accuracy is set to 0.75, 0.8, 0.85, 0.9, and 0.95 on the datasets MNIST, FEMNIST, and CIFAR10, respectively, and the training time required for SAEHFL to achieve the target accuracy is verified. Fig.11 , Fig.12 and Fig.13 , when the target accuracy of the MNIST dataset is 0.9 and 0.95, the target accuracy of the FEMNIST dataset is 0.8 and 0.85, and the target accuracy of the CIFAR10 dataset is 0.7, the time of SAFedAvg is blank. This is because after a fixed run of 400 epochs, SAFedAvg cannot reach the target accuracy. On the MNIST dataset, when the target accuracy is 0.9, the training time of SAEHFL is 564.39s, and the training time of SFedAvg, SAFedAvg, Random, FedProx, and HFLCHFL are 707.3s, 918.38s, 782.66s, 1620.36s, and 656.58s, respectively. Therefore, the training time of SAEHFL is reduced by 142.91s, 353.99s, 218.27s, 1055.97s, and 92.19s, respectively. This is because SAEHFL groups devices with similar data distribution, which allows the training and model aggregation in each group to converge faster and reduces the number of communication rounds within the group. At the same time, SAEHFL only needs one specific device in each group to communicate with the server, replacing the communication between devices and the server with communication between devices, reducing the communication links between devices and the server, increasing the communication bandwidth of each communication device, and improving communication efficiency.

[0167] In this embodiment, the number of communication epochs required for SAEHFL to reach the target accuracy is compared with SFedAvg, SAFedAvg, Random, FedProx, and HFLCHFL on the data sets MNIST, FEMNIST, and CIFAR10. According to the results in Table 2, the target accuracy is set to 0.75, 0.8, 0.85, 0.9, and 0.95 under the MNIST data set. When the target accuracy is 0.9, SAEHFL requires 66 epochs, but SFedAvg, SAFedAvg, Random, FedProx, and HFLCHFL require 62, 543, 113, 127, and 71, respectively. The reason why SAEHFL takes more epochs to reach the target accuracy than SFedAvg is that SAEHFL uses semi-asynchronous communication between ES and CS, which slows down the convergence speed and requires more communication epochs to aggregate more update information and improve model accuracy. However, the number of epochs that SAEHFL reaches the target accuracy is lower than that of the other four methods, because SAEHFL adopts group training aggregation, first aggregating local models within the group, and then aggregating models across groups at the global level, reducing the excessive impact of a single device on the global model, making the aggregated model more balanced and stable, alleviating the convergence fluctuations caused by data distribution differences, and to a certain extent avoiding the problem of slow convergence or falling into local optimality under NonIID data. On the other hand, SAEHFL uses full device participation for training, avoiding the decrease in accuracy caused by information loss.

[0168] Table 2:

[0169]

[0170] The implementation of the SAEHFL algorithm includes three steps: performing clustering grouping according to the resource heterogeneity and data heterogeneity of the MD to ensure that the training completion time of different groups is close while minimizing the data distribution differences within the group. The MDs within the group are aggregated in a ring order to reduce the communication links between the MD and the ES and increase the communication bandwidth of each link. SAEHFL adopts synchronous communication between different groups within the ES, asynchronous communication between different ESs, and semi-asynchronous communication between ES and CS, which not only ensures the convergence speed of the model but also reduces the waiting time. Because, in order to verify the effectiveness of these three methods, the model accuracy and training time of SAEHFL are compared with SAEHFL-Randomization, SAEHFL-All Communications, and SAEHFL-Synchronous training for 400 epochs. The results are as follows Fig.14 , Fig.15 , Fig.16 , Fig.17 , Fig.18 , Fig.19 As shown. According to the results, the model accuracy obtained by training SAEHFL-Synchronous for 400 epochs is slightly higher than that of other methods, but the time required is more than that of other methods. The training time of SAEHFL-Randomization is shorter than that of SAEHFL-AllCommunications and SAEHFL-Synchronous, but its accuracy is lower. Although the accuracy of SAEHFL is lower than that of SAEHFL-Synchronous, its accuracy is still at a high level, and the reduction in its training time is much greater than the difference in accuracy. In summary, SAEHFL can improve the training speed of the model while ensuring the accuracy of the model, with less overhead.

[0171] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, a person skilled in the art should understand that the specific implementation modes of the present invention can still be modified or replaced by equivalents, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be included in the scope of protection of the claims of the present invention.

Claims

1. An efficient hierarchical federated learning method for heterogeneous resources and data, characterized in that: include: Build the HFL framework, initialize the global model parameters inside the cloud server CS in the HFL framework, and send the global model parameters to the edge server ES; The edge server ES sends the initial global model to the mobile devices MD participating in the training; Based on the periodic data of the mobile device MD processing the sample data, the time for each mobile device MD to perform local model training is obtained; Based on the transmission rate at which each mobile device MD uploads the local updated parameter model to the edge server ES, the transmission time when each mobile device MD transmits the parameter model to the edge server ES is obtained; and then the training time of the local model in each training iteration is obtained; The edge server ES uses the K-means algorithm to group the mobile devices MD within the service range and sort them according to computing power; Each mobile device MD passes the trained model to the next device it is connected to. The device continues to train the model using its local data set and aggregates the received model after the training is completed. When the training is completed on the last device in each group, the device will synchronously upload the locally trained and aggregated final model to the edge server ES for the first layer of aggregation. Between the edge server ES and the cloud server CS, each edge server ES uploads part of the aggregated model to the buffer in the cloud, and determines whether the number of models in the buffer reaches the threshold. If so, a global model aggregation operation is performed.

2. The efficient hierarchical federated learning method for heterogeneous resources and data according to claim 1, characterized in that: The HFL framework includes a cloud server CS, M edge servers and N mobile devices ; During operation, each edge server ES is located on a base station covering a local area, serving mobile devices MD within its coverage, and each mobile device MD can only be connected to one edge server ES.

3. The efficient hierarchical federated learning method for heterogeneous resources and data according to claim 2, characterized in that: The periodic data of processing sample data by the mobile device MD is used to obtain the time for each mobile device MD to perform local model training, specifically: ; in, The number of CPU cycles required to process one data sample for the mobile device MD, The number of CPU cycles required to perform one round of local training; The CPU frequency allocated by the mobile device MD for local training.

4. The efficient hierarchical federated learning method for heterogeneous resources and data according to claim 3, characterized in that: The transmission time when each mobile device MD transmits the parameter model to the edge server ES is obtained based on the transmission rate at which each mobile device MD uploads the local updated parameter model to the edge server ES, specifically: In the HFL system, the mobile device MD is connected to the edge server ES through a wireless network. The mobile device MD updates the local model Sent to the edge server ES to which it is connected; Based on frequency division multiplexing FDMA, the entire communication bandwidth is divided into several frequency bands. The mobile device MD always occupies this frequency band during the communication process, so that all mobile devices MD occupy different bandwidth resources for communication at the same time; therefore, the transmission rate at which each mobile device MD uploads the local updated parameter model to the edge server ES in the HFL system is: ; in, is the channel bandwidth, Indicates the number of packets of bandwidth, represents the gain of the channel, represents the transmission power between the mobile device MD and the edge server, represents noise; therefore, the transmission time when the mobile device MD transmits the parameter model to the edge server ES is: ; in, Indicates the parameter size of the local model, in bits; is the transmission rate; The time for the local model to be trained in each training iteration is specifically: ; in, Calculate time locally; In the HFL model training scenario, the communication between the mobile device MD and the edge server ES adopts the synchronous transmission method. In each training iteration, the completion time of the local model training is the longest among all the group training completion times, that is, ; in, is the number of groups.

5. The efficient hierarchical federated learning method for heterogeneous resources and data according to claim 4, characterized in that: The edge server ES uses the K-means algorithm to group the mobile devices MD within the service range and sort them according to computing power, specifically: using the communication between the mobile devices MD instead of the communication between the mobile devices MD and the edge server ES; pre-grouping the mobile devices MD according to computing power, grouping devices with smaller differences in computing power into the same group, and then using the K-means algorithm to select devices with smaller data distribution similarity from each group, and regroup them into one group to obtain the final grouping result.

6. The efficient hierarchical federated learning method for heterogeneous resources and data according to claim 4, characterized in that: The mobile devices MD are pre-grouped according to computing power, and devices with smaller computing power differences are grouped into the same group. Then, the K-means algorithm is used to select devices with smaller data distribution similarity from each group and regroup them into one group to obtain the final grouping result. Specifically, in the initialization stage of group training, all devices under each edge server ES are regarded as a whole, and mobile devices MD with different computing power and similar data distribution are grouped into one group. The number of groups is k, and the goal of grouping is to make each group have the smallest data distribution difference. The data distribution and computing power of each device are represented by a feature vector: Let The data distribution feature vector of each device is: , which can usually be expressed by the category distribution or statistical characteristics of data samples; Set up The computing power of a device is ; Represent the characteristics of each mobile device MD as a vector ,in, is a weight factor used to balance the impact of data distribution characteristics and computing power characteristics; The K-means goal is to minimize the sum of the squared distances from a device to the center of its cluster: ; in, is the index of the device, is the cluster index, Indicates A clustered set of devices, It is The center vector of each cluster, No. The center vectors of clusters are first randomly selected Device feature vectors are used as initial cluster centers ; For each device , calculate its distance to each cluster center, and assign the device to the cluster center closest to it: ; in and are two different cluster center vectors, arrive The distance ratio arrive The distance is closer, for each cluster , calculate the new cluster center as the average value of all MD feature vectors in the cluster: ; Repeat the above steps until the cluster center no longer changes significantly or the maximum number of iterations is reached, and a group of devices with similar data distribution and balanced computing power is obtained.

7. The efficient hierarchical federated learning method for heterogeneous resources and data according to claim 6, characterized in that: After the mobile devices MD are grouped, the method further includes: the mobile device MD performs local training to update the gradient, and transmits the model gradient to the next connected mobile device MD for pre-aggregation in the ring sequence group until the last device in each group, specifically: Each mobile device has its own specific user data set , whose size is ; Mobile devices participate in HFL collaborative training to optimize the global model; Among them, the optimization goal of local training is: ; in, Indicates the device In local dataset The loss function during the training process is: is the first samples; Represents the sum of loss functions of all devices; In HFL, each edge server communicates with mobile devices within its coverage area, represented as , the size of each group is ; The data set of mobile devices contained in each edge server is represented as ; The MDs in each edge server group are grouped according to their computing power and data distribution characteristics, so that MDs with similar computing power and small feature distribution differences are evenly divided into different groups, expressed as , the size of each group is ; When the local model training reaches the preset number of iterations, all selected nodes within the service range of each edge server will upload the gradient update and aggregate all received gradient updates to obtain a new local gradient; Obtain the optimal global model to minimize the global dataset The global loss function trained on ; where the global dataset is represented as , the size is , is the sum of the data sets of all devices within the service range of the edge server ES; when the edge server ES completes the aggregation, the gradient aggregated by the edge server ES is uploaded to the cloud server CS for aggregation, and finally the global gradient is obtained; the optimization target of the cloud server CS is expressed as: ; in, is the model gradient, is a single sample, is the data sample size for each group, is the data sample size of all groups; the parameter updating method is the gradient descent method, and the updating process of the model parameters is: ; in, represents the training round, is the learning rate when the local model parameters are updated; When training the model in the HFL system, it is necessary to ensure that the training time does not exceed the required task deadline; the goal of this model training is to minimize the delay of model training and improve the training accuracy of the model without affecting the convergence performance of the model in the environment of model hierarchical aggregation; minimizing the loss and training delay in the model training process is expressed as: ; The training delay of the model cannot exceed the defined deadlines, specifically: ; in, Indicates the iteration round in the training process; (1); (2); Formula 1 indicates that after all mobile devices MD under each edge server ES are regrouped, each mobile device MD can only be divided into one group; Formula 2 indicates that all mobile devices MD under each ES should participate in the group training.

8. The efficient hierarchical federated learning method for heterogeneous resources and data according to claim 7 is characterized in that: Each mobile device MD passes the trained model to the next device to which it is connected. The device continues to train the model using its local data set and aggregates the received model after the training is completed. When the training is completed on the last device in each group, the device synchronously uploads the locally trained and aggregated final model to the edge server ES for the first layer of aggregation. Specifically: The training steps of each mobile device MD in each round include: (1) The time required for the mobile device MD to train its own model locally The data are sent to the edge server ES, and then the edge server ES groups the devices according to time and sorts the mobile devices MD in each group according to the predetermined requirements; (2) The mobile device MD receives the global model from the edge server ES and performs local training to update the model; (3) The mobile device MD transmits the updated model to the connected mobile device MD based on the pre-set training sequence; (4) After the current mobile device MD receives the local training model sent by its previous mobile device MD, it uses its local data set to continue local training on the received model; when the model is transmitted to a location without a successor device, the device uploads the locally updated model to the edge server ES for synchronization aggregation.

Citation Information

Patent Citations

  • Network topology construction method and system in hierarchical federated learning scene

    CN114650227A

  • Resource allocation and aggregation optimization method and device for hierarchical federated learning system

    CN117076132A