A layered federated learning method, device, equipment and storage medium

CN118863090BActive Publication Date: 2026-09-08TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410809397.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2026-09-08
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

[0004]需要注意的是,在客户端动态参与的设备-边缘-云上的分层联邦学习场景中,客户端可能无法参与每一轮全局训练,这表明在每一轮全局训练中都需要重新评估关于客户端选择和客户端与边缘的关联决策,如此存在决策开销较大的技术问题

Benefits of technology

[0028] In embodiments of the present invention, a first approach identifies clients that may participate in future model training offline, while a second approach rapidly explores suitable alternative clients online when any client selected in the first approach cannot participate in model training. In this way, a phased decision-making strategy can optimize client selection and the association between clients and edge servers, thereby reducing decision-making overhead and improving decision-making efficiency, ultimately achieving seamless hierarchical federated learning in dynamic network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118863090B_ABST
    Figure CN118863090B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a layered federated learning method, device, equipment and storage medium. The method comprises: clustering clients associated with an edge server based on attribute information of the clients; before model training starts, performing a first scheme: selecting a group of long-term clients according to the probability of each client participating in each round of global iteration learning in an offline manner, and assigning them to suitable edge servers for subsequent global model aggregation; after the model training starts, performing a second scheme: if there is an offline long-term client among the long-term clients under the edge server, recruiting a short-term client from the class in which the offline long-term client is located in the corresponding clustering result, and assigning it to a suitable edge server for subsequent global model aggregation. In this way, the selection of clients, the association between clients and edge servers can be optimized, and the purpose of reducing decision overhead can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of federated learning technology, and in particular to a hierarchical federated learning method, apparatus, device, and storage medium. Background Technology

[0002] The proliferation of smart devices has led to the generation of massive amounts of data, driving the rise of AI-based applications throughout the Internet of Things (IoT) ecosystem. Traditional machine learning methods typically involve collecting data from distributed IoT devices and aggregating it at a central location for model training. This paradigm shift from centralized to distributed machine learning methods has seen federated learning emerge as one of the most compelling approaches.

[0003] However, federated learning relies on frequent model transfers between clients and cloud servers, which can incur significant overhead (e.g., communication latency and energy consumption). Furthermore, reducing the frequency of model aggregation prolongs the training duration to achieve optimal convergence, further increasing network resource consumption. In recent years, hierarchical federated learning has provided a recognized framework to address these challenges—introducing an intermediate layer into the federated learning model training architecture for intermediate model aggregation. For example, an edge server layer frequently synchronizes model parameters from its associated clients and exchanges model updates with the cloud server at a lower frequency.

[0004] It is important to note that in a hierarchical federated learning scenario where the client dynamically participates in the device-edge-cloud architecture, the client may not be able to participate in every round of global training. This means that decisions regarding client selection and the association between the client and the edge need to be re-evaluated in each round of global training, which presents a technical problem of high decision-making overhead. Summary of the Invention

[0005] Embodiments of the present invention provide a hierarchical federated learning method, apparatus, device, and storage medium.

[0006] In a first aspect, embodiments of the present invention provide a hierarchical federated learning method, the method comprising:

[0007] Based on the past training participation records of each client, calculate the probability of each client participating in each round of global iterative learning;

[0008] For each edge server, cluster the clients that it can associate with based on the attribute information of the clients it can associate with;

[0009] Before model training begins, the first approach is implemented: a group of long-term clients are selected offline based on the probability of each client participating in each round of global iterative learning, and these long-term clients are assigned to appropriate edge servers for subsequent global model aggregation.

[0010] After model training begins, the second approach is implemented: for each edge server, if there are disconnected long-term clients among its subordinate long-term clients, short-term clients are recruited online from the class where the disconnected long-term clients belong in the corresponding clustering results, and the short-term clients are assigned to appropriate edge servers for subsequent global model aggregation.

[0011] In some possible implementations of the first aspect, the probability of each client participating in each round of global iterative learning is calculated based on the client's past training participation records, including:

[0012] Based on historical statistics of the client's performance in previous rounds of global iterative learning, the probability of the client participating in each round of global iterative learning is calculated.

[0013] In some possible implementations of the first aspect, clustering of the clients to which it can be associated is performed based on the attribute information of the clients to which it can be associated, including:

[0014] Based on the attribute information of the clients that can be associated with it, the cosine similarity between the attribute information of any two clients among the clients that can be associated with it is calculated, and the clients that can be associated with it are clustered in this way.

[0015] In some possible implementations of the first aspect, the attribute information includes: data distribution, latency and energy consumed in participating in a round of edge iteration.

[0016] In some possible implementations of the first aspect, the clustering algorithm used is the DBSCAN clustering algorithm.

[0017] In some possible implementations of the first aspect, a group of long-term clients is selected offline based on the probability of each client participating in each round of global iterative learning, and these long-term clients are assigned to appropriate edge servers for subsequent global model aggregation, including:

[0018] In an offline manner, a group of long-term clients is initially selected based on the probability of each client participating in each round of global iterative learning. The association combination of long-term clients and edge servers that minimizes the first system cost objective function used before model training begins is greedily selected. The first system cost objective function is obtained by decomposing the system cost objective function that incorporates latency and energy consumption. The selection of long-term clients is iteratively optimized to obtain the optimal solution of the first system cost objective function. The long-term clients are then assigned to appropriate edge servers for subsequent global model aggregation.

[0019] In some possible implementations of the first aspect, for each edge server, if there are offline long-term clients among its subordinate long-term clients, short-term clients are recruited online from the class of the offline long-term clients in the corresponding clustering results, and the short-term clients are assigned to appropriate edge servers for subsequent global model aggregation, including:

[0020] For each edge server, if there are disconnected long-term clients among its subordinate long-term clients, short-term clients are recruited online from the cluster of the disconnected long-term clients in the corresponding clustering results. The association combination of short-term clients and edge servers is greedily selected to minimize the second system cost objective function used after the model training begins. The second system cost objective function is obtained by decomposing the system cost objective function that integrates latency and energy consumption. Finally, the optimal solution of the second system cost objective function is obtained, and short-term clients are assigned to appropriate edge servers for subsequent global model aggregation.

[0021] Secondly, embodiments of the present invention provide a hierarchical federated learning apparatus, the apparatus comprising:

[0022] The calculation module is used to calculate the probability of each client participating in each round of global iterative learning based on the past training participation records of each client.

[0023] The clustering module is used to cluster the clients that can be associated with each edge server based on the attribute information of the clients that can be associated with it.

[0024] The first execution module is used to execute the first scheme before model training begins: select a group of long-term clients offline based on the probability of each client participating in each round of global iterative learning, and assign the long-term clients to appropriate edge servers for subsequent global model aggregation;

[0025] The second execution module is used to execute the second scheme after the model training begins: for each edge server, if there are disconnected long-term clients among its subordinate long-term clients, short-term clients are recruited online from the class where the disconnected long-term clients are located in the corresponding clustering results, and the short-term clients are assigned to suitable edge servers for subsequent global model aggregation.

[0026] Thirdly, embodiments of the present invention provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0027] Fourthly, embodiments of the present invention provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described above.

[0028] In embodiments of the present invention, a first approach identifies clients that may participate in future model training offline, while a second approach rapidly explores suitable alternative clients online when any client selected in the first approach cannot participate in model training. In this way, a phased decision-making strategy can optimize client selection and the association between clients and edge servers, thereby reducing decision-making overhead and improving decision-making efficiency, ultimately achieving seamless hierarchical federated learning in dynamic network environments.

[0029] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0030] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0031] Figure 1 A flowchart of a hierarchical federated learning method provided by an embodiment of the present invention is shown;

[0032] Figure 2 The illustration shows a specific application diagram of a hierarchical federated learning method provided by an embodiment of the present invention;

[0033] Figure 3 This diagram illustrates a comparison of accuracy between the phased decision-making method for client-side dynamic threats provided by the hierarchical federated learning method of the present invention and existing baseline methods.

[0034] Figure 4 This diagram illustrates a comparison of the phased decision-making method for client-side dynamic threats in the hierarchical federated learning method provided by an embodiment of the present invention with existing baseline methods in terms of training time, energy consumption, system cost, and runtime.

[0035] Figure 5 A structural diagram of a hierarchical federated learning device provided by an embodiment of the present invention is shown;

[0036] Figure 6 A structural diagram of an exemplary electronic device capable of implementing an embodiment of the present invention is shown. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] Furthermore, the term "and / or" in this invention is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this invention generally indicates that the preceding and following related objects have an "or" relationship.

[0039] To address the technical problems encountered in the background art, embodiments of the present invention provide a hierarchical federated learning method, apparatus, device, and storage medium. The following detailed description, in conjunction with the accompanying drawings, of specific embodiments of the hierarchical federated learning method, apparatus, device, and storage medium provided by the present invention.

[0040] Figure 1 A flowchart of a hierarchical federated learning method provided by an embodiment of the present invention is shown, as follows: Figure 1 As shown, the hierarchical federated learning method 100 may include the following steps:

[0041] S110: Calculate the probability of each client participating in each round of global iterative learning based on the past training participation records of each client.

[0042] In some embodiments, the probability of a client participating in each round of global iterative learning can be calculated based on historical statistical data of the client performing model training tasks in previous rounds of global iterative learning.

[0043] S120: For each edge server, cluster the clients that can be associated with it based on the attribute information of the clients that can be associated with it.

[0044] In some embodiments, the cosine similarity between the attribute information of any two clients among the clients that can be associated with can be calculated based on the attribute information of the clients that can be associated with (e.g., data distribution, latency and energy consumed in a round of edge iteration, etc.), and the DBSCAN clustering algorithm is used to cluster the clients that can be associated with.

[0045] S130, before model training begins, execute the first scheme: select a group of long-term clients offline based on the probability of each client participating in each round of global iterative learning, and assign the long-term clients to appropriate edge servers for subsequent global model aggregation.

[0046] In some embodiments, a set of long-term clients can be initially selected offline based on the probability of each client participating in each round of global iterative learning. The association combination of long-term clients and edge servers that minimizes the first system cost objective function used before model training begins is greedily selected. The first system cost objective function is obtained by splitting the system cost objective function that incorporates latency and energy consumption. The selection of long-term clients is iteratively optimized to obtain the optimal solution of the first system cost objective function. In this way, long-term clients are assigned to appropriate edge servers for subsequent global model aggregation.

[0047] S140, after model training begins, execute the second scheme: for each edge server, if there are disconnected long-term clients among its subordinate long-term clients, recruit short-term clients online from the class where the disconnected long-term clients are located in the corresponding clustering results, and assign the short-term clients to suitable edge servers for subsequent global model aggregation.

[0048] In some embodiments, for each edge server, if there are disconnected long-term clients among its subordinate long-term clients, short-term clients are recruited online from the cluster of the disconnected long-term clients in the corresponding clustering results. The association combination of short-term clients and edge servers that minimizes the second system cost objective function used after the start of model training is greedily selected. The second system cost objective function is obtained by splitting the system cost objective function that integrates latency and energy consumption. Finally, the optimal solution of the second system cost objective function is obtained, so as to allocate short-term clients to suitable edge servers for subsequent global model aggregation.

[0049] To facilitate further understanding, the above content will be described in detail below with reference to specific embodiments:

[0050] like Figure 2 As shown, the hierarchical federated learning method provided in this embodiment of the invention optimizes client selection and the association between clients and edge servers through a phased decision-making strategy. This minimizes the total system cost (including latency and energy consumption) while maintaining satisfactory model accuracy, even with dynamic client participation. This strategy employs two complementary schemes (i.e., a first scheme and a second scheme) to achieve asynchronous execution at different time scales, improving decision-making time efficiency and ensuring seamless model training. The specific network framework, optimization objectives, and method steps are shown below:

[0051] Network framework and optimization goals:

[0052] Assume the environment contains a cloud server acting as a central aggregator and multiple edge servers. As an intermediary aggregator, multiple clients Each customer n i Described as a quad array [local dataset] CPU frequency v i Transmit power q i Participation status ξ i The system achieves joint minimization of training latency (including computation and communication latency) and energy consumption by optimizing client selection and the client-edge correlation matrix A.

[0053] As an example, the number of experiments is M=100, each experiment contains 1=30 global iterations, the dataset is MNIST, the training model is a convolutional neural network, and there is 1 cloud server, S=4 edge servers, and N=100 clients in the environment.

[0054] Method and steps:

[0055] (1) Assume that the binary variable ξ of the client participates in each round of global iterative learning. i The probability of obedience is p i Bernoulli distribution, which will be based on the client n i The historical statistical data from the previous 100 global iterations used to perform model training tasks were used to estimate each p. i For example, at the end of the Kth observation rolling window, p i The estimated value is calculated using the following formula:

[0056]

[0057] in, p represents the value at the end of the Kth observation rolling window. i Estimated value; In the g-th global iteration, ξ represents... i The observed value; τ represents the length of the observation scrolling window; K represents the total number of windows.

[0058] (2) Calculate the values ​​of each edge server s j The client vector [y] i T i,j E i,j Cosine similarity between and similarity matrix Among them, y i Indicates client n i Data distribution; Ti,j Indicates client n i The latency consumed in participating in a round of edge iteration; E i,j Indicates client n i The energy consumed in participating in one round of edge iteration. Given a similarity threshold ψ of a neighborhood. min Density threshold P of the neighborhood min The DBSCAN clustering algorithm was used to cluster each edge server s j The clients are clustered so that long-term disconnected clients can be replaced during actual model training, ensuring smooth training execution.

[0059] (3) The system cost objective function that integrates time delay and energy consumption Decomposed into the first system cost objective function Second system cost objective function These two methods are used respectively before and after model training begins. Indicates time delay; Indicates energy consumption; λ t , λ e These represent the corresponding weights. Here, we introduce the concept of system continuity, defined as:

[0060]

[0061] Define the objective function Define the objective function

[0062] (4) Before the actual model training begins, execute the first scheme: offline, select a group of long-term clients based on the probability of each client participating in each round of global iterative learning, that is, a group of clients with higher probabilities, and greedily select the ones that make the first system cost objective function The minimum combination of long-term clients and edge servers is used to iteratively optimize the selection of long-term clients, ultimately yielding the first system cost objective function. The optimal solution A′ is obtained, and then the finally selected long-term clients are assigned to appropriate edge servers for subsequent global model aggregation.

[0063] (5) After the actual model training begins, execute the second scheme: for ξ i For long-term clients with a value of 1, directly select and assign an edge server consistent with the first scheme; for ξ i Long-term clients with a value of 0, i.e., long-term disconnected clients, are clustered based on the clustering results corresponding to their respective edge servers. The client with the highest cosine similarity to the client belonging to the same class is selected as the short-term client, and the client is greedily selected to maximize the cost objective function of the second system. The minimum short-term client and edge server combination yields the second system cost objective function. The approximate optimal solution A″ is obtained, and then short-term clients are assigned to appropriate edge servers for subsequent global model aggregation.

[0064] The following is combined with Figures 3-4 This demonstrates a performance comparison between the phased decision-making method for client-side dynamic threats provided in the hierarchical federated learning method of this invention and other baseline methods. Figure 3 curves and Figure 4 The bar charts (a)-(c) show that the phased decision-making method is slightly lower than the single-stage optimal decision-making method in terms of model accuracy and system cost control, but better than other baseline methods. Figure 4 As shown in the bar chart (d), the phased decision-making method performs exceptionally well in terms of runtime, second only to the method of randomly selecting clients. This indicates that the phased decision-making method proposed here can efficiently make client selection and client-to-edge server association decisions that are beneficial to system cost control and model convergence, and has excellent time efficiency.

[0065] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0066] The above is an introduction to the method embodiments. The following describes the solution of the present invention further through device embodiments.

[0067] Figure 5 A structural diagram of a hierarchical federated learning device provided by an embodiment of the present invention is shown, as follows: Figure 5 As shown, the hierarchical federated learning device 500 may include:

[0068] The calculation module 510 is used to calculate the probability of each client participating in each round of global iterative learning based on the past training participation records of each client.

[0069] Clustering module 520 is used to cluster the clients that can be associated with each edge server based on the attribute information of the clients that can be associated with it.

[0070] The first execution module 530 is used to execute the first scheme before model training begins: select a group of long-term clients offline based on the probability of each client participating in each round of global iterative learning, and assign the long-term clients to appropriate edge servers for subsequent global model aggregation.

[0071] The second execution module 540 is used to execute the second scheme after the model training begins: for each edge server, if there are disconnected long-term clients among its subordinate long-term clients, then short-term clients are recruited online from the class where the disconnected long-term clients are located in the corresponding clustering results, and the short-term clients are assigned to suitable edge servers for subsequent global model aggregation.

[0072] In some embodiments, the calculation module 510 is specifically used for:

[0073] Based on historical statistics of the client's performance in previous rounds of global iterative learning, the probability of the client participating in each round of global iterative learning is calculated.

[0074] In some embodiments, the clustering module 520 is specifically used for:

[0075] Based on the attribute information of the clients that can be associated with it, the cosine similarity between the attribute information of any two clients among the clients that can be associated with it is calculated, and the clients that can be associated with it are clustered in this way.

[0076] In some embodiments, the first execution module 530 is specifically used for:

[0077] In an offline manner, a group of long-term clients is initially selected based on the probability of each client participating in each round of global iterative learning. The association combination of long-term clients and edge servers that minimizes the first system cost objective function used before model training begins is greedily selected. The first system cost objective function is obtained by decomposing the system cost objective function that incorporates latency and energy consumption. The selection of long-term clients is iteratively optimized to obtain the optimal solution of the first system cost objective function. The long-term clients are then assigned to appropriate edge servers for subsequent global model aggregation.

[0078] In some embodiments, the second execution module 540 is specifically used for:

[0079] For each edge server, if there are disconnected long-term clients among its subordinate long-term clients, short-term clients are recruited online from the cluster of the disconnected long-term clients in the corresponding clustering results. The association combination of short-term clients and edge servers is greedily selected to minimize the second system cost objective function used after the model training begins. The second system cost objective function is obtained by decomposing the system cost objective function that integrates latency and energy consumption. Finally, the optimal solution of the second system cost objective function is obtained, and short-term clients are assigned to appropriate edge servers for subsequent global model aggregation.

[0080] Understandable, Figure 5 Each module / unit in the hierarchical federated learning device 500 shown has the ability to implement Figure 1 The functions of each step in the hierarchical federated learning method 100 shown, and the corresponding technical effects they achieve, will not be elaborated here for the sake of brevity.

[0081] Figure 6 A structural diagram of an exemplary electronic device capable of implementing embodiments of the present invention is shown. Electronic device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 600 may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown in this invention, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0082] like Figure 6 As shown, the electronic device 600 may include a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0083] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0084] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer program product, including a computer program tangibly contained in a computer-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform method 100 by any other suitable means (e.g., by means of firmware).

[0085] The various embodiments described above in this invention can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), payload programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0086] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0087] In the context of this invention, a computer-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of computer-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0088] It should be noted that the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute method 100 and achieve the corresponding technical effects achieved by the embodiments of the present invention in executing the method. For the sake of brevity, these will not be elaborated here.

[0089] In addition, the present invention also provides a computer program product, which includes a computer program that implements method 100 when executed by a processor.

[0090] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this invention does not impose any limitations on them.

[0091] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A hierarchical federated learning method, characterized in that, The method includes: Based on the past training participation records of each client, calculate the probability of each client participating in each round of global iterative learning; For each edge server, cluster the clients that it can associate with based on the attribute information of the clients it can associate with; Before model training begins, the first approach is implemented: a group of long-term clients are selected offline based on the probability of each client participating in each round of global iterative learning, and these long-term clients are assigned to appropriate edge servers for subsequent global model aggregation. After model training begins, the second approach is implemented: for each edge server, if there are disconnected long-term clients among its subordinate long-term clients, short-term clients are recruited online from the class where the disconnected long-term clients belong in the corresponding clustering results, and the short-term clients are assigned to appropriate edge servers for subsequent global model aggregation.

2. The method according to claim 1, characterized in that, The calculation of the probability of each client participating in each round of global iterative learning, based on the client's past training participation records, includes: Based on historical statistics of the client's performance in previous rounds of global iterative learning, the probability of the client participating in each round of global iterative learning is calculated.

3. The method according to claim 1, characterized in that, The clustering of clients that can be associated with a client based on its attribute information includes: Based on the attribute information of the clients that can be associated with it, the cosine similarity between the attribute information of any two clients among the clients that can be associated with it is calculated, and the clients that can be associated with it are clustered in this way.

4. The method according to claim 3, characterized in that, The attribute information includes: data distribution, latency and energy consumed in participating in a round of edge iteration.

5. The method according to claim 3, characterized in that, The clustering algorithm used is the DBSCAN clustering algorithm.

6. The method according to claim 1, characterized in that, The step of selecting a group of long-term clients offline based on the probability of each client participating in each round of global iterative learning, and assigning these long-term clients to appropriate edge servers for subsequent global model aggregation, includes: In an offline manner, a group of long-term clients is initially selected based on the probability of each client participating in each round of global iterative learning. The association combination of long-term clients and edge servers that minimizes the first system cost objective function used before model training begins is greedily selected. The first system cost objective function is obtained by decomposing the system cost objective function that incorporates latency and energy consumption. The selection of long-term clients is iteratively optimized to obtain the optimal solution of the first system cost objective function. The long-term clients are then assigned to appropriate edge servers for subsequent global model aggregation.

7. The method according to claim 6, characterized in that, For each edge server, if any of its long-term clients have disconnected, short-term clients are recruited online from the cluster of the disconnected long-term clients in the corresponding clustering results. These short-term clients are then assigned to suitable edge servers for subsequent global model aggregation, including: For each edge server, if there are disconnected long-term clients among its subordinate long-term clients, short-term clients are recruited online from the cluster of the disconnected long-term clients in the corresponding clustering results. The association combination of short-term clients and edge servers is greedily selected to minimize the second system cost objective function used after the model training begins. The second system cost objective function is obtained by decomposing the system cost objective function that integrates latency and energy consumption. Finally, the optimal solution of the second system cost objective function is obtained, and short-term clients are assigned to appropriate edge servers for subsequent global model aggregation.

8. A hierarchical federated learning device, characterized in that, The device includes: The calculation module is used to calculate the probability of each client participating in each round of global iterative learning based on the past training participation records of each client. The clustering module is used to cluster the clients that can be associated with each edge server based on the attribute information of the clients that can be associated with it. The first execution module is used to execute the first scheme before model training begins: select a group of long-term clients offline based on the probability of each client participating in each round of global iterative learning, and assign the long-term clients to appropriate edge servers for subsequent global model aggregation; The second execution module is used to execute the second scheme after the model training begins: for each edge server, if there are disconnected long-term clients among its subordinate long-term clients, short-term clients are recruited online from the class where the disconnected long-term clients are located in the corresponding clustering results, and the short-term clients are assigned to suitable edge servers for subsequent global model aggregation.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Federal mutual learning model training method for non-independent identically distributed data

    CN114091667A

  • Federal learning privacy protection method and system based on homomorphic encryption

    CN117640253A