A Network Function Computing Encoding Scheme for Hierarchical Federated Learning

By adopting a hybrid network structure and network function computing coding scheme with multi-edge server connection in cloud edge collaborative federated learning, the problems of high communication costs and poor scalability in traditional solutions are solved, and more efficient network function computing speed and better scalability are achieved.

CN118869142BActive Publication Date: 2025-05-27SUN YAT SEN UNIVERSITY SHENZHEN +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410827971.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2025-05-27
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

The existing cloud-edge collaborative federated learning based on tree transmission structure fails to fully utilize the computing power of edge servers, resulting in higher communication costs and narrow application scope of federated learning solutions based on network encoding and poor expansion.

Method used

A network function computing coding scheme applied to hierarchical federated learning is proposed, using a hybrid network structure model for multi-edge server connection, combining network function computing technology and greedy algorithms, making full use of the computing power of edge servers and improving the computing rate of network function.

Benefits of technology

The hybrid network structure and network function calculation and coding scheme connected by multi-edge servers improve the transmission rate and convergence speed of local aggregation, and enhance the scalability and lossless transmission capabilities of the scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118869142B_ABST
    Figure CN118869142B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of cloud communication technology, and specifically discloses a network function computing coding scheme applied to hierarchical federated learning, which consists of three parts: a hybrid network structure model connected by multiple edge servers; an encoding transmission scheme for user-edge servers; and the global aggregation and local aggregation processes of the model. The scheme makes full use of the characteristics of federated learning function computing, combines network function computing technology and the idea of greedy algorithm, and makes full use of the computing power of edge servers to improve the network function computing rate, thereby reducing the overall communication cost of the system. Even when the finite field is small, the scheme can effectively improve the network function computing rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud communication, and particularly to a network function computing coding scheme applied to hierarchical federated learning. Background Art

[0002] 1. Hierarchical Federated Learning Based on Cloud-Edge-Device Collaboration

[0003] Traditional cloud-centric machine learning requires collecting data from all parties and sending all the data to a cloud-based server or data center for processing and training. However, uploading all this data to the central server will consume a large amount of network bandwidth, and the transmission delay of data interaction is high. It is not suitable for scenarios with low transmission delay requirements and there is also a risk of privacy leakage. In order to enable artificial intelligence systems to more efficiently collaborate on data training while meeting data privacy, security, and regulatory requirements, Google proposed the federated learning framework in 2016. In federated learning, after training the model locally, the user device uploads the model update parameters instead of the local data to the central server. This means that the process of data storage and modeling training is migrated to the user device for execution. The user device only needs to upload the model update parameters to the central server, and then the central server performs weighted aggregation to obtain the training result, thus effectively protecting the privacy of user data.

[0004] Since federated networks may often involve thousands of devices participating in training and each training requires multiple rounds of communication, and resources such as network bandwidth are limited, reducing communication costs has become a key concern. Currently, most methods for reducing communication costs focus on two aspects: one is to reduce the total number of rounds of uploading data; the other is to reduce the data uploaded in each round. These methods rarely consider optimizing the communication network structure of federated learning. The traditional federated learning structure is a two-layer network based on the cloud. As Figure 1 shown, a large number of mobile devices communicate directly with the cloud center server. However, the mobile devices and the central server are often far apart, and the communication of the wide area network requires a large amount of network bandwidth and has serious delays. People introduced edge computing in federated learning and proposed a three-layer heterogeneous federated learning network for collaborative training based on cloud-edge-device, as Figure 2 shown, to reduce communication costs by transferring part of the computational workload to the edge server.

[0005] 2. Network Coding and Network Function Computing

[0006] Network function computing is a field closely related to network coding. Network Coding is a technology for information transmission and data storage. Intermediate nodes do not simply store and forward data, but allow intermediate nodes to process the information before sending it during data transmission, thereby improving the efficiency and reliability of the network.

[0007] Next, the implementation of network coding will be described using a typical single-source and two-destination butterfly network in network coding. As Figure 3 and Figure 4 shown, assume a communication network G = (V, E), where V describes the set of nodes in G and E describes the set of edges in G. S is the source, and C 1 and C 2 are both destination nodes, and the capacity of the channel is 1. The main task to be achieved by this network is to send the information w 1 and w 2 from the source node S to the destination nodes C 1 and C 2 . Figure 3 For the traditional router to store and forward information process, since the capacity of the channel is 1, after the edge server 3 receives w 1 and w 2 , it takes two time slots to send out the message, so C 1 and C 2 receive w 1 and w 2 simultaneously in two time slots. By using network coding technology as Figure 4 shown, the edge server 3 can perform an exclusive OR operation on the transmitted w 1 and w 2 to get w 1 ⊕w 2 , and transmit the operation result w 1 ⊕w 2 to node 4, and this process only takes one time slot. C 1 receives w 1 from node 1 and w 1 ⊕w 2 from node 4, and can obtain w 1 ⊕w 1 ⊕w 2 =w 2 . Similarly, C v can also obtain w 1 , and C 1 and C 2 receive w 1 and w 2 simultaneously in one time slot. Therefore, using network coding can achieve the maximum flow-minimum cut bound.

[0008] The above coding scheme is for a determined network topology, while random network coding is a network coding scheme for an unknown topology. The destination divides the data to be sent into several data packets and then encodes each data packet. The encoding process involves selecting a set of random coefficients and multiplying them with the original data packets to generate the encoded data packets. When the edge servers of random network coding receive multiple encoded data packets, they linearly combine these data packets and then send the generated combined packets. When the receiving end receives a sufficient number of encoded data packets, the original data packets can be restored by solving a system of linear equations.

[0009] The network function calculation problem is, on the premise that the edge servers have the network coding function, for the case where in some networks the central server is interested in the correlation functions of the source data rather than the original data, such as wireless sensor networks. A wireless sensor network consists of nodes with sensing, wireless communication, and computing capabilities. Such a network not only has to complete the task of sensing the environment but also, through a series of messages passed between nodes and the computing on the nodes, transfer the functions of the relevant data to the designated receiving nodes. For example, in environmental monitoring, the relevant statistical values of the temperature sensor readings may be the average, median, and mode of the temperature; in a temperature alarm network, the data of interest may be the maximum value of the temperature readings. In this case, the edge servers not only store and forward the collected data but also play a dual role of computing and communication. The network function calculation problem is an extended network coding problem. Traditional network coding needs to restore the original information of the source at the destination and can be considered as finding the identity function of the source at the destination. The uplink communication network of federated learning based on cloud-edge-end collaboration also has this function calculation feature. In the classic federated averaging algorithm, after the local training of the users is completed, they upload the updated parameters of the model to the central server. What the central server needs is the arithmetic sum of the user data rather than the original data. Therefore, the edge servers can play a computing role.

[0010] 3. Federated Learning Based on Network Coding

[0011] As the combination of coding technology and federated learning becomes closer and closer, network coding has also been introduced into federated learning transmission, mainly including deterministic network coding and random network coding. The NC-FLs (Network Coding-Federated Learning System) proposed based on deterministic network coding studied the coding scheme in a two-user butterfly network. This coding scheme sends the data out in a combined manner after dividing it into two parts. Finally, the central server decodes to obtain the original data after receiving enough messages. The uplink process makes use of the network function calculation technology, and the central server only needs to solve the arithmetic sum of the source data. For example Figure 5 、 Figure 6The following are the downlink and uplink processes of federated learning based on network coding respectively:

[0012] The FedNC (Federated Learning-Network Coding) coding scheme proposed based on random network coding. The idea is to mix the information of local models by performing random linear combinations on the original data packets before further aggregation during upload. FedNC improves the performance of traditional FL in several important aspects, including security, throughput, and robustness.

[0013] Disadvantages of the existing technologies:

[0014] 1. Disadvantages of the existing cloud-edge-end collaborative federated learning based on tree-like transmission structure

[0015] Currently, most of the cloud-edge-end collaborative federated learning considers the tree-like network structure, that is, only considering the situation where each end node is only connected to one edge server, and the edge server is then connected to the cloud, lacking the consideration of optimizing more complex transmission structures, such as the situation where users can connect to multiple edge servers. Therefore, the traditional scheme does not fully utilize the computing power of edge servers, resulting in higher communication costs.

[0016] 2. Disadvantages of the existing federated learning based on network coding

[0017] Currently, the NC-FLs scheme based on deterministic network coding only studies the coded transmission based on the butterfly network. This coding scheme is only suitable for the case of two users and lacks the rate analysis of the coding scheme from the perspective of information theory;

[0018] Currently, the FedNC scheme based on random network coding requires a relatively large finite field to ensure the decoding success rate of the central server.

[0019] In summary, the application scope of the existing federated learning schemes based on network coding is relatively narrow and the scalability is poor. Summary of the Invention

[0020] Aiming at the problem that the single network structure in the heterogeneous federated learning of cloud-edge-end collaboration is difficult to meet the actual user needs, the present invention aims to propose a hybrid network structure with better scalability and multiple edge server connections based on the traditional tree-like network structure, and propose a corresponding network coding scheme. The scheme makes full use of the characteristics of federated learning function calculation, combines network function calculation technology and the idea of greedy algorithm, and makes full use of the computing power of edge servers to improve the function calculation rate of the network, thereby reducing the overall communication cost of the system. Even when the finite field is small, the scheme can effectively improve the function calculation rate of the network.

[0021] The technical solution adopted by the present invention is as follows:

[0022] A network function computing coding scheme applied to hierarchical federated learning, comprising the following three components:

[0023] P1. A hybrid network structure model connected by multiple edge servers;

[0024] P2. An encoding transmission scheme for user-edge servers;

[0025] P3. The global aggregation and local aggregation processes of the model.

[0026] Preferably, P1 is a three-layer cloud-edge-end collaborative network structure in which the user side is connected to one or more edge servers, and one or more edge servers are connected to the cloud.

[0027] Preferably, the network structure is specifically:

[0028] In the network, at the top layer is 1 central server or central servers, which perform data aggregation on the parameters in the system to further perform federated learning calculations; in the middle layer are r edge servers m i , 1 ≤ i ≤ r, or edge servers, which perform data aggregation operations and forwarding; at the bottom layer are s users σ i , 1 ≤ i ≤ s, each user randomly connects to a different edge server and uploads its own parameter data, and the number of edge servers connected by user σ i is denoted as C i .

[0029] Preferably, for each user σ i , 1 ≤ i ≤ s, the generated parameter data x i are all k symbols, and each symbol is included in the alphabet A = {0, 1, 2, …, q - 1}: x i = (x i,1 , x i,2 , …, x i,k ), x i,j ∈ A.

[0030] Preferably, P2 includes the following steps:

[0031] S2.1. Equally divide the parameter data of each user;

[0032] For each user σ i , 1 ≤ i ≤ s, the generated parameter data x i are all k symbols, and the parameter data x i generated by each user is equally divided into r parts, that is, the number of edge servers; in actual application scenarios, k is much larger than r, and k is approximately regarded as divisible by r; the parameter data x iAfter evenly dividing it into r parts, each part of the parameter data contains symbols, which are arranged in the original data order as the first part, the second part,..., the r-th part; between the user and the edge server, the data is in parts, that is symbols, and is uploaded and aggregated as a whole;

[0033] S2.2. Edge server initialization settings;

[0034] Set the allocated data volume of edge server m i to be l i , 1 ≤ i ≤ r, and the initial values of l i are all 1;

[0035] Set the retrieval index of edge server m i to be d i , 1 ≤ i ≤ r, and the initial values of d i are all i;

[0036] Start cycling from the (i + 1) % r part in each round, where % represents the modulo operation;

[0037] S2.3. Sort and select the edge server for the data to be allocated in this round of decision;

[0038] Define the following two rules, and sort the edge servers according to the allocated data volume l i of each edge server m i , 1 ≤ i ≤ r:

[0039] If l a < l b , then the sorting of m a is more forward;

[0040] If l a = l b and a < b, then the sorting of m a is more forward;

[0041] Traverse the edge servers from front to back according to the above sorting, and judge whether the data of the users connected to the edge server m i has been transmitted completely;

[0042] If not, select m i , if so, continue to traverse the next edge server in the sorting, and re-execute the judgment until an edge server is selected;

[0043] S2.4. Retrieve the underlying user data connected to the edge server selected in S2.3, and generate an uplink coding transmission scheme through the greedy strategy;

[0044] For the edge server m selected in S2.3i , retrieve the data to be transmitted currently among the underlying users it is connected to, with a total of r parts; for each part of the data, solve how many nodes still have not transmitted this part of the data; according to m i 's current retrieval index d i , traverse and retrieve the data from the d i -th part to the ((d i + r - 1) % r)-th part, where % represents the modulo operation;

[0045] If the j-th part of the data has the largest number of nodes to be transmitted, then the currently untransmitted j-th part among these node data is allocated to be transmitted to the edge server m i , and m i aggregates these data; if there is a situation where the number of nodes to be transmitted for the a-th part of the data is the same as that for the b-th part of the data, then the detection order of the two in the traversal retrieval should be compared. If a is detected earlier than b, then the a-th part should be selected to be allocated to be transmitted to the edge server m i ; finally, update m i 's current retrieval index d i to (j + 1) % r, where % represents the modulo operation;

[0046] S2.5. Update the amount of allocated data of the edge server selected in S2.3;

[0047] For the edge server m i selected in S2.3, up to now, it has been allocated to perform b i rounds of data aggregation, b i > 1; denote the number of users in the j-th round of aggregation as a j , 1 < j < b i , that is, arithmetic sum operations have been performed on a j portions of parameter data; according to information theory, the alphabet size of each symbol of the data generated in the j-th round of aggregation is a j (q - 1) + 1, and each portion of parameter data is still symbols; after b i rounds of aggregation, that is, arithmetic sum operations, up to now, the data size that m i needs to be transmitted to the central server is:

[0048]

[0049] Since both k and r are constants, so in each subsequent iteration, update l according to the following formula i :

[0050]

[0051] to measure the edge server m iThe allocated data volume;

[0052] Calculate and update m i The corresponding l i ;

[0053] S2.6. Loop through S2.3 - S2.5 until all the data of the underlying users has been allocated and transmitted to the corresponding edge servers, then end the loop; thus, obtain the uplink coding transmission scheme between the underlying users and the edge servers, denoted as Greedy - NFC, Greedy Algorithm for Network Function Computing;

[0054] Before the start of training, the central server generates the coding scheme according to the network topology. During each round of training, the user only needs to upload the model update parameters to the edge server according to the corresponding transmission matrix, and there is no need to traverse the above process again.

[0055] Preferably, P3 includes the following steps:

[0056] S3.1. Assume that the user uploads data to the edge server every τ 1 rounds, and the edge server aggregates τ 2 times and then uploads it to the cloud server, that is, the user conducts a global aggregation only after training τ 1 τ 2 times; the initial value of the aggregation round number τ of the edge server is 0; the user uses the gradient parameters received from the edge server or the central server for local training, and after every τ 1 rounds of local training, the data needs to be processed and sent to the edge server;

[0057] Using the federated averaging algorithm, the objective function of the central server:

[0058]

[0059] where N is the number of clients participating in the training, D i is the dataset of user i, and |D i | is the data volume of the data D i , is the sum of all user datasets;

[0060] After the user finishes training τ 1 rounds, the obtained gradient parameters w i (k) need to be multiplied by |D i | to get |D i |×w i (k), and then |D i |×w i(k) Sent to the intermediate layer nodes according to the previously mentioned Greedy-NFC coding scheme; in particular, if the amount of data trained by users is equal, the original gradient parameters are directly transmitted without processing the data, obtaining:

[0061]

[0062] S3.2. The user uploads to the edge server according to the Greedy-NFC method. After the edge server receives the data, it performs arithmetic sum operations on the same part of the data; updates the aggregation count τ + 1 of the edge server;

[0063] S3.3. If τ < τ 2 , then the aggregated data is fed back to the user, i.e., local aggregation; since the user receives data from multiple edge servers, it is necessary to calculate the federated average of the same part of the data:

[0064]

[0065] where w (j) (k + 1) refers to the updated parameter of the j-th part, N j refers to the number of users aggregated in the j-th part, |D j | refers to the sum of the data of N j users aggregated in the j-th part; if the amount of data trained by users is equal, then:

[0066]

[0067] S3.4. If τ = τ 2 , then the aggregated data is transmitted to the central server, i.e., global aggregation, and the central server calculates the federated average of each part of the data:

[0068]

[0069] Then the updated gradient is sent back to the user, and the user does not need to process the received data; similarly, if the amount of data trained by users is equal, then:

[0070]

[0071] At the same time, the edge server restarts calculating the aggregation count: τ = 0;

[0072] S3.5. Iterate S3.1 → S3.2 → S3.3 → S3.4 until the training requirements are met.

[0073] The present invention has the following characteristics and advantages:

[0074] 1. Improve the transmission rate

[0075] The traditional tree structure does not consider the situation where users connect to multiple edge servers, and the traditional communication process requires the restoration of original data on the central server. The network based on multiple edge servers considered in the present invention and the coding scheme proposed in combination with network function calculation can achieve computing while transmitting, and effectively improve the network function calculation rate by using the greedy idea.

[0076] 2. Improve the convergence speed of local aggregation

[0077] Since users can connect to the data of multiple edge servers simultaneously, each edge server can aggregate more users during local aggregation. Thus, the aggregated data can be fed back to the user side before the next round of training, which can improve the convergence speed of training to a certain extent.

[0078] 3. Scalability

[0079] Since this communication scheme starts from the transmission structure, it is independent of federated learning. At the same time, compared with the coding scheme of traditional NC-FLs which is only applicable to the two-user case, the scheme we proposed can be extended to any hierarchical federated learning with cloud-edge-end collaboration.

[0080] 4. Lossless transmission

[0081] The coding scheme designed in the present invention is lossless, so it will not cause loss of accuracy. Brief description of the drawings

[0082] Figure 1 is a schematic diagram of federated learning with a traditional two-layer network structure;

[0083] Figure 2 is a schematic diagram of a hierarchical federated learning structure with cloud-edge-end collaboration;

[0084] Figure 3 is a schematic diagram of the communication process of a traditional butterfly network;

[0085] Figure 4 is a schematic diagram of the communication process of a butterfly network based on network coding;

[0086] Figure 5 is a schematic diagram of the downlink process of federated learning based on network coding;

[0087] Figure 6 is a schematic diagram of the uplink process of federated learning based on network coding;

[0088] Figure 7 is a schematic diagram of a hierarchical federated learning structure with multi-edge service connection;

[0089] Figure 8 is an example of the coding scheme of the embodiment of the present invention;

[0090] Figure 9 It is the transmission scheme for splitting the data in the embodiment of the present invention into 2 parts;

[0091] Figure 10 It is the transmission scheme for splitting the data in the embodiment of the present invention into 3 parts. Detailed implementation manners

[0092] The global aggregation uplink coding transmission scheme and the local aggregation downlink feedback scheme applied to the hybrid network structure model for multi-edge server connection proposed by the present invention combine network function computing technology and improve the network function computing rate. The user data is evenly split into the same number of parts as the number of edge servers, and in each round, the edge server with the smallest amount of allocated data is selected according to the greedy algorithm. The edge server starts to circularly retrieve and greedily aggregate the data parts of the most users from the next round of the previous round of aggregation nodes. These designs for balancing the transmission volume from the edge server to the top layer node are the keys for the present invention to improve the network function computing rate.

[0093] 1. Generation process of user-edge server coding scheme

[0094] Consider Figure 8 the network structure in: 1 top layer node, 3 edge servers, 3 bottom layer users, and each user can be connected to two intermediate layer nodes. Assume that the parameter data x i to be transmitted by the bottom layer users are all 3 symbols from the alphabet A = {0, 1}. Assume that the amount of data trained by each user is equal.

[0095] S1: Evenly split the parameter data of each user;

[0096] Since the number of 3 edge servers r = 3, therefore, for users σ 1 , σ 2 , σ 3 split the data x 1 , x 2 , x 3 into 3 equal parts, obtaining (x 11 , x 12 , x 13 ), (x 21 , x 22 , x 23 ), (x 31 , x 32 , x 33 ).

[0097] S2: Initialization;

[0098] Set the amount of allocated data for each edge server: l 1 = l 2 = l 3= 1. Set the retrieval index for each edge server: d 1 = 1, d 2 = 2, d 3 = 3.

[0099] Start loop S3 - S5;

[0100] Loop 1:

[0101] S3: Sort and select the edge server for the data to be allocated: m 1 →m 2 →m 3 (Because l 1 = l 2 = l 3 , but 1 < 2 < 3). So m 1 is the edge server for the data to be allocated in this round of decision-making.

[0102] S4: m 1 The current retrieval index d 1 = 1, so start retrieving from the first part, retrieve the data to be transmitted currently allocated to the underlying users σ 1 connected to m 1 , σ 2 . The retrieval results are as follows: The data to be transmitted and allocated in the first part is x 11 +x 21 , the data to be transmitted and allocated in the second part is x 12 +x 22 , the data to be transmitted and allocated in the third part is x 13 +x 23 . Since the amount of data to be transmitted and allocated is 2 for all, according to the detection order of the three in the traversal retrieval, select the data in the first part to be allocated and transmitted to the edge server m 1 . m 1 Aggregates the data (finds the arithmetic sum) to get x 11 +x 21 . The retrieval index of m 1 is updated to d 1 = 2.

[0103] S5: According to the formula b 1 = 1, a 1 = 2, q = 2, update the amount of allocated data of m 1 to: l 1 = 2*(2 - 1)+1 = 3.

[0104] Loop 2:

[0105] S3: Sort and select the edge server for the data to be allocated: m 2 →m 3 →m1 (Since l 2 = l 3 = 0 < l 1 = 3, but 2 < 3). Therefore, m 2 is the edge server for the data to be allocated in this round of decision-making.

[0106] S4: m 2 The current retrieval index d 2 = 2. Therefore, starting from the second part, retrieve the data to be transmitted currently allocated to the underlying users σ 2 connected to m 1 , σ 3 The retrieval results are as follows: the data to be transmitted and allocated in the second part is x 12 + x 32 , the data to be transmitted and allocated in the third part is x 13 + x 33 , and the data to be transmitted and allocated in the first part is x 31 . Since the amounts of data to be transmitted and allocated in the second and third parts are both 2, according to the detection order in the traversal retrieval, select the data in the second part to be allocated and transmitted to the edge server m 2 . m 2 Aggregates the data (finds the arithmetic sum) to get x 12 + x 32 . The retrieval index of m 2 is updated to d 2 = 3.

[0107] S5: According to the formula b 2 = 1, a 1 = 2, q = 2, update the amount of allocated data of m 2 to: l 2 = 2*(2 - 1)+1 = 3.

[0108] Loop 3:

[0109] S3: Sort and select the edge server for the data to be allocated: m 3 → m 1 → m 2 (Since l 3 = 0 < l 1 = l 2 = 3). Therefore, m 3 is the edge server for the data to be allocated in this round of decision-making.

[0110] S4: m 3 The current retrieval index d 3 = 3. Therefore, starting from the third part, retrieve the data to be transmitted currently allocated to the underlying users σ 3 connected to m 2 , σ3 The data to be allocated for transmission currently is retrieved in the following order: The data to be allocated for transmission in the third part is x 23 +x 33 , the data to be allocated for transmission in the first part is x 31 , the data to be allocated for transmission in the second part is x 22 . The data to be allocated for transmission in the third part is the most, so the data in the third part is selected to be allocated for transmission to the edge server m 3 . m 3 Aggregate the data (find the arithmetic sum) to get x 23 +x 33 . m 3 's retrieval index is updated to d 3 =1.

[0111] S5: According to the formula b 3 =1, a 1 =2, q = 2, update the allocated data volume of m 3 to: l 3 =2*(2 - 1)+1 = 3.

[0112] Loop 4:

[0113] S3: Sort and select the edge server for the data to be allocated: m 1 →m 2 →m 3 (because l 1 =l 2 =l 3 , but 1 < 2 < 3). So m 1 is the edge server for the data to be allocated in this round of decision-making.

[0114] S4: The current retrieval index d 1 of m 1 =2, so starting from the second part, retrieve the underlying users σ 2 connected to m 1 、σ 2 for the data to be allocated for transmission currently. The retrieval results are as follows: The data to be allocated for transmission in the second part is x 22 , the data to be allocated for transmission in the third part is x 13 . Since the data volumes to be allocated for transmission in the second part and the third part are both 1, according to the detection order in the traversal retrieval, select the data in the second part to be allocated for transmission to the edge server m 1 . Since there is only one piece of data, m 1 does not need to aggregate the data again. m 1 's retrieval index is updated to d 1 =3.

[0115] S5: According to the formula b 1 = 2, a 1 = 2, a 2 = 1, q = 2, update m 1 The allocated data volume of is: l 1 = [2*(2 - 1)+1]*[1*(2 - 1)+1]= 6.

[0116] Loop 5:

[0117] Similarly to the above loop, the result is: Select the data of the 3rd part (x 13 ) is allocated and transmitted to the edge server m 2 , and the data x 22 is directly fed back to the user σ 3 .

[0118] Loop 6:

[0119] Similarly to the above loop, the result is: Select the data of the 1st part (x 31 ) is allocated and transmitted to the edge server m 3 .

[0120] S6: At this time, all the data of the user has been uploaded, and the edge servers m 1 , m 2 , m 3 obtain the aggregated data respectively as:

[0121] The above is the generation process of the user-to-middle-layer node coding scheme.

[0122] 2. Global aggregation and local aggregation process

[0123] S1: The user uses the gradient parameters received from the edge server or the central server for local training. After every τ 1 rounds of local training, the data needs to be processed and sent to the edge server; using the federated averaging algorithm, the objective function of the central server:

[0124]

[0125] where N is the number of clients participating in the training, D i is the dataset of user i, |D i | is the data volume of the data D i , is the sum of all user datasets;

[0126] After the user has trained for τ 1 rounds, the obtained gradient parameters w i (k) and |D i|Multiply to get| D i |× w i (k), and then | D i |× w i (k) is sent to the edge server according to the aforementioned Greedy - NFC coding scheme; in particular, if the amount of data trained by users is equal, the original gradient parameters are directly transmitted without data processing, obtaining:

[0127]

[0128] S2: Users upload to the edge server according to the Greedy - NFC method. After the edge server receives the data, it performs arithmetic sum operations on the same part of the data; update the aggregation times τ + 1 of the edge server;

[0129] S3: If τ < τ 2 Then the aggregated data is fed back to the user. User σ 1 receives the data fed back from the edge server m 1 , m 2 The same part of the data is averaged according to to obtain new parameters: User σ 2 3 , σ 3 Similarly.

[0130] S4: If τ = τ 2 Then the aggregated data is transmitted to the central server. Similarly, the same part of the data is averaged according to to obtain and then transmitted back to the user side to start a new round of training.

[0131] S5: Iterate S1 → S2 → S3 → S4 until the training requirements are met.

[0132] 3. Efficiency comparison between this scheme and traditional schemes

[0133] First, elaborate on the calculation method of the network function calculation rate: If a coding scheme enables the central server to calculate the k symbol functions f(x 1,j , x 2,j , …, x s,j ),(j = 1, 2, …, k) of s users without error, and each edge transmits at most n symbols from the alphabet B, then we call the network function calculation rate of this (k, n) coding scheme: Here we assume A = B, then the network function calculation rate is simplified to: The physical meaning is that the target function can be calculated k times by using the network n times.

[0134] Compared with the traditional solution, the encoding scheme for the uplink in the present invention improves the network function calculation rate. Since users need to send sufficient information, we do not consider the transmission optimization from users to edge servers, but only need to consider the optimization of the communication bottleneck from edge servers to the central server. The advantages of this encoding scheme will be intuitively illustrated below by comparing it with the simple encoding schemes of Figure 9 and Figure 10 .

[0135] In the traditional tree structure, each user is only connected to one edge server. Therefore, the entire data needs to be sent to the edge server. Even after the edge server performs calculations, it still needs to transmit k symbols to the central server, and the network function calculation rate is 1. If we consider the multi-edge server connection network encoding scheme shown in Figure 9 (assuming that q in the alphabet A = {0, 1,..., q - 1} is large enough): Each user divides the data into two equal parts (because the user has two sending links), and the size of each part of the data after division is symbols, which are sent to the edge server through two links respectively. Due to the "barrel effect", the overall communication efficiency of the network is limited by the link that needs to transmit the most data. It can be seen that the communication bottleneck is on the link from edge server m 2 to the central server ρ, because two parts of the data need to be sent respectively, that is, k symbols. Compared with the tree structure, the communication cost is not reduced, and the network function calculation rate is still 1. Another encoding scheme is the one proposed in the present invention (see the specific embodiments). The data generated by each user is divided into three parts, and each part of the data has symbols. As shown in Figure 10 , after encoding, the amount of data sent from each edge server to the top-level node is two parts of data, that is, the data that needs to be sent is . According to , the network function calculation rate is obtained as

[0136] The embodiments described above are only used to describe the preferred mode of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A network function computing coding scheme for hierarchical federated learning, characterized in that: It includes the following three components: P1, hybrid network structure model with multiple edge servers connected; P2, encoding and transmission scheme between user and edge server; P3, global aggregation and local aggregation process of the model; The steps in P2 are as follows: S2.

1. Equally divide the parameter data of each user; Per user , 1 Generated parameter data Both symbols, each with user-generated parameter data Divided equally Part, that is, the number of edge servers; Greater than , quilt divisibility; S2.2, edge server initialization settings; S2.3, sort and select the edge server to which the data to be allocated in this round of decision; S2.4, retrieve the user data connected to the edge server selected in S2.3, and generate an uplink coding transmission scheme through a greedy strategy; For S2.3 selected edge servers , retrieve the data currently to be allocated for transmission among the users to which it is connected, a total of For each part of the data, find out how many nodes have not yet transmitted the data; according to The current search index , will be from Part to Part of the data is traversed and retrieved, Represents the remainder operation; If Some of the data to be transmitted is the largest, so the node data that has not yet been transmitted Some are allocated for transmission to edge servers ,Depend on Aggregate these data; if the Partial data and If the number of partial data to be transmitted is the same, the order of their retrieval in the traversal search should be compared. Compare First check out, you should select Some are allocated for transmission to edge servers ; Finally, The current search index Updated to ,in Represents the remainder operation; S2.5, updating the allocated data volume of the edge server selected in S2.3; For S2.3 selected edge servers , which has been allocated so far Round data aggregation, ; The number of users aggregated in rounds is recorded as , 1< < , that is, Arithmetic and operations are performed on the parameter data; According to information theory, The alphabet size of each symbol of the data generated by round aggregation is , each parameter data is still symbols; go through Round aggregation, that is, after arithmetic and operation, up to now The size of data that needs to be transmitted to the central server is: ; because , are all constants, so each subsequent iteration is updated according to the following formula : ; To measure edge servers The amount of allocated data; Calculate and update The corresponding ; S2.6, loop S2.3-S2.5 until the data of all underlying users have been allocated and transmitted to the corresponding edge servers, then end the loop; At this point, the uplink coded transmission scheme between the underlying user and the edge server is obtained, denoted as Greedy-NFC, Greedy Algorithm for Network Function Computing; The steps in P3 are as follows: S3.

1. Users every The data is uploaded to the edge server once in a round, and the edge server aggregates After uploading to the cloud server, the user training Global aggregation is performed only after the edge server aggregation rounds. The initial value is 0; the user uses the gradient parameters received from the edge server or the central server to perform local training. After the round, the data needs to be processed and sent to the edge server; Using the federated average algorithm, the central server's objective function is: ; in is the number of clients participating in the training, Is a user Datasets, The data The amount of data, is the sum of all user data sets; User training completed After the round, the gradient parameters need to be obtained and Multiply to get ,Again According to the Greedy-NFC encoding scheme mentioned above, it is sent to the middle layer node; if the amount of data trained by the user is equal, the original gradient parameters are directly transmitted without processing the data, and the result is: ; S3.2, the user uploads the data to the edge server according to the Greedy-NFC method. After receiving the data, the edge server performs arithmetic and operation on the same part of the data; Number of times the edge server aggregation is updated ; S3.3 If , the aggregated data is fed back to the user, that is, local aggregation; since the user receives data from multiple edge servers, it is necessary to calculate the federated average of the same part of the data: ; in Refers to Some update parameters, Refers to The number of partially aggregated users, Refers to Partially polymerized The sum of the data of users; if the amount of training data of users is equal, then: ; S3.4 If , the aggregated data is transmitted to the central server, i.e. global aggregation, and the central server calculates the federated average of each part of the data: ; The updated gradient is then transmitted back to the user, who does not need to process the data after receiving it. Similarly, if the amount of data trained by the user is equal, then: ; At the same time, the edge server starts calculating the aggregation times again: ; S3.5 Iteration , until the training requirements are met.

2. The solution according to claim 1, characterized in that: P1 connects one or more edge servers for each user, and each edge server is connected to a three-layer cloud-edge collaborative network structure with a central server.

3. The solution according to claim 2, characterized in that: The network structure is specifically as follows: In the network, the top layer is a central server that performs data aggregation on the parameters of users in the system to further perform federated learning calculations; the middle layer is Edge Servers , , performs some data aggregation operations and forwarding; the bottom layer is Users , , each user randomly connects to a different edge server and uploads his own parameter data to the edge server. The number of connected edge servers is recorded as .

4. The solution according to claim 3, characterized in that: Each user , 1 , the parameter data generated Both symbols, each of which is contained in the alphabet Inside: , .

5. The solution according to claim 1, characterized in that: P2 includes the following steps: S2.

1. Equally divide the parameter data of each user; Per user , Generated parameter data Both symbols, each with user-generated parameter data Divided equally Part, that is, the number of edge servers; Greater than , quilt divisibility; Each user's parameter data Divide equally After the copy, each part of the parameter data contains symbols, arranged in the order of the original data. Part, Part, ..., Part; between the user and the edge server, the data is in parts, that is, symbols, upload and aggregate operations as a whole; S2.2, edge server initialization settings; Setting up the Edge Server The amount of data allocated for , , The initial value is 1; Setting up the Edge Server The search index is , , The initial values ​​are ; Each round from Part starts the loop, where % represents the remainder operation; S2.3, sort and select the edge server to which the data to be allocated in this round of decision; Define the following two rules, according to each edge server , Amount of allocated data To sort the edge servers: like ,but The order of is higher; like and ,but The order of is higher; According to the above order, traverse the edge servers from front to back and determine the edge server Whether the data of the connected user has been transmitted; If no, select ,If so, continue to traverse the next edge server in the sorting and re-execute the judgment until an edge server is selected; S2.4, retrieve the user data connected to the edge server selected in S2.3, and generate an uplink coding transmission scheme through a greedy strategy; S2.5, updating the allocated data volume of the edge server selected in S2.3; S2.6, loop S2.3-S2.5 until the data of all underlying users have been allocated and transmitted to the corresponding edge servers, then end the loop; At this point, the uplink coded transmission scheme between the underlying user and the edge server is obtained, denoted as Greedy-NFC, Greedy Algorithm for Network Function Computing; Before training begins, the central server generates a coding scheme based on the network topology. In each round of training, the user only needs to upload the model update parameters to the edge server according to the corresponding transmission matrix, without having to traverse the above process again.

Citation Information

Patent Citations

  • Edge calculation and resource optimization method based on federated learning

    CN113791895A

  • System and method for distributed learning of wireless edge dynamics

    CN114930347A