A Federated Learning System, Method, Computer Device, and Storage Medium

By segmenting the model training task in the federated learning system, edge terminal devices are responsible for feature extraction, and edge servers are responsible for model parameter update and federated aggregation, which solves the problem of traditional federated learning methods' dependence on the central server and poor risk resistance, and achieves the improvement of data privacy and risk resistance.

CN114462577BActive Publication Date: 2025-06-17STATE GRID INFORMATION & TELECOMM BRANCH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210113913.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-30
Publication Date
2025-06-17
Estimated Expiration
2042-01-30

AI Technical Summary

Technical Problem

Traditional federated learning methods have strong dependence on central servers, poor risk resistance, and risk of privacy leakage.

Method used

Adopting a decentralized federated learning system, by segmenting model training tasks between edge servers and edge terminal devices, edge terminal devices are responsible for feature extraction, and edge servers are responsible for circular iterative updates and federated aggregation of model parameters.

Benefits of technology

It protects the data privacy of edge terminal devices, reduces the computing resource consumption and storage pressure of edge terminal devices, and improves the risk resistance of federated learning systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462577B_ABST
    Figure CN114462577B_ABST
Patent Text Reader

Abstract

The present invention discloses a federated learning system, method, computer device and storage medium. By utilizing the low coupling between the networks inside the neural network model, the model splitting technology is adopted to split the model training into the feature extraction part of the edge terminal device and the parameter iterative update part of the edge server, so that the edge server and the edge terminal device jointly implement the model training. The edge terminal device is only responsible for the feature extraction part of the data, and the edge server is responsible for the cyclic iterative update, sharing and federated aggregation of the model parameters, protecting the data privacy of the edge terminal device and reducing the consumption of computing resources and storage pressure of the edge terminal device. In the federated aggregation stage, the edge server replaces the central server to achieve decentralization, solves the problems of strong dependence on the central server and poor anti-risk ability in the traditional federated learning method, and improves the anti-risk ability of the federated learning system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a federated learning system, method, computer device, and storage medium. Background Art

[0002] With the popularization of "Internet of Everything", the power grid system has realized the integration of energy flow and information flow through widely distributed intelligent terminals and sensing devices.

[0003] However, traditional big data analysis and learning methods need to upload the data information collected by edge terminal devices to a central server for centralized learning, inevitably facing the risk of privacy leakage. Therefore, the federated learning method implemented based on the central server emerged as the times require.

[0004] However, traditional federated learning methods still need to train a complete model in edge terminal devices to obtain model parameters, and upload the model parameters to the central server for model integration to update the model parameters, which is prone to problems such as strong dependence on the central server and poor anti-risk ability. Summary of the Invention

[0005] The present invention provides a federated learning system, method, computer device, and storage medium to solve the problems of strong dependence on the central server and poor anti-risk ability of traditional federated learning methods, and provide a decentralized federated learning system to protect the data privacy of edge terminal devices while improving the anti-risk ability of the system.

[0006] According to one aspect of the present invention, a federated learning system is provided, including: a plurality of edge servers and a plurality of edge terminal devices; each edge server is connected to at least one edge terminal device;

[0007] Any edge server is used to establish two completely identical initial global models, send the pre-trained feature extraction network corresponding to one of the initial global models to the edge terminal devices, and send the initial fully connected network in the other initial global model to other edge servers;

[0008] The edge terminal device is used to input a private data set into the pre-trained feature extraction network to obtain a feature vector, and send the feature vector to the edge server;

[0009] The edge server, as an edge computing node, is used to perform parameter training on the received initial fully connected network based on the feature vector to obtain an edge training model, and synchronize the edge computing model to an aggregation node, where the aggregation node is determined from each of the edge computing nodes based on a preset policy;

[0010] The aggregation node is used to aggregate the edge training models to obtain an aggregated global model. If the aggregated global model does not converge, the initial fully connected network in the initial global model is updated based on the aggregated global model, and the next aggregation node is determined. The updated fully connected network and the identification information of the next aggregation node are synchronized to each edge computing node. If the aggregated global model converges, the aggregated global model is sent to each edge server.

[0011] According to another aspect of the present invention, a federated learning method is provided, which is applied to an edge server in a federated learning system. The method includes:

[0012] Establish two exactly the same initial global models;

[0013] Send the pre-trained feature extraction network corresponding to one of the initial global models to the edge terminal device;

[0014] Send the initial fully connected network in the other initial global model to other edge servers in the federated learning system.

[0015] According to another aspect of the present invention, a federated learning method is provided, which is applied to an edge terminal device in a federated learning system. The method includes:

[0016] Obtain a private data set and receive the pre-trained feature extraction network sent by an edge server in the federated learning system;

[0017] Input the private data set into the pre-trained feature extraction network to obtain feature vectors;

[0018] Send the feature vectors to the edge server.

[0019] According to another aspect of the present invention, a federated learning method is provided, which is applied to an edge server as an edge computing node in a federated learning system. The method includes:

[0020] Receive feature vectors sent by at least one edge terminal device in the federated learning system;

[0021] Perform parameter training on the received initial fully connected network based on the feature vectors to obtain an edge training model;

[0022] Synchronize the edge computing model to the aggregation node, and the aggregation node is determined from each edge computing node in the federated learning system based on a preset policy.

[0023] According to another aspect of the present invention, a federated learning method is provided, which is applied to an edge server as an aggregation node in a federated learning system. The method includes:

[0024] Receive the edge training models sent by each edge computing node in the federated learning system;

[0025] Perform model aggregation on each of the edge training models to obtain an aggregated global model;

[0026] If the aggregated global model does not converge, update the initial fully connected network in the initial global model based on the aggregated global model, determine the next aggregation node, and synchronize the updated fully connected network and the identification information of the next aggregation node to each of the edge computing nodes;

[0027] If the aggregated global model converges, send the aggregated global model to each edge server in the federated learning system.

[0028] According to another aspect of the present invention, there is provided a computer device, which includes:

[0029] At least one processor; and

[0030] A memory communicatively connected to the at least one processor; wherein,

[0031] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the federated learning method according to any embodiment of the present invention.

[0032] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the federated learning method according to any embodiment of the present invention when executed.

[0033] The technical solution of the embodiments of the present invention discloses a federated learning system, method, computer device and storage medium. By utilizing the low coupling between each network inside the neural network model, the model splitting technology is adopted to split the model training into the feature extraction part of the edge terminal device and the parameter iterative update part of the edge server, enabling the edge server and the edge terminal device to jointly implement model training. The edge terminal device is only responsible for the feature extraction part of the data, and the edge server is responsible for the cyclic iterative update, sharing and federated aggregation of the model parameters, protecting the data privacy of the edge terminal device and reducing the computing resource consumption and storage pressure of the edge terminal device. In the federated aggregation stage, the edge server replaces the central server to achieve decentralization, solves the problems of strong dependence on the central server and poor anti-risk ability in the traditional federated learning method, and improves the anti-risk ability of the federated learning system.

[0034] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 is a schematic structural diagram of a federated learning system provided according to Embodiment 1 of the present invention;

[0037] Figure 2 is a flowchart of a federated learning method provided according to Embodiment 2 of the present invention;

[0038] Figure 3 is a flowchart of a federated learning method provided according to Embodiment 3 of the present invention;

[0039] Figure 4 is a flowchart of a federated learning method provided according to Embodiment 4 of the present invention;

[0040] Figure 5 is a flowchart of a federated learning method provided according to Embodiment 5 of the present invention;

[0041] Figure 6 is a schematic structural diagram of a federated learning device provided according to Embodiment 6 of the present invention;

[0042] Figure 7 is a schematic structural diagram of a federated learning device provided according to Embodiment 7 of the present invention;

[0043] Figure 8 is a schematic structural diagram of a federated learning device provided according to Embodiment 8 of the present invention;

[0044] Figure 9 is a schematic structural diagram of a federated learning device provided according to Embodiment 9 of the present invention;

[0045] Figure 10 is a schematic structural diagram of a computer device for implementing the federated learning method of the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0047] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0048] Embodiment 1

[0049] Figure 1 For Embodiment 1 of the present invention, a schematic structural diagram of a federated learning system is provided. This embodiment is applicable to the situation of performing federated learning on a model. As Figure 1 shown, the system includes: a plurality of edge servers 110 and a plurality of edge terminal devices 120; each edge server 110 is connected to at least one edge terminal device 120;

[0050] Any edge server 110 is used to establish two completely identical initial global models, send the pre-trained feature extraction network corresponding to one of the initial global models to each edge terminal device, and send the initial fully connected network in the other initial global model to other edge servers;

[0051] The edge terminal device 120 is used to input the private data set into the pre-trained feature extraction network to obtain feature vectors, and send the feature vectors to the edge server;

[0052] The edge server 110, as an edge computing node, is used to perform parameter training on the received initial fully connected network based on the feature vectors to obtain an edge training model, and synchronize the edge computing model to the aggregation node, and the aggregation node is determined from each edge computing node based on a preset policy;

[0053] The aggregation node is used to perform model aggregation on each edge training model to obtain an aggregated global model. If the aggregated global model has not converged, the initial fully connected network in the initial global model is updated based on the aggregated global model, and the next aggregation node is determined, and the updated fully connected network and the identification information of the next aggregation node are synchronized to each edge computing node; if the aggregated global model has converged, the aggregated global model is sent to each edge terminal device.

[0054] Among them, the edge terminal device 120 is the underlying device in the federated learning system. The edge terminal device 120 can be integrated with a data acquisition module, a communication module, a storage module, or a computing module. Exemplarily, in the power grid system, the edge terminal device 120 can be a smart meter, a smart sensing device, etc. The edge terminal device 120 is connected to the edge server 110, and can transmit the data collected by the edge terminal device 120 or the data obtained through calculation or processing to the edge server 110.

[0055] Compared with the edge terminal device 120, the edge server 110 has more storage and computing resources and is usually deployed on the edge side of the network. In the present invention, while acting as a computing server, the edge server also has the opportunity to act as a central server, responsible for completing the model aggregation and update sent from other edge servers.

[0056] Specifically, an edge server can be arbitrarily selected from the federated learning system according to user requirements or to adapt to application scenario requirements. Two completely identical initial global models are established through the edge server 110. The initial global model is set according to actual requirements. In this embodiment, the architecture and function of the initial global model are not limited. Due to the low coupling between the internal networks of the neural network model, the initial global model can be divided into two relatively independent networks: an initial feature extraction network and an initial fully connected network. For the same reason, the pre-trained global model can also be divided into two relatively independent networks: a pre-trained feature extraction network and a pre-trained fully connected network. The edge server sends the pre-trained feature extraction network obtained after pre-training and model splitting of one of the initial global models to each edge terminal device.

[0057] After receiving the pre-trained feature extraction network sent by the edge server, the edge terminal device 120 is used to input the private data set into the pre-trained feature extraction network to obtain a feature vector, and send the feature vector to the edge server 110.

[0058] It should be noted that in this embodiment, the pre-trained feature extraction network in the edge terminal device is a frozen network, that is, its model parameters will not be updated during the federated process. It should also be noted that the private data set refers to the private data sets collected by each edge terminal device, which is only used for feature extraction in the current edge terminal device, does not need to be shared with other edge terminal devices, and does not need to be aggregated to the server, protecting the data privacy and data security of each edge terminal device.

[0059] In addition, the edge server 110 sends the initial fully connected network obtained after splitting another initial global model to other edge servers in the federated learning system, so that each edge server in the federated learning system has the same initial fully connected network and the feature vectors sent by the edge terminal devices. Thus, as an edge computing node, the edge server 110 inputs the feature vectors into the received initial fully connected network for model training to obtain an edge training model, and synchronizes the edge computing model to the aggregation node. It should be noted that the edge training model is a local model obtained by the edge server based on the private data set.

[0060] Among them, the aggregation node is an edge server selected and determined from all edge servers acting as edge computing nodes based on a preset policy.

[0061] It should be noted that the meaning of the edge computing node synchronizing the edge computing model to the aggregation node is that for an edge computing node that is not the aggregation node, it sends the edge computing model to the aggregation node. For an edge computing node that is the aggregation node and already has its own edge computing model, it only needs to receive the edge computing models sent by other edge computing nodes.

[0062] After the above steps, the aggregation node receives an edge training model separately trained by each edge computing node, and performs model aggregation on each edge computing model to obtain an aggregated global model. The first aggregated global model needs to be trained through iterative loops to converge. Therefore, if the aggregated global model has not converged, the parameters of the initial fully connected network in the initial global model are updated based on the parameters of the aggregated global model, and the next aggregation node is determined. The updated fully connected network and the identification information of the next aggregation node are synchronized to each edge computing node, so that the edge computing node performs parameter training on the received initial fully connected network based on the feature vectors to obtain an edge training model, and synchronizes the edge computing model to the next aggregation node. The above iterative process is executed cyclically until the aggregated global model converges, and a complete aggregated global model is obtained after the model training process. The aggregated global model is sent to each edge server.

[0063] The technical solution of this embodiment utilizes the low coupling between the networks inside the neural network model. Based on the model segmentation technology, the model training process is split into a feature extraction part on the edge terminal device and a parameter iterative update part on the edge server, enabling the edge server and the edge terminal device to jointly implement model training. The edge terminal device is only responsible for the feature extraction part of the data, and the edge server is responsible for tasks that consume a large amount of computing resources, such as the establishment of the initial global model, the cyclic iterative update, sharing, and federated aggregation of model parameters; it protects the data privacy of the edge terminal device and at the same time reduces the computing resource consumption and storage pressure of the edge terminal device; it is particularly suitable for systems with limited computing and storage capabilities of edge terminal devices, such as power grid systems.

[0064] In the federated aggregation stage, the edge server replaces the central server, and each edge computing node has the opportunity to act as an aggregation node to aggregate the edge training models obtained by training each edge computing node, achieving a decentralized effect, solving the problems of strong dependence on the central server and poor risk resistance in traditional federated learning methods, and improving the risk resistance of the federated learning system.

[0065] Optionally, send the pre-trained feature extraction network corresponding to one of the initial global models to the edge terminal device;

[0066] Pre-train one of the initial global models based on the public dataset to obtain a pre-trained global model;

[0067] Perform a model segmentation operation on the pre-trained global model to obtain a pre-trained feature extraction network; the pre-trained feature extraction network is the part of the pre-trained global model except for the pre-trained fully connected network;

[0068] Send the pre-trained feature extraction network to the edge terminal device.

[0069] Among them, the public dataset refers to the dataset that is mutually public or shared among edge terminals; the pre-trained global model obtained by pre-training one of the initial global models based on the public dataset has certain feature extraction capabilities and can perform feature extraction work.

[0070] Intuitively, the model segmentation operation is to segment a complete global neural network model at a certain level according to certain rules. The effectiveness of the neural network model lies in the low coupling between the internal layers of the network, that is, each hidden layer can be executed independently with the output of the previous layer as its input. Therefore, the pre-trained global model obtained by pre-training the initial global model can be subjected to a model segmentation operation to obtain a pre-trained feature extraction network, and the pre-trained feature extraction network is sent to the edge terminal device so that the edge terminal device can use the pre-trained feature extraction network to perform feature extraction of the dataset.

[0071] Since the network structures of global models established based on different types or functions are different, the parts for model segmentation may be different. The main purpose of the model segmentation operation is to split the fully connected network that needs to be iteratively optimized for parameters from the feature extraction network that does not require iterative optimization.

[0072] Exemplarily, for the time series prediction LSTM model for predicting power load, it is split into two parts along the last layer of the LSTM: the LSTM network and the fully connected layer. The LSTM network can abstract key features and patterns from the user's previous records, while the fully connected layer can remap the learned distributed features to the sample label space.

[0073] Optionally, sending the fully connected network of another initial global model to other edge servers includes:

[0074] Performing a model segmentation operation on the initial global model to obtain an initial fully connected network;

[0075] Sending the initial fully connected network to other edge servers.

[0076] Specifically, performing a model segmentation operation on the initial global model based on the same model segmentation operation as the pre-trained global model to obtain an initial fully connected network and an initial feature extraction network, and sending the initial fully connected network to other edge servers so that each edge server stores the same initial fully connected network, thereby enabling network parameter training according to their respective data sets.

[0077] Optionally, determining the next aggregation node includes:

[0078] Obtaining the comprehensive index of each edge computing node;

[0079] Determining the edge computing node corresponding to the highest comprehensive index as the next aggregation node;

[0080] Wherein, the comprehensive index includes at least one of the following: the performance index of the edge computing node, the evaluation index of the edge training model corresponding to the edge computing node, and the deviation degree between the edge training model and the aggregated global model.

[0081] Specifically, the aggregation node determines the edge computing node with the best performance based on the performance of the edge training models obtained by each edge computing node through parameter training of the received initial fully connected network based on the feature vector. This performance is quantified by the comprehensive index of each edge computing node, and the comprehensive index of the edge computing node includes at least one of the following: the performance index of the edge computing node, the evaluation index of the edge training model corresponding to the edge computing node, and the deviation degree between the edge training model and the aggregated global model.

[0082] Exemplarily, the performance indices of the edge computing nodes may include: software performance indices and hardware performance indices. The software performance indices, for example, may be throughput and computing speed, etc.; the hardware performance indices may be, for example, the CPU size and the storage space size, etc. The evaluation indices of the edge training model corresponding to the edge computing nodes may be the accuracy and recall rate of the edge training model, or the amount of private data samples, etc. The deviation degree between the edge training model and the aggregated global model refers to the deviation degree between the edge training model obtained by the edge computing node and the aggregated global model aggregated by the aggregation node. The comprehensive index may be any one of the performance indices of the edge computing nodes, the evaluation indices of the edge training model corresponding to the edge computing nodes, and the deviation degree between the edge training model and the aggregated global model, etc., or the weighted sum of each index.

[0083] Embodiment 2

[0084] Figure 2 The present invention provides a flowchart of a federated learning method for Embodiment 2. This embodiment is applicable to the situation of performing federated learning on a model. This method can be executed by a federated learning device, which can be implemented in the form of hardware and / or software. The federated learning device can be configured in the terminal server in the federated learning system provided in Embodiment 1. As Figure 2 shown, the method includes:

[0085] S210. Establish two completely identical initial global models.

[0086] Among them, the initial global model is set according to actual requirements. This embodiment does not limit the architecture and functions of the initial global model, etc.

[0087] S220. Send the pre-trained feature extraction network corresponding to one of the initial global models to the edge terminal devices.

[0088] Due to the low coupling between the networks inside the neural network model, the pre-trained global model can be divided into two relatively independent networks: the pre-trained feature extraction network and the pre-trained fully connected network. The pre-trained feature extraction network obtained after pre-training and model splitting of one of the initial global models is sent to each edge terminal device, so that each edge server has the same pre-trained feature extraction network to perform feature extraction of the data set.

[0089] S230. Send the initial fully connected network in the other initial global model to other edge servers in the federated learning system.

[0090] For the same reason, another initial global model can be split into two relatively independent networks: an initial feature extraction network and an initial fully connected network. The initial fully connected network is sent to other edge servers in the federated learning system, so that each edge server has the same initial fully connected network for training network parameters.

[0091] The technical solution of this embodiment can build two identical initial global models; send the pre-trained feature extraction network corresponding to one of the initial global models to the edge terminal device; send the initial fully connected network in the other initial global model to other edge servers in the federated learning system. By taking advantage of the low coupling between the networks inside the neural network model and based on the model splitting technology, the model can be split into a feature extraction part and a parameter iterative update part. The edge terminal device executes the feature extraction part, and the edge server executes the parameter iterative update part to jointly implement model training.

[0092] Embodiment III

[0093] Figure 3 The flowchart of a federated learning method provided in Embodiment III of the present invention is applicable to the situation of performing federated learning on a model. This method can be executed by a federated learning device, which can be implemented in the form of hardware and / or software, and the federated learning device can be configured in the edge terminal device in the federated learning system provided in Embodiment I. As Figure 3 shown, the method includes:

[0094] S310. Obtain a private dataset and receive the pre-trained feature extraction network sent by the edge server in the federated learning system.

[0095] The pre-trained feature extraction network is a network for feature extraction obtained by splitting the pre-trained initial global model.

[0096] The private dataset refers to the private datasets collected by each edge terminal device, which is only used for feature extraction in the current edge terminal device, and does not need to be shared with other edge terminal devices or aggregated to the server.

[0097] S320. Input the private dataset into the pre-trained feature extraction network to obtain feature vectors.

[0098] Specifically, the pre-trained feature extraction network has a certain feature extraction ability. By inputting the private dataset into the pre-trained feature extraction network to obtain feature vectors, the data privacy and data security of each edge terminal device are protected.

[0099] S330. Send the feature vectors to the edge server.

[0100] Specifically, the edge terminal device is only responsible for the feature extraction part and sends the feature vector to the edge server. The edge server performs the process of parameter iterative training, reducing the consumption of computing resources and storage pressure of the edge terminal device.

[0101] The technical solution of this embodiment is to obtain a private data set and receive a pre-trained feature extraction network sent by the edge server in the federated learning system; input the private data set into the pre-trained feature extraction network to obtain a feature vector; send the feature vector to the edge server, protecting the data privacy and data security of each edge terminal device and reducing the consumption of computing resources and storage pressure of the edge terminal device; it is especially suitable for systems with limited computing and storage capabilities of edge terminal devices, such as power grid systems.

[0102] Embodiment 4

[0103] Figure 4 This is a flowchart of a federated learning method provided by Embodiment 4 of the present invention. This embodiment is applicable to the situation of performing federated learning on a model. This method can be executed by a federated learning device, which can be implemented in the form of hardware and / or software. The federated learning device can be configured in the edge server serving as an edge computing node in the federated learning system provided by Embodiment 1. As Figure 4 shown, the method includes:

[0104] S410. Receive the feature vectors sent by at least one edge terminal device in the federated learning system.

[0105] Among them, the feature vector is obtained by the feature extraction network in the edge terminal device.

[0106] S420. Perform parameter training on the received initial fully connected network based on the feature vector to obtain an edge training model.

[0107] Specifically, input the feature vector into the initial fully connected network for parameter training to obtain an edge training model, which is a local model trained by the feature vectors extracted by some edge terminal devices based on the private data set.

[0108] S430. Synchronize the edge computing model to the aggregation node, and the aggregation node is determined from the edge computing nodes of the federated learning system based on a preset policy.

[0109] Specifically, the meaning of the edge computing node synchronizing the edge computing model to the aggregation node means that for an edge computing node that is not the aggregation node, send the edge computing model to the aggregation node. For an edge computing node that is the aggregation node and already has its own edge computing model, only need to receive the edge computing models sent by other edge computing nodes.

[0110] Exemplarily, the preset policy may be to elect an aggregation node from each edge computing node based on the performance index of the edge computing node. Among them, the performance index of the edge computing node may include: hardware performance index, such as CPU size and storage space size, etc.; software performance index, such as throughput and computing speed, etc.; or the aggregation node may be determined from each of the edge computing nodes based on a random algorithm.

[0111] The technical solution of this embodiment is to receive the feature vectors sent by at least one edge terminal device in the federated learning system; perform parameter training on the received initial fully connected network based on the feature vectors to obtain an edge training model; synchronize the edge computing model to the aggregation node, and the aggregation node is determined from each edge computing node in the federated learning system based on a preset policy, and perform complex parameter iterative update work in the edge server acting as the edge computing node, reducing the computing resource consumption and storage pressure of the edge terminal device; it is especially applicable to systems where the computing and storage capabilities of edge terminal devices are limited, such as power grid systems.

[0112] Embodiment Five

[0113] Figure 5 The flowchart of a federated learning method provided in Embodiment Five of the present invention is applicable to the situation of performing federated learning on a model. This method can be executed by a federated learning device, which can be implemented in the form of hardware and / or software, and the federated learning device can be configured in the edge server acting as the aggregation node in the federated learning system provided in Embodiment One. As Figure 5 shown, the method includes:

[0114] S510. Receive the edge training models sent by each edge computing node in the federated learning system.

[0115] Specifically, after the relevant connections are established among the edge computing nodes in the federated learning system and the aggregation node is determined from each edge computing node, the aggregation node receives the edge training models sent by each edge computing node.

[0116] S520. Perform model aggregation on each edge training model to obtain an aggregated global model.

[0117] Specifically, any existing aggregation method can be used for model aggregation, such as the optimal model selection method, the selection voting method, or the method of aggregating parameters based on weights. This embodiment does not limit this.

[0118] S530. If the aggregated global model has not converged, update the initial fully connected network in the initial global model based on the aggregated global model, determine the next aggregation node, and synchronize the updated fully connected network and the identification information of the next aggregation node to each edge computing node.

[0119] Among them, the identification information of the next aggregation node may be the identification information of the server serving as the next aggregation node.

[0120] Specifically, after the aggregation node aggregates the edge training models to obtain an aggregated global model, it determines whether the aggregated global model converges. If not, it updates the network parameters of the initial fully connected network in the initial global model based on the model parameters of the aggregated global model. In addition, it determines the next aggregation node from each edge computing node, obtains the identification information of the next aggregation node, and synchronizes the updated fully connected network and the identification information of the next aggregation node to each edge computing node.

[0121] Exemplarily, the method for determining the next aggregation node may be to determine the edge computing node with the best performance or model training performance as the next aggregation node. For example, determine the edge computing nodes corresponding to the highest comprehensive index as the next aggregation node. The comprehensive index may include at least one of the following: the performance index of the edge computing node, the evaluation index of the edge training model corresponding to the edge computing node, and the deviation degree between the edge training model and the aggregated global model.

[0122] S540. If the aggregated global model converges, send the aggregated global model to each edge server in the federated learning system.

[0123] Specifically, if the aggregated global model converges, it means that the training of the aggregated global model is complete, and the aggregated global model is sent to each edge server in the federated learning system. Each edge service can store the aggregated global model or distribute it to the terminal device.

[0124] The technical solution of this embodiment receives the edge training models sent by each edge computing node in the federated learning system; aggregates the edge training models to obtain an aggregated global model; if the aggregated global model does not converge, updates the initial fully connected network in the initial global model based on the aggregated global model, determines the next aggregation node, and synchronizes the updated fully connected network and the identification information of the next aggregation node to each edge computing node; if the aggregated global model converges, sends the aggregated global model to each edge server in the federated learning system, enabling each edge computing node to have the opportunity to become the aggregation center. The edge server replaces the central server to perform complex model loop iteration, update, and aggregation, realizing decentralization, solving the problems of strong dependence on the central server and poor anti-risk ability in traditional federated learning methods, and improving the anti-risk ability of the federated learning system.

[0125] Embodiment Six

[0126] Figure 6 This is a schematic structural diagram of a federated learning device provided in Embodiment Six of the present invention. AsFigure 6 As shown in the figure, the device can be configured in the terminal server in the federated learning system provided in the first embodiment. The device includes: a model establishment module 610, a first sending module 620, and a second sending module 630;

[0127] Among them, the model establishment module 610 is used to establish two completely identical initial global models;

[0128] The first sending module 620 is used to send the pre-trained feature extraction network corresponding to one of the initial global models to the edge terminal device;

[0129] The second sending module 630 is used to send the initial fully connected network in the other initial global model to other edge servers in the federated learning system.

[0130] The federated learning device provided by the embodiment of the present invention can execute the federated learning method provided by the first or second embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0131] Embodiment Seven

[0132] Figure 7 It is a schematic structural diagram of a federated learning device provided by the seventh embodiment of the present invention. As Figure 7 shown in the figure, the device can be configured in the edge terminal device in the federated learning system provided in the first embodiment. The device includes: a data acquisition module 710, a feature extraction module 720, and a sending module 730;

[0133] Among them, the data acquisition module 710 is used to acquire a private data set and receive the pre-trained feature extraction network sent by the edge server in the federated learning system;

[0134] The feature extraction module 720 is used to input the private data set into the pre-trained feature extraction network to obtain a feature vector;

[0135] The sending module 730 is used to send the feature vector to the edge server.

[0136] The federated learning device provided by the embodiment of the present invention can execute the federated learning method provided by the first or third embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0137] Embodiment Eight

[0138] Figure 8 It is a schematic structural diagram of a federated learning device provided by the eighth embodiment of the present invention. As Figure 8As shown in the figure, the device can be configured in the edge server serving as an edge computing node in the federated learning system provided in the first embodiment. The device includes: a receiving module 810, a parameter training module 820, and a synchronization module 830;

[0139] Among them, the receiving module 810 is used to receive the feature vectors sent by at least one edge terminal device in the federated learning system.

[0140] The parameter training module 820 is used to perform parameter training on the received initial fully connected network based on the feature vectors to obtain an edge training model.

[0141] The synchronization module 830 is used to synchronize the edge computing model to the aggregation node, and the aggregation node is determined from each edge computing node in the federated learning system based on a preset policy.

[0142] The federated learning device provided in the embodiment of the present invention can execute the federated learning method provided in the first or fourth embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0143] Embodiment Nine

[0144] Figure 9 As shown in the figure, it is a schematic structural diagram of a federated learning device provided in the ninth embodiment of the present invention. As Figure 9 shown, the device can be configured in the edge server serving as an aggregation node in the federated learning system provided in the first embodiment. The device includes: a receiving module 910, a model aggregation module 920, an update module 930, and a sending module 940;

[0145] Among them, the receiving module 910 is used to receive the edge training models sent by each edge computing node in the federated learning system;

[0146] The model aggregation module 920 is used to perform model aggregation on each edge training model to obtain an aggregated global model;

[0147] The update module 930 is used to, if the aggregated global model has not converged, update the initial fully connected network in the initial global model based on the aggregated global model, and determine the next aggregation node, and synchronize the updated fully connected network and the identification information of the next aggregation node to each edge computing node.

[0148] The sending module 940 is used to, if the aggregated global model has converged, send the aggregated global model to each edge server in the federated learning system.

[0149] The federated learning device provided in the embodiment of the present invention can execute the federated learning method provided in the first or fifth embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0150] Embodiment Ten

[0151] Figure 10 FIG. shows a schematic structural diagram of a computer device 10 that can be used to implement the embodiments of the present invention. The computer device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The computer device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0152] As Figure 10 shown, the computer device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the computer device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0153] Multiple components in the computer device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the computer device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0154] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the federated learning method.

[0155] In some embodiments, the federated learning method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the computer device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the federated learning method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the federated learning method by any other suitable means (e.g., by means of firmware).

[0156] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0157] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0158] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0159] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0160] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0161] A computing system may include edge terminal devices and a server. The edge terminal devices and the server are generally far from each other and usually interact via a communication network. The relationship between the edge terminal devices and the server is generated by computer programs running on corresponding computers and having an edge terminal device-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0162] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0163] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A federated learning system, characterized in that, Including: A plurality of edge servers and a plurality of edge terminal devices; each edge server is connected to at least one edge terminal device; Any edge server is used to establish two identical initial global models, send the pre-training feature extraction network corresponding to one of the initial global models to the edge terminal device, and send the initial fully connected network in the other initial global model to other edge servers; The edge terminal device is used to input the private data set into the pre-training feature extraction network to obtain a feature vector, and send the feature vector to the edge server; The edge server serves as an edge computing node, and is used to perform parameter training on the received initial fully connected network based on the feature vector to obtain an edge training model, and synchronize the edge computing model to the aggregation node, and the aggregation node is determined from each of the edge computing nodes based on a preset policy; The aggregation node is used to perform model aggregation on each of the edge training models to obtain an aggregated global model. If the aggregated global model does not converge, update the initial fully connected network in the initial global model based on the aggregated global model, and determine the next aggregation node, and synchronize the updated fully connected network and the identification information of the next aggregation node to each of the edge computing nodes; If the aggregated global model converges, send the aggregated global model to each of the edge servers.

2. The system according to claim 1, characterized in that, Sending the pre-training feature extraction network corresponding to one of the initial global models to the edge terminal device; Pre-training one of the initial global models based on a public data set to obtain a pre-trained global model; Performing a model splitting operation on the pre-trained global model to obtain a pre-training feature extraction network; the pre-training feature extraction network is the part of the pre-trained global model other than the pre-trained fully connected network; Sending the pre-training feature extraction network to the edge terminal device.

3. The system according to claim 2, characterized in that, The sending the fully connected network of the other initial global model to other edge servers includes: Performing the model splitting operation on the initial global model to obtain an initial fully connected network; Sending the initial fully connected network to other edge servers.

4. The system according to claim 1, characterized in that, The determining the next aggregation node includes: Obtaining the comprehensive index of each of the edge computing nodes; Determining the edge computing node corresponding to the highest comprehensive index as the next aggregation node; Wherein, the comprehensive index includes at least one of the following: the performance index of the edge computing node, the evaluation index of the edge training model corresponding to the edge computing node, and the deviation degree between the edge training model and the aggregated global model.

5. A federated learning method, characterized in that, An edge server applied to the federated learning system according to any one of claims 1-4, the method includes: Establishing two identical initial global models; Sending the pre-training feature extraction network corresponding to one of the initial global models to the edge terminal device; Sending the initial fully connected network in the other initial global model to other edge servers in the federated learning system.

6. A federated learning method, characterized in that, An edge terminal device applied to the federated learning system according to any one of claims 1-4, the method includes: Obtain a private dataset and receive a pre-trained feature extraction network sent by an edge server in a federated learning system; Input the private dataset into the pre-trained feature extraction network to obtain feature vectors; Send the feature vectors to the edge server.

7. A federated learning method, characterized in that, An edge server applied as an edge computing node in the federated learning system according to any one of claims 1-4, the method comprising: Receive feature vectors sent by at least one edge terminal device in the federated learning system; Perform parameter training on the received initial fully connected network based on the feature vectors to obtain an edge training model; Synchronize the edge computing model to an aggregation node, where the aggregation node is determined from each of the edge computing nodes in the federated learning system based on a preset policy.

8. A federated learning method, characterized in that, An edge server applied as an aggregation node in the federated learning system according to any one of claims 1-4, the method comprising: Receive edge training models sent by each edge computing node in the federated learning system; Perform model aggregation on each of the edge training models to obtain an aggregated global model; If the aggregated global model has not converged, update the initial fully connected network in the initial global model based on the aggregated global model, and determine the next aggregation node, and synchronize the updated fully connected network and the identification information of the next aggregation node to each of the edge computing nodes; If the aggregated global model has converged, send the aggregated global model to each edge server in the federated learning system.

9. A computer device, characterized in that, The computer device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the federated learning method according to any one of claims 5-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the federated learning method according to any one of claims 5-8 when executed.

Citation Information

Patent Citations

  • Hierarchical federated learning method and device based on asynchronous communication, terminal equipment and storage medium

    CN112532451A

  • Federal learning method and device, equipment and storage medium

    CN113516250A