Federated Learning Platform, Method, Device and Storage Medium Based on Distributed Architecture
The distributed federated learning architecture addresses data transmission and compatibility issues by using a main node to manage data aggregation and distribution, enhancing stability and efficiency in federated learning systems.
Patent Information
- Application Number
- CN202011280929.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-11-16
AI Technical Summary
The existing federated learning framework has problems such as long data transmission chains, many transmission nodes, poor universality and poor compatibility of machine learning algorithms. Especially when the public network is unstable, the data transmission failure rate is high, and the data interaction pressure is high in point-to-point mode.
The federated learning platform based on distributed architecture is adopted, and the master-slave architecture is used. The master node is responsible for the aggregation and distribution of data, and the child node is responsible for the data loading and model training. In each round of training, the master node only needs to request the child node once, ensure data security through public-private key encryption, and transfer data interaction through the master node.
It reduces the failure rate of public network requests, improves the stability and efficiency of the system, reduces the amount of data transmission, enhances the universality of the system and the compatibility of machine learning algorithms, and improves the security and reliability of data interaction.
Smart Images

Figure CN113923225B_ABST
Abstract
Description
Background Art
[0002] In existing federated learning frameworks, a peer node-based (similar to p2p) framework is usually adopted. After each client finishes processing, it sends the data to all other clients, or the data is passed sequentially among the clients (ring). However, the above-mentioned peer node-based federated learning framework has at least the following technical problems:
[0003] (1) Long data transmission chain: In a ring architecture, data needs to be transmitted node by node in sequence. When the network of the public network is unstable, the failure rate of data transmission is relatively high.
[0004] (2) Many transmission nodes: If a broadcast architecture is adopted, since each node needs to send data to all other nodes, the data transmission volume is large.
[0005] (3) Poor generality. For example, when multiple customers want to perform training on a recognized trusted platform, the point-to-point mode cannot be realized.
[0006] (4) In the point-to-point mode, the compatibility of machine learning algorithms is poor. Especially for algorithms with features not on the same client or hybrid algorithms, it is necessary to send messages bidirectionally for multiple interactive negotiations, resulting in a large data interaction pressure.
[0007] It should be noted that the information disclosed in the above background art section is only used to strengthen the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0008] The purpose of the present disclosure is to provide a federated learning platform, method, device, and storage medium based on a distributed architecture, which at least overcome the problems of large data transmission pressure and large interaction volume in the related art to a certain extent.
[0009] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0010] According to one aspect of the present disclosure, a federated learning platform based on a distributed architecture is provided, including: sub-node devices, which acquire partial feature data for performing federated learning, and the partial feature data of all sub-node devices, when aggregated, includes all feature data for federated learning; a master node device, connected to the sub-node devices, for receiving a request instruction for federated learning from the sub-node devices, acquiring partial feature data sent by a first type of sub-node devices in the sub-node devices according to the request instruction, and distributing the partial feature data to a second type of sub-node devices in the sub-node devices according to the request instruction, and the second type of sub-node devices perform partial federated learning in federated learning according to the received partial feature data.
[0011] In one embodiment of the present disclosure, data interaction between any two of the sub-node devices is relayed through the master node device.
[0012] In one embodiment of the present disclosure, the sub-node device determines a corresponding public-private key pair according to the request instruction, and feeds back the public key to the master node device. The sub-node device encrypts the partial feature data to be sent using the private key, and forwards the encrypted partial feature data to the master node device.
[0013] In one embodiment of the present disclosure, the sub-node device performs partial federated learning based on partial feature data, and the training of the partial federated learning includes at least one of horizontal training, vertical training, and transfer training.
[0014] In one embodiment of the present disclosure, the master node device is further configured to obtain the progress of the sub-node device in performing partial federated learning, and obtain the output result of the partial federated learning.
[0015] In one embodiment of the present disclosure, the master node device is further configured to obtain the running state information of the sub-node device in performing partial federated learning, and the running state information includes the data interaction volume and / or the resource occupancy rate.
[0016] In one embodiment of the present disclosure, the federated learning platform based on a distributed architecture further includes: a metadata database connected to the sub-node device. The metadata database stores metadata, and the sub-node device obtains feature data by preprocessing the metadata.
[0017] In one embodiment of the present disclosure, the master node device is further configured to determine a federated learning training task, allocate partial federated learning training tasks to the sub-node devices according to the federated learning training task, and perform interaction of training data among the sub-node devices according to the federated learning training task.
[0018] In one embodiment of the present disclosure, the metadata database includes at least one of a Redis database, an Hbase database, an Elasticsearch database, and an Oracle database.
[0019] According to another aspect of the present disclosure, a federated learning method based on a distributed architecture is provided, including: obtaining partial feature data for performing federated learning sent by sub-node devices, and the partial feature data of all the sub-node devices contains all the feature data of the federated learning after aggregation; receiving a request instruction for federated learning sent by the sub-node devices; obtaining the partial feature data sent by the first type of sub-node devices in the sub-node devices according to the request instruction, and distributing the partial feature data to the second type of sub-node devices in the sub-node devices according to the request instruction for the second type of sub-node devices to perform partial federated learning in the federated learning according to the received partial feature data, where the sub-node devices are the first type of sub-node devices or the second type of sub-node devices.
[0020] In an embodiment of the present disclosure, the federated learning method based on a distributed architecture further includes: obtaining a data interaction request and a data packet sent by the sub-node devices; parsing the data interaction request to determine a target sub-node device; forwarding the data packet to the target sub-node device, where the target sub-node device is the first type of sub-node devices or the second type of sub-node devices.
[0021] In an embodiment of the present disclosure, the federated learning method based on a distributed architecture further includes: receiving a public key sent by the sub-node devices, where the sub-node devices store a private key that matches the public key; receiving the partial feature data encrypted by the sub-node devices using the private key; and performing parsing processing on the encrypted partial feature data using the public key.
[0022] In an embodiment of the present disclosure, the federated learning method based on a distributed architecture further includes: receiving an intermediate result of an encrypted gradient value and a loss value calculated by the sub-node devices; determining a gradient value according to the intermediate result of the encrypted gradient value and the loss value; and feeding back the gradient value to the sub-node devices for the sub-node devices to update the local machine learning model according to the gradient value.
[0023] According to still another aspect of the present disclosure, an electronic device is provided, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the federated learning method based on a distributed architecture as described in any one of the above via executing the executable instructions.
[0024] According to yet another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the federated learning method based on a distributed architecture as described in any one of the above is implemented.
[0025] The federated learning solution provided by the embodiments of the present disclosure is based on a distributed master-slave architecture. The master node is responsible for operations such as cluster maintenance, data aggregation and distribution, and the slave nodes are responsible for operations such as data loading and model training. In each round of training, the master node only needs to request each slave node once, reducing the number of public network requests and thus reducing the failure rate. At the same time, the master node can always sense whether the data transmitted by the slave nodes is successful, and even if it fails, it can be solved by retrying.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0028] Figure 1A A schematic diagram showing the structure of a federated learning system in an embodiment of the present disclosure;
[0029] Figure 1B A schematic diagram showing another structure of a federated learning system in an embodiment of the present disclosure;
[0030] Figure 2 A schematic diagram showing the architecture of a federated learning platform based on a distributed architecture in an embodiment of the present disclosure;
[0031] Figure 3 A schematic diagram showing another federated learning method based on a distributed architecture in an embodiment of the present disclosure;
[0032] Figure 4 A schematic diagram showing another federated learning method based on a distributed architecture in an embodiment of the present disclosure;
[0033] Figure 5 A schematic diagram showing another federated learning method based on a distributed architecture in an embodiment of the present disclosure;
[0034] Figure 6 A schematic diagram showing another federated learning method based on a distributed architecture in an embodiment of the present disclosure;
[0035] Figure 7 A schematic diagram showing an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0037] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0038] The solution provided by this application is based on a distributed master-slave architecture. The master node is responsible for operations such as cluster maintenance, data aggregation and distribution, and the slave nodes are responsible for operations such as data loading and model training. In each round of training, the master node only needs to request each slave node once, reducing the number of public network requests and thus reducing the failure rate. At the same time, the master node can always sense whether the data transmitted by the slave nodes is successful, and even if it fails, it can be resolved by retrying.
[0039] For different data sets, federated learning is divided into horizontal federated learning, vertical federated learning, and federated transfer learning.
[0040] (1) In horizontal federated learning, when there is a large overlap in user features between two data sets but a small overlap in users, we divide the data sets horizontally (i.e., along the user dimension) and extract the part of the data where the user features of both parties are the same but the users are not completely the same for training. This method is called horizontal federated learning. For example, there are two banks in different regions. Their user groups come from their respective regions and have little intersection. However, their businesses are very similar, so the user features recorded are the same. At this time, we can use horizontal federated learning to build a joint model. For instance, when a single user uses an Android phone, the model parameters are continuously updated locally and uploaded to the Android cloud, enabling the data owners with the same feature dimensions to build a joint model.
[0041] (2) In the case where there is a large overlap in users but a small overlap in user features between two datasets, we split the datasets vertically (i.e., along the feature dimension) and select the part of the data where the users are the same but the user features are not completely the same for training. This method is called vertical federated learning. For example, there are two different institutions, one is a bank in a certain place and the other is an e-commerce company in the same place. Their user groups are very likely to include most of the residents in that area, so the intersection of users is large. However, since the bank records the income and expenditure behaviors and credit ratings of users, while the e-commerce company keeps the browsing and purchase histories of users, the intersection of their user features is small. Vertical federated learning is to aggregate these different features in an encrypted state to enhance the model's ability. Currently, many machine learning models such as logistic regression models, tree structure models, and neural network models have gradually been proven to be able to be built on this federated system.
[0042] (3) In the case where there is a small overlap in both users and user features between two datasets, instead of splitting the data, transfer learning is used to overcome the situation of insufficient data or labels. This method is called federated transfer learning. For example, there are two different institutions, one is a bank located in China and the other is an e-commerce company located in the United States. Due to geographical restrictions, the intersection of the user groups of these two institutions is very small. At the same time, due to the different types of institutions, only a small part of their data features overlap. In this case, to conduct effective federated learning, it is necessary to introduce transfer learning to solve the problems of small unilateral data scale and few label samples, thereby improving the effect of the model.
[0043] The solution provided in the embodiments of this application involves technologies such as federated learning and platform architecture, and will be specifically described through the following embodiments.
[0044] Figure 1A Shows a flowchart of a federated learning platform based on a distributed architecture in an embodiment of the present disclosure. The method provided in the embodiments of the present disclosure can be executed by any electronic device with computing and processing capabilities, such as Figure 1A the first terminal 104 shown (or such as Figure 1A the second terminal 106 shown) and / or the server cluster 102. In the following illustrative examples, the first terminal 104 (or such as Figure 1A the second terminal 106 shown) is used as the execution subject for example.
[0045] Figure 1A Shows a schematic structural diagram of a federated learning system in an embodiment of the present disclosure, including a server cluster 102 and multiple first terminals 104 (or such as Figure 1A the second terminal 106 shown).
[0046] The first terminal 104 (or such as Figure 1AThe second terminal 106) shown may be a mobile terminal such as a mobile phone, a game console, a tablet computer, an e-book reader, smart glasses, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a smart home device, an AR (Augmented Reality) device, a VR (Virtual Reality) device, etc. Alternatively, the first terminal 104 (or as Figure 1A the second terminal 106) shown may also be a personal computer (PC), such as a laptop and a desktop computer, etc.
[0047] Among them, an application program for providing federated learning may be installed in the first terminal 104 (or the second terminal 106 as Figure 1A shown).
[0048] The first terminal 104 (or the second terminal 106 as Figure 1A shown) is connected to the server cluster 102 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0049] The first type of private data 108 of the first terminal 104 and the second type of private data 110 of the second terminal 106 cannot directly perform interactive processing 112, but after performing encrypted sample alignment processing 114 between the first terminal 104 and the second terminal 106, data interaction can be carried out.
[0050] As Figure 1B shown, after the encrypted sample alignment processing 114, a first aligned data end 118 and a second aligned data end 116 are obtained, and federated learning between the first aligned data end 118 and the second aligned data end 116 is carried out through the assisting end 120. Encrypted interaction 124 of intermediate results is carried out between the first aligned data end 118 and the second aligned data end 116. The assisting end 120 distributes a public key 122 to the first aligned data end 118, the first aligned data end 118 sends encrypted aggregated gradients and losses 126 to the assisting end 120, the assisting end 120 sends an instruction 128 to update the models of both sides to the first aligned data end 118, and the second aligned data end 116 and the assisting end 120 also perform the above processing.
[0051] The server cluster 102 is a single server, or consists of several servers, or is a virtualization platform, or is a cloud computing service center. The server cluster 102 is used to provide background services for the application program that provides federated learning. Optionally, the server cluster 102 undertakes the main computing work, and the first terminal 104 (or as Figure 1AThe second terminal 106 shown undertakes secondary computing tasks; or, the server cluster 102 undertakes secondary computing tasks, and the first terminal 104 (or the second terminal 106 shown as Figure 1A undertakes primary computing tasks; or, the first terminal 104 (or the second terminal 106 shown as Figure 1A and the server cluster 102 perform collaborative computing using a distributed computing architecture.
[0052] In some alternative embodiments, the server cluster 102 is used to store federated learning information, such as images to be detected, a reference image library, and images for which detection has been completed.
[0053] Optionally, the clients of the application programs installed in different first terminals 104 (or the second terminal 106 shown as Figure 1A are the same, or the clients of the application programs installed on two first terminals 104 (or the second terminal 106 shown as Figure 1A are clients of the same type of application program on different control system platforms. Depending on the differences in the terminal platforms, the specific forms of the clients of the application programs can also be different. For example, the clients of the application programs can be mobile phone clients, PC clients, or World Wide Web (Web) clients, etc.
[0054] Those skilled in the art can be aware that the number of the above-mentioned first terminals 104 (or the second terminal 106 shown as Figure 1A can be more or less. For example, there can be only one such terminal, or there can be dozens or hundreds of such terminals, or even more. The embodiments of the present application do not limit the number and device types of the terminals.
[0055] Optionally, the system can further include a management device, which is connected to the server cluster 102 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0056] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0057] Next, each step of the federated learning platform based on the distributed architecture in this exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.
[0058] As Figure 2 shown, the federated learning platform based on the distributed architecture includes: a sub-node device 212 that obtains partial feature data for federated learning, and the partial feature data of all sub-node devices 212 contains all the feature data for federated learning after aggregation; a master-node device 202 that is connected to the sub-node device 212, is used to receive a request instruction for federated learning from the sub-node device 212, obtains the partial feature data sent by the first type of sub-node device 212 in the sub-node device 212 according to the request instruction, and distributes the partial feature data to the second type of sub-node device 212 in the sub-node device 212 according to the request instruction, and the second type of sub-node device 212 performs partial federated learning in the federated learning according to the received partial feature data.
[0059] In the above embodiments, in the federated learning platform based on the master-slave architecture, each sub-node device 212 only interacts with the master-node device 202. The master-node device 202 performs data aggregation and distribution, and only needs to request each sub-node device 212 once in each round of training. At the same time, the master-node device 202 can statistically analyze the network and other resource conditions of each sub-node device 212 in real time and actively retry in case of failure.
[0060] Furthermore, compared with the peer-to-peer federated learning platform in the prior art, the federated learning platform based on the distributed architecture of the present disclosure has improvements in terms of stability, efficiency, etc., and is also more convenient than other modes in subsequent distributed expansion and data statistics.
[0061] In an embodiment of the present disclosure, the data interaction between any two sub-node devices 212 is relayed through the master-node device 202.
[0062] In an embodiment of the present disclosure, the sub-node device 212 determines the corresponding public-private key pair according to the request instruction, and feeds back the public key to the master-node device 202. The sub-node device 212 encrypts the partial feature data to be sent with the private key and forwards the encrypted partial feature data to the master-node device 202.
[0063] In an embodiment of the present disclosure, the sub-node device 212 determines the corresponding public-private key pair according to the request instruction and feeds back the public key to the master-node device 202.
[0064] In the above embodiments, the sub-node device 212 is mainly responsible for local data training, model generation and storage, etc. The sub-node device 212 can perform distributed federated learning or multi-level multi-threaded inference, and perform data summarization and aggregation through the master-node device 202. The sub-node device 212 further improves the efficiency, reliability and security of federated learning by detecting the results of data plugging, data monitoring, leakage alarms and data access and reporting them to the master-node device 202 in real time.
[0065] In an embodiment of the present disclosure, the sub-node device 212 performs partial federated learning based on partial feature data, and the training of the partial federated learning includes at least one of horizontal training, vertical training and transfer training.
[0066] In the above embodiments, the sub-node device 212 performs partial federated learning based on partial feature data. The feature data is distributed among multiple clients or in a hybrid algorithm, and no additional steps are required. Multiple sub-node devices 212 cooperate with each other to implement federated learning, and the master-node device 202 coordinates the federated learning process.
[0067] In one embodiment of the present disclosure, the master node device 202 is further configured to obtain the progress of the partial federated learning executed by the slave node device 212, and obtain the output result of the partial federated learning.
[0068] In the above embodiment, the master node device 202 is further configured to obtain the progress of the partial federated learning executed by the slave node device 212, and obtain the output result of the partial federated learning, so as to improve the efficiency of the federated learning. Additionally, if the response time of the output result of the federated learning is too long, the master node device 202 may instruct the slave node device 212 to re-upload the output result or re-perform the partial federated learning.
[0069] In one embodiment of the present disclosure, the master node device 202 is further configured to obtain the running status information of the partial federated learning executed by the slave node device 212, and the running status information includes the data interaction volume and / or the resource occupancy rate.
[0070] In the above embodiment, the master node device 202 is further configured to obtain the running status information of the partial federated learning executed by the slave node device 212, and the running status information includes the data interaction volume and / or the resource occupancy rate, to monitor the running pressure of the federated learning platform, so as to reduce the possibility of the federated learning platform crashing.
[0071] In one embodiment of the present disclosure, the federated learning platform based on a distributed architecture further includes: a metadata database, connected to the slave node device 212, and the metadata database stores metadata, and the slave node device 212 obtains feature data by preprocessing the metadata.
[0072] In the above embodiment, the metadata database includes at least one of a Redis database, an Hbase database, an Elasticsearch database, and an Oracle database. Specifically:
[0073] (1) A Redis (Remote Dictionary Server) server (such as Figure 2 204A or 204B shown), that is, a remote dictionary service, is an open-source log-type, Key-Value database written in ANSI C language, supporting the network, and can be based on memory or persistent, and provides APIs in multiple languages.
[0074] (2) An HBase–Hadoop Database (such as Figure 2 206A or 206B shown) is a highly reliable, high-performance, column-oriented, scalable distributed storage system, and a large-scale structured storage cluster can be built on low-cost PC Servers using HBase technology.
[0075] (3) ES (Elasticsearch) database (such as 208A or 208B shown in Figure 2 ) is an open-source distributed search engine. At the same time, the ES database is also a distributed document database, where each field can be indexed, and the data of each field can be searched, and it can be horizontally extended to hundreds of servers for storing and processing PB-level data.
[0076] The ES database is designed for high availability and scalability, and can store, search, and analyze a large amount of data in a very short time. It is usually used as the core engine in complex search scenarios.
[0077] (4) Oracle database (such as 210A or 210B shown in Figure 2 ), as a general database system, it has complete data management functions. As a relational database, it is a complete relational product. As a distributed database, it realizes distributed processing functions. In an embodiment of the present disclosure, the master node device 202 is further configured to determine a federated learning training task, and allocate part of the federated learning training tasks to the slave node device 212 according to the federated learning training task, and perform interaction of training data among the slave node devices 212 according to the federated learning training task.
[0078] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0079] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, method, or program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0080] Figure 3 Shows a flowchart of a federated learning platform based on a distributed architecture in an embodiment of the present disclosure. The method provided by the embodiment of the present disclosure can be executed by any electronic device with computing and processing capabilities.
[0081] As shown in Figure 3 , when an electronic device executes a federated learning platform based on a distributed architecture, it includes the following steps:
[0082] Step S302: Obtain partial feature data sent by the child node devices for federated learning. After aggregating the partial feature data of all the child node devices, it contains all the feature data for the federated learning.
[0083] Step S304: Receive a request instruction for federated learning sent by the child node devices.
[0084] Step S306: Obtain partial feature data sent by the first type of child node devices among the child node devices according to the request instruction.
[0085] Step S308: Distribute the partial feature data to the second type of child node devices among the child node devices according to the request instruction for the second type of child node devices to perform partial federated learning in the federated learning based on the received partial feature data. The child node devices are the first type of child node devices or the second type of child node devices.
[0086] In the above embodiment, in the federated learning platform based on the master - slave architecture, each child node device only interacts with the master node device. The master node device performs data aggregation and distribution, and only needs to request each child node device once in each round of training. At the same time, the master node device can real - time statistically analyze the network and other resource conditions of each child node device and actively retry in case of failure.
[0087] Furthermore, compared with the peer - to - peer federated learning platform in the prior art, the federated learning platform based on the distributed architecture in the present disclosure has improvements in many aspects such as stability and efficiency, and is also more convenient than other modes in subsequent distributed expansion and data statistics.
[0088] In Figure 3 Based on the method steps shown, as Figure 4 shown, the federated learning method based on the distributed architecture further includes:
[0089] Step S402: Obtain a data interaction request and a data packet sent by the child node devices.
[0090] Step S404: Analyze the data interaction request to determine the target child node device.
[0091] Step S406: Forward the data packet to the target child node device. The target child node device is the first type of child node device or the second type of child node device.
[0092] In the above embodiment, the master node device is used to forward the data packets between the child node devices, and the master node device is used to forward the interaction data between the child node devices, which is beneficial to improving the data security and reliability of the child node devices.
[0093] InFigure 3 Based on the method steps shown, such as Figure 5 shown, the federated learning method based on a distributed architecture further includes:
[0094] Step S502, receiving the public key sent by the child node device, where the child node device stores a private key that matches the public key,
[0095] Step S504, receiving the partial feature data encrypted by the child node device using the private key.
[0096] Step S506, parsing and processing the encrypted partial feature data using the public key.
[0097] In the above embodiment, by generating a public-private key pair in the child node device and sending the public key to the master node device after the child node device stores the private key, the master node device can perform data packet verification and decryption processing.
[0098] In Figure 3 Based on the method steps shown, such as Figure 6 shown, the federated learning method based on a distributed architecture further includes:
[0099] Step S602, receiving the intermediate result of the encrypted gradient value and the loss value calculated by the child node device.
[0100] Step S604, determining the gradient value according to the intermediate result of the encrypted gradient value and the loss value.
[0101] Step S606, feeding back the gradient value to the child node device for the child node device to update the local machine learning model according to the gradient value.
[0102] In the above embodiment, steps S602, S604, and S606 are iterated until the loss function converges, thus completing the entire training process. During the sample alignment and model training process, the partial feature data of the child node device is retained locally, and data interaction during training does not cause data privacy leakage.
[0103] Furthermore, by feeding back the gradient value to the child node device for the child node device to update the local machine learning model according to the gradient value, the update of the machine learning model is realized, further improving the accuracy and reliability of federated learning.
[0104] Next, the electronic device 700 according to this embodiment of the present invention will be described with reference to Figure 7 The electronic device 700 shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention. Figure 7
[0105] As shown Figure 7 in FIG. 700, the electronic device 700 takes the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: at least one of the above-mentioned processing units 710, at least one of the above-mentioned storage units 720, and a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710).
[0106] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 710, so that the processing unit 710 executes the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification. For example, the processing unit 710 can execute all the steps of the federated learning based on the distributed architecture of the present disclosure, as well as other steps defined in the federated learning platform based on the distributed architecture of the present disclosure.
[0107] The storage unit 720 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 7201 and / or a cache storage unit 7202, and may further include a read-only storage unit (ROM) 7203.
[0108] The storage unit 720 may further include a program / utilities 7204 having a set (at least one) of program modules 7205. Such program modules 7205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.
[0109] The bus 730 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0110] The electronic device 700 can also communicate with one or more external devices 740 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device, and / or communicate with any device (such as a router, a modem, etc.) that enables the electronic device 700 to communicate with one or more other computing devices. Such communication can be carried out through the input / output (I / O) interface 750. Moreover, the electronic device 700 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 760. As shown in the figure, the network adapter 760 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0111] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0112] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which there is a program product capable of implementing the above method of this specification. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "exemplary method" section of this specification.
[0113] A program product for implementing the above method according to an embodiment of the present invention is described. It can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0114] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0115] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, and the like, or any suitable combination of the foregoing.
[0116] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0117] It should be noted that although several modules or units of a device for action execution are mentioned in the foregoing detailed description, such a division is not mandatory. In fact, according to embodiments of the present disclosure, the features and functions of two or more of the above-mentioned modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by a plurality of modules or units.
[0118] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in that specific order, or that all of the shown steps must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0119] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0120] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A federated learning platform based on a distributed architecture, characterized in that, Including: Sub-node devices, which obtain partial training data for federated learning. The partial training data of all the sub-node devices, when aggregated, includes all the training data for the federated learning; A master node device, connected to the sub-node devices, for determining a federated learning training task, allocating partial federated learning training tasks to the sub-node devices according to the federated learning training task, and performing interaction of training data among the sub-node devices according to the federated learning training task; Receiving a request instruction for federated learning of the sub-node devices, obtaining partial training data sent by a first type of sub-node devices among the sub-node devices according to the request instruction, and distributing the partial training data to a second type of sub-node devices among the sub-node devices according to the request instruction. The second type of sub-node devices perform partial federated learning in the federated learning according to the received partial training data. The partial federated learning performed by the second type of sub-node devices includes training a model according to the received partial training data, and data aggregation and distribution are performed by the master node device; Data interaction between any two of the sub-node devices is relayed through the master node device. The master node device statistically monitors the network and other resource conditions of each sub-node device in real time and actively retries in case of failure; The master node device receives the intermediate results of encrypted gradient values and loss values calculated by the sub-node devices; determines the gradient values according to the intermediate results of the encrypted gradient values and the loss values; feeds back the gradient values to the sub-node devices; and the sub-node devices update the local machine learning model according to the gradient values.
2. The federated learning platform based on a distributed architecture according to claim 1, wherein The sub-node devices determine corresponding public-private key pairs according to the request instruction, and feedback the public keys to the master node device. The sub-node devices encrypt the partial training data to be sent with the private keys and forward the encrypted partial training data to the master node device.
3. The federated learning platform based on a distributed architecture according to claim 1, wherein The sub-node devices perform partial federated learning according to the partial training data. The training of the partial federated learning includes at least one of horizontal training, vertical training, and transfer training.
4. The federated learning platform based on a distributed architecture according to claim 1, wherein The master node device is further used for obtaining the progress of the sub-node devices in performing the partial federated learning, and obtaining the output results of the partial federated learning.
5. The federated learning platform based on a distributed architecture according to claim 1, wherein The master node device is further used for obtaining the running state information of the sub-node devices in performing the partial federated learning, and the running state information includes data interaction volume and / or resource occupancy rate.
6. The federated learning platform based on a distributed architecture according to claim 1, wherein It further includes: A metadata database, connected to the sub-node devices. The metadata database stores metadata, and the sub-node devices obtain the training data by preprocessing the metadata.
7. The federated learning platform based on a distributed architecture according to claim 6, wherein the meta database includes at least one of a Redis database, an Hbase database, an Elasticsearch database, and an Oracle database.
8. A federated learning method based on a distributed architecture, characterized in that, The method is executed by a master node device, which performs data aggregation and distribution, is used to determine a federated learning training task, allocates part of the federated learning training task to slave node devices according to the federated learning training task, and performs interaction of training data among the slave node devices according to the federated learning training task; the method includes: Obtaining partial training data sent by slave node devices for federated learning, and the partial training data of all the slave node devices includes all the training data of the federated learning after aggregation; Receiving a request instruction for federated learning sent by the slave node device; Obtaining partial training data sent by a first type of slave node device among the slave node devices according to the request instruction, and distributing the partial training data to a second type of slave node device among the slave node devices according to the request instruction, so that the second type of slave node device performs part of the federated learning in the federated learning according to the received partial training data, and the part of the federated learning performed by the second type of slave node device includes training a model according to the received partial training data, and the slave node device is the first type of slave node device or the second type of slave node device; Data interaction between any two of the slave node devices is relayed by the master node device, and the master node device real-time statistics the network and other resource conditions of each slave node device, and actively retries in case of failure; Receiving intermediate results of encrypted gradient values and loss values calculated by the slave node device; Determining a gradient value according to the intermediate results of the encrypted gradient values and the loss values; Feeding back the gradient value to the slave node device for the slave node device to update the local machine learning model according to the gradient value.
9. The federated learning method based on a distributed architecture according to claim 8, wherein Further includes: Obtaining a data interaction request and a data packet sent by the slave node device; Parsing the data interaction request to determine a target slave node device; Forwarding the data packet to the target slave node device, and the target slave node device is the first type of slave node device or the second type of slave node device.
10. The federated learning method based on a distributed architecture according to claim 8 or 9, characterized in that, Further includes: Receiving a public key sent by the slave node device, and the slave node device stores a private key matching the public key; Receiving partially encrypted training data sent by the slave node device through the private key; Performing parsing processing on the encrypted partial training data through the public key.
11. An electronic device, characterized in that, Includes: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the federated learning method based on a distributed architecture according to any one of claims 8 to 10 by executing the executable instructions.
12. A computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the federated learning method based on a distributed architecture according to any one of claims 8 to 10.
Citation Information
Patent Citations
Model updating method and device based on longitudinal federation learning, equipment and medium
CN111325352A
Data security sharing method, system and device
CN111901309A