Model training method, center node device and computer program product
By introducing a trusted execution environment (TEE) into federated learning, establishing a secure channel between the central node device and the edge node device, solving the computing and network load problems, and achieving efficient model training and data protection.
Patent Information
- Application Number
- CN202410111584.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-07-29
AI Technical Summary
The existing federated learning methods have a great burden on computing and network load, which is difficult to deploy in large-scale production environments, and are poorly scalable, making them unable to effectively protect data privacy and security.
The Trusted Execution Environment (TEE) is used to realize federated learning, by establishing a secure channel between the central node device and the edge node device, and using TEE's data protection mechanism to perform plain text calculations, reducing computing resource requirements and improving training efficiency.
Significantly reduces computing load, improves model training efficiency, ensures data privacy and security, and is suitable for large-scale production environments.
Smart Images

Figure CN120387530A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer processing, and more particularly, to a model training method based on federated learning, a central node device, and a computer program product. Background Art
[0002] With the development of computer technology, more and more artificial intelligence technologies have emerged and been widely applied. At the same time, the privacy and security of data have become increasingly important and received more and more attention. As a kind of machine learning technology, federated learning allows a group of organizations or groups within the same organization to train and improve a shared global machine learning model in a collaborative and iterative manner. Federated learning enables members within a group to build a general and powerful machine learning model without sharing data, thus solving key problems such as data privacy, data security, data access rights, and heterogeneous data access. Currently, federated learning has been widely applied in many industries including telecommunications and the Internet of Things due to its secure protection of data privacy and the mechanism of joint modeling. Summary of the Invention
[0003] Embodiments of the present disclosure provide a model training method based on federated learning, a central node device, and a computer program product.
[0004] According to a first aspect of the present disclosure, there is provided a model training method for federated learning, which is executed by a central node device having a trusted execution environment (TEE). The method includes: in response to receiving configuration information, communicating with a remote verification device to verify the TEE of the central node device, where the configuration information includes a partitioning method of training sample data for model training; in response to the verification of the TEE passing, establishing a secure channel with one or more edge node devices, where each of the one or more edge node devices has a TEE; and iteratively performing the following operations until a predetermined condition is met: selecting at least one edge node device for training the model from the one or more edge node devices; updating global model parameters based on training results received via the secure channel from the selected at least one edge node device, where the selected at least one edge node device trains corresponding models according to the partitioning method, and the training results correspond to the partitioning method; and sending the updated global model parameters to the one or more edge node devices via the secure channel.
[0005] According to a second aspect of the present disclosure, there is provided a central node device having a trusted execution environment (TEE). The central node device includes: at least one processor; and a memory coupled to the at least one processor and having instructions stored thereon, which when executed by the at least one processor cause the central node device to perform actions including: in response to receiving configuration information, communicating with a remote verification device to verify the TEE of the central node device, the configuration information including a partitioning manner of training sample data for model training; in response to successful verification of the TEE, establishing a secure channel with one or more edge node devices, wherein each of the one or more edge node devices has a TEE; and iteratively performing the following operations until a predetermined condition is met: selecting at least one edge node device for training a model from the one or more edge node devices; updating global model parameters based on training results received via the secure channel from the selected at least one edge node device, wherein the selected at least one edge node device trains a corresponding model according to the partitioning manner and the training results correspond to the partitioning manner; and sending the updated global model parameters to the one or more edge node devices via the secure channel.
[0006] According to a third aspect of the present disclosure, there is provided a computer program product tangibly stored on a non - volatile computer - readable medium and including machine - executable instructions which, when executed, cause the machine to perform the steps of the method in the first aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above - mentioned and other objects, features, and advantages of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0008] Figure 1 A schematic diagram of an example system in which embodiments of the present disclosure can be implemented is shown;
[0009] Figure 2 A flowchart showing a method for model training based on federated learning according to an embodiment of the present disclosure;
[0010] Figure 3 A detailed block diagram showing a model training system based on federated learning according to an embodiment of the present disclosure;
[0011] Figure 4 A signaling diagram showing the establishment of a secure channel between a central node device and an edge node device according to an embodiment of the present disclosure;
[0012] Figure 5Shows a schematic diagram of the working process of the model training system when the partitioning method includes the horizontal federated learning method according to an embodiment of the present disclosure;
[0013] Figure 6 Shows a schematic diagram of the working process of the model training system when the partitioning method includes the vertical federated learning method according to an embodiment of the present disclosure;
[0014] Figure 7 Schematically shows a simplified block diagram of a device suitable for implementing an example embodiment of the present disclosure.
[0015] In the respective drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0017] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0018] Traditional machine learning processes collect data from various devices and gather the collected data at a central server to train various types of models such as neural network models. However, this method inevitably includes data transmission between various devices and the central server, which limits the ability of the model to learn in real time. In addition, since data is often sensitive, collecting data across different devices for any purpose inevitably raises data security and privacy issues.
[0019] In contrast, Federated Learning (FL), as a type of machine learning technology, is a method of training and updating a model at local devices using local data. Then, the local devices send these locally trained models or model parameters from the local devices to a central server. The central server performs an aggregation operation (e.g., averaging weights, etc.) on the multiple models or model parameters received from multiple local devices, and then sends the combined and improved global model or global model parameters back to all local devices for the local devices to update the local models.
[0020] Due to its secure protection of data privacy and the mechanism of joint modeling, federated learning has been widely applied in multiple industries. To protect the data security and privacy of all parties involved, currently, homomorphic encryption or secret sharing multi-party computation is usually used in the security protocol layer of open-source federated learning.
[0021] For example, in industrial applications, there is both homomorphic encryption based on a central processing unit (CPU) and homomorphic encryption accelerated by a graphics processing unit (GPU). However, these solutions usually result in a very large computational load and a long latency, leading to low computational and training efficiency. For the secret sharing method, since this method generates a large number of network data packets, the network communication will be very busy. Due to these reasons, current federated learning methods actually serve security at the cost of performance and are still difficult to deploy in large-scale production environments. In addition, for secret sharing or homomorphic encryption, it was originally designed for computations between two parties. When the number of parties involved increases, these solutions cannot scale well and the performance will drop sharply. Therefore, it is desirable to provide a federated learning-based model training method that can reduce the computational load, reduce the network load, and improve the training efficiency.
[0022] Accordingly, to at least address the above problems and other potential problems, embodiments of the present disclosure propose a model training method based on federated learning, which is executed by a central node device having a trusted execution environment (TEE). The method includes: in response to receiving configuration information, communicating with a remote verification device to verify the TEE of the central node device, where the configuration information includes a partitioning method of training sample data for model training; in response to successful verification of the TEE, establishing a secure channel with one or more edge node devices, where each of the one or more edge node devices has a TEE; and iteratively performing the following operations until a predetermined condition is met: selecting at least one edge node device for training the model from the one or more edge node devices; updating global model parameters based on training results received from the selected at least one edge node device via the secure channel, where the selected at least one edge node device trains the corresponding model according to the partitioning method, and the training results correspond to the partitioning method; and sending the updated global model parameters to the one or more edge node devices via the secure channel.
[0023] The model training method and system based on federated learning according to embodiments of the present disclosure implement a brand-new architecture based on federated learning by leveraging a trusted execution environment (TEE) to train a model. The model training method based on federated learning according to embodiments of the present disclosure does not rely on the security protocol layer in the current standard implementation of federated learning, but rather fully utilizes the advantages of the data protection mechanism in the TEE and can perform plaintext calculations in a secure environment, thereby being able to significantly save computing resources, reduce the computing load, and improve the efficiency of model training.
[0024] Embodiments of the present disclosure will be described in further detail below with reference to the accompanying drawings, where Figure 1 FIG. shows a schematic diagram of an example system 100 in which embodiments of the present disclosure can be implemented.
[0025] The example system 100 can implement the model training method based on federated learning according to embodiments of the present disclosure. The system 100 includes a central node device 110 and one or more edge node devices 120-1, 120-2... 120-N, where N is a positive integer greater than or equal to 1. Each edge node device 120-i (1 ≤ i ≤ N) may include a corresponding model 122-i and sample data 124-i for training the model 122-i.
[0026] The central node device 110 and one or more edge node devices 120-1, 120-2... 120-N can implement the training for the global model by iteratively training the corresponding models. In each iteration operation, the central node device 110 can select at least one edge node device 120-s (1 ≤ s ≤ t; t is the number of edge node devices selected for training the model during the current iteration operation) from one or more edge node devices 120-1, 120-2... 120-N to train the corresponding model. Correspondingly, the selected edge node device 120-s can use its respective sample data 124-s to train the corresponding model 122-s, and can send the training result to the central node device 110. In each iteration operation, the central node device 110 processes the received training results to obtain updated global model parameters. The central node device 110 can further send the updated global model parameters to each edge node device 120-i among the edge node devices 120-1, 120-2... 120-N. Each edge node device 120-i can update the parameters of the local model based on the received global model parameters, so as to obtain the trained local model during each iteration operation.
[0027] The central node device 110 and one or more edge node devices 120-1, 120-2... 120-N can repeatedly perform the above iterative operation until a predetermined condition is met. In some embodiments, the predetermined condition may include but is not limited to: the number of times of performing the iterative operation reaches a predetermined number or the training accuracy of the global model reaches an accuracy threshold, etc.
[0028] In some embodiments, the central node device 110 and each edge node device 120-i each have a trusted execution environment TEE. Moreover, the central node device 110 and each edge node device 120-i each perform operations related to model training in the TEE area. For example, for each edge node device 120-i, the corresponding model 122-i and sample data 124-i are both stored in the TEE area of the edge node device 120-i. For the central node device 110, it performs operations related to model training in the TEE area. For example, the central node device 110 can perform the model training method according to the embodiments of the present disclosure in the TEE area.
[0029] In addition, in some embodiments, the central node device 110 and each edge node device 120-i are connected through a secure channel 130-i, and the secure channel 130-i is a secure channel supported by the TEE, so as to further ensure the security of the data calculated in plaintext. Although Figure 1It is shown that there is a corresponding secure channel between each edge node device 120-i and the central node device 110. However, it can be understood that the central node device 110 and multiple edge node devices can share a secure channel, and the present disclosure does not limit this.
[0030] In some embodiments, the central node device 110 can communicate with a remote verification device to verify the TEE of the central node device 110 in response to receiving configuration information. In some embodiments, the configuration information includes the partitioning method of the training sample data for model training and the information of the remote verification device for verifying the TEE of the central node device 110. The central node device 110 can establish a secure channel 130-i between the central node device 110 and one or more edge node devices 120-i in response to the successful verification of the TEE. In some embodiments, each of the one or more edge node devices has a TEE. The central node device 110 can receive various types of data including training results from the corresponding edge node device 120-i or send data to the edge node device 120-i through the secure channel 130-i.
[0031] In some embodiments, during each iteration operation, the edge node device 120-s selected for training the model can train the corresponding model 122-s according to the partitioning method, and the training results sent by the edge node device 120-s to the central node device 110 correspond to the partitioning method.
[0032] For example, the partitioning method can include horizontal federated learning and vertical federated learning. In the case where the partitioning method includes horizontal federated learning, the training results sent by the selected edge node device 120-s for training the model to the central node device 110 include the model parameters obtained by training the model 122-s in the edge node device 120-s, such as the weights of the model. In the case where the partitioning method includes vertical federated learning, the training results sent by the selected edge node device 120-s for training the model to the central node device 110 include the intermediate results obtained by training the model 122-s in the edge node device 120-s. In some embodiments, the intermediate results include information related to the gradients of the model 122-s in the edge node device 120-s. The central node device 110 can obtain updated global model parameters based on the training results from the edge node device 120-s in each iteration operation.
[0033] In addition, in some embodiments, the configuration information may also include structural information of the model stored in the edge node device 120-i. The central node device 110 may send at least the model structure information included in the configuration information as edge node configuration information to each edge node device 120-i. Each edge node device 120-i may configure the local model 122-i based on the structural information of the model in the edge node configuration information. Moreover, the structure of the model deployed in the edge node devices 120-1, 120-2...120-N corresponds to the division method in the configuration information. Specifically, in the case where the division method includes a horizontal federated learning method, the structures of the multiple models respectively deployed in the edge node devices 120-1, 120-2...120-N are the same as each other. In the case where the division method includes a vertical federated learning method, the structures of at least some of the multiple models respectively deployed in the edge node devices 120-1, 120-2...120-N may be different.
[0034] In some embodiments, the central node device 110 may be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframe computers, edge computing nodes, computing devices in cloud environments, and the like. The present disclosure does not limit the specific type of the central node device 110. The edge node devices 120-i may include but are not limited to personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multi-processor systems, consumer electronics, wearable electronic devices, smart home devices, and combinations of any of the above systems or devices. The present disclosure does not limit the specific type of edge node devices.
[0035] The model training method based on federated learning according to the embodiments of the present disclosure does not rely on the security protocol layer in the current standard federated learning implementation, but fully utilizes the advantages of the data protection mechanism in TEE and can perform plaintext calculations in a secure environment, thereby significantly saving computing resources and reducing computing load while improving the efficiency of model training.
[0036] Combined with the above Figure 1 A block diagram of an example system 100 is depicted in which embodiments of the present disclosure can be implemented. Figure 2 Describe a model training method based on federated learning according to an embodiment of the present disclosure. Figure 2 Flowchart of the model training method 200 based on federated learning according to an embodiment of the present disclosure is shown. Figure 1The example system 100 shown is used to describe the actions involved in method 200. For example, in some embodiments, method 200 may be executed by the central node device 110. It should be understood that method 200 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this regard.
[0037] At block 201, the central node device 110 may communicate with a remote attestation device in response to receiving configuration information to attest the TEE of the central node device 110. In some embodiments, the configuration information includes the partitioning method of the training sample data for model training.
[0038] The central node device 110 may have a TEE. The TEE is a secure area inside the hardware, which utilizes the characteristics of the hardware to ensure the confidentiality and integrity of the code and data in the TEE area, thereby resisting attacks from untrusted software. According to the model training method of the embodiments of the present disclosure, the data and code related to model training are stored in the TEE area of the central node device 110 by utilizing the security characteristics of the TEE, so that it is possible to calculate data in plaintext without further performing encryption operations, thereby saving computing resources and reducing the computing load.
[0039] The configuration information may include information related to model training input by the user. In some embodiments, the configuration information may include, but is not limited to, one or more of the following: the number of edge node devices participating in model training; the device addresses of the edge node devices participating in model training; information about the remote attestation device for TEE attestation of each edge node device (e.g., the address of the remote attestation device); information about the remote attestation device for TEE attestation of the central node device (e.g., the address of the remote attestation device); the partitioning method of the training sample data for model training; the update method used by the central node device when updating the global model parameters (e.g., the FedAvg method, etc.); the structural information of the model at each edge node device, etc.
[0040] In some embodiments, the central node device 110 may communicate with a remote attestation device in response to receiving configuration information to attest the TEE of the central node device 110. The attestation process for the TEE can ensure that the code runs securely in the corresponding device. The attestation process may include, but is not limited to, the central node device 110 sending proof information representing the TEE state, such as measurement values, to the remote attestation device. The remote attestation device may verify the proof information to determine whether the TEE can pass the attestation. For example, the remote attestation device may match the received measurement values with reference values, and when it is determined that the measurement values match the reference values, the remote attestation device may determine that the TEE passes the attestation.
[0041] In block 202, the central node device 110 may establish a secure channel with one or more edge node devices in response to the remote verification device passing the verification of the TEE. In some embodiments, one or more of the edge node devices each have a TEE. In some embodiments, there may be a corresponding secure channel between each edge node device 120-i and the central node device 110. Alternatively, the central node device 110 may also share a secure channel with multiple edge node devices.
[0042] In some embodiments, after the central node device 110 determines that the remote verification device has passed the verification of the TEE, the central node device 110 may establish a secure channel 130-i with the edge node device 120-i. In some embodiments, the central node device 110 may send edge node configuration information to the edge node device 120-i. The central node device 110 may send some or all of the configuration information received at block 201 as the edge node configuration information to the edge node device 120-i. In some embodiments, the edge node configuration information at least includes at least one of the following: information about the remote verification device for performing TEE verification with the corresponding edge node device 120-i (e.g., the address of the remote verification device); the partitioning method of the training sample data for model training; and the structural information of the model at the corresponding edge node device 120-i, etc. After receiving the edge node configuration information, the edge node device 120-i may communicate with the corresponding remote verification device according to the information about the remote verification device for performing TEE verification on each edge node device included in the edge node configuration information to perform TEE verification. The specific verification method may be similar to the implementation method of the central node device and the remote verification device performing TEE verification described above. For the sake of brevity, it will not be elaborated here.
[0043] In some embodiments, the remote verification device for the central node device 110 and the remote verification device for the edge node device 120-i may be the same or different, and the present disclosure does not limit this. In addition, the remote verification device for the edge node device 120-i and the remote verification device for the edge node device 120-(i + 1) may be the same or different, and the present disclosure does not limit this.
[0044] After determining that the verification of the TEE for the edge node device 120-i has passed, the central node device 110 may send a request to establish a secure channel to the edge node device 120-i. When the edge node device 120-i determines that the conditions for establishing a secure channel are met, it may send an acknowledgment message to the central node device 110. After receiving the acknowledgment message from the edge node device 120-i, the central node device 110 establishes a secure channel 130-i with the edge node device 120-i.
[0045] After establishing secure channels with one or more edge node devices participating in model training, the central node device 110 may iteratively perform the operations at block 203-205 until the predetermined condition at 206 is met.
[0046] In some embodiments, each edge node device 120-i may configure the local model 122-i based on the structural information of the model in the received edge node configuration information. And the structures of the models deployed in the edge node devices 120-1, 120-2... 120-N correspond to the partitioning method. Specifically, in the case where the partitioning method includes horizontal federated learning, the structures of the multiple models respectively deployed in the edge node devices 120-1, 120-2... 120-N are the same. In the case where the partitioning method includes vertical federated learning, the structures of at least some of the multiple models respectively deployed in the edge node devices 120-1, 120-2... 120-N may be different.
[0047] During the iterative operation, at block 203, the central node device 110 may select at least one edge node device, such as 120-s (1 ≤ s ≤ t; t is the number of edge node devices selected for model training during the current iterative operation), from the one or more edge node devices 120-1, 120-2... 120-N participating in model training for model training during the current iterative operation. Correspondingly, the selected edge node device 120-s may train the corresponding model 122-s using the sample data 124-s stored in the local TEE area. The selected edge node device 120-s may send the training result obtained after training the model to the central node device 110 through the secure channel established with the central node device 110.
[0048] In some embodiments, the selected edge node device 120-s for model training may train the corresponding model according to the partitioning method of the training sample data included in the edge node configuration information, and the training result of the model training corresponds to the partitioning method.
[0049] In some embodiments, the partitioning methods of the training sample data may include horizontal federated learning and vertical federated learning. In the case where the partitioning method includes horizontal federated learning, the selected edge node device 120-s for training the model may train the model 122-s based on the locally stored training sample data 124-s according to the horizontal federated learning method, and the training result sent to the central node device 110 includes the model parameters obtained by training the model 122-s, for example, the weights of the model. In the case where the partitioning method includes horizontal federated learning, the models 122-i at each of the edge node devices 120-i in the edge node devices 120-1, 120-2... 120-N have the same structure.
[0050] In the case where the partitioning method includes vertical federated learning, the selected edge node device 120-s for training the model trains the model 122-s based on the locally stored training sample data 124-s according to the vertical federated learning method. Since the locally stored training sample data 124-s is not complete sample data (for example, there may be a situation where there is no sample label), the training result sent by the edge node device 120-s to the central node device 110 includes the intermediate result obtained by training the model 122-s in the edge node device 120-s. In some embodiments, the intermediate result includes information related to the gradient of the model 122-s in the selected edge node device 120-s for training the model.
[0051] In block 204, the central node device 110 may update the global model parameters based on the training results received from at least one selected edge node device for training the model through a secure channel. In some embodiments, the central node device 110 may process the training results based on the update method (such as the FedAvg method, etc.) included in the configuration information when updating the global model parameters to update the global model parameters.
[0052] In block 205, the central node device 110 may send the updated global model parameters to each of the edge node devices 120-i in the edge node devices 120-1, 120-2... 120-N. Each edge node device 120-i may update the local model based on the received global model parameters, thereby obtaining the trained local model in the current iteration operation process.
[0053] At block 206, the central node device 110 determines whether a predetermined condition is satisfied. In some embodiments, the predetermined condition may include, but is not limited to, the number of times of performing iterative operations reaching a predetermined number or the training accuracy of the global model reaching an accuracy threshold, etc. When it is determined that the predetermined condition is satisfied, the central node device 110 terminates the model training, as shown in block 207. When it is determined that the predetermined condition is not yet satisfied, the central node device 110 returns to block 203 and continues to execute the operations in blocks 203 to 205 until the predetermined condition is satisfied. The specific execution process for blocks 203 to 205 can be understood with reference to the above description. For the sake of brevity, it will not be elaborated here.
[0054] In some embodiments, the model 122-i and the training sample data 124-i of each edge node device 120-i among the edge node devices 120-1, 120-2... 120-N are stored in the TEE area of the edge node device 120-i. And the central node device 110 processes the training results in the TEE area, which will be described in detail below.
[0055] According to the model training method based on federated learning according to an embodiment of the present disclosure, by performing model training in the TEE area and transmitting data in a secure channel supported by the TEE, it is possible not to rely on the security protocol layer in the current standard federated learning implementation, and fully utilize the advantages of the data protection mechanism in the TEE, and be able to perform plaintext calculations in a secure environment, thereby being able to significantly save computing resources, reduce the computing load, and at the same time improve the efficiency of model training.
[0056] The following will be combined with Figure 3 Describe the federated learning-based model training system 300 according to an embodiment of the present disclosure. Figure 3 Show a detailed block diagram of the federated learning-based model training system according to an embodiment of the present disclosure. Figure 3 Show Figure 1 The detailed block diagrams of the respective node devices in Figure 3 In Figure 1 The same elements as those in Figure 3 As shown, the central node device 110 includes a control module 112, an input / output (I / O) module 114, and a custom operator module 116. In some embodiments, the control module 112, the input / output (I / O) module 114, and the custom operator module 116 are all stored in the TEE area of the central node device 110.
[0057] In Figure 3Among them, each edge node device 120-i (1 ≤ i ≤ N; N is a positive integer greater than or equal to 1) participating in model training includes an input / output (I / O) module 121-i, a model 122-i, and a custom operator module 123-i. It should be understood that although not shown in Figure 3 it, sample data 124-i for training the corresponding model 122-i (as Figure 1 schematically shown) is also stored in each edge node device 120-i (for example, in the TEE area of the edge node device 120-i). The input / output (I / O) module 1212-i, the model 122-i, and the custom operator module 123-i are also stored in the TEE area of the edge node device 120-i.
[0058] In some embodiments, the control module 112 in the central node device 110 is used to control the overall operation process of model training. After receiving the configuration information input by the user, the control module 112 analyzes and extracts the configuration information, and sends the extracted configuration information to the input / output module 114 in the central node device 110. The control module 112 is also used to process the training results received from the edge node devices 120-i according to the parameter update method in the received configuration information to obtain updated global model parameters.
[0059] In some embodiments, the configuration information may include, but is not limited to, one or more of the following: the number of edge node devices participating in model training; the device addresses of the edge node devices participating in model training; the information of the remote verification device for TEE verification with each edge node device (for example, the address of the remote verification device); the information of the remote verification device for TEE verification with the central node device (for example, the address of the remote verification device); the partitioning method of the training sample data for model training; the update method (such as the FedAvg method, etc.) used by the central node device to update the global model parameters; the structural information of the models at each edge node device, etc.
[0060] In some embodiments, the input / output module 114 in the central node device 110 receives the extracted configuration information from the control module 112, and in response to receiving this configuration information, the input / output module 114 can communicate with the remote verification device according to the information of the remote verification device for TEE verification with the central node device 110 included in the configuration information to verify the TEE of the central node device 110. The specific verification process can refer to the above description and will not be elaborated here.
[0061] When the input / output module 114 determines that the verification for the TEE of the central node device 110 has passed, the edge node configuration information can be sent to each edge node device 120-i based on the device addresses of the edge node devices participating in model training included in the configuration information. Specifically, the input / output module 114 can send the edge node configuration information to the input / output module 121-i of each edge node device 120-i. The edge node configuration information sent to each edge node device 120-i can include all or at least a part of the configuration information in the configuration information received by the input / output module 114 from the control module 112. The edge node configuration information sent to the input / output module 121-i of each edge node device 120-i includes at least one of the following: information about the remote verification device for TEE verification of the corresponding edge node device 120-i (e.g., the address of the remote verification device); the partitioning method of the training sample data for model training; and the structure information of the model at the corresponding edge node device 120-i, etc.
[0062] Since the edge node configuration information can include the information about the remote verification device for TEE verification of the corresponding edge node device, after the input / output module 121-i of each edge node device 120-i receives the edge node configuration information, it can communicate with the corresponding remote verification device according to the information about the remote verification device for TEE verification included in the edge node configuration information, so as to realize the verification of the TEE. After determining that the verification of its TEE has passed, the input / output module 121-i of the edge node device 120-i with TEE verification passed can send a verification passed notification message to the central node device 110 to establish a secure channel between the central node device 110 and the edge node device with TEE verification passed.
[0063] During the process of establishing a secure channel, the central node device 110 (e.g., the input / output device 114 of the central node device 110) may send a request to establish a secure channel to each edge node device 120-i (e.g., the input / output module 121-i) that has passed the TEE verification. When the edge node device 120-i that receives the request to establish a secure channel determines that the conditions for establishing the secure channel are met, it may send confirmation information to the central node device 110 through the input / output module 121-i. In response to this confirmation information, the central node device 110 establishes a secure channel 130-i between the central node device 110 and the corresponding edge node device 120-i. This secure channel 130-i is used for data transmission between the central node device 110 and the edge node device 120-i. For example, the training result obtained after the edge node device 120-i trains the model 122-i can be transmitted via the secure channel 130-i to be sent to the central node device 110.
[0064] Continuing to refer to Figure 3 , Figure 3 the central node device 110 in also includes a custom operator module 116. In some embodiments, the custom operator module 116 creates an operator for the neural network model of the central node device 110 based on the operator of the neural network model, and stores the created operator in the TEE area of the central node device 110. In some embodiments, the custom operator module 116 may modify the syntax of the operator to the syntax supported by the TEE without changing the function of the operator based on the operator of the neural network model, for example, modifying it from the C language to the syntax supported by the TEE. The custom operator module 116 saves the created operator in the TEE area of the central node device 110 for the control module 112 to call during the process of processing the training result to update the parameters. In some embodiments, when the partitioning method of the training sample data in the configuration information indicates the vertical federated learning method, the central node module 110 may call the created operator to process the training result, thereby obtaining the parameters of the global model.
[0065] Similarly, as Figure 3 shown, each edge node device 120-i also has a corresponding custom operator module 123-i. The custom operator module 123-i creates an operator for the neural network model 122-i of the edge node device 120-i based on the operator of the corresponding neural network model 122-i, and stores the created operator in the TEE area of the edge node device 120-i. The edge node device 120-i may call the operator created by the custom operator module 123-i during the process of training its neural network model 122-i.
[0066] The above combination Figure 3 describes a block diagram of a system for model training according to an embodiment of the present disclosure. Figure 4 shows a signaling diagram for establishing a secure channel between a central node device and an edge node device according to an embodiment of the present disclosure. Figure 4 The order of the respective steps in Figure 4 is not strictly defined, and those skilled in the art can understand that
[0067] The central node device 110 may receive configuration information input by a user at 411. The configuration information may include, but is not limited to, one or more of the following: the number of edge node devices participating in model training; the device addresses of the edge node devices participating in model training; information about remote verification devices for performing TEE verification with each edge node device (e.g., the address of the remote verification device); information about remote verification devices for performing TEE verification with the central node device (e.g., the address of the remote verification device); the partitioning method of training sample data for model training; the update method used by the central node device when updating global model parameters (e.g., the FedAvg method, etc.); and the structural information of the model at each edge node device, etc.
[0068] In response to receiving the configuration information, the central node device 110 (e.g., Figure 3 the input / output module 114 in
[0069] ) may communicate with the remote verification device for verifying the TEE indicated in the configuration information to verify the TEE in the central node device 110, as shown in block 412. Figure 3 After determining that the verification of the TEE is passed, the central node device 110 (e.g.,
[0070] ) may send all or at least part of the configuration information in the configuration information as edge node configuration information to each edge node device 120-i among the multiple edge node devices participating in model training at 413. For example, the input / output module 121-i of each edge node device 120-i may receive the edge node configuration information from the central node device 110. The edge node configuration information may include at least one or more of the following: information about remote verification devices for verifying the TEE of the corresponding edge node device (e.g., the address of the remote verification device), the partitioning method of training sample data for model training, and the structural information of the model at the corresponding edge node device 120-i, etc.In response to receiving the edge node configuration information, the edge node device 120-i can verify the TEE in the edge node device 120-i by communicating with the remote verification device indicated in the edge node configuration information, as shown at 414.
[0071] After determining that the verification for the TEE is passed, the edge node device 120-i can send the information indicating the verification passed to the central node device 110 at 415. For example, the input / output module 114 in the central node device 110 can receive the verification passed information to establish a secure channel with the edge node device 120-i subsequently.
[0072] At 416, the central node device 110 can send a request to establish a secure channel to the edge node device that has passed the TEE verification. The received edge node device can determine whether the conditions for establishing the secure channel are met, and when determining that the conditions for establishing the secure channel are met, send a confirmation message to the central node device 110 at 417, and the confirmation message indicates confirmation of establishing a secure channel between the central node device 110 and the edge node device.
[0073] In response to receiving the confirmation message from the edge node device, the central node device 110 establishes a secure channel 130-i for communicating with the edge node device at 418. In some embodiments, the secure channel is a TEE-supported secure channel 130-i for data communication between the central node device 110 and the edge node device.
[0074] In some embodiments, the central node device 110 can establish a separate secure channel with each edge node device 120-i. Alternatively, the central node device 110 can share a secure channel with multiple edge node devices, and the present disclosure does not limit this.
[0075] After establishing a secure channel with each of the edge node devices 120-1, 120-2... 120-N in the edge node devices 120-1, 120-2... 120-N, the central node device 110 can execute the training process of the global model together with the edge node devices 120-1, 120-2... 120-N participating in the model training. In some embodiments, the central node device 110 can iteratively execute the steps in block 203 to block 205 in the method flow chart in Figure 2 until the predetermined condition in block 206 is met. The specific iterative operation process can refer to the description in the above in combination with Figure 2 For the sake of brevity, it will not be elaborated here.
[0076] The above has described in combination with Figure 4 the signaling diagram for establishing a secure channel between the central node device and the edge node device according to the embodiments of the present disclosure. The following will be described in combination withFigure 5 and Figure 6 A schematic diagram describing the working process of the model training system according to the partitioning method of the training sample data in the configuration information.
[0077] Figure 5 It shows a schematic diagram of the working process of the model training system when the partitioning method includes the horizontal federated learning method according to an embodiment of the present disclosure. As Figure 5 shown, when the partitioning method of the training sample data in the configuration information includes the horizontal federated learning method, the sample data included in each edge node device participating in the model training is complete, that is, the training sample data included in each edge node device includes the feature x and the label y of the sample. As Figure 5 shown, the edge node device 120-1 includes the training sample data D1, D2... Dn, and Dj (1≤j≤n) includes the feature x and the corresponding label y. The edge node device 120-2 includes the training sample data D(n+1), D(n+2)... D(n+m), and Dk ((n+1)≤k≤(n+m)) includes the feature x and the corresponding label y. The edge node device 120-N includes the training sample data D(n+m+1), D(n+m+2)... D(n+m+q), and Du ((n+m+1)≤u≤(n+m+q)) includes the feature x and the corresponding label y. It can be understood that although Figure 5 it is shown that each sample data includes three features x1, x2, and x3, this is only exemplary, and the training sample data included in each edge node device may include any number of features.
[0078] The edge node device 120-i can configure the model 122-i stored in the local TEE area according to the model structure information included in the edge node configuration information received from the central node device 110. In some embodiments, when the partitioning method includes the horizontal federated learning method, the models in each edge node device have the same structure, and the sample data used for training the model in each edge node device has the same dimension.
[0079] In some embodiments, during the process of model training, after establishing a secure channel between the central node device 110 and the edge node devices 120-1, 120-2... 120-N participating in the model training, the central node device 110 can perform the iterative operations described below with the edge node devices 120-1, 120-2... 120-N to achieve the training of the global model.
[0080] Specifically, in each iteration operation, the central node device 110 can select at least one edge node device 120-s (1 ≤ s ≤ t; t is the number of edge node devices selected for training the model in this iteration operation) from the edge node devices 120-1, 120-2... 120-N for training the model, so as to train the model 122-s stored locally. The selected edge node device 120-s for training the model can use various appropriate methods, utilize the training sample data stored in the local TEE area, and train the local model by invoking the operators in the custom operator module 123-s. The selected edge node device 120-s can use the trained model parameters (e.g., the weights of the model) as the training result, and send them to the central node device 110 via the input / output module 121-s through the secure channel 130-s.
[0081] In addition, the central node device 110 (e.g., the input / output module 114 of the central node device 110) receives the model parameter w from the edge node device 120-s through the secure channel 130-s s , and this model parameter w s represents the parameters of the model in the corresponding edge node device. The central node device 110 (e.g., the control module 112 of the central node device 110) processes the received model parameters w1, w2... w t (t is the number of edge node devices selected in this iteration operation) according to the update method (e.g., FedAvg method, etc.) specified in the configuration information for updating the global model parameters, so as to obtain the updated global model parameter W in this iteration operation, and send the updated global model parameter W to each edge node device 120-i participating in model training among the edge node devices 120-1, 120-2... 120-N participating in model training. The edge node device 120-i updates the model stored in the local TEE area based on the updated global model parameter W, so as to obtain the updated model of each edge node device in this iteration operation.
[0082] It can be understood that the above process can be iteratively executed until the training process meets a predetermined condition (e.g., reaching the number of iterations or the training accuracy of the global model reaching the accuracy threshold, etc.) and terminates.
[0083] In some embodiments, the central node device 110 may execute the method for model training according to the embodiments of the present disclosure in the TEE area. Moreover, the edge node device 120-i stores the corresponding model and training sample data in the local TEE area and executes the method for model training according to the embodiments of the present disclosure in the TEE area. In addition, the data transmission between the central node device 110 and the edge node device 120-i is performed via a secure channel supported by the TEE. Therefore, it is possible not to rely on the security protocol layer in the current standard implementation of federated learning, but to make full use of the advantages of the data protection mechanism in the TEE, and be able to perform plaintext calculations in a secure environment, thereby being able to significantly save computing resources, reduce the computing load, and improve the efficiency of model training.
[0084] Figure 6 FIG. shows a schematic diagram of the working process of the model training system when the partitioning method includes the vertical federated learning method according to the embodiments of the present disclosure. As Figure 6 shown, when the partitioning method for the training sample data in the configuration information includes the vertical federated learning method, the sample data included in each edge node device participating in model training is not complete, that is, the training sample data included in each edge node device does not necessarily include both the feature x and the label y of the sample. For example, one or more edge node devices may include partial features and not include the label. For another example, one or more edge node devices may include the label but only include partial features.
[0085] As Figure 6 shown, the training sample data included in the edge node device 120-1 includes features x1, x2, and x3 and does not include the label y of the data. The training sample data included in the edge node device 120-2 includes features x1, x4, x6, x7, and the label y. The training sample data included in the edge node device 120-2 includes features x3, x4, x5, and x7, but does not include the label y of the data. It can be understood that Figure 6 the examples of the training sample data in
[0086] In some embodiments, the edge node device 120-i may configure the model 122-i stored in the local TEE area according to the structure information of the model included in the edge node configuration information received from the central node device 110. In some embodiments, when the partitioning method includes the vertical federated learning method, since the training sample data may have different dimensions, the models in each edge node device may have different structures, as Figure 6 schematically shown.
[0087] In some embodiments, during the process of model training, after establishing a secure channel between the central node device 110 and the edge node devices 120-1, 120-2, ……, 120-N participating in the model training, the central node device 110 can perform the iterative operations described below with the edge node devices 120-1, 120-2, ……, 120-N to implement the training of the global model.
[0088] Specifically, in each iterative operation, the central node device 110 selects at least one edge node device 120-s (1 ≤ s ≤ t; t is the number of edge node devices selected for training the model in this iterative operation) from the edge node devices 120-1, 120-2, ……, 120-N for training the locally stored model 122-s. In some embodiments, during each iterative operation, the central node device 110 can send an alignment request to each edge node device 120-s participating in this iterative operation to determine the training sample data with a common serial number among the edge node devices selected for training the model. Accordingly, the determined training sample data with a common serial number is aligned among the edge node devices selected for training the model. As Figure 6 shown, in response to receiving the alignment request from the central node device 110, the edge node devices selected for model training in this iterative operation perform the alignment operation of the training sample data with a common serial number, so that the training sample data with the same serial number can be aligned. For example, as Figure 6 shown, the training sample data with serial numbers 1, 2 to n are respectively aligned. It can be understood that the alignment of the training sample data performed by the edge node device can include logical alignment.
[0089] During each iterative operation, the edge node device 120-s selected for model training can adopt various appropriate methods, utilize the training sample data stored in the local TEE area, and train the corresponding model by calling the operator in the custom operator module 123-s. Since the training sample data stored in the local TEE area has different dimensions and does not necessarily have labels, when the division method includes a vertical federated learning method, the training result obtained by the edge node device for training the model is an intermediate result of the training process. In some embodiments, the intermediate result includes information related to the gradient of the model 122-s in the edge node device 120-s. In some embodiments, the intermediate result includes an element associated with the serial number of the training sample data, for example, a gradient-related information element associated with the serial number of the training sample data. The edge node device 120-s can send the intermediate result to the central node device 110 through the input / output module 121-s and via the secure channel 130-s.
[0090] In each iteration, the central node device 110 (eg, the input / output module 114 of the central node device 110) receives the intermediate result R from the edge node device 120-i through the secure channel 130-s. s , the intermediate result R s Characterizes the gradient-related information of the model in the corresponding edge node device. The central node device 110 (e.g., the control module 112 of the central node device 110) updates the received intermediate results R1, R2, ... R according to the update method used to update the model parameters specified in the configuration information (e.g., FedAvg method, etc.). t (t is the number of edge node devices selected in this iterative operation) to obtain the updated global model parameter W. In some embodiments, the intermediate results R1, R2...R t Each intermediate result R in i The central node device 110 can obtain the serial number of each training sample data and the label corresponding to the sample data of the serial number in advance. Figure 6 As shown in , the central node device 110 can pre-acquire the following information: the label corresponding to the training sample data with a serial number of 1 is category b, the label corresponding to the training sample data with a serial number of 2 is category a, and so on.
[0091] In each iteration operation, the central node device 110 (e.g., the control module 112 of the central node device 110) can, based on the serial number and label of the training sample data, and according to the model parameter update method set in the configuration information, call the corresponding operator stored in the TEE area of the central node device 110 to process the received intermediate results R1, R2…R t to obtain the updated global model parameter W. The central node device 110 can send the updated global model parameter W as the global model parameter to each edge node device 120-i participating in model training among the edge node devices 120-1, 120-2……120-N participating in model training. The edge node device 120-i updates the model stored in the local TEE area based on the updated global model parameter W, so as to obtain the updated model in each edge node device in this iteration operation.
[0092] It can be understood that the above process can be iteratively executed until the training process meets a predetermined condition (e.g., reaching the number of iterations or the training accuracy of the global model reaching an accuracy threshold, etc.) and terminates.
[0093] In some embodiments, the central node device 110 executes the method for model training according to the embodiments of the present disclosure in the TEE area. And, the edge node device 120-i stores the corresponding model and training sample data in the local TEE area to execute the method for model training according to the embodiments of the present disclosure in the TEE area. In addition, the data transmission between the central node device 110 and the edge node device 120-i is carried out via a secure channel supported by the TEE. Therefore, it can not rely on the security protocol layer in the current standard implementation of federated learning, but make full use of the advantages of the data protection mechanism in the TEE, and can perform plaintext calculations in a secure environment, so as to be able to significantly save computing resources, reduce the computing load while improving the efficiency of model training.
[0094] Figure 7FIG. 0 shows a schematic block diagram of an exemplary device 700 that can be used to implement embodiments of the present disclosure. The central node device and / or edge node device according to embodiments of the present disclosure can both be implemented using device 700. As shown in the figure, device 700 includes a processing unit (e.g., central processing unit CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or computer program instructions loaded from storage unit 708 into random access memory (RAM) 703. In RAM 703, various programs and data required for the operation of device 700 can also be stored. CPU 701, ROM 702, and RAM 703 are connected to each other via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.
[0095] Multiple components in device 700 are connected to I / O interface 705, including: an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, optical disc, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0096] Each of the processes and processes described above, such as method 200, can be executed by processing unit 701. For example, in some embodiments, method 200 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more actions of method 200 described above can be performed.
[0097] The present disclosure can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.
[0098] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0099] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0100] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0101] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.
[0102] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is created that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0103] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0104] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0105] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A model training method based on federated learning, which is executed by a central node device with a trusted execution environment (TEE). The method includes: Responding to receiving configuration information, communicating with a remote verification device to verify the TEE of the central node device, where the configuration information includes the partitioning method of training sample data for model training; Responding to the verification of the TEE passing, establishing a secure channel with one or more edge node devices, where each of the one or more edge node devices has a TEE; And Iteratively performing the following operations until a predetermined condition is met: Selecting at least one edge node device for training the model from the one or more edge node devices; Updating global model parameters based on the training results received from the selected at least one edge node device through the secure channel, where the selected at least one edge node device trains the corresponding model according to the partitioning method, and the training results correspond to the partitioning method; And Sending the updated global model parameters to the one or more edge node devices through the secure channel.
2. The method according to claim 1, wherein establishing a secure channel with one or more edge node devices includes: Sending edge node configuration information to the one or more edge node devices, where the edge node configuration information includes information about the remote verification device for TEE verification with the one or more edge node devices; Responding to each of the one or more edge node devices passing TEE verification, sending a request to establish a secure channel to the one or more edge node devices; And Responding to the confirmation from the one or more edge node devices, establishing the secure channel between the central node device and the one or more edge node devices.
3. The method according to claim 1, wherein the configuration information further includes structure information of the model for training deployed in each of the one or more edge node devices corresponding to the partitioning method.
4. The method according to claim 1, wherein in response to the partitioning method including horizontal federated learning, the training results include model parameters obtained by training the model in each of the selected at least one edge node device.
5. The method according to claim 4, further including: Processing the received model parameters according to the model parameter update method in the configuration information to obtain the updated global model parameters; And Sending the updated global model parameters to each of the one or more edge node devices through the secure channel.
6. The method according to claim 1, wherein in response to the partitioning method including vertical federated learning, the training results include intermediate results of training the model in each of the selected at least one edge node device.
7. The method according to claim 6, wherein the intermediate result includes elements associated with the serial numbers of the training sample data.
8. The method according to claim 6, further comprising: Constructing operators used in the model training process; And Storing the constructed operators in the TEE of the central node device.
9. The method according to claim 7, further comprising: Based on the serial numbers and labels of the training sample data, and according to the model parameter update method in the configuration information, calling the corresponding operators stored in the TEE to process the intermediate result to obtain the updated global model parameters; And Sending the updated global model parameters to each of the one or more edge node devices via the secure channel.
10. The method according to claim 1, wherein the models respectively corresponding to the one or more edge node devices are stored in the TEEs of the corresponding edge node devices.
11. A central node device having a trusted execution environment TEE, the central node device comprising: At least one processor; And A memory coupled to the at least one processor and having instructions stored thereon, the instructions causing the central node device to perform actions when executed by the at least one processor, the actions including: In response to receiving configuration information, communicating with a remote verification device to verify the TEE of the central node device, the configuration information including the partitioning method of the training sample data for model training; In response to the verification of the TEE passing, establishing a secure channel with one or more edge node devices, wherein each of the one or more edge node devices has a TEE; and Iteratively performing the following operations until a predetermined condition is met: Selecting at least one edge node device for training the model from the one or more edge node devices; Updating the global model parameters based on the training results received from the selected at least one edge node device via the secure channel, wherein the selected at least one edge node device trains the corresponding model according to the partitioning method, and the training results correspond to the partitioning method; and Sending the updated global model parameters to the one or more edge node devices via the secure channel.
12. The central node device according to claim 11, wherein establishing a secure channel with one or more edge node devices includes: Sending edge node configuration information to the one or more edge node devices, the edge node configuration information including information about the remote verification device for TEE verification of the one or more edge node devices; In response to each of the one or more edge node devices passing TEE verification, sending a request to establish a secure channel to the one or more edge node devices; And In response to confirmations from the one or more edge node devices, establishing the secure channel between the central node device and the one or more edge node devices.
13. The central node device according to claim 11, wherein the configuration information further includes structure information of a model for training deployed in each of the one or more edge node devices corresponding to the partitioning method.
14. The central node device according to claim 11, wherein in response to the partitioning method including a horizontal federated learning method, the training result includes model parameters obtained by training the model in each of the at least one selected edge node devices.
15. The central node device according to claim 14, the instruction further causes the central node device to perform an action when executed by the at least one processor, the action including: Processing the received model parameters according to the model parameter update method in the configuration information to obtain the updated global model parameters; And Sending the updated global model parameters to each of the one or more edge node devices via the secure channel.
16. The central node device according to claim 11, wherein in response to the partitioning method including a vertical federated learning method, the training result includes intermediate results of training the model in each of the at least one selected edge node devices.
17. The central node device according to claim 16, the instruction further causes the central node device to perform an action when executed by the at least one processor, the action including: Constructing operators used in the model training process; And Storing the constructed operators in the TEE of the central node device.
18. The central node device according to claim 17, the instruction further causes the central node device to perform an action when executed by the at least one processor, the action including: Based on the serial number and label of the training sample data, and according to the model parameter update method in the configuration information, calling the corresponding operator stored in the TEE to process the intermediate results to obtain the updated global model parameters; And Sending the updated global model parameters to each of the one or more edge node devices via the secure channel.
19. The central node device according to claim 11, wherein the models respectively corresponding to the one or more edge node devices are stored in the TEE of the corresponding edge node devices.
20. A computer program product, the computer program product being tangibly stored on a non - volatile computer - readable medium and including machine - executable instructions that, when executed, cause a central node device to perform the following steps: In response to receiving configuration information, communicate with a remote verification device to verify the trusted execution environment TEE of the central node device, the configuration information including a partitioning method of training sample data for model training; In response to successful verification of the TEE, establish a secure channel with one or more edge node devices, where each of the one or more edge node devices has a TEE; And Iteratively perform the following operations until a predetermined condition is met: Select at least one edge node device from the one or more edge node devices for training the model; Update the global model parameters based on the training results received from the selected at least one edge node device through the secure channel, where the selected at least one edge node device trains the corresponding model respectively according to the partitioning method, and the training results correspond to the partitioning method; And Send the updated global model parameters to the one or more edge node devices through the secure channel.