A model training method and a node device

By encrypting and averaging model parameters directly among nodes in a decentralized federated learning system, the method safeguards user privacy and enhances security by eliminating the need for a central server, addressing data breaches in federated learning.

CN114004265BActive Publication Date: 2025-07-15HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010670085.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-13
Publication Date
2025-07-15
Estimated Expiration
2040-07-13

AI Technical Summary

Technical Problem

In federated learning, the training data of node devices is easily leaked by attackers through central control devices, resulting in security risks to user privacy data.

Method used

The encryption method is used to directly calculate parameters and values between node devices to avoid assistance from central control devices, and to realize a decentralized architecture, and to protect the weight value of node devices from being leaked through the ElGamal encryption algorithm and voting mechanism.

Benefits of technology

It effectively protects the privacy data of node devices, avoids attackers from obtaining training data by collecting weight values, and enhances the robustness and security of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114004265B_ABST
    Figure CN114004265B_ABST
Patent Text Reader

Abstract

A model training method and a node device are applied to federated learning in the field of artificial intelligence. This method is applied to a federated learning system, which at least includes a first node device and N second node devices. The method includes: the first node device obtains a first training data set; then inputs the first training data set into a first model for training to determine a first parameter value of the target parameter of the first model; encrypts the first parameter value to obtain a first encrypted representation of the first parameter value; receives a second encrypted representation of the second parameter value of the target parameter sent by each second node device; finally, calculates a parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representation, and the parameter sum value is used to update the first model, so as to obtain an updated second model. The node device directly calculates the parameter sum value of all nodes and does not know the weight value of each node device, thereby protecting the privacy data of each node device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a model training method and node device. Background Art

[0002] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0003] Federated learning (FL) is an emerging basic technology of artificial intelligence. Federated learning trains training data based on a decentralized machine learning framework. The training data is scattered on different node devices, and these node devices cooperate with each other to train the model together.

[0004] The system architecture of federated learning includes a central control device (such as a cloud server) and multiple node devices. The central control device first sends a machine learning network architecture and a set of initialized parameters to each node device. Each node device uses local data (such as photos in the local photo library) to train the network architecture, and then each node device uploads the locally trained parameters to the central control device. The central control device collects the parameters uploaded by each node device, calculates the average value of these parameters in proportion, and the average value is the result of the calculation. The central control device sends the result to each node device, and each node device is trained again based on the result. Through multiple iterative training, the finally selected parameters can make the training model accuracy reach the system's predetermined requirements.

[0005] During the federated learning training process, each node device needs to send the parameters obtained from each training to the central control device. During this process, attackers have an opportunity to collect the parameters of each node device, or they can attack the central control device to obtain the parameters of each node device, and use the parameters to extract the information of the training data used by the node device (such as photos in the photo library), thereby leading to the leakage of user privacy data. There are great security risks to user privacy data. Summary of the invention

[0006] Embodiments of the present application provide a model training method and a node device. Among them, this method is applied to a federated learning system, which includes a first node device and N second node devices, where N is an integer greater than or equal to 1. For example, N is 1, 2, 3, etc. When N is 1, the federated learning system includes 2 node devices. When N is 2, the federated learning system includes 3 node devices. When N is 3, the federated learning system includes 4 node devices. The number of node devices included in the federated learning system is not limited. Among them, a first model to be trained (such as a neural network) and initial parameters (such as all 0s) are preset in both the first node device and the N second node devices. Alternatively, the first node device and the second node devices download the first model (such as a neural network) and initial parameters (such as all 0s) from the server to the local side. The first node device and the second node devices train the first models on their respective local sides. This method can be applied to each node device in the system, and the actions performed by each node device are similar. In the embodiments of the present application, a node device in the system (such as the first node device) is used as the execution subject for illustration.

[0007] In a first aspect, an embodiment of the present application provides a model training method, which is applied to a first node device. The first node device obtains a first training data set; the first training data set is a set of user data stored at the local end of the first node device, and the user data refers to data generated based on user behavior. The user data can be user input data received by the first node device through an input device. For example, the user data can be text input by the user received by the mobile phone through the display screen, etc.; then, the first training data set is input into the first model for training to determine the first parameter value of the target parameter of the first model; further, the first parameter value is encrypted to obtain the first encrypted representation of the first parameter value. The first parameter value is a set of parameter values generated by the first node device, or it can also be one of the set of parameter values. For example, the first parameter value is the weight value corresponding to the directed line of each neuron in the hidden layer; further, receive the second encrypted representation of the second parameter value of the target parameter sent by each second node device. The second parameter value is obtained after the second node device inputs the second training data set into the first model for training; it should be noted that, in order to distinguish the parameter values obtained by the first node device during training from the parameter values obtained by the second node device during training, the parameter values obtained by the first node device during training are called the first parameter values, and the parameter values obtained by the second node device during training are called the second parameter values. In order to distinguish the encrypted representation of the parameter values generated by the first node device from the encrypted representation of the parameter values generated by the second node device, the encrypted representation of the parameter values generated by the first node device is called the first encrypted representation, and the encrypted representation of the parameter values generated by the second node device is called the second encrypted representation; that is, the parameter values (second parameter values) obtained by each of the N second node devices during training can be different, and the encrypted representations (second encrypted representations) generated by each second node device can also be different. The second parameter value is a set of parameter values corresponding to each node or, it can also be one of the set of parameter values. For example, if N is 2, the first node device receives 2 second encrypted representations sent by 2 second node devices; finally, the first node device calculates the parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representation. The parameter sum value is used to update the first model, thereby obtaining the updated second model.

[0008] The model training method provided in this embodiment can be applied to the scenario of training a model with private data. The federated learning system architecture includes a first node device and N second node devices. Among them, both the first node device and the N second node devices will use the data sets stored locally to train the first model. Each node device trains the first model according to its own local training data set, obtains the weight values obtained by training the first model for each of its own, then encrypts the weight values, and broadcasts the encrypted weight values. As a result, each node device will receive the encrypted weight values broadcast by other nodes. Finally, each node device further directly calculates the parameter sum value of all the weight values according to the encrypted weight values of all the nodes, and then each node device updates the first model according to the parameter sum value. Since each node device directly calculates the parameter sum value of all the nodes and does not know the weight values of each node device, the weights of each node device will not be leaked, preventing attackers from collecting the weights of each node device to extract the training data of the node device, thus protecting the private data of each node device. And the parameter sum value of all the nodes is calculated by the cooperation of multiple node devices without the assistance of a central control device, realizing a decentralized architecture, preventing attackers from attacking the central control device to obtain the parameters of each node device, thus protecting the private data of each node device from being leaked.

[0009] In a possible implementation, the first parameter value has L bit positions, where L is an integer greater than or equal to 2; encrypting the first parameter value to obtain the first encrypted representation of the first parameter value may specifically include: the first node device generates a first private key for each of the L bit positions of the first parameter value, and the first private key may be a random number; then, the first node device may generate a first public key according to the first private key and receive N second public keys sent by the N second node devices. The second public key is generated by the second node device according to a second private key, and the first public key and the second public key are discrete logarithms; finally, the first node device encrypts the first parameter value according to the first public key and the N second public keys to obtain the first encrypted representation of the first parameter value. It should be noted that, in order to distinguish the public key generated by the first node device from the public key generated by the second node device, the public key generated by the first node device is called the first public key, and the public key generated by the second node device is called the second public key. The N public keys (second public keys) generated by the N node devices may be different. In this example, the first parameter value is in binary representation, and a first private key is generated for each of the L bit positions, increasing the complexity of the first encrypted representation.

[0010] In a possible implementation, the second parameter value has L bits, where L is an integer greater than or equal to 2; calculating the parameter sum value of the first parameter value and the second parameter value based on the first encrypted representation and the second encrypted representation may specifically include: the first node device first determines the sum value of the bit values of the t-th bit in the first encrypted representation and the bit value of the t-th bit in the second encrypted representation, where t takes each integer from 1 to L, that is, calculates the sum value of the bit values of the two parameter values on the first bit, the sum value of the bit values of the two parameter values on the second bit, until the sum value of the bit values of the two parameter values on the L-th bit; then, calculates the parameter sum value of the first parameter value and the second parameter value based on the sum value of the bit values corresponding to each bit among the L bits.

[0011] In this embodiment, based on the voting mechanism, each bit of the weight value of each node device is equivalent to a candidate in the voting mechanism. Each node device is equivalent to a voter in the voting mechanism. At this time, each weight value is a binary string, and the bit value on each bit is 0 or 1. The 0 or 1 on each bit is equivalent to a ballot in the voting mechanism. In the voting mechanism, directly calculate the sum of all votes of each candidate, but do not disclose the votes cast by each voter. It is equivalent to each node device further directly calculating the parameter sum value of all weight values based on the encrypted weight values of all nodes, and then each node device updates the first model according to the parameter sum value. Since each node device directly calculates the parameter sum value of all nodes and does not know the weight value of each node device, nor discloses the weight value of each node device, it prevents attackers from obtaining the training data of the node device by collecting the weight values of each node device, thus protecting the privacy data of each node device.

[0012] In a possible implementation, the federated learning system further includes a third node device, and the method further includes: the first node device receives the third public key broadcast by the third node device, where the third public key is calculated by the third node device based on the generated private key; then, the first node device calculates the correlation relationship of the first private key, the second private key, and the third private key based on the first public key, the second public key, and the third public key; further, the first node device may calculate the third encrypted representation of the third parameter value corresponding to the third node device according to the correlation relationship; finally, the first node device calculates the parameter sum value of the first parameter value, the second parameter value, and the third parameter value based on the first encrypted representation, the second encrypted representation, and the third encrypted representation. In this example, if one of the multiple node devices (such as the third node device) drops out of the line and does not broadcast the encrypted representation of the weight value, then other node devices can restore the encrypted representation of the weight of the dropped-out node device according to the public key broadcast by the dropped-out node device before. In this example, even if a node device drops out of the line, it will not affect the calculation of the parameter sum value of the weight value, enhancing the robustness of the federated learning system.

[0013] In a possible implementation, the method further includes: the first node device determines a target parameter value according to the first data volume of the first training data set, the second data volume of the second training data set, and the parameter and the value, and the target parameter value is used to determine the second model.

[0014] In a possible implementation, the first data volume is the same as the second data volume. In this example, the data volumes of the training data sets of each node device are the same, and the average value of the weight values can be directly calculated without calculating the data volume, saving the computing power of the node device.

[0015] In a possible implementation, the second data volume is c×M; where M is the data magnitude and c is a factor, and c is greater than 0; the first node device receives the second data volume sent by the second node device, and the second data volume is represented by c. Representing the data volume as c×M simplifies the representation of the data volume of the training data, expands the flexibility that each node device can have different data volumes, and also gives the parameter values of the local neural network contributed by each node device the flexibility of different importance.

[0016] In a possible implementation, the second model can be an image classification model, the first training data set is an image data set, and each image in the image data set has a corresponding classification label. In this example, the image data set used for training at the local end will not be leaked, protecting the user's privacy data.

[0017] In a possible implementation, the second model is a text prediction model, and the first training data set includes multiple text records; where each text record includes the first text input by the user and the second text selected by the user. In this example, the multiple text records used for training at the local end will not be leaked, protecting the user's privacy data.

[0018] In a possible implementation, the method further includes: the first node device obtains first data; inputs the first data into a second model, and outputs result data through the second model; if the difference between the result data and first target data is greater than a threshold value, broadcasts a parameter adjustment message, where the parameter adjustment message is used to notify the second node device to continue training the second model; if the difference between the result data and first target data is less than or equal to the threshold value, determines the second model according to parameter sum values. In this example, since the parameter sum values are obtained based on the joint training of multiple node devices, that is, both the first node device and N second node devices contribute to the training of the first model (such as a neural network), the second model is tested using the data at the local end of the first node device to check whether the second model meets the system requirements of the first node device. If the difference between the result data and first target data is greater than the threshold value, broadcasts a parameter adjustment message to perform the next round of model parameter adjustment, so that the second model can meet the system requirements of the first node device.

[0019] In a possible implementation, inputting the first data into the second model and outputting result data through the second model may specifically include: inputting a first image into the second model and outputting a classification result through the second model; if the difference between the classification result and the category information of the first image is greater than the threshold value, performs the step of broadcasting parameter adjustment; if the difference between the classification result and the category information of the first image is not greater than the threshold value, performs the step of determining the second model according to parameter sum values.

[0020] In a possible implementation, the first data is third text, and inputting the first data into the second model and outputting result data through the second model may include: inputting the third text into the second model and outputting associated text for the third text through the second model; if the difference between the associated text of the third text output by the second model and the fourth text selected by the user is greater than the threshold value, performs the step of broadcasting parameter adjustment; if the difference between the associated text of the third text output by the second model and the fourth text selected by the user is not greater than the threshold value, performs the step of determining the second model according to parameter sum values.

[0021] Second aspect, an embodiment of the present application provides a method for image classification, which is applied to a first node device. The first node device belongs to a federated learning system, and the federated learning system at least further includes a second node device. The method includes: the first node device obtains a first image; the first node device inputs the first image into an image classification model, and outputs a classification result through the image classification model. The image classification model is: the first node device inputs the obtained first training data set into a first model for training. The first training data set includes multiple images, and each image has a corresponding classification label; determines a first parameter value of the target parameter of the first model; encrypts the first parameter value to obtain a first encrypted representation of the first parameter value; receives a second encrypted representation of a second parameter value of the target parameter sent by the second node device. The second parameter value is obtained after the second node device inputs a second training data set into the first model for training; calculates a parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representation, and is obtained according to the parameter sum value.

[0022] Third aspect, an embodiment of the present application provides a method for determining associated text, which is applied to a first node device. The first node device belongs to a federated learning system, and the federated learning system at least further includes a second node device. The method includes: the first node device receives text input by a user through an input device; the first node device inputs the text into a text prediction model, and outputs associated text of the text through the text prediction model. The text prediction model is: the first node device inputs the obtained first training data set into a first model for training. The first training data set includes multiple text records; wherein, each text record includes first text input by the user and second text selected by the user; determines a first parameter value of the target parameter of the first model; encrypts the first parameter value to obtain a first encrypted representation of the first parameter value; the first node device also receives second encrypted representations of second parameter values of the target parameter sent by N second node devices. The second parameter value is obtained after the second node device inputs a second training data set into the first model for training; calculates a parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representation, and is obtained according to the parameter sum value.

[0023] Fourth aspect, an embodiment of the present application provides a node device, which has the functions executed by the first node device in the first aspect, the second aspect, or the third aspect described above; these functions can be implemented by hardware or by hardware executing corresponding software; the hardware or software includes one or more modules corresponding to the above functions.

[0024] Fifth aspect, an embodiment of the present application provides a node device, including: a processor, the processor is coupled to at least one memory, and the processor is configured to read a computer program stored in the at least one memory, so that the node device executes the method described in the first aspect or the second aspect or the third aspect above.

[0025] Sixth aspect, an embodiment of the present application provides a computer-readable medium, the computer-readable storage medium is configured to store a computer program, and when the computer program runs on a computer, the computer is caused to execute the method described in the first aspect or the second aspect or the third aspect above.

[0026] Seventh aspect, the present application provides a chip system, the chip system includes a processor, configured to support a first node device to implement the functions involved in the first aspect or the second aspect or the third aspect above. In a possible design, the chip system further includes a memory, and the memory is configured to store necessary program instructions and data of the first node device. The chip system may be composed of chips or may include chips and other discrete devices. Description of the Drawings

[0027] Figure 1 It is a schematic diagram of an artificial intelligence main body framework in an embodiment of the present application;

[0028] Figure 2 It is a schematic diagram of the system framework of traditional federated learning;

[0029] Figure 3 It is a schematic diagram of another example of the system framework of federated learning in an embodiment of the present application;

[0030] Figure 4 It is a schematic diagram of a neural network in an embodiment of the present application;

[0031] Figure 5 It is a schematic flowchart of an example of a model training method in an embodiment of the present application;

[0032] Figure 6 It is a schematic flowchart of another example of a model training method in an embodiment of the present application;

[0033] Figure 7 It is a schematic diagram of the principle of encrypting parameter values using a voting mechanism in an embodiment of the present application;

[0034] Figure 8 It is a schematic structural diagram of an example of a node device in an embodiment of the present application;

[0035] Figure 9 It is a schematic structural diagram of a chip in an embodiment of the present application;

[0036] Figure 10 Schematic structural diagram of another example of the node device in the embodiment of the present application;

[0037] Figure 11 Schematic structural diagram of another example of the node device in the embodiment of the present application. Detailed implementation manners

[0038] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Terms such as "first" and "second" in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices.

[0039] Figure 1 Shows a schematic diagram of an artificial intelligence main framework, which describes the overall working process of an artificial intelligence system and is applicable to the general requirements in the field of artificial intelligence.

[0040] Next, the above artificial intelligence main framework will be elaborated from two dimensions of the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis).

[0041] The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom".

[0042] The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.

[0043] (1) Infrastructure:

[0044] The infrastructure provides computing power support for the AI system, enables communication with the external world, and is supported by the basic platform. It communicates with the external world through sensors; the computing power is provided by intelligent chips (such as hardware acceleration chips like CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external world to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.

[0045] (2) Data

[0046] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.

[0047] (3) Data Processing

[0048] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0049] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on the data.

[0050] Reasoning refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formal information to perform machine thinking and solve problems according to the reasoning control strategy. The typical function is search and matching.

[0051] Decision-making refers to the process of making decisions after the intelligent information is reasoned, and usually provides functions such as classification, sorting, prediction, etc.

[0052] (4) General Capabilities

[0053] After the data undergoes the above-mentioned data processing, some general capabilities can be formed further based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0054] (5) Intelligent Products and Industry Applications

[0055] Intelligent products and industry applications refer to the products and applications of the AI system in various fields. It is the encapsulation of the overall AI solution, productizes the intelligent information decision-making, and realizes the landing application. Its application fields mainly include: intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, safe city, intelligent terminal, etc.

[0056] The embodiments of the present application mainly relate to the machine learning (or deep learning) content in the above-mentioned part (3). The present application relates to the federated learning method in machine learning. Federated learning (or joint learning) is a decentralized machine learning framework. The main difference between federated learning and traditional machine learning lies in that in traditional machine learning, the training data is concentrated in a database, and the training device generates a target model based on the training data maintained in the database. In federated learning, the training data is scattered on different node devices, each node device has its own training data, and there is no data exchange between the nodes. Through the cooperation of these node devices, machine learning training is jointly carried out.

[0057] The present application provides a model training method, which is applied to a joint learning system. The main difference between the system framework of joint learning in the present application and the system framework of joint learning in the traditional method is that: the system framework of joint learning in the traditional method includes multiple node devices and a central control device. The system framework of joint learning in the present application includes multiple node devices and does not include a central control device.

[0058] Exemplarily, please refer to Figure 2 As shown, in traditional joint learning, the central control device 101 sends a machine learning network architecture (such as a neural network) and a set of initial weight values to each node device 102. After receiving, each node device 102 uses the local data to train the neural network, obtains the trained parameter values, and then uploads the parameter values and the data volume of the training data to the central control device 101. The central control device 101 calculates the average value according to the proportion according to the parameter values and the data volume of the training data uploaded by each node device 102 collected by the central control device 101, as shown in the following formula (1). The average value W' is the result of this calculation.

[0059]

[0060] Among them, k is the number of node devices, W k is a set of weight values trained by the kth node device, and n k is the data volume of the training data of the kth node device. Then, the central control device 101 transmits the result W' back to each node device 102. This process needs to be repeated multiple times so that the finally selected parameter values can make the accuracy of the training model reach the system's predetermined requirements. In the traditional method, each node device sends its own trained weight values to the central control device. In this process, an attacker may attack the central control device to collect the weight values of each node device, and each node device will disclose its own weight values, resulting in the leakage of training data.

[0061] Please refer to Figure 3As shown, in the system framework of federated learning provided in this application, the system framework includes multiple node devices 301, and the multiple node devices 301 can be communicatively connected to each other. Each node device 301 can interact through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof. In this application, each node device 301 is both a storage device for training data sets and an execution device for training models. Optionally, each node device 301 can also be a data acquisition device for acquiring training data. For example, if the training data set is an image data set, the images can be images acquired by the node device 301 (such as a mobile phone) through a camera. In this application, the parameters in the model to be trained are calculated through federated learning, and a decentralized approach is achieved based on encryption methods to protect the parameter values trained by each node device 301. The core of this federated learning collaboration is to calculate the average value of the parameter values provided by each node device 301. Since the data volumes of each node device 301 are public, the core of federated learning is to calculate the sum value of a parameter value. The calculation of this sum value means that each node device directly calculates the parameter sum value of all weight values based on the encrypted weight values of all nodes. What each node device 301 finally obtains is the sum value of the parameter values (or called the parameter sum value). One node device 301 does not know the parameter values trained by other node devices 301, so the weight values of each node device are not leaked, protecting the privacy data of each node device. And the parameter sum values of all nodes are calculated through the cooperation of multiple node devices, without the assistance of a central control device, achieving a decentralized architecture and preventing attackers from attacking the central control device to obtain the parameters of each node device, thus protecting the privacy data of each node device from being leaked.

[0062] In this application, the node device can be a terminal device (or also called a user device), or a server. Alternatively, among the multiple node devices, some node devices are terminal devices and some are servers. Among them, the terminal device can represent any computing device. For example, the terminal device can be a smart phone, a tablet computer, a personal computer, a computer workstation, a smart camera, a vehicle terminal, a terminal in autonomous driving, a terminal in assisted driving, a terminal in intelligent healthcare, a terminal in industrial Internet of Things, etc. The server can be a file server, a data server, an application server, etc. In the embodiments of this application, the multiple node devices are described by taking terminal devices as an example. For example, the multiple node devices can all be mobile phones as an example.

[0063] Exemplarily, a first model to be trained (which may be an initial network architecture) and initial parameter values (such as all 0s) of the first model can be preset in each node device. The first model can be a neural network. Among them, the neural network includes but is not limited to deep neural network, convolutional neural network, artificial neural network, or recurrent neural network, etc. As for which type of machine learning architecture the first model is can be determined according to the application scenario of the model. For example, if the first model is applied to the scenario of image classification, the first model can be a convolutional neural network architecture specialized in image classification. If the first model can be applied to the scenario of text analysis, the first model can be a recurrent neural network specialized in text analysis. In the embodiments of this application, the first model can be exemplified by a deep neural network.

[0064] Exemplarily, please refer to Figure 4 As shown, a deep neural network (DNN) can be understood as a neural network with many hidden layers. Divided by the positions of different layers in the DNN, the neural network inside the DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers (or called hidden layers, or intermediate layers), and the layers are fully connected. The work of each layer in the hidden layer of the deep neural network can be expressed by the mathematical expression Formula (2). Among them, is the input vector, is the output vector, b is the bias, f() represents the activation function, and W is the weight vector. Each value in the weight vector W represents the weight value of a neuron in this layer of the neural network. The vector W determines the space transformation from the input space to the output space described above, that is, the weight W of each layer controls how to transform the space.

[0065] A neuron in the hidden layer receives multiple inputs, performs calculations inside the neuron and then outputs. The directed line from the input layer to the hidden layer represents signal conduction, and each line has a weight value. There are multiple neurons in the hidden layer. The neurons perform an inner product calculation on the input and the weights on the directed lines, perform calculations through an activation function, and add a bias configured for the neuron, which is the output of the neuron, and then output through the output layer. Exemplarily, taking the output of a neuron as an example, the output of the neuron can be expressed as: h1 = f(x1w1 + x2w2) + b, Formula (3). Among them, h1 is the output of the neuron, f() is the activation function, b is the bias, x1 and x2 are the input values, and w1 and w2 are the weight values. It can be understood that x1 and x2 are in the above formula (2) The input values therein, w1 and w2 are the weight values in W in the above formula (2). It should be noted that Figure 4 It is only an exemplary illustration for convenience. In actual applications, the hidden layer includes multiple layers.

[0066] The neural network training process refers to continuously adjusting the weight values (or the derivative values of the weight values) in the neural network for a set of training data (or training data set) and a neural network structure, so that the output of the deep neural network is as close as possible to the value that is truly desired to be predicted. Therefore, the weight of each layer of the neural network can be updated by comparing the predicted value of the current network with the truly desired target value and according to the difference between the two. For example, if the predicted value of the network is high, adjust the weight to make it predict lower, and continuously adjust until the neural network can predict the truly desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0067] In the embodiments of the present application, during the training process of the model, aiming at the weight value (such as w1) corresponding to the directed line of each neuron in the hidden layer (refer to the above formula 3), calculate the total value of the weight values of each node device, so as not to disclose the weight value of each node device and protect the training data in each node device from being leaked. In this embodiment, the system framework of federated learning includes k node devices, and k is an integer greater than or equal to 2. For the convenience of description, in this example, the first node device and the second node device are taken as examples for description. Among them, the first model to be trained (such as a neural network) and initial parameters (such as all 0) are pre-set in both the first node device and the second node device. Alternatively, the first node device and the second node device download the first model (such as a neural network) and initial parameters (such as all 0) from the server to the local side. The first node device and the second node device can be described by taking a mobile phone as an example. Among them, the actions performed by the first node device and the second node device are similar, and the second node device can be understood by referring to the first node device. In this embodiment, the method for training the model is described from the side of the first node device with the first node device as the execution subject.

[0068] Please refer to Figure 5 As shown, the present application provides an embodiment of a model training method, including:

[0069] Step 501, the first node device obtains the first training data set.

[0070] The first node device obtains a first training data set for training a first model from the local side, where the data in the first training data set has corresponding class labels. The second node device obtains a second training data set for training the first model from the local side.

[0071] Exemplarily, the first training data set is a set of user data stored at the local side of the first node device, and the user data refers to the data generated based on the user's behavior. The user data is related to the user. For example, the user data can be the data input by the user received by the first node device through an input device, such as the user data can be the text input by the user received by the mobile phone through the display screen. Or, the user data can be an image captured by the mobile phone through a camera. The image can be a portrait image or a landscape image, etc. Or, the user data is an image downloaded by the mobile phone through a browser. The user data has privacy and is data that the user does not want to be leaked.

[0072] In one example, the first training data set takes the picture data set as an example. The first node device obtains a first picture data set from the local photo library, and the images in the first picture data set have corresponding class labels. The images in the first picture data set can be images captured by a camera, or the images can also be images downloaded by the first node device through a browser and then stored at the local side.

[0073] In another example, the first training data set takes multiple text records as an example, where each text record includes the first text input by the user and the second text selected by the user. The text in this example does not limit the representation form of the language. For example, it can be Chinese or it can also be English, and the text does not limit the number of characters. For example, the text can be a single character or a phrase. The first text is the text input by the user received by the mobile phone through an input device (such as the display screen), and the second text is the text selected by the user from multiple texts recommended by the first model.

[0074] Step 502: The first node device inputs the first training data set into the first model for training to determine the first parameter value (such as the first weight value) of the target parameter of the first model.

[0075] The first model can be understood as an initial neural network architecture. The target parameter can be a weight, or it can be the derivative of a weight. The target parameter is described by taking the weight as an example, and the parameter value is described by taking the weight value as an example.

[0076] The first node device determines a set of weight values, and the set of weight values is a weight vector (such as W in Equation (2)). Each value in the weight vector represents the weight value of a neuron in this layer of the neural network. The first weight value (such as w1 in Equation (3)) is at least one weight value in the set of weight values.

[0077] It is understandable that the first weight value in this step is the weight value that has been determined according to the loss function after multiple iterations of training and is relatively optimal for local training.

[0078] Step 503: The first node device encrypts the first parameter value to obtain a first encrypted representation of the first parameter value.

[0079] The first node device encrypts the first parameter value through the ElGamal encryption method to obtain a first encrypted representation of the first parameter value, and broadcasts the first encrypted representation. The purpose of broadcasting is to enable the second node device to receive the first encrypted representation. Optionally, the first node device can also broadcast the data volume of the first training dataset.

[0080] Optionally, the first parameter value is in binary representation and has L bit positions. Wherein, L is an integer greater than or equal to 2. A first private key is generated for each of the L bit positions of the first parameter value, and the first private key can be a random value. The first node device generates a first public key according to the first private key and sends the first public key to the second node device. The first public key and the second public key are discrete logarithms. At the same time, the first node device will receive the second public key sent by the second node device. Further, the first node device can encrypt the first parameter value according to the first public key and the second public key to obtain a first encrypted representation of the first parameter value.

[0081] Step 504: The first node device receives a second encrypted representation of the second parameter value of the target parameter sent by the second node device, and the second parameter value is obtained after the second node device inputs the second training dataset into the first model for training.

[0082] The first node device receives the second encrypted representation of the second parameter value broadcast by the second node device. The second parameter value is in binary representation and has L bit positions.

[0083] It should be noted that there is no timing limit for Step 503 and Step 504. Optionally, Step 504 can also be before Step 503.

[0084] Step 505: The first node device calculates the parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representation, and the parameter sum value is used to determine the second model.

[0085] First, the first node device determines the bitwise sum value of the t-th bit in the first encrypted representation and the bit value of the t-th bit in the second encrypted representation, where t sequentially takes each integer from 1 to L. That is, calculate the bitwise sum value of the two parameter values at the first bit position, the bitwise sum value of the two parameter values at the second bit position, until the bitwise sum value of the two parameter values at the L-th bit position.

[0086] Then, calculate the parameter sum value (weight sum value) of the first parameter value and the second parameter value according to the bitwise sum value corresponding to each bit in the L bits.

[0087] Step 506: The first node device inputs the first data into the second model and outputs the result data through the second model.

[0088] The first node device inputs the local data into the second model and outputs the result data through the second model. This output result is used to test the second model to check whether the second model meets the system requirements of the first node device. Since the parameter sum value obtained in the above step 505 is based on the joint training of two node devices, that is, both the first node device and the second node device contribute to the training of the neural network. By testing the second model with the local data of the first node device, it is tested whether the second model meets the system requirements of the first node device.

[0089] Step 507: Determine whether the difference between the result data and the target data is greater than the threshold. If the difference between the result data and the target data is greater than the threshold, then return to step 502, broadcast the parameter adjustment message, perform the next round of model parameter adjustment, update the first model in step 501 to the second model, that is, input the first training data set into the second model and train the second model to determine the weight value of the second model. The parameter adjustment message is used to notify the second node device that the next round of parameter adjustment training is required.

[0090] If the difference between the result data and the target data is less than or equal to the threshold, execute step 508. Optionally, if the difference between the result data and the target data is less than or equal to the threshold and the first node device has not received the parameter adjustment message sent by the second node device, execute step 508.

[0091] For example, input the first image into the second model and output the classification result through the second model. The first image can be an image in the first training data set. If the difference between the classification result and the category information of the first image is greater than the threshold, then broadcast the parameter adjustment message. The parameter adjustment message is used to notify the second node device to continue training the second model; if the difference between the classification result and the category information of the first image is less than or equal to the threshold, then determine the second model according to the parameter sum value.

[0092] For another example, the first node device inputs the third text into the second model and outputs the associated text for the third text through the second model. The third text can be the text received in real time by the first node device through the input device, or the third text can also be the data in the first training dataset. If the difference between the associated text output through the second model and the fourth text selected by the user is greater than the threshold, a parameter adjustment message is broadcast, and the parameter adjustment message continues to train the second model. If the difference between the associated text and the fourth text selected by the user is less than or equal to the threshold, the second model is determined according to the parameter sum value.

[0093] Step 508: Determine the second model according to the parameter sum value.

[0094] Calculate the weighted average value according to the parameter sum value as shown in the following formula (4), and then calculate the second model according to the weighted average value.

[0095] (w′1n1 + w′2n2 +... + w′ k n k ) / (n1 + n2 +... + n k ), formula (4). Where k is the number of node devices, n k is the data volume broadcast by the k-th node device, and w′ k is: the weighted sum value of the k node devices determined for w k .

[0096] In the first case, the data volume broadcast by each node device is the same, and the average value of the weight values can be determined according to the following formula (5).

[0097] (w′1n1 + w′2n2 +... + w′ k n k ) / (n1 + n2 +... + n k ) = (w′1 + w′2 +... + w′ k ) / k, formula (5).

[0098] In this case, the data volume of the training dataset of each node device is the same, and the average value of the weight values can be directly calculated without calculating the data volume, saving the computing power of the node device.

[0099] In the second case, the data volume of the training dataset of the i-th node device is represented by c i ×M. Where M is the order of magnitude of the data volume, and this order of magnitude is recognized by both the first node device and the second node device. For example, M = 1000 or M = 10000, etc. c i is the factor of the i-th node device. When the i-th node device broadcasts the data volume, it can only broadcast this c i to simplify the representation of the data volume of the training data.

[0100] In this case, the weighted average value is calculated by formula (4), and the data volume is represented as c i ×M, which simplifies the representation of the data volume of the training data, expands the flexibility that each node device can have different data volumes, and also gives the parameter values of the local neural network contributed by each node device the flexibility of different importance.

[0101] It should be noted that each weight value in the above formula (5) is the parameter sum value calculated in step 505. Each weight value in formula (5) can be calculated in one batch, and the calculation of each weight value has no constraint of sequence.

[0102] The model training method provided in this embodiment can be applied to the scenario of training a model with private data. The federated learning system architecture includes a first node device and a second node device. Among them, both the first node device and the second node device will use the data set stored locally to train the first model. Each node device trains the first model according to its own local training data set, obtains the weight value obtained by training the first model for each, then encrypts the weight value, and broadcasts the encrypted weight value. As a result, each node device will receive the encrypted weight values broadcast by other nodes. Finally, each node device further directly calculates the parameter sum value of all weight values according to the encrypted weight values of all nodes, and then each node device updates the first model according to the parameter sum value. Since each node device directly calculates the parameter sum values of all nodes and does not know the weight values of each node device, the weights of each node device will not be leaked, preventing attackers from collecting the weights of each node device to extract the training data of the node device, thereby protecting the private data of each node device. And the parameter sum values of all nodes are calculated through the cooperation of multiple node devices without the assistance of a central control device, realizing a decentralized architecture and preventing attackers from attacking the central control device to obtain the parameters of each node device, thereby protecting the private data of each node device from being leaked.

[0103] Please refer to Figure 6 As shown, another embodiment of a model training method is provided in this application. In this embodiment, a voting mechanism based on a cryptographic method is used to encrypt the weight values of all node devices and calculate the sum value of the weight values of all node devices. First, the parameters involved in this embodiment are described. In this embodiment, the parameter values are encrypted by ElGamal encryption, and the ElGamal encryption algorithm includes parameters (G, g, p). Among them, G is a multiplicative group with a size of q, q is a prime number, p = rq + 1, P is a prime number, g is a primitive root of p, and r is a coefficient. The system framework of federated learning in this embodiment includes k node devices. For the convenience of description, in this embodiment, 4 node devices are taken as an example, and N is usedi Represents the i-th node device, where i is an integer greater than or equal to 1 and less than or equal to k. The k node devices are sorted according to an index value, and the k node devices know this sorting from each other. For example, the first node device (denoted as N1), the second node device (denoted as N2), the third node device (denoted as N3), and the fourth node device (denoted as N4).

[0104] Step 601: Each node device obtains a training data set from its local side.

[0105] For example, the first node device obtains the first training data set for training the first model from the storage device on its local side. At the same time, the second node device obtains the second training data set for training the first model from its local side. The third node device obtains the third training data set for training the first model from its local side. The fourth node device obtains the fourth training data set for training the first model from its local side.

[0106] This step can be understood by referring to Figure 5 Step 501 in the corresponding embodiment.

[0107] Step 602: Each node device inputs the training data set on its local side into the first model for training to determine the parameter values of the target parameters of the first model.

[0108] For example, the first node device inputs the first training data set on its local side into the first model for training to determine the first weight value (denoted as A) of the first model. At the same time, the second node device inputs the second training data set on its local side into the first model for training to determine the second weight value (denoted as B) of the first model. The third node device inputs the third training data set on its local side into the first model for training to determine the third weight value (denoted as C) of the first model. The fourth node device inputs the fourth training data set on its local side into the first model for training to determine the fourth weight value (denoted as D) of the first model.

[0109] This step can be understood by referring to Figure 5 Step 502 in the corresponding embodiment.

[0110] Step 603: Each of the node devices generates a private key for each of the L bits of its respective parameter value.

[0111] The weight values of each of the node devices are represented as binary numbers, and the length of the binary number is L.

[0112] Each of the k node devices determines a parameter value (i.e., weight value), for a total of k parameter values. L is the length of the binary value of the maximum value among the k parameter values. Among them, 2 L<q. For example, the first parameter value is 1 (binary representation: 01), the second parameter value is 2 (binary representation: 10), the third parameter value is 3 (binary representation: 11), and the fourth parameter value is 4 (binary representation: 100). The fourth parameter value is the maximum value, and the length of the binary value of the fourth parameter value is 3, so the L is 3.

[0113] The i-th node device N i The binary representation of the weight value of is: a i,1 a i,2... a i,t... a i,L . Where t represents the t-th bit.

[0114] For example, the first weight value is represented in binary as: a 1,1 a 1,2... a 1,t... a 1,L . The second weight value is represented in binary as a 2,1 a 2,2... a 2,t... a 2,L . The third weight value is represented in binary as a 3,1 a 3,2... a 3,t... a 3,L . The fourth weight value is represented in binary as a 4,1 a 4,2... a 4,t... a 4,L .

[0115] Step 604: Each node device generates a public key according to its own private key and broadcasts it. At the same time, it receives the public keys sent by other node devices.

[0116] The i-th node device N i Generates a random number x for each bit, uses the random number x as the private key, and calculates the public key according to the private key. The public key is: Equation (5). It can be understood as a part of the ElGamel ciphertext. Where i represents the i-th node device, t represents the t-th bit, and t is an integer greater than or equal to 1 and less than or equal to L.

[0117] Each node device broadcasts its own public key, so that other node devices can receive the public key.

[0118] Exemplarily, N1 broadcasts the first public key, and N2, N3, and N4 can all receive the first public key. N2 broadcasts the second public key, and N1, N3, and N4 can all receive the second public key. N3 broadcasts the third public key, and N1, N2, and N4 can all receive the third public key. N4 broadcasts the fourth public key, and N1, N2, and N3 can receive the fourth public key.

[0119] Optionally, when each node device broadcasts its public key, it also broadcasts a zero-knowledge proof zkp(x i,t ), where zkp(x i,t ) indicates that the i-th node device has the information of x i,t , but zkp(x i,t ) does not disclose the information of x i,t .

[0120] Step 605: Each node device encrypts its respective parameter value according to the public keys of all node devices, and each node device obtains an encrypted representation of its respective parameter value.

[0121] First, after each node device receives the public keys broadcast by other node devices, it further determines the association relationship between the private keys generated by each node device according to all the public keys. This association relationship is represented by the following relational expression (6):

[0122] Expression (6). Where j represents the j-th node device, represents the association relationship between the private keys.

[0123] Then, each node device obtains the encrypted representation of the parameter value according to the above expressions (5) and (6) as: Expression (7). Where a i,t represents the bit value of the t-th bit of the i-th node device, and a i,t is 0 or 1.

[0124] Exemplarily, the first encrypted representation of the first parameter value of the first node device is:

[0125] The second encrypted representation of the second parameter value of the second node device is:

[0126] The third encrypted representation of the third parameter value of the third node device is:

[0127] The third encrypted representation of the third parameter value of the fourth node device is:

[0128] Finally, each node device broadcasts its respective encrypted representation so that other node devices can receive the encrypted representation broadcast by this node device. Optionally, each node device also broadcasts a zero-knowledge proof zkp(a i,t ). Where zkp(a i,t ) can prove that a i,t is a bit value.

[0129] Step 606: Each node device determines the sum of the bit values on the t-th bit of the k encrypted representations, where t ranges over each integer from 1 to L.

[0130] First, after each node device receives the encrypted representations broadcast by other node devices, it calculates the product of the encrypted representations of all node devices (such as calculating the product of V 1,t , V 2,t , V 3,t and V 4,t ), denoted as Equation (8).

[0131] Then, taking 4 node devices as an example, calculate the following formula through the above Equation (6):

[0132]

[0133] Obtain: Equation (9).

[0134] Finally, according to the above Equation (9), use Shank’s baby-step giant-step algorithm to calculate the sum of the bit values on the t-th bit: a 1,t +a 2,t +a 3,t +a 4,t .

[0135] Step 607: Each node device calculates the sum of the parameters of the weights of the k node devices according to the sum of the bit values corresponding to each bit among the L bits.

[0136] Exemplarily, taking 4 node devices as an example, and taking each weight value as 3 bits, the sum of the parameters of the weight values can be obtained as follows:

[0137] A + B + C + D = 2 2 ×(a 1,1 +a 2,1 +a 3,1 +a 4,1 ) + 2 1 ×(a 1,2 +a 2,2 +a 3,2 +a 4,2 ) + 2 0 ×(a 1,3 +a 2,3 +a 3,3 +a 4,3 )

[0138] It can be understood that, please refer to Figure 7Each weight value is represented in binary. L bits are equivalent to L candidates in the voting mechanism. Each node device is equivalent to a voter in the voting mechanism. At this time, each weight value is a binary string, and the bit value on each bit is 0 or 1. The 0 or 1 on each bit is equivalent to a ballot in the voting mechanism. Calculating the bit value and value of the weight value of the t-th bit is equivalent to counting the ballots for the t-th candidate. In the voting mechanism, the sum of all the votes for each candidate is directly calculated, but the votes cast by each voter are not disclosed. And through the adopted cryptographic algorithm and zero-knowledge proof, the voting information can be verified, that is, the bit value and value (such as a 1,t +a 2,t +a 3,t +a 4,t ) in this embodiment can be verified. For example, if the value on a certain bit is incorrect, that is, a value other than 0 or 1, this voting mechanism can detect it to reduce the probability of incorrect parameter sum values, thereby improving the accuracy of the model.

[0139] Step 608: For the i-th node device, the i-th node device inputs the first data into the second model and outputs result data through the second model.

[0140] Exemplarily, the first node device inputs the first data at the local end into the second model and outputs the first result data through the second model. The second node device inputs the second data at the local end into the second model and outputs the first result data through the second model. The third node device inputs the third data at the local end into the second model and outputs the third result data through the second model. The fourth node device inputs the fourth data at the local end into the second model and outputs the fourth result data through the second model.

[0141] If the difference between the result data and the target data is greater than the threshold, return to step 602, broadcast a parameter adjustment message, and perform the next round of model parameter adjustment. The first model in step 601 is updated to the second model, that is, the first training dataset is input into the second model to train the second model, thereby determining the weights of the second model. This parameter adjustment message is used to notify other node devices that the next round of parameter adjustment training is required.

[0142] If the difference between the result data and the target data is less than or equal to the threshold, execute step 609. Determine the second model according to the parameter sum value.

[0143] This step can be understood by referring to Figure 5 step 506 in the corresponding embodiment.

[0144] Step 609: For the i-th node device, the i-th node device determines the second model according to the parameter sum value.

[0145] This step is understood in conjunction with Figure 5 step 507 in the corresponding example, which will not be elaborated here.

[0146] Optionally, among the k node devices in step 605, one node device does not broadcast its encrypted representation. For example, the third node device drops offline, and the first node device, the second node device, and the fourth node device do not receive the third encrypted representation broadcast by the third node device.

[0147] For each node device in N1, N2, and N4, calculate the third encrypted representation according to the public key received in step 604.

[0148] If j < 3, then calculate and broadcast according to Equation (6)

[0149] For example, the first node device calculates and broadcasts Equation (10). The second node device calculates and broadcasts Equation (11).

[0150] If j > 3, then calculate and broadcast according to Equation (6)

[0151] For example, the fourth node device calculates and broadcasts Equation (12).

[0152] The first node device multiplies the above Equation (10), Equation (11), and Equation (12) to obtain the third encrypted representation. Similarly, the second node device and the fourth node device obtain the third encrypted representation.

[0153] Then, in step 606, for each node device, calculate the product of the encrypted representations of each node device (such as calculating the product of V 1,t , V 2,t , V 3,t and V 4,t ) to obtain: Equation (9).

[0154] According to the above Equation (9), use Shank’s baby-step giant-step algorithm to calculate the bit value and value of the weight value of the t-th bit: a 1,t +a 2,t +a 3,t +a 4,t .

[0155] Finally, execute step 607.

[0156] In this example, if one of the multiple node devices goes offline and the offline node device does not broadcast the encrypted representation of the weight value, then the other node devices can recover the encrypted representation of the weight of the offline node device based on the public key previously broadcast by the offline node device. In this example, even if a node device goes offline, it will not affect the parameters and values for calculating the weight value, enhancing the robustness of the federated learning system.

[0157] The model training method provided in this embodiment can be applied to the scenario of training a model with user data (such as privacy data). The federated learning system architecture includes multiple node devices. Among them, each of the multiple node devices will use the dataset stored locally to train the first model. Each node device trains the first model according to its respective local training dataset to obtain its own weight value for the first model, and then encrypts the weight value based on the voting mechanism of the ElGamal encryption algorithm. And broadcasts the encrypted weight value, so that each node device will receive the encrypted weight values broadcast by other nodes. Based on this voting mechanism, each bit of the weight values of each node device is equivalent to a candidate in the voting mechanism. Each node device is equivalent to a voter in the voting mechanism. At this time, each weight value is a binary string, and the bit value at each bit is 0 or 1. The 0 or 1 at each bit is equivalent to a ballot in the voting mechanism. In the voting mechanism, the sum of all the votes of each candidate is directly calculated, but the votes cast by each voter are not leaked. Equivalent to each node device further directly calculating the parameters and values of all the weight values according to the encrypted weight values of all the nodes, and then each node device updates the first model according to the parameters and values. Since each node device directly calculates the parameters and values of all the nodes and does not know the weight values of each node device, nor will it leak the weight values of each node device, it prevents attackers from obtaining the training data of the node devices by collecting the weight values of each node device, thus protecting the privacy data of each node device. And the parameters and values of all the nodes are calculated by the cooperation of multiple node devices without the assistance of a central control device, realizing a decentralized architecture, preventing attackers from attacking the central control device to obtain the weight values of each node device, thus protecting the privacy data of each node device from being leaked.

[0158] In an application scenario, the system framework of federated learning includes mobile phone A, mobile phone B, and mobile phone C, which can be connected through the same WiFi. The first model (such as a neural network) is trained through the federated learning of the user data on the local sides of mobile phone A, mobile phone B, and mobile phone C respectively. The first model can be an application in the mobile phone. For example, this application is used to classify the pictures in the mobile phone. This application can be pre-loaded in the mobile phone, or this application can also be downloaded by the user to the local side of the mobile phone according to the needs. For example, this application is a neural network architecture and has an initial weight value. For mobile phone A, there are multiple images stored in the local side of mobile phone A, and these multiple images may include the user's portrait photos, ID photos, landscape photos, or the photos may include the user's privacy data (such as address, ID number, etc.). These user data are not expected to be leaked. Through the model training method in the above method embodiments, the neural network can be trained on the local side without leaking its own weight value, protecting the user data in the mobile phone from being leaked. The photos used for training in the photo library on the local side of mobile phone A are the photos that have been pre-classified by the user, and each photo has a classification label. Mobile phone A prompts the user whether to update this application (or it can also automatically update without prompting the user). After the user determines to update, mobile phone A obtains the image data set with classification labels on the local side, trains the first model (neural network), obtains the weight value, and obtains the encrypted representation of this weight value and broadcasts it. At the same time, mobile phone B and mobile phone C will also broadcast the encrypted representations of their respective weight values, so that each node device will receive the encrypted representations of the weight values broadcast by other nodes. Finally, each node device further directly calculates the parameter sum value of all weight values according to the encrypted representations of the weight values of all nodes (without decrypting each weight value), and then each node device updates the first model according to this parameter sum value. Since each node device directly calculates the parameter sum value of all nodes and does not know the weight value of each node device, the weight value of each node device is avoided from being leaked, ensuring the security of the user's privacy data.

[0159] For another example, the first model can also be a text prediction model (or a keyboard prediction model), and the first model can be an input method application. Each time the user inputs text, mobile phone A receives the text input by the user through the display screen, and the input method application recommends multiple associated texts after the current text to the user. The user selects a target text from the multiple associated texts. Mobile phone A locally stores multiple text records; each text record includes the first text input by the user and the second text selected by the user. The multiple text records are used as the training data of the first model. Mobile phone A inputs the training data into the first model to obtain parameter values, and then encrypts the parameter values to obtain an encrypted representation of the parameter values. Through the model training method in the present application, a second model (text prediction model) is obtained, and during the model training process, the leakage of user data is avoided.

[0160] In another application scenario, there are two data owners (i.e., enterprise A and enterprise B). Enterprise A has a first business system, and enterprise B has a second business system. This scenario can be extended to a scenario including multiple data owners. For example, enterprise A and B want to jointly train a machine learning model, and their business systems respectively have relevant data (i.e., training datasets) of their own users. In addition, enterprise A and enterprise B also have label data for the relevant data. For data privacy protection and security considerations, the first business system (such as the first node device) and the second business system (such as the second node device) can apply the model training method provided in the above method embodiments to train the machine learning model to protect user data.

[0161] Of course, other scenarios in the present application that can be used for protecting the local training data of node devices in federated learning are not limited by the above examples, and the application scenarios are not listed one by one here.

[0162] The above describes the method of model training. Next, the application scenarios of the machine learning model obtained by the methods provided by the corresponding embodiments through Figure 5 and Figure 6 are illustrated by examples.

[0163] Exemplarily, an embodiment of the present application provides an image classification method. This method is applied to the first node device. After the first node device jointly trains the first model, the first model on the local side is updated to a second model, and the second model is used for classifying images.

[0164] First, the first node device obtains a first image. The first image can be an image stored in the local storage device, or the first image can also be an image captured by the camera of the first node device in real time. For example, the image is a photo of the user.

[0165] Then, the first node device inputs the first image into the image classification model, and the classification result is output through the image classification model. For example, the classification result is "person", and the first image is stored in the "person" folder. In this example, the input image is classified by the image classification model.

[0166] Exemplarily, an embodiment of the present application provides a method for determining associated text, which is applied to the first node device. The first node device updates the first model at the local end to the second model after jointly training the first model, and the second model is used to predict text. For example, the second model can be an input method application installed on a mobile phone.

[0167] The first node device receives the text input by the user through the input device.

[0168] For example, the first node device receives "mine" input by the user through the display screen.

[0169] The first node device inputs the text into the text prediction model, and the associated text of the text is output through the text prediction model. The first recommended associated text of this application is "address", and the second recommended associated text is "name".

[0170] The first node device receives the target text selected by the user through the display screen. For example, the target text is "address".

[0171] Then, the first node device inputs the text "address" selected by the user into the text prediction model, and the associated text of the text is output through the text prediction model. For example, the associated text of "address" is "Shenzhen". In this example, the associated text of the input text is predicted by the text prediction model.

[0172] The above describes a model training method. Next, the node device 800 to which this method is applied is described. This node device can be any node device in the federated learning system. Please refer to Figure 8 As shown, this node device includes a processing module 801 and a transceiver module 802; where

[0173] The processing module 801 is used to obtain the first training dataset;

[0174] Input the obtained first training dataset into the first model for training, and determine the first parameter value of the target parameter of the first model;

[0175] Encrypt the first parameter value to obtain the first encrypted representation of the first parameter value;

[0176] A transceiver module 802, configured to receive a second encrypted representation of the second parameter value of the target parameter sent by each of the second node devices, where the second parameter value is obtained after the second node devices input a second training data set into the first model for training;

[0177] The processing module 801 is further configured to calculate a parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representation, where the parameter sum value is used to update the first model, so as to obtain an updated second model.

[0178] Further, the processing module 801 is configured to execute Figure 5 Steps 501 - 503, steps 505 - 508 in the corresponding method embodiment. For specific reference, please refer to Figure 5 The specific descriptions in steps 501 - 503, steps 505 - 508 in the corresponding method embodiment. The transceiver module 802 is configured to execute Figure 5 Step 504 in the corresponding method embodiment. For specific reference, please refer to Figure 5 The specific description in step 504 in the corresponding method embodiment.

[0179] The processing module 801 is further configured to execute Figure 6 Steps 601 - 603 and steps 605 - 609 in the corresponding method embodiment. For specific reference, please refer to Figure 5 The specific descriptions in steps 601 - 603 and steps 605 - 609 in the corresponding method embodiment. The transceiver module 802 is configured to execute Figure 6 Step 604 in the corresponding method embodiment. For specific reference, please refer to Figure 6 The description in step 604 in the corresponding method embodiment.

[0180] In one implementation, the processing module 801 may be a processing device, and the functions of the processing device may be partially or fully implemented by software.

[0181] Optionally, the functions of the processing device may be partially or fully implemented by software. In this case, the processing device may include a memory and a processor, where the memory is used to store a computer program, and the processor reads and executes the computer program stored in the memory to execute the corresponding processing and / or steps in any method embodiment.

[0182] Optionally, the processing device may only include a processor. The memory for storing the computer program is located outside the processing device, and the processor is connected to the memory through a circuit / wire to read and execute the computer program stored in the memory.

[0183] Optionally, the processing device may be one or more chips, or one or more integrated circuits.

[0184] Exemplarily, an embodiment of the present application provides a chip structure. Please refer to Figure 9 As described above, the chip includes:

[0185] The chip can be manifested as a neural network processor NPU. The NPU is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 903. The arithmetic circuit 903 is controlled by the controller 904 to extract matrix data from the memory and perform multiplication operations.

[0186] In some implementations, the arithmetic circuit 903 internally includes multiple processing units (Process Engine, PE). In some implementations, the arithmetic circuit 903 is a two-dimensional systolic array. The arithmetic circuit 903 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 903 is a general matrix processor.

[0187] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 902 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 901 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 908 (accumulator).

[0188] The unified memory 906 is used to store input data and output data. The weight data is directly transported to the weight memory 902 through the storage unit access controller 905 (Direct Memory Access Controller, DMAC). The input data is also transported to the unified memory 906 through the DMAC.

[0189] The BIU is the Bus Interface Unit, that is, the bus interface unit 910, which is used for the interaction between the AXI bus and the DMAC and the instruction fetch memory 909 (Instruction Fetch Buffer).

[0190] The bus interface unit 910 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch memory 909 to obtain instructions from the external memory, and is also used for the storage unit access controller 905 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0191] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 906, or transfer the weight data to the weight memory 902, or transfer the input data to the input memory 901.

[0192] The vector computing unit 907 has multiple arithmetic processing units. When needed, it further processes the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / FC layer network calculations in neural networks, such as Pooling, Batch Normalization, Local Response Normalization, etc.

[0193] In some implementations, the vector computing unit 907 can store the processed output vector into the unified buffer 906. For example, the vector computing unit 907 can apply a non-linear function to the output of the arithmetic circuit 903, such as a vector of accumulated values, to generate activation values. In some implementations, the vector computing unit 907 generates normalized values, combined values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 903, for example, for use in subsequent layers in a neural network.

[0194] The instruction fetch buffer 909 connected to the controller 904 is used to store the instructions used by the controller 904;

[0195] The unified memory 906, the input memory 901, the weight memory 902, and the instruction fetch memory 909 are all On-Chip memories. The external memory is private to this NPU hardware architecture.

[0196] Among them, Figure 4 The operations of each layer in the shown neural network can be executed by the arithmetic circuit 903 or the vector computing unit 907.

[0197] The arithmetic circuit 903 or the vector computing unit 907 calculates a parameter value (such as a first parameter value), and the main CPU is used to read the computer program stored in the at least one memory, so that the node device executes the method described in the above method embodiments.

[0198] Please refer to Figure 10As shown in the figure, the embodiments of the present application further provide another node device, which may be the terminal device 1000. For example, the node device may be a smart phone, a tablet computer, a personal computer, a computer workstation, a smart camera, a vehicle-mounted terminal, a terminal in unmanned driving, a terminal in assisted driving, a terminal in intelligent healthcare, a terminal in industrial Internet of Things, etc.

[0199] As Figure 10 shown in the figure, the terminal device 1000 includes a processor 1001, a transceiver 1002, and a memory 1003. Among them, the processor 1001, the transceiver 1002, and the memory 1003 can communicate with each other through an internal connection path to transmit control signals and / or data signals. The memory 1003 is used to store computer programs, and the processor 1001 is used to call and run the computer programs from the memory 1003 to control the transceiver 1002 to send and receive signals. Optionally, the terminal device 1000 may further include an antenna. The transceiver 1002 transmits or receives wireless signals through the antenna. Optionally, the processor 1001 and the memory 1003 may be integrated into a processing device, and the processor 1001 is used to execute the program code stored in the memory 1003 to implement the above functions.

[0200] Optionally, the memory 1003 may also be integrated in the processor 1001. Or, the memory 1003 is independent of the processor 1001, that is, located outside the processor 1001.

[0201] The processor 1001 may be used to execute the actions implemented by the node device described in the previous method embodiments. For specific reference, see Figure 5 or Figure 6 the description in the corresponding method embodiments, which will not be elaborated here. The memory 1003 is used to implement the storage function. Optionally, the transceiver 1002 may be used to perform the function of broadcasting information (such as public keys, encrypted representations, etc.) executed by the node device to other node devices.

[0202] Figure 8 The processing and / or operations performed by the processing module 801 in Figure 10 can be implemented by the processor 1001 shown in

[0203] . For specific details, please refer to the detailed description of the method embodiments, which will not be elaborated here.

[0204] In addition, in order to make the functions of the terminal device more complete, the terminal device 1000 may further include one or more of an input unit 1006, a display unit 1007, a camera 1005, an audio circuit 1008, etc. The audio circuit may further include a speaker 10082, a microphone 10084, etc.

[0205] Please refer to Figure 11 As shown, the embodiment of the present application further provides another node device, which is a server 1100. For example, the server may be a file server, a data server, an application server, etc.

[0206] Please refer to Figure 11 As shown, the embodiment of the present application further provides a server 1100. Figure 11 FIG. is a schematic structural diagram of a server provided by the embodiment of the present application. The server 1100 may vary greatly due to different configurations or performances, and may include one or more processors 1122 and a memory 1132, and one or more readable storage media 1130 (such as one or more mass storage devices) for storing application programs 1142 or data 1144. Among them, the memory 1132 and the readable storage media 1130 may be transient storage or persistent storage. The program stored in the readable storage media 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the processor 1122 may be configured to communicate with the readable storage media 1130 and execute a series of instruction operations in the readable storage media 1130 on the server 1100.

[0207] The server 1100 may further include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141.

[0208] In the embodiment of the present application, the processor is used to read the computer program stored in the at least one memory, so that the server executes Figure 5 and Figure 6 the method steps executed by the server in the corresponding embodiment. For specific details, please refer to the description in the method embodiment, which will not be elaborated here.

[0209] It can be understood that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed through the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0210] The embodiments of the present application also provide a computer-readable storage medium. A program is stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the steps performed by the node device in the method described in the foregoing Figure 5 and Figure 6 illustrated embodiments.

[0211] The embodiments of the present application also provide a computer program product. When it runs on a computer, it causes the computer to execute the steps performed by the node device in the method described in the foregoing Figure 5 and Figure 6 illustrated embodiments.

[0212] The embodiments of the present application also provide a circuit system. The circuit system includes a processing circuit configured to execute the steps performed by the node device in the method described in the foregoing Figure 5 and Figure 6 illustrated embodiments.

[0213] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0214] In the several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0215] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0216] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist physically separately for each unit, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0217] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0218] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present application.

Claims

1. A model training method, characterized in that, Applied to a federated learning system, the federated learning system includes a first node device and N second node devices, where N is an integer greater than or equal to 1. The method is applied to the first node device and includes: Obtain a first training data set; Input the first training data set into a first model for training to determine a first parameter value of the target parameter of the first model; Encrypt the first parameter value to obtain a first encrypted representation of the first parameter value; Receive a second encrypted representation of the second parameter value of the target parameter sent by each second node device, where the second parameter value is obtained after the second node device inputs a second training data set into the first model for training; Calculate a parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representation, where the parameter sum value is used to update the first model, thereby obtaining an updated second model; The first parameter value has L bits, where L is an integer greater than or equal to 2. The encrypting the first parameter value to obtain a first encrypted representation of the first parameter value includes: Generate a first private key for each of the L bits of the first parameter value; Generate a first public key according to the first private key and receive a second public key sent by each second node device, where the second public key is generated by the second node device according to a second private key; Encrypt the first parameter value according to the first public key and the second public key to obtain a first encrypted representation of the first parameter value.

2. The method according to claim 1, wherein The second parameter value has L bits; The calculating a parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representation includes: Determine a sum value of the bit values of the t-th bit in the first encrypted representation and the bit value of the t-th bit in the second encrypted representation, where t takes each integer from 1 to L; Calculate a parameter sum value of the first parameter value and the second parameter value according to the sum value of the bit values corresponding to each of the L bits.

3. The method according to claim 1, wherein The federated learning system further includes a third node device, and the method further includes: Receive a third public key broadcast by the third node device, where the third public key is calculated by the third node device according to a generated third private key; Calculate the correlation relationship of the first private key, the second private key and the third private key according to the first public key, the second public key and the third public key; Calculate a third encrypted representation of a third parameter value corresponding to the third node device according to the correlation relationship; The calculating a parameter sum value of the first parameter and the second parameter according to the encrypted first parameter value and the encrypted second parameter value includes: Calculate a parameter sum value of the first parameter value, the second parameter value and the third parameter value according to the first encrypted representation, the second encrypted representation and the third encrypted representation.

4. The method according to claim 1, characterized in that, The method further includes: Determine a target parameter value according to a first data volume of the first training data set, a second data volume of the second training data set, and the parameter and value, where the target parameter value is used to determine the second model.

5. The method according to claim 4, characterized in that The first data volume is the same as the second data volume.

6. The method according to claim 4, wherein The second data volume is c×M; where M is the data order of magnitude and c is a factor, and c is greater than 0; The method further includes: Receive the second data volume sent by the second node device, where the second data volume is represented by c.

7. The method according to any one of claims 1-6, characterized in that The second model is an image classification model, the first training data set is an image data set, and each image in the image data set has a corresponding classification label.

8. The method according to any one of claims 1 to 6, characterized in that The second model is a text prediction model, and the first training data set includes multiple text records; where each text record includes a first text input by the user and a second text selected by the user.

9. The method according to claim 1, characterized in that, The method further includes: Obtain first data; Input the first data into the second model, and output result data through the second model; If the difference between the result data and the first target data is greater than a threshold, broadcast a parameter adjustment message, where the parameter adjustment message is used to notify the second node device to continue training the second model; If the difference between the result data and the first target data is less than or equal to the threshold, determine the second model according to the parameter and value.

10. The method according to claim 9, characterized in that, The inputting the first data into the second model and outputting result data through the second model includes; Input a first image into the second model and output a classification result through the second model; If the difference between the classification result and the category information of the first image is greater than the threshold, perform the step of broadcasting parameter adjustment; If the difference between the classification result and the category information of the first image is not greater than the threshold, perform the step of determining the second model according to the parameter and value.

11. The method according to claim 9, characterized in that, The first data is a third text, and the inputting the first data into the second model and outputting result data through the second model includes: Input the third text into the second model and output an associated text for the third text through the second model; If the difference between the associated text of the third text output by the second model and the fourth text selected by the user is greater than the threshold, perform the step of broadcasting parameter adjustment; If the difference between the associated text of the third text output by the second model and the fourth text selected by the user is not greater than the threshold, perform the step of determining the second model according to the parameter and value.

12. A method for image classification, characterized in that, Applied to a first node device, the first node device belongs to a federated learning system, and the federated learning system further includes at least N second node devices, and the method includes: The first node device obtains a first image; The first node device inputs the first image into an image classification model, and outputs a classification result through the image classification model. The image classification model is as follows: The first node device inputs the obtained first training data set into a first model for training. The first training data set includes multiple images, and each of the images has a corresponding classification label; determine a first parameter value of the target parameter of the first model; encrypt the first parameter value to obtain a first encrypted representation of the first parameter value; receive second encrypted representations of second parameter values of the target parameter sent by N second node devices, where the second parameter values are obtained after the second node devices input a second training data set into the first model for training; calculate a parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representations, and obtain it according to the parameter sum value; The first parameter value has L bit positions, where L is an integer greater than or equal to 2; the encrypting the first parameter value to obtain a first encrypted representation of the first parameter value includes: generating a first private key for each of the L bit positions of the first parameter value; generating a first public key according to the first private key, and receiving second public keys sent by each of the second node devices, where the second public keys are generated by the second node devices according to second private keys; encrypting the first parameter value according to the first public key and the second public keys to obtain a first encrypted representation of the first parameter value.

13. A method for determining associated text, characterized in that, Applied to a first node device, the first node device belongs to a federated learning system, and the federated learning system further includes at least N second node devices. The method includes: The first node device receives text input by a user through an input device; The first node device inputs the text into a text prediction model, and outputs associated text of the text through the text prediction model. The text prediction model is as follows: The first node device inputs the obtained first training data set into a first model for training. The first training data set includes multiple text records; where each of the text records includes first text input by a user and second text selected by the user; determine a first parameter value of the target parameter of the first model; encrypt the first parameter value to obtain a first encrypted representation of the first parameter value; receive second encrypted representations of second parameter values of the target parameter sent by N second node devices, where the second parameter values are obtained after the second node devices input a second training data set into the first model for training; calculate a parameter sum value of the first parameter value and the second parameter value according to the first encrypted representation and the second encrypted representations, and obtain it according to the parameter sum value; The first parameter value has L bit positions, where L is an integer greater than or equal to 2; encrypting the first parameter value to obtain a first encrypted representation of the first parameter value includes: generating a first private key for each of the L bit positions of the first parameter value; generating a first public key based on the first private key, and receiving a second public key sent by each of the second node devices, where the second public key is generated by the second node device based on a second private key; encrypting the first parameter value based on the first public key and the second public key to obtain a first encrypted representation of the first parameter value.

14. A node device, characterized in that, Comprising: Comprising a processor, the processor being coupled to at least one memory, the processor being configured to read a computer program stored in the at least one memory, such that the node device executes the method according to any one of claims 1 to 11, or such that the node device executes the method according to claim 12, or such that the node device executes the method according to claim 13.

15. A computer-readable medium, characterized in that, The computer-readable storage medium is used to store a computer program, which, when run on a computer, causes the computer to execute the method according to any one of claims 1 to 11, or causes the computer to execute the method according to claim 12, or causes the computer to execute the method according to claim 13.

Citation Information

Patent Citations

  • Model training method based on federated learning

    CN110955907A

  • Federated learning system and method based on block chain

    CN111212110A