A method, device, storage medium, and program product for joint training of a model

By conducting joint training of the model between the passive and active parties of vertical federated learning and using homomorphic encryption technology to process the gradient, the problems of low gradient computing efficiency and communication time in vertical federated learning are solved, and more efficient model training is achieved.

CN112818374BActive Publication Date: 2025-06-10WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110230932.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-02
Publication Date
2025-06-10
Estimated Expiration
2041-03-02

AI Technical Summary

Technical Problem

In vertical federated learning, gradient calculation can only be completed by the active or passive party, resulting in low model training efficiency, and a large amount of data is required for each gradient calculation, resulting in a long communication time, especially in scenarios with large data volume and multiple participants.

Method used

By conducting joint training of the model between the passive and the active party, the encryption gradient is generated using homomorphic encryption technology, and the encryption gradient is sent to the coordinator for decryption and update when the synchronization conditions are met, reducing the data transmission amount and communication time.

Benefits of technology

It shortens the training time of the model, improves the training efficiency of the model, reduces the transmitted data and communication time, and is suitable for scenarios with large data volume and multiple participants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112818374B_ABST
    Figure CN112818374B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device, computer-readable storage medium and program product for joint training of a model, which is applied to the passive party in vertical federated learning. The passive party and the active party in vertical federated learning respectively use their own feature data for model training. The method includes: obtaining the second ciphertext training result sent by the active party; obtaining the first ciphertext training result and the number of trained rounds; determining the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result; when the number of trained rounds meets the synchronization condition, sending the first encrypted gradient to the coordinator so that the coordinator determines the first decrypted gradient based on the first encrypted gradient; receiving the first decrypted gradient sent by the coordinator, and updating its own training model based on the first decrypted gradient to obtain an updated training model. By improving the interaction process of data during model training in vertical federated learning, the training time of the model can be shortened and the training efficiency of the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and relates to, but is not limited to, a method, device, storage medium and program product for jointly training a model. Background Art

[0002] In recent years, regression models have been widely used to solve various problems. When the training data of the regression model is vertically distributed among various data parties, since the feature data owned by each data party may involve privacy, in order to avoid the leakage of private data, each data party can perform joint training of the model through vertical federated learning.

[0003] However, in the related art when performing vertical federated modeling, the gradient calculation can only be completed by one of the active party or the passive party, resulting in a low training efficiency of the model; and for each calculated gradient, it needs to be sent to the coordinator, resulting in a large amount of transmitted data and a long communication time. Especially in the scenarios of large amounts of data and multiple participants, the communication time even exceeds the calculation time, seriously affecting the training efficiency of the model. Summary of the Invention

[0004] Embodiments of the present application provide a method, device, equipment, computer-readable storage medium and computer program product for jointly training a model, which can shorten the training time of the model and improve the training efficiency of the model.

[0005] The technical solution of the embodiments of the present application is implemented as follows:

[0006] Embodiments of the present application provide a method for jointly training a model, which is applied to the passive party of vertical federated learning. The passive party and the active party of vertical federated learning respectively use their own feature data for model training. The method includes:

[0007] Obtain the second ciphertext training result sent by the active party;

[0008] Obtain the first ciphertext training result and the number of trained rounds;

[0009] Determine the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result;

[0010] When the number of trained rounds meets the synchronization condition, send the first encrypted gradient to the coordinator, so that the coordinator determines the first decrypted gradient based on the first encrypted gradient;

[0011] Receive the first decrypted gradient sent by the coordinator, and update the own training model based on the first decrypted gradient to obtain an updated training model.

[0012] An embodiment of the present application provides a method for joint training of a model, which is applied to a coordinator in vertical federated learning. The method includes:

[0013] Generate a public key for encryption and a private key for decryption;

[0014] Send the public key to a passive party and an active party for model training respectively, so that the passive party and the active party determine a first encrypted gradient and a second encrypted gradient respectively based on the public key;

[0015] Receive the first encrypted gradient sent by the passive party and the second encrypted gradient sent by the active party;

[0016] Decrypt the first encrypted gradient and the second encrypted gradient respectively based on the private key to obtain a first decrypted gradient and a second decrypted gradient;

[0017] Send the first decrypted gradient and the second decrypted gradient to the passive party and the active party respectively, so that the passive party and the active party update their respective training models based on the first decrypted gradient and the second decrypted gradient respectively.

[0018] An embodiment of the present application provides a joint training device for a model, which is applied to a passive party in vertical federated learning. The passive party and the active party in vertical federated learning use their own feature data for model training respectively. The device includes:

[0019] A first acquisition module, configured to acquire a second ciphertext training result sent by the active party;

[0020] A second acquisition module, configured to acquire a first ciphertext training result and the number of trained rounds;

[0021] A first determination module, configured to determine a first encrypted gradient based on the first ciphertext training result and the second ciphertext training result;

[0022] A first sending module, configured to send the first encrypted gradient to the coordinator when the number of trained rounds meets the synchronization condition, so that the coordinator determines a first decrypted gradient based on the first encrypted gradient;

[0023] A first receiving module, configured to receive the first decrypted gradient sent by the coordinator;

[0024] A first update module, configured to update its own training model based on the first decrypted gradient to obtain an updated training model.

[0025] An embodiment of the present application provides a joint training device for a model, which is applied to a coordinator in vertical federated learning. The device includes:

[0026] A generation module, configured to generate a public key for encryption and a private key for decryption;

[0027] A second sending module, configured to send the public key to a passive party and an active party for model training respectively, so that the passive party and the active party respectively determine a first encrypted gradient and a second encrypted gradient based on the public key;

[0028] A second receiving module, configured to receive the first encrypted gradient sent by the passive party and the second encrypted gradient sent by the active party;

[0029] A decryption module, configured to decrypt the first encrypted gradient and the second encrypted gradient respectively based on the private key to obtain a first decrypted gradient and a second decrypted gradient;

[0030] A third sending module, configured to send the first decrypted gradient and the second decrypted gradient to the passive party and the active party respectively, so that the passive party and the active party respectively update their training models based on the first decrypted gradient and the second decrypted gradient.

[0031] An embodiment of the present application provides a device for joint training of a model, and the device includes:

[0032] A memory, configured to store executable instructions;

[0033] A processor, configured to implement the method provided by the embodiment of the present application when executing the executable instructions stored in the memory.

[0034] An embodiment of the present application provides a computer-readable storage medium, on which executable instructions are stored, and when executed by a processor, the method provided by the embodiment of the present application is implemented.

[0035] An embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method provided by the embodiment of the present application is implemented.

[0036] The embodiment of the present application has the following beneficial effects:

[0037] The joint training method of the model provided by the embodiment of the present application is applied to the passive party of vertical federated learning. The passive party and the active party of vertical federated learning respectively use their own feature data for model training. In this way, the original single active party training model is improved to be that the passive party and the active party respectively use their own feature data for model training, which can shorten the training time of the model and improve the training efficiency of the model. When the passive party conducts model training, it first obtains the second ciphertext training result sent by the active party, then obtains the first ciphertext training result and the number of trained rounds, determines the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result, and when the number of trained rounds meets the synchronization condition, sends the determined first encrypted gradient to the coordinator of vertical federated learning, so that the coordinator decrypts the first encrypted gradient to obtain the first decrypted gradient, and then the coordinator sends the decrypted first decrypted gradient to the passive party; the passive party updates its own training model based on the first decrypted gradient, thereby obtaining the updated training model and completing one synchronous update of the model. In this way, the first encrypted gradient will be sent to the coordinator for one synchronous update only when the number of trained rounds meets the synchronization condition, and the first encrypted gradient will not be sent to the coordinator when the synchronization condition is not met, and the passive party continues to use its own feature data for model training, so as to reduce the amount of data transmitted, reduce communication time, further shorten the training time of the model, and improve the training efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic diagram of a network architecture of the joint training method of the model provided by the embodiment of the present application;

[0039] Figure 2 It is a schematic diagram of the composition structure of the joint training device of the model provided by the embodiment of the present application;

[0040] Figure 3 It is a schematic diagram of an implementation process of the joint training method of the model provided by the embodiment of the present application;

[0041] Figure 4 It is another schematic diagram of an implementation process of the joint training method of the model provided by the embodiment of the present application;

[0042] Figure 5 It is still another schematic diagram of an implementation process of the joint training method of the model provided by the embodiment of the present application;

[0043] Figure 6 It is yet another schematic diagram of an implementation process of the joint training method of the model provided by the embodiment of the present application;

[0044] Figure 7 It is still another schematic diagram of an implementation process of the joint training method of the model provided by the embodiment of the present application;

[0045] Figure 8A Schematic diagram of the network architecture for linear regression interaction in the longitudinal model in the related art;

[0046] Figure 8B Schematic diagram of the process for linear regression interaction in the longitudinal model in the related art;

[0047] Figure 9A Schematic diagram of the network architecture for linear regression interaction in the longitudinal model provided by the embodiments of the present application;

[0048] Figure 9B Schematic diagram of the process for linear regression interaction in the longitudinal model provided by the embodiments of the present application;

[0049] Figure 10 Schematic diagram of the process for logistic regression interaction in the longitudinal model provided by the embodiments of the present application. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0051] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0052] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0054] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0055] 1) Vertical Federated Learning. When there is a large overlap in users between two datasets but a small overlap in user features, on the premise of ensuring information security during big data exchange, protecting terminal data and personal data privacy, and ensuring legality and compliance, the dataset is sliced vertically (i.e., by feature dimension), and a part of the data that is the same for both parties' users but the user features are not completely the same is taken for training in machine learning.

[0056] 2) Homomorphic Encryption. Homomorphic encryption is a cryptographic technique based on the computational complexity theory of mathematical problems. Processing the homomorphically encrypted data gives an output, and decrypting this output results in the same result as processing the unencrypted original data using the same method.

[0057] 3) Logistic Regression. A machine learning method used to solve binary classification (0 or 1) problems and estimate the probability of something.

[0058] 4) Linear Regression. A regression analysis that models the relationship between one or more independent variables and a dependent variable using a least squares function called a linear regression equation. Both Logistic Regression and Linear Regression are a type of generalized linear model. Logistic Regression assumes that the dependent variable follows a Bernoulli distribution, while Linear Regression assumes that the dependent variable follows a Gaussian distribution.

[0059] The following describes the exemplary application of the device implementing the embodiments of the present application. The device provided by the embodiments of the present application can be implemented as a joint training device for a model. Below, the exemplary applications covered when the device is implemented as a joint training device for a model will be described.

[0060] Figure 1 A schematic diagram of a network architecture for a joint training method of a model provided by an embodiment of the present application, as Figure 1 shown. In this network architecture, it includes a joint training device for a model and Network 400. The joint training device for a model includes a passive party 100, an active party 200, and a coordinator 300. To support an exemplary application, the passive party 100, active party 200, and coordinator 300 included in the joint training device for a model are devices capable of joint training and supporting interaction, which can be servers or devices such as desktop computers, laptop computers, mobile phones, and tablet computers. The passive party 100, active party 200, and coordinator 300 are interconnected through Network 300. Network 300 can be a wide area network, a local area network, or a combination of both, and uses wireless or wired links to achieve data transmission.

[0061] In the embodiments of the present application, the passive party 100 is a data provider, the active party 200 is a data provider with labeled data, and the coordinator 300 is a third party for joint training. The passive party 100 and the active party 200 need to perform vertical federated model training without disclosing the labeled data of the active party 200 and the feature data of both parties. When performing joint training of the model, the coordinator 300 generates a public key and a private key, and sends the public key to the passive party 100 and the active party 200. The passive party 100 and the active party 200 respectively use their own feature data for model training. When the number of training rounds of each participating party meets the corresponding synchronization condition, a gradient synchronization is performed. The following takes the model joint training process of the passive party 100 as an example for illustration.

[0062] The passive party 100 first obtains the second ciphertext training result from the active party 200, and the second ciphertext training result is obtained by the active party 200 based on its own feature data and the public key. Then the passive party 100 obtains the first ciphertext training result based on its own feature data and the public key, and obtains the number of trained rounds. Then the passive party 100 determines the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result. When it is determined that the number of trained rounds meets the synchronization condition, the passive party 100 sends the first encrypted gradient to the coordinator 300. The coordinator 300 performs homomorphic decryption on the first encrypted gradient based on the private key to obtain the first decrypted gradient, and sends the first decrypted gradient to the passive party 100. After receiving the first decrypted gradient, the passive party 100 updates its training model according to the first decrypted gradient to obtain the updated training model, thereby completing a synchronous update of the model. At the same time, the active party 200 also uses its own feature data for joint training of the model. In this way, by improving the original single active party training model to enable the passive party and the active party to respectively use their own feature data for model training, the training time of the model is shortened and the training efficiency of the model is improved. Moreover, the first encrypted gradient is sent to the coordinator for a synchronous update only when the number of trained rounds meets the synchronization condition, and the first encrypted gradient is not sent to the coordinator when the synchronization condition is not met. The passive party continues to use its own feature data for model training, so that the amount of data transmitted can be reduced, the communication time can be reduced, and further the training time of the model can be shortened and the training efficiency of the model can be improved.

[0063] The device provided in the embodiments of the present application can be implemented in a hardware or a combination of software and hardware manner. The following describes various exemplary implementations of the device provided in the embodiments of the present application.

[0064] See Figure 2 , Figure 2 which is a schematic structural diagram of the composition of the joint training device of the model provided in the embodiments of the present application. In the embodiments of the present application, the joint training device 10 of the model is shown by taking the passive party 100 device as an example. According to Figure 2The exemplary structure of the joint training device 10 of the shown model is presented. Other exemplary structures of the joint training device 10 can be foreseen. Therefore, the structures described herein should not be regarded as restrictive. For example, some components described below can be omitted, or components not described below can be added to meet the special requirements of certain applications.

[0065] Figure 2 The joint training device 10 of the shown model includes: at least one processor 110, a memory 140, at least one network interface 120, and a user interface 130. Each component in the joint training device 10 of the model is coupled together through a bus system 150. It can be understood that the bus system 150 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 150 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 150.

[0066] The user interface 130 may include a display, a keyboard, a mouse, a touchpad, and a touch screen, etc.

[0067] The memory 140 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory). The volatile memory can be a random access memory (RAM, Random Acces s Memory). The memory 140 described in the embodiments of the present application is intended to include any suitable type of memory.

[0068] The memory 140 in the embodiments of the present application is capable of storing data to support the operation of the joint training device 10 of the model. Examples of these data include: any computer programs for operating on the joint training device 10 of the model, such as an operating system and application programs. Among them, the operating system contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs can include various application programs.

[0069] As an example of the method provided in the embodiments of the present application being implemented in software, the method provided in the embodiments of the present application can be directly embodied as a combination of software modules executed by the processor 110. The software modules can be located in a storage medium. The storage medium is located in the memory 140. The processor 110 reads the executable instructions included in the software modules in the memory 140 and combines the necessary hardware (for example, including the processor 110 and other components connected to the bus 150) to complete the method provided in the embodiments of the present application.

[0070] As an example, the processor 110 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0071] The exemplary application and implementation of the device provided in the embodiments of the present application will be combined to illustrate the joint training method of the model provided in the embodiments of the present application.

[0072] Figure 3 It is a schematic diagram of an implementation process of the joint training method of the model provided in the embodiments of the present application, which is applied to Figure 1 the passive party in the network architecture shown. In the embodiments of the present application, the passive party and the active party of vertical federated learning respectively use their own feature data for model training. The following will be combined with Figure 3 the steps shown to illustrate the method provided in the embodiments of the present application.

[0073] Step S301, obtain the second ciphertext training result sent by the active party.

[0074] In the embodiments of the present application, when performing joint training of the model, the coordinator generates a public key for encryption and a private key for decryption. In one implementation, the public key is used for homomorphic encryption, and the private key is used for homomorphic decryption. The coordinator sends the public key to the passive party and the active party. The passive party and the active party respectively use their own feature data for model training. In the embodiments of the present application, the process of joint training of the model is described with the passive party as the execution subject. The passive party obtains the second ciphertext training result from the active party.

[0075] The active party obtaining the second ciphertext training result can be implemented as follows: The active party obtains its own feature data and training model. After initializing the training model, it obtains the initial parameters of the training model, inputs its own feature data into the training model to obtain the second plaintext training result. For example, represent the initial parameters of the training model as w G , represent the feature data of the active party itself as x G , represent the label data of the active party as y. The second plaintext training result based on the linear regression model can be expressed as w G x G -y. Then, using the public key received from the coordinator, perform homomorphic encryption on the second plaintext training result to obtain the second ciphertext training result. Use [[ ]] to represent the value after homomorphic encryption. Then the second ciphertext training result can be expressed as [[w G x G -y]]. The active party sends the second ciphertext training result [[w G xG -y]] is sent to the passive party.

[0076] In some other embodiments, training can also be performed based on a logistic regression model. In this case, the second plaintext training result based on the logistic regression model can be expressed as Then the second ciphertext training result can be expressed as The active party sends the second ciphertext training result to the passive party.

[0077] Step S302: Obtain the first ciphertext training result and the number of trained rounds.

[0078] The number of trained rounds here is the number of times the passive party has trained its training model. Each time it is trained, the number of trained rounds is incremented by 1.

[0079] The passive party obtains its own feature data and training model. After initializing the training model, it obtains the initial parameters of the training model, and inputs its own feature data into the training model to obtain the first plaintext training result. For example, represent the parameters of the training model as w H , and represent the feature data of the passive party itself as x H , then the first plaintext training result based on the linear regression model can be expressed as w H x H . Then the passive party uses the public key received from the coordinator to perform homomorphic encryption on the first plaintext training result w H x H to obtain the first ciphertext training result. Similarly, use [[ ]] to represent the value after homomorphic encryption. Then the first ciphertext training result can be expressed as [[w H x H .

[0080] It should be noted here that since the active party also performs model training, after the passive party obtains the first ciphertext training result [[w H x H , it also sends this first ciphertext training result [[w H x H to the active party.

[0081] Correspondingly, when training is performed based on a logistic regression model, at this time, the first plaintext training result based on the logistic regression model can be expressed as Then the first ciphertext training result can be expressed as The passive party sends the first ciphertext training result to the active party.

[0082] Since the data sent from the active party to the passive party is encrypted, and the data sent from the passive party to the active party is also encrypted, and neither the active party nor the passive party has the key for homomorphic decryption, it is impossible to obtain the feature data of the other party and the parameter information of the training model of the other party, thus protecting data privacy and avoiding information leakage.

[0083] Step S303: Determine a first encrypted gradient based on the first encrypted training result and the second encrypted training result.

[0084] The passive party calculates the encrypted residual value using the first encrypted training result and the second encrypted training result. Since both the first encrypted training result and the second encrypted training result are encrypted, the calculated encrypted residual value is also a ciphertext, denoted as [[di]].

[0085] In the embodiment of the present application, when calculating the encrypted residual value [[di]], it can be calculated based on a linear model, and the encrypted residual value [[di]] can be determined according to the sum of the first encrypted training result and the second encrypted training result. For example, the encrypted residual value [[di]] calculated by a linear regression model can be expressed as [[w H x H + [[w G x G - y]], and the encrypted residual value [[di]] calculated by a logistic regression model can be expressed as

[0086] After obtaining the encrypted residual value [[di]], the passive party uses its own feature data x and the encrypted residual value [[di]] to calculate the first encrypted gradient, denoted as [[di]]x H .

[0087] Here, the second encrypted gradient determined by the active party based on the first encrypted training result and the second encrypted training result is denoted as [[di]]x G .

[0088] In the embodiment of the present application, since the passive party and the active party respectively use their own feature data for model training, determine their respective encrypted residual values, and thus determine their respective encrypted gradients, compared with the method in the related art where only a single participating party determines the encrypted residual value and then determines the encrypted gradient, the asynchronous training of multiple parties including the active party and the passive party can shorten the training time of the model and improve the training efficiency of the model.

[0089] Step S304: When the number of trained rounds meets the synchronization condition, send the first encrypted gradient to the coordinator, so that the coordinator determines a first decrypted gradient based on the first encrypted gradient.

[0090] The passive party determines whether the synchronization condition is met based on the number of trained rounds obtained in step S302. When the synchronization condition is met, it is determined that gradient synchronization needs to be performed in this round. When performing gradient synchronization, the passive party calculates the first encrypted gradient [[di]]x H and sends it to the coordinator. When the synchronization condition is not met, the passive party updates its own training model according to the first encrypted gradient to obtain an updated training model, and then returns to step S302 to continue asynchronous training locally.

[0091] Similarly, when the number of trained rounds of the active party meets its own synchronization condition, the active party calculates the second encrypted gradient [[di]]x G and sends it to the coordinator.

[0092] Since in the embodiments of the present application, the passive party and the active party perform a synchronous update only after meeting their respective synchronization conditions, and do not send the first encrypted gradient to the coordinator when the synchronization condition is not met, the passive party and the active party continue to asynchronously update their respective models using their own feature data, thereby reducing the amount of data transmitted, reducing communication time consumption, and then shortening the training time of the model and improving the training efficiency of the model.

[0093] Step S305: Receive the first decrypted gradient sent by the coordinator.

[0094] After receiving the first encrypted gradient [[di]]x H and the second encrypted gradient [[di]]x G , the coordinator uses the pre-generated private key for decryption to decrypt the first encrypted gradient [[di]]x H and the second encrypted gradient [[di]]x G to obtain the first decrypted gradient and the second decrypted gradient, and then sends the first decrypted gradient and the second decrypted gradient to the passive party and the active party respectively. The decryption here can be homomorphic decryption.

[0095] Step S306: Update the own training model based on the first decrypted gradient to obtain an updated training model.

[0096] After receiving the first decrypted gradient, the passive party updates the parameter w H of its own training model, thereby realizing the update of its own training model to obtain an updated training model.

[0097] On the other side, after receiving the second decrypted gradient sent by the coordinator, the active party updates the parameter w GUpdate it to update its own training model and obtain the updated training model. Thus, the passive party and the active party complete a synchronous training. Repeat the above training process until a trained model is obtained.

[0098] The method for jointly training a model provided by an embodiment of this application is applied to the passive party in vertical federated learning. The passive party and the active party in vertical federated learning respectively use their own feature data for model training. The method includes: obtaining the second ciphertext training result sent by the active party; obtaining the first ciphertext training result and the number of trained rounds; determining the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result; when the number of trained rounds meets the synchronization condition, sending the first encrypted gradient to the coordinator so that the coordinator determines the first decrypted gradient based on the first encrypted gradient; receiving the first decrypted gradient sent by the coordinator and updating its own training model based on the first decrypted gradient to obtain the updated training model. In this way, the original single active-party training model is improved to be that the passive party and the active party respectively use their own feature data for model training, which can shorten the training time of the model and improve the training efficiency of the model; and the first encrypted gradient will only be sent to the coordinator for a synchronous update when the number of trained rounds meets the synchronization condition, which can reduce the amount of data transmitted, shorten the communication time, and can further shorten the training time of the model and improve the training efficiency of the model.

[0099] Based on the foregoing embodiments, another embodiment of this application provides a method for jointly training a model. Figure 4 It is another schematic diagram of the implementation process of the method for jointly training a model provided by an embodiment of this application, and is applied to Figure 1 the passive party in the network architecture shown. In the embodiment of this application, the passive party and the active party in vertical federated learning respectively use their own feature data for model training. As Figure 4 shown, the method for jointly training the model includes the following steps:

[0100] Step S401, obtain the second ciphertext training result sent by the active party.

[0101] In the embodiment of this application, steps S401 to S403, and steps S405 to S407 respectively correspond to steps S301 to S306 in the Figure 3 shown embodiment one by one. For the implementation manners and effects of steps S401 to S403 and steps S405 to S407, please refer to the descriptions of steps S301 to S306.

[0102] Step S402, obtain the first ciphertext training result and the number of trained rounds.

[0103] Step S403: Determine the first encrypted gradient based on the first encrypted text training result and the second encrypted text training result.

[0104] After each round of training, the passive party and the active party perform a gradient synchronization once. Although this can shorten the training time of the model and improve the training efficiency compared to the single-party training method where only the active party or only the passive party trains the model, due to differences in the computing power and the size of the sample data of the passive party and the active party, generally they will not complete the training at the same time. Moreover, performing a gradient synchronization every time a round of training is completed causes the active party, the passive party, and the coordinator to interact back and forth multiple times, and a large amount of communication data increases the communication burden and greatly increases the communication time.

[0105] In the embodiments of the present application, the passive party and the active party can first perform asynchronous training locally multiple times respectively and then perform a gradient synchronization once, so that the number of interactions is reduced, thereby reducing the communication burden and communication time, further shortening the training time of the model, and improving the training efficiency of the model. In the embodiments of the present application, the following steps are used to achieve performing a gradient synchronization once after each party has trained multiple times respectively.

[0106] Step S404: Determine whether the number of trained rounds satisfies the synchronization condition.

[0107] When the number of trained rounds satisfies the synchronization condition, it indicates that this round of training requires the coordinator to synchronize the encrypted gradients of both the passive party and the active party once, and at this time, step S405 is entered; when the number of trained rounds does not satisfy the synchronization condition, it indicates that this round of training does not require synchronization of the encrypted gradients, and at this time, step S408 is entered.

[0108] One implementation method for determining whether the number of trained rounds satisfies the synchronization condition is: determining whether the number of trained rounds satisfies the synchronization condition based on a first preset threshold; when the number of trained rounds can be divided evenly by the first preset threshold, it is determined that the number of trained rounds satisfies the synchronization condition; when the number of trained rounds cannot be divided evenly by the first preset threshold, it is determined that the number of trained rounds does not satisfy the synchronization condition.

[0109] The passive party pre-sets a first preset threshold, and performs a gradient synchronization once when the number of rounds trained by the passive party (different from the number of trained rounds obtained in step S402) reaches the first preset threshold after the last gradient synchronization. Here, the first preset threshold can be any integer value, for example, set to 10. The passive party conducts training, and after every 10 rounds of training, synchronizes the first encrypted gradient with the second encrypted gradient obtained by the active party.

[0110] It should be noted that, like the passive party, the active party can also train multiple rounds by itself. When the number of trained rounds is divisible by the third preset threshold, gradient synchronization is performed. Therefore, the second encrypted gradient sent by the active party to the coordinator for gradient synchronization can be the second encrypted gradient obtained when the number of rounds trained by the active party after the last gradient synchronization (different from the number of trained rounds used to divide the third preset threshold) reaches the third preset threshold. The third preset threshold here can also take any integer value.

[0111] In some embodiments, the values of the first preset threshold and the third preset threshold can be determined based on the duration of one round of training for the passive party and the active party respectively, ensuring that the time difference between the first time when the passive party sends the first encrypted gradient to the coordinator and the second time when the active party sends the second encrypted gradient to the coordinator is less than the preset duration, so as to avoid either side waiting too long. For example, if the passive party takes 10 s (seconds) for one round of training and the active party takes 15 s for one round of training, then any common multiple of 10 and 15 can be used to set the first preset threshold and the third preset threshold. For example, based on the common multiple 150, the first preset threshold is set to 15 and the third preset threshold is set to 10. After the passive party and the active party train by themselves for 150 s, that is, after the passive party trains by itself for 15 rounds, it sends the first encrypted gradient to the coordinator, and after the active party trains by itself for 10 rounds, it sends the second encrypted gradient to the coordinator for gradient synchronization.

[0112] Step S405: Send the first encrypted gradient to the coordinator so that the coordinator determines the first decrypted gradient based on the first encrypted gradient.

[0113] Step S406: Receive the first decrypted gradient sent by the coordinator.

[0114] In some embodiments, the coordinator can optimize the first decrypted gradient and then send the optimized first decrypted gradient to the passive party, which can improve the calculation accuracy, enable the training model to converge as soon as possible, and further shorten the training time.

[0115] Step S407: Update the own training model based on the first decrypted gradient to obtain an updated training model.

[0116] So far, the passive party and the active party complete one round of synchronous training, return to step S401, and repeat the above training process until a trained model is obtained.

[0117] Step S408: Optimize the first encrypted gradient based on a preset step size to obtain an optimized first encrypted gradient.

[0118] Here, the preset step size is the increment for adjusting the first encrypted gradient after each round of training preset.

[0119] When optimizing the first encrypted gradient based on a preset step size, one implementation is to add the preset step size to the first encrypted gradient to obtain the optimized first encrypted gradient. Another implementation is to multiply the preset step size by the first encrypted gradient to obtain the optimized first encrypted gradient. Of course, there can also be other optimization methods, which are not limited in the embodiments of this application.

[0120] Step S409: Update its own training model based on the optimized first encrypted gradient to obtain an updated training model.

[0121] After obtaining the optimized first encrypted gradient, the passive party updates the parameters w of its own training model H to achieve local update of its own training model and obtain an updated training model.

[0122] At this point, the passive party completes an asynchronous update, returns to step S402, and continues to train the updated training model locally until the number of trained rounds meets the synchronization condition.

[0123] The method for jointly training the model provided by the embodiments of this application is applied to the passive party in vertical federated learning. The passive party and the active party in vertical federated learning respectively use their own feature data for model training, improving the original single active party training model to be that the passive party and the active party respectively use their own feature data for model training, which can shorten the training time of the model and improve the training efficiency of the model. And in the method provided by the embodiments of this application, the passive party and the active party in vertical federated learning first train locally for multiple rounds respectively. When the number of trained rounds meets the synchronization condition, they perform a gradient synchronization of their respective encrypted gradients. When the number of trained rounds does not meet the synchronization condition, they continue to train the model locally, so as to reduce the number of interactions between each participating party and the coordinator, reduce the amount of data transmitted, reduce the communication burden, and reduce the communication time, which can further shorten the training time of the model and improve the training efficiency of the model.

[0124] In some embodiments, the above Figure 3 "obtaining the first ciphertext training result" in step S302 or Figure 4 step S402 in the illustrated embodiment can be implemented as the following steps:

[0125] Step S3021: Receive the public key sent by the coordinator.

[0126] When jointly training the model, the coordinator generates a public key for homomorphic encryption and a private key for homomorphic decryption, and then sends the public key to the passive party and the active party, enabling the passive party and the active party to jointly train the model using the public key while ensuring the privacy of their respective data.

[0127] Step S3022: Obtain the sample data for joint training and its own training model.

[0128] Since the active party and the passive party conduct joint training, it is required that their respective sample data come from the same samples, that is, the samples of the active party's sample data and the passive party's sample data are unified. That is to say, the identifiers (denoted as id) of the sample data are the same. Based on this, the active party and the passive party encrypt the ids of their respective feature data in advance to obtain the encrypted id of the active party and the encrypted id of the passive party, and take the intersection of the encrypted id of the active party and the encrypted id of the passive party to complete the common sample screening.

[0129] In one implementation, obtaining the sample data for joint training can be implemented as: obtaining the identifier of the current round of training samples from the active party; based on the identifier, screening out the data corresponding to the identifier from its own feature data; and determining the data corresponding to the identifier as the sample data for joint training.

[0130] During each synchronous training, the active party randomly selects a part of the samples from the common samples for the current synchronous training. The active party sends the identifiers (i.e., ids) of the randomly selected samples to the active party training module and the passive party training module, so that they select the corresponding data from their own data tables based on the ids as the sample data for the current round of joint training.

[0131] Step S3023: Input the sample data into the training model for training to obtain the first training result.

[0132] In the first round of training, after the passive party obtains its own training model, it initializes it to obtain the parameter w of the training model H , and inputs the sample data x H into the initialized training model to obtain the first training result. Here, the first training result is in plaintext, that is, the first plaintext training result w H x H .

[0133] Step S3024: Encrypt the first training result based on the public key to obtain the first ciphertext training result.

[0134] In the embodiments of the present application, the first plaintext training result w H x HPerform encryption, such as any one of additive homomorphic encryption, multiplicative homomorphic encryption, hybrid multiplicative homomorphic encryption, subtractive homomorphic encryption, divisive homomorphic encryption, algebraic homomorphic encryption (also known as fully homomorphic encryption), or arithmetic homomorphic encryption. Here, the fully homomorphic encryption means that the encryption function satisfies both additive homomorphism and multiplicative homomorphism. In the embodiments of the present application, additive homomorphic encryption is used to encrypt the first plaintext training result w H x H Performing encryption can improve the operation efficiency compared with other methods such as fully homomorphic encryption.

[0135] In some embodiments, the above Figure 3 In the embodiment shown in step S303 or Figure 4 In the embodiment shown in step S403, "determining the first encryption gradient based on the first ciphertext training result and the second ciphertext training result" can be implemented as the following steps:

[0136] Step S3031: Perform regression analysis on the first ciphertext training result and the second ciphertext training result to obtain an encrypted residual value.

[0137] Here, the passive party can perform linear regression analysis on the first ciphertext training result [[w H x H and the second ciphertext training result [[w G x G -y]], or can perform logistic regression analysis on the first ciphertext training result [[w H x H and the second ciphertext training result [[w G x G -y]]. Since both the first ciphertext training result [[w H x H and the second ciphertext training result [[w G x G -y]] are encrypted, the calculated encrypted residual value is also a ciphertext, denoted as [[di]].

[0138] When calculating the encrypted residual value [[di]] using a linear regression model, [[di]] can be expressed as [[w H x H +[[w G x G -y]], and when calculating the encrypted residual value [[di]] using a logistic regression model, [[di]] can be expressed as

[0139] Step S3032: Determine the first encryption gradient based on the data corresponding to the identifier and the encrypted residual value.

[0140] After obtaining the encrypted residual value [[di]], the passive party uses the sample data x H and the encrypted residual value [[di]] to calculate the first encrypted gradient, denoted as [[di]]x H . Then, the first encrypted gradient is sent to the coordinator for multi-party synchronization.

[0141] In the embodiment of the present application, the passive party determines the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result. When the number of trained rounds does not meet the synchronization condition and asynchronous training is performed, the local training model is updated. The first ciphertext training results obtained each time are different, that is, when asynchronous training is performed, the first ciphertext training result changes in each round of training, and the second ciphertext training result does not change; when the number of trained rounds meets the synchronization condition and synchronous training is performed, the passive party re-obtains the second ciphertext training result once, and both the first ciphertext training result and the second ciphertext training result change in this round of training. In this way, only after each synchronization, the active party and the passive party will send their respective ciphertext training results to each other, thereby reducing the amount of data transmitted, shortening the communication time, further shortening the training time of the model, and improving the training efficiency of the model.

[0142] Based on the foregoing embodiments, the embodiment of the present application further provides a method for jointly training a model, Figure 5 which is another schematic diagram of the implementation process of the method for jointly training a model provided by the embodiment of the present application, and is applied to Figure 1 the passive party in the network architecture shown. In the embodiment of the present application, the passive party and the active party in vertical federated learning respectively use their own feature data for model training. As Figure 5 shown, the method for jointly training the model includes the following steps:

[0143] Step S501, obtain the second ciphertext training result sent by the active party.

[0144] In the embodiment of the present application, steps S501 to S505 and step S509 respectively correspond to Figure 3 steps S301 to S305 in the embodiment shown. For the implementation process and effects of steps S501 to S505 and step S508, please refer to the descriptions of steps S301 to S305.

[0145] Step S502, obtain the first ciphertext training result and the number of trained rounds.

[0146] Step S503, determine the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result.

[0147] Step S504, when the number of trained rounds meets the synchronization condition, send the first encrypted gradient to the coordinator so that the coordinator determines the first decrypted gradient based on the first encrypted gradient.

[0148] Step S505, receive the first decrypted gradient sent by the coordinator.

[0149] Since it is a joint training of multiple participants, the passive party alone cannot determine whether the training is completed. Based on this, the coordinator determines whether the model converges based on the first encrypted gradient, the second encrypted gradient, and the preset gradient norm threshold sent by the passive party and the active party. Or, the coordinator can also determine whether the model converges based on the number of trained rounds of the passive party and the active party and the preset number of rounds threshold. Or, the coordinator can also determine whether the model converges based on the first loss value and the second loss value sent by the passive party and the active party.

[0150] Step S506, receive the convergence information sent by the coordinator.

[0151] Here, the convergence information is determined by the coordinator based on the first decrypted gradient and the preset gradient norm threshold; or, the convergence information is determined by the coordinator based on the number of trained rounds and the preset number of rounds threshold; or, the convergence information is determined by the coordinator based on the first loss value sent by itself and the second loss value sent by the active party.

[0152] In some embodiments, when the convergence information is determined by the coordinator based on the first loss value sent by itself and the second loss value sent by the active party, before step S506, the method further includes:

[0153] Step S51, determine the first loss value of this round of training based on the public key and the first training result.

[0154] Here, the passive party uses the public key to homomorphically encrypt the first training result to obtain the first loss value of this round of training, denoted as

[0155] Similarly, the active party also uses the public key to homomorphically encrypt the second training result to obtain the second loss value of this round of training, denoted as The second training result here is obtained by the active party receiving the public key sent by the coordinator, obtaining the sample data for joint training and its own training model, and inputting the sample data into the training model for training.

[0156] Step S52, send the first loss value to the coordinator so that the coordinator determines the loss value of this round of training based on the first loss value and the second loss value sent by the active party, and determines the convergence information of this round of training based on the loss value.

[0157] The coordinator can use the sum of the first loss value and the second loss value as the loss value for this round of training, that is Determine the convergence information for this round of training according to L. For the process of the coordinator determining the convergence information, refer to the description in step S606 below.

[0158] Step S507, determine whether the convergence information indicates convergence or the training is completed.

[0159] When the convergence information indicates convergence or the training is completed, it means that the passive party does not need to continue training, and at this time, enter step S508; when the convergence information does not indicate convergence and does not indicate that the training is completed, it means that the training is not completed, and at this time, enter step S509 to update the model.

[0160] Step S508, determine the updated training model as the trained target model.

[0161] When it is determined that there is no need to continue training, the obtained training model is the trained target model. Thus, the joint training process of the model is completed, and the obtained target model is applicable to both the passive party and the active party.

[0162] Step S509, update its own training model based on the first decryption gradient to obtain the updated training model.

[0163] After the update is completed, return to step S501 to continue the next round of training until the trained target model is obtained.

[0164] In the method for jointly training a model provided in the embodiments of the present application, the passive party determines whether to continue training according to the convergence information sent by the coordinator. When the convergence information indicates convergence or the training is completed, the trained target model is determined, thereby completing the joint training of the model and obtaining a model applicable to both the passive party and the active party.

[0165] Based on the foregoing embodiments, the embodiments of the present application further provide a method for jointly training a model, Figure 6 which is another schematic implementation flow diagram of the method for jointly training a model provided in the embodiments of the present application, and is applied to Figure 1 the coordinator in the network architecture shown. As Figure 6 shown, the method for jointly training a model includes the following steps:

[0166] Step S601, generate a public key for encryption and a private key for decryption.

[0167] In the embodiments of the present application, the public key generated by the coordinator can be the public key for homomorphic encryption, and the generated private key can be the private key for homomorphic decryption. Through the public key, each participating party does not need to send its own private data to the other party or the coordinator, which can protect the privacy of the data of each participating party.

[0168] Homomorphic encryption is a cryptographic technology based on the computational complexity theory of mathematical problems. Processing the data encrypted by homomorphic encryption obtains an output, and decrypting this output results in the same result as processing the unencrypted original data using the same method. The generated public key can be used for any one of additive homomorphic encryption, multiplicative homomorphic encryption, hybrid multiplicative homomorphic encryption, subtractive homomorphic encryption, divisive homomorphic encryption, algebraic homomorphic encryption (also known as fully homomorphic encryption), and arithmetic homomorphic encryption. Here, the fully homomorphic encryption means that the encryption function satisfies both additive homomorphism and multiplicative homomorphism.

[0169] In some embodiments, the public key generated by the coordinator can be used for additive homomorphic encryption, and the generated private key can be used for additive homomorphic decryption, or the public key generated by the coordinator can be used for multiplicative homomorphic encryption, and the generated private key can be used for multiplicative homomorphic decryption, or the public key generated by the coordinator can be used for fully homomorphic encryption, and the generated private key can be used for fully homomorphic decryption. Compared with using fully homomorphic encryption and decryption, using additive homomorphic encryption and decryption during encryption and decryption can improve the operation efficiency.

[0170] Step S602: Send the public key to the passive party and the active party for model training respectively, so that the passive party and the active party respectively determine the first encrypted gradient and the second encrypted gradient based on the public key.

[0171] Here, the active party and the passive party are different participating parties for joint model training. After receiving the public key, the passive party and the active party respectively use their own feature data for model training.

[0172] Step S603: Receive the first encrypted gradient sent by the passive party and the second encrypted gradient sent by the active party.

[0173] Here, the first encrypted gradient and the second encrypted gradient can be the encrypted gradients respectively sent by the passive party and the active party to the coordinator after one round of training. At this time, the steps executed by the coordinator correspond to Figure 3 the steps executed by the passive party in the embodiment shown. The first encrypted gradient and the second encrypted gradient can also be the encrypted gradients respectively sent by the passive party and the active party to the coordinator when their respective numbers of trained rounds have reached the corresponding first preset threshold and third preset threshold after they have trained themselves for multiple rounds. At this time, the steps executed by the coordinator correspond to Figure 4 the steps executed by the passive party in the embodiment shown.

[0174] Step S604: Decrypt the first encrypted gradient and the second encrypted gradient respectively based on the private key to obtain a first decrypted gradient and a second decrypted gradient.

[0175] Step S605: Send the first decrypted gradient and the second decrypted gradient to the passive party and the active party respectively, so that the passive party and the active party update their respective training models based on the first decrypted gradient and the second decrypted gradient respectively.

[0176] The coordinator decrypts the first encrypted gradient and the second encrypted gradient using the private key generated in step S601 to obtain a first decrypted gradient and a second decrypted gradient. Then, the first decrypted gradient is sent to the passive party so that the passive party updates the model trained by itself according to the first decrypted gradient. At the same time, the second decrypted gradient is sent to the active party so that the active party updates the model trained by itself according to the second decrypted gradient.

[0177] In some embodiments, after the coordinator obtains the first decrypted gradient and the second decrypted gradient, the first decrypted gradient and the second decrypted gradient can be optimized respectively and then sent to the corresponding passive party and active party, which can improve the calculation accuracy, make the training model converge as soon as possible, and further shorten the training time.

[0178] One implementation of optimizing the first encrypted gradient and the second encrypted gradient respectively can be: obtaining a first adjustment coefficient and a second adjustment coefficient; optimizing the first encrypted gradient based on the first adjustment coefficient to obtain an optimized first encrypted gradient; optimizing the second encrypted gradient based on the second adjustment coefficient to obtain an optimized second encrypted gradient. Here, the first adjustment coefficient and the second adjustment coefficient are preset coefficients for adjusting the first decrypted gradient and the second decrypted gradient after each round of synchronization. Among them, the first adjustment coefficient and the second adjustment coefficient can be the same or different. When implementing the optimization, the adjustment coefficient (including the first adjustment coefficient and the second adjustment coefficient) can be added to the corresponding decrypted gradient (including the first decrypted gradient and the second decrypted gradient) to obtain the optimized decrypted gradient, or the adjustment coefficient can be multiplied by the decrypted gradient to obtain the optimized decrypted gradient. Of course, there can also be other optimization methods, which are not limited in the embodiments of the present application.

[0179] The joint training method of the model provided by the embodiment of the present application is applied to the coordinator of vertical federated learning. The method includes: generating a public key for encryption and a private key for decryption; sending the public key to the passive party and the active party for model training respectively, so that the passive party and the active party determine the first encrypted gradient and the second encrypted gradient respectively based on the public key; receiving the first encrypted gradient sent by the passive party and the second encrypted gradient sent by the active party; decrypting the first encrypted gradient and the second encrypted gradient respectively based on the private key to obtain the first decrypted gradient and the second decrypted gradient; sending the first decrypted gradient and the second decrypted gradient to the passive party and the active party respectively, so that the passive party and the active party update their respective training models based on the first decrypted gradient and the second decrypted gradient. By improving the interaction process of data during model training in vertical federated learning and performing synchronization once when the synchronization condition is met, the number of interactions between each participating party and the coordinator can be reduced, the amount of data transmitted can be reduced, the communication burden can be alleviated, and the communication time-consuming can be reduced. Therefore, the training time of the model can be shortened and the training efficiency of the model can be improved.

[0180] In some embodiments, after Figure 6 step S605 of the embodiment shown, the coordinator can further determine whether to continue training the model after this round of training. When the coordinator determines that the model has converged or the coordinator determines that the training has been completed, the passive party and the active party can be notified to end the training. Based on this, after the above step S605, the method may further include the following steps:

[0181] Step S606, obtaining convergence information.

[0182] In one implementation, the collaboration party can determine whether the model has converged based on the decrypted gradient. At this time, obtaining the convergence information can be implemented as: based on the received first decrypted gradient and second decrypted gradient, calculating the sum of the decrypted gradients, and determining the difference between the sum of the decrypted gradients and a preset gradient norm threshold; determining whether the difference is less than a second preset threshold; when the difference is less than the second preset threshold, determining the convergence information as converged; when the difference is greater than or equal to the second preset threshold, determining the convergence information as not converged.

[0183] In one implementation, the collaboration party can determine whether the model has converged based on the number of training rounds of the passive party. At this time, obtaining the convergence information can be implemented as: obtaining the number of training rounds of the passive party; determining whether the number of training rounds is greater than a preset number of rounds threshold; when the number of training rounds is greater than the preset number of rounds threshold, determining the convergence information as training completed; when the number of training rounds is less than or equal to the preset number of rounds threshold, determining the convergence information as not completed training.

[0184] In one implementation, the collaborating party can also determine whether the model has converged based on the number of training rounds of the initiating party. In this case, obtaining the convergence information can be implemented as follows: obtaining the number of training rounds of the initiating party; determining whether the number of training rounds is greater than a preset round threshold; when the number of training rounds is greater than the preset round threshold, determining the convergence information as training completed; when the number of training rounds is less than or equal to the preset round threshold, determining the convergence information as training not completed.

[0185] Here, when determining whether the model has converged based on the passive party and the initiating party, the preset round thresholds set can be the same or different.

[0186] Or obtaining the number of training rounds of the passive party; or determining the loss value of this round of training based on the first loss value sent by the passive party and the second loss value sent by the initiating party;

[0187] When the difference is less than a second preset threshold, or when the loss value is less than a preset loss threshold, determining the convergence information as converged; or when the number of training rounds is greater than the preset round threshold, determining the convergence information as training completed.

[0188] In one implementation, the collaborating party can determine whether the model has converged based on the loss value. In this case, obtaining the convergence information can be implemented as follows: receiving the first loss value sent by the passive party and the second loss value sent by the initiating party; adding the first loss value and the second loss value to obtain the loss value of this round of training; determining whether the loss value is less than a preset loss threshold; when the loss value is less than the preset loss threshold, determining the convergence information as converged; when the loss value is greater than or equal to the preset loss threshold, determining the convergence information as not converged.

[0189] Step S607, sending the convergence information to the passive party and the initiating party.

[0190] The collaborating party sends the convergence information to the passive party and the initiating party to inform the passive party and the initiating party to continue training or end training.

[0191] The various ways of obtaining the convergence information provided in the embodiments of the present application can implement the determination of model convergence in different scenarios, increasing the flexibility of applicable scenarios.

[0192] Based on the foregoing embodiments, the embodiments of the present application further provide a method for jointly training a model, Figure 7 which is another schematic flowchart of the implementation process of the method for jointly training a model provided in the embodiments of the present application, and is applied to Figure 1 the network architecture shown, as Figure 7 shown, the method for jointly training a model includes the following steps:

[0193] Step S701, the coordinator generates a public key for encryption and a private key for decryption.

[0194] In the embodiments of the present application, the coordinator can generate a public key for homomorphic encryption, so that each participating party does not need to send its own private data to the other party or the coordinator, and can protect the privacy of the data of each participating party.

[0195] Step S702, the coordinator sends the public key to the passive party and the active party for model training respectively.

[0196] Step S703, the active party obtains the common sample identifiers of itself and the passive party based on vertical federated learning.

[0197] Here, the identifier can be the id value of the sample.

[0198] Step S704, the active party obtains the sample size of this round of training.

[0199] Here, the sample size can be a randomly determined value, and the sample size of each round of training can be different.

[0200] Step S705, the active party screens out the corresponding number of identifiers from the common sample identifiers based on the sample size.

[0201] Here, when the active party screens, it can randomly screen among the common sample identifiers, or screen in a preset manner (such as in order).

[0202] Step S706, the active party determines the screened corresponding number of identifiers as the sample identifiers for this round of training.

[0203] Step S707, the active party sends the sample identifiers for this round of training to the passive party.

[0204] Step S708, the active party obtains the second ciphertext training result and the second number of trained rounds.

[0205] The second number of trained rounds here is the number of rounds that the active party has trained itself. Each time the active party trains, the second number of trained rounds increases by 1.

[0206] In the embodiments of the present application, for the active party to obtain the second ciphertext training result, it can be realized as follows: based on the identifier, screen out the data corresponding to the identifier from its own feature data; determine the data corresponding to the identifier as the second sample data for joint training; input the second sample data into its own training model for training to obtain a second training result; encrypt the second training result based on the public key to obtain the second ciphertext training result.

[0207] Step S709, the active party sends the second ciphertext training result to the passive party.

[0208] Step S710, the passive party obtains the first ciphertext training result and the first number of trained rounds.

[0209] The first number of trained rounds here is the number of trained rounds in the above embodiment. Each time the passive party trains, the first number of trained rounds increases by 1. Since the active party and the passive party train separately, the first number of trained rounds and the second number of trained rounds are generally not equal.

[0210] In the embodiment of the present application, the passive party obtaining the first ciphertext training result can be implemented as: based on the identifier, screening out the data corresponding to the identifier from its own feature data; determining the data corresponding to the identifier as the first sample data for joint training; inputting the first sample data into its own training model for training to obtain a first training result; encrypting the first training result based on the public key to obtain the first ciphertext training result.

[0211] Step S711, the passive party sends the first ciphertext training result to the active party.

[0212] Here, the order of step S708 and step S710 is not limited.

[0213] Step S712, the passive party determines a first encrypted gradient based on the first ciphertext training result and the second ciphertext training result.

[0214] Step S713, the passive party determines whether the first number of trained rounds meets the synchronization condition.

[0215] In one implementation, it can be determined whether the first number of trained rounds meets the synchronization condition based on a first preset threshold. When the first number of trained rounds is divisible by the first preset threshold, it is determined that gradient synchronization is required, and at this time, step S716 is entered; when the first number of trained rounds is not divisible by the first preset threshold, step S714 is entered.

[0216] Step S714, the passive party optimizes the first encrypted gradient based on a first preset step size to obtain an optimized first encrypted gradient.

[0217] Here, the first preset step size is the increment for adjusting the first encrypted gradient after each round of training preset.

[0218] Step S715, the passive party updates its own training model based on the optimized first encrypted gradient to obtain an updated training model.

[0219] Here, after step S715 is executed, the process returns to step S710 to continue the next round of training process. Here, when step S712 is executed again in the next round of training, if the updated second ciphertext training result is not received, the second ciphertext training result from the previous round of training is used to determine the first encrypted gradient.

[0220] Step S716, the passive party sends the first encrypted gradient to the coordinator.

[0221] Here, after step S716, the process enters step S722.

[0222] Step S717, the active party determines a second encrypted gradient based on the first ciphertext training result and the second ciphertext training result.

[0223] Step S718, the active party determines whether the second number of trained rounds satisfies the synchronization condition.

[0224] In one implementation, it can be determined whether the second number of trained rounds satisfies the synchronization condition based on a third preset threshold. When the second number of trained rounds is divisible by the third preset threshold, it is determined that gradient synchronization is required, and at this time, the process enters step S721; when the second number of trained rounds is not divisible by the first preset threshold, the process enters step S719.

[0225] Step S719, the active party optimizes the second encrypted gradient based on a second preset step size to obtain an optimized second encrypted gradient.

[0226] Here, the second preset step size is the increment for adjusting the second encrypted gradient after each round of training as preset.

[0227] Step S720, the active party updates its training model based on the optimized second encrypted gradient to obtain an updated training model.

[0228] Here, after step S720 is executed, the process returns to step S708 to continue the next round of training process. Here, when step S717 is executed again in the next round of training, if the updated first ciphertext training result is not received, the first ciphertext training result from the previous round of training is used to determine the second encrypted gradient.

[0229] Step S721, the passive party sends the second encrypted gradient to the coordinator.

[0230] Step S722, the coordinator determines a first decrypted gradient based on the first encrypted gradient and determines a second decrypted gradient based on the second encrypted gradient.

[0231] The coordinator decrypts the first encrypted gradient and the second encrypted gradient respectively using the private key for homomorphic decryption to obtain the first decrypted gradient and the second decrypted gradient.

[0232] Step S723: The coordinator sends the first decrypted gradient to the passive party and the second decrypted gradient to the active party.

[0233] Step S724: The coordinator determines the difference between the sum of the decrypted gradients and the preset gradient norm threshold.

[0234] Here, the sum of the decrypted gradients is the sum of the first decrypted gradient and the second decrypted gradient.

[0235] Step S725: The coordinator determines whether the difference is less than the second preset threshold.

[0236] When the difference is less than the second preset threshold, go to step S726; when the difference is greater than or equal to the second preset threshold, it is determined that convergence has not occurred, and return to step S704 to continue training.

[0237] Step S726: The coordinator determines the convergence information as converged.

[0238] Step S727: The coordinator sends the convergence information to the active party and the passive party.

[0239] After determining convergence, the coordinator notifies the active party and the passive party that there is no need to continue training.

[0240] Step S728: The active party determines the updated training model as the trained target model.

[0241] Step S729: The passive party determines the updated training model as the trained target model.

[0242] In the model joint training method provided by the embodiments of the present application, the passive party and the active party simultaneously use their own feature data for model training. Compared with the process in the related art where only the active party or only the passive party calculates the gradient for model training, the multi-party synchronous training of the active party and the passive party can shorten the training time of the model and improve the training efficiency of the model; and after the passive party and the active party each train multiple rounds separately, then synchronize their encrypted gradients once, and then the cooperation party determines whether the model has converged. When convergence has not occurred, continue training; when convergence has occurred, the active party and the passive party determine the updated training model as the trained target model, which can not only reduce the number of interactions between each participating party and the coordinator, but also reduce the communication burden and communication time consumption, thereby further shortening the training time of the model and improving the training efficiency of the model.

[0243] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0244] In a three - party vertical federated learning scenario, for example, set the Arbiter party (Party A) (corresponding to the coordinator in the above text), the Host party (Party H) (corresponding to the passive party in the above text), and the Guest party (Party G) (corresponding to the active party in the above text). The label provider (i.e., the active party) Party G has data labels, the data provider (i.e., the passive party) Party H has some feature data that is not in Party G's data, and Party A is a third - party as the coordinator. Party H and Party G need to model and predict (including linear models such as logistic regression and linear regression) without disclosing Party G's label information and the feature data of both parties. A scenario that requires vertical federated modeling is: Party G is an insurance sales company. Party G wants to predict the price of auto insurance policies that potential customers are willing to buy. Then the policy price is y, and the participating party Party H may be a certain auto brand. Party G and Party H are not willing to directly interact with each other's data, so vertical federated modeling is required.

[0245] In practical applications, communication latency often becomes one of the biggest efficiency bottlenecks in federated learning modeling. Especially in scenarios with large amounts of data and multiple participating parties, the communication latency may be higher than the computation latency. Taking the above example, it may be that either Party G or Party H has a device configuration that cannot communicate quickly, or the amount of information (sample size and feature data) in each communication is huge, resulting in a large number of communication interactions and low modeling efficiency.

[0246] The following introduces the vertical model regression interaction process in related technologies. Figure 8A It is a schematic diagram of the network architecture for vertical model linear regression interaction in related technologies. Figure 8B It is a schematic diagram of the process for vertical model regression interaction in related technologies. Figure 8A The participating parties included are the three parties A, G, and H. Party H represents the data provider that does not share data with Party G.

[0247] Premise settings: Party H and Party G complete the screening of common samples through the encrypted ID intersection. In the following training sessions, it is default that the same ID value is used each time. Party A and Party H participate in the training simultaneously and interact with Party G.

[0248] Step S801, Party A generates a public key and transmits it to Party H and Party G.

[0249] The public key here is the public key in the above text, and this public key is used for homomorphic encryption.

[0250] Step S802, Party G determines the amount of training data for each time within each round and sends it to Party H and Party G.

[0251] In this round of training, let x G represent the sample features on the G side (i.e., sample data), and x H represent the sample features on the H side.

[0252] Step S803: Parties H and G respectively initialize their local models and calculate local intermediate calculation results.

[0253] Here, after Party H initializes its local model, the parameters of the model on the H side are denoted as w H , then the local intermediate calculation result on the H side calculated according to the parameters of the model on the H side and the sample features is w H x H , and this result is the predicted value of each sample on the H side.

[0254] After Party G initializes its local model, the parameters of the model on the G side are denoted as w G , then the local intermediate calculation result on the G side calculated according to the parameters of the model on the G side and the sample features is w G x G , and this result is the predicted value of each sample on the G side.

[0255] Step S804: Party H encrypts its intermediate calculation result using homomorphic encryption technology to generate an encrypted intermediate calculation result, and sends this encrypted intermediate calculation result to Party G.

[0256] Here, Party H performs homomorphic encryption using the public key. Denote the value after using homomorphic encryption with [[ ]], the encrypted intermediate calculation result of Party H is denoted as [[w H x H , and sends this encrypted intermediate calculation result [[w H x H to Party G.

[0257] Step S805: Party G combines the encrypted intermediate calculation result sent by Party H to calculate the encrypted residual value [[di]], and Party G sends [[di]] to Party H.

[0258] Here, the encrypted intermediate calculation result of Party G is denoted as [[w G x G -y]], Party G uses [[w H x H sent by Party H and its own [[w G x G -y]], the encrypted residual value [[di]] calculated based on linear regression can be expressed as [[w H x H + [[w G x G -y]], where y is the label provided by Party G.

[0259] Since Party G does not have the private key corresponding to the public key for decryption, it cannot decrypt this value, thus avoiding the data leakage of Party H.

[0260] Step S806: Party G and Party H respectively calculate the encrypted local gradients using their own intermediate calculation results and the encrypted residual values [[di]], and send them to Party A.

[0261] Here, the encrypted local gradients are the first encrypted gradient and the second encrypted gradient in the above text. When Party A uses the preset loss threshold to judge convergence or not, Party G also needs to send the encrypted loss value to Party A.

[0262] Calculate the encrypted loss value L of this round of training H For

[0263] Similarly, the intermediate calculation result encrypted by Party G is expressed as [[w G x G , and calculate the encrypted loss value of this round of training.

[0264] Calculate the encrypted loss value L of this round of training H For

[0265] Similarly, calculate the encrypted loss value L of this round of training G For

[0266] The encrypted local gradient calculated by Party G The encrypted local gradient calculated by Party H

[0267] Step S807: Party A uses the private key to decrypt the encrypted local gradient, and performs optimization processing on the decrypted local gradient, and sends the processing results to Party H and Party G respectively. Party A judges whether to converge according to the preset gradient norm threshold or the preset loss threshold, and sends the obtained convergence information to Party H and Party G.

[0268] Here, performing optimization processing on the decrypted local gradient can be multiplying the decrypted local gradient by the update step size.

[0269] The convergence judgment criterion here is: at the end of each round of training, calculate the sum of the gradient norms of all Party G and Party H And compare it with the preset gradient norm threshold. If the sum of the gradient norms is less than the preset gradient norm threshold, it is considered that the model converges; if the sum of the gradient norms is greater than or equal to the preset gradient norm threshold, it is considered that the model does not converge, and enter the next round of training.

[0270] Alternatively, the convergence judgment criterion can also be: calculate the sum of the loss values on both the G side and the H side, and use a preset loss threshold to judge whether to converge. If the sum of the loss values is less than the preset loss threshold, it is considered that the model converges; if the sum of the loss values is greater than or equal to the preset loss threshold, continue the next round of training.

[0271] Step S808, the H and G parties update the local model parameters.

[0272] Repeat steps S803 to S808 until all test data has been used.

[0273] Repeat steps S802 to S808 until the model converges or reaches the maximum number of model training rounds.

[0274] Here, the maximum number of model training rounds is the preset round threshold in the above text.

[0275] In the related art, through steps S801 to S808, the H and G parties train some linear regression model parameters. During the whole process, neither party discloses its own data and model parameter information, and at the same time, Party A cannot know the data information of the H and G parties. In the related art, all gradient calculations are completed by one party under the interaction mechanism of vertical linear model modeling. Due to the limitation of the data interaction mechanism in the vertical regression scenario of the federated learning system design, it is not convenient to realize the asynchronous update of each participating party.

[0276] To solve this problem, the embodiment of the present application proposes an asynchronous update idea, that is, each participating party updates the gradient locally, and synchronizes the gradient changes every n rounds. The training optimization scheme can improve the training efficiency and reduce the overall training time.

[0277] The embodiment of the present application transforms the data interaction process of vertical federated learning linear model modeling in the related art. It changes from the original unified calculation of encrypted gradients by Party G to the calculation of gradients by Party H and Party G respectively for their own parties. On the overall process, the residual sent by Party G to Party H is changed to its own encrypted intermediate calculation result. Therefore, Party H can calculate the residual by itself by combining its own intermediate calculation result, and then calculate the gradient. This change creates conditions for asynchronous update while reducing the operation of encrypted data. By using the new process for asynchronous update, the number of data interactions can be reduced and the communication time can be compressed.

[0278] Next, the vertical model regression interaction process in the embodiment of the present application is introduced. Figure 9A It is a schematic diagram of the network architecture for vertical model linear regression interaction provided by the embodiment of the present application. Figure 9B It is a schematic diagram of the process for vertical model linear regression interaction provided by the embodiment of the present application. Similar to Figure 8A the same, Figure 9A the participating parties included are Party A, Party G, and Party H. Party H represents the data provider that does not share data with Party G.

[0279] Premise setting: Party H and Party G complete the screening of common samples through the intersection of encrypted IDs. In the following training sessions, it is default that the same ID value is used each time. Party A and Party H participate in the training simultaneously and interact with Party G.

[0280] By each party calculating the residual d respectively, it is possible to achieve synchronous gradient calculation for each party.

[0281] Step S901: Party A generates a public key and transmits it to Party H and Party G.

[0282] Step S902: Party G determines the amount of training data for each time within each round and sends it to Party H and Party G.

[0283] Step S903: Party H and Party G initialize their local models respectively and calculate the local intermediate calculation results.

[0284] Here, after Party H initializes its local model, the parameters of the model on the H side are denoted as w H , then according to the parameters of the model on the H side and the sample features, the local intermediate calculation result on the H side is calculated as w H x H , and this result is the predicted value of each sample on the H side.

[0285] After Party G initializes its local model, the parameters of the model on the G side are denoted as w G , then according to the parameters of the model on the G side and the sample features, the local intermediate calculation result on the G side is calculated as w G x G , and this result is the predicted value of each sample on the G side.

[0286] Step S904: Party H encrypts its intermediate calculation result using homomorphic encryption technology to generate the encrypted intermediate calculation result of Party H, and sends this encrypted intermediate calculation result to Party G. Party G encrypts its intermediate calculation result using homomorphic encryption technology to generate the encrypted intermediate calculation result of Party G, and sends this encrypted intermediate calculation result to Party H.

[0287] Here, the encrypted intermediate calculation result of Party H is denoted as [[w H x H , and this encrypted intermediate calculation result [[w H x H is sent to Party G; the encrypted intermediate calculation result of Party G is denoted as [[w G x G , and this encrypted intermediate calculation result [[w G x G is sent to Party H.

[0288] Step S905, Party G combines the intermediate calculation results sent by Party H to calculate the encrypted residual value [[di]], and Party H combines the intermediate results sent by Party G to calculate the encrypted residual value [[di]].

[0289] Here, the encrypted intermediate calculation result of Party G is expressed as [[w G x G -y]], and Party G uses [[w H x H sent by Party H and its own [[w G x G -y]]. The encrypted residual value [[di]] calculated based on linear regression can be expressed as [[w H x H + [[w G x G -y]], where y is the label provided by Party G.

[0290] Similarly, the encrypted intermediate calculation result of Party H is expressed as [[w H x H , and Party H uses its own [[w H x H and [[w G x G -y]] sent by Party G. The encrypted residual value [[di]] calculated based on linear regression can be expressed as [[w H x H + [[w G x G -y]], where y is the label provided by Party G.

[0291] Since Party G does not have the private key corresponding to the public key for decryption, it cannot decrypt this value, thus avoiding the data leakage of Party H.

[0292] Step S906, Party G and Party H respectively calculate the encrypted local gradients [[di]]x G and [[di]]x H using their own intermediate calculation results and encrypted residual values.

[0293] In some embodiments, [[di]]x G and [[di]]x H can be optimized respectively based on the first preset step size and the second preset step size.

[0294] Note that this step is performed synchronously by both parties in Step S905. Compared with the old scheme where Party G calculates the gradient uniformly, it avoids the delay waiting caused by possible network communication delays and accelerates the training efficiency.

[0295] Step S907, Party G and Party H send the encrypted local gradient [[di]]x to Party A.

[0296] When using the loss value for convergence judgment, Party G and Party H also send their own loss values L G and L H to Party A, and Party A combines the loss values by itself as L = L G + L H .

[0297] Party A decrypts the gradient using the private key (i.e., the private key in the above text) and sends it to each participating party respectively. Party A decides whether to converge based on the gradient norm or the loss value and notifies Party G and Party H.

[0298] In the embodiments of the present application, the calculation formula of L G the loss value is: L H the calculation formula of the loss value is:

[0299] Step S908, Party H and Party G update the local model parameters.

[0300] In the asynchronous gradient rounds: Repeat steps S905 to S906.

[0301] In the synchronous gradient rounds (multiples of n): Repeat steps S903 to S908.

[0302] Repeat steps S902 to S908 until the model converges or reaches the maximum number of model training rounds.

[0303] Figure 10 This is the schematic diagram of the process of vertical model logistic regression interaction provided by the embodiments of the present application. The difference from the linear regression interaction method is only that the regression model for calculating the residual value di is different, and the other steps are the same. For details, please refer to Figure 9A and Figure 9B the descriptions in the embodiments shown.

[0304] The embodiments of the present application describe an asynchronous update mechanism for linear model modeling under the vertical framework of federated learning. Aiming at the defect that the interaction mechanism of the solution in the related art cannot be conveniently updated asynchronously, by improving the interaction mechanism of the existing model, the time and encryption calculation that each participant needs to wait for transmission in communication are reduced, and at the same time, conditions are created for asynchronous update. After a specific number of local updates, interactions are carried out, effectively and controllably reducing the interactions of the modeling parties and confidential calculations while not excessively sacrificing the reliability of the resulting model, and improving the overall modeling efficiency. Moreover, in the framework of federated learning, the communication cost is high during training. By setting the same number of update rounds, compared with the existing linear regression training solution, the solution of the embodiments of the present application can also reduce the communication times of each participant, thereby reducing the calculation operations on the encrypted arrays and further improving the overall training efficiency.

[0305] The following continues to describe the exemplary structure of the software module implementation of the joint training device of the model provided by the embodiments of the present application. In some embodiments, as Figure 2 shown, the joint training device 111 of the model stored in the memory 140 is applied to the passive party of vertical federated learning. The passive party and the active party of vertical federated learning respectively use their own feature data for model training. The software modules in the joint training device 111 of the model may include:

[0306] The first acquisition module 112 is used to acquire the second ciphertext training result sent by the active party;

[0307] The second acquisition model 113 is used to acquire the first ciphertext training result and the number of trained rounds;

[0308] The first determination module 114 is used to determine the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result;

[0309] The first sending module 115 is used to send the first encrypted gradient to the coordinator when the number of trained rounds meets the synchronization condition, so that the coordinator determines the first decrypted gradient based on the first encrypted gradient;

[0310] The first receiving module 116 is used to receive the first decrypted gradient sent by the coordinator;

[0311] The first update module 117 is used to update its own training model based on the first decrypted gradient to obtain an updated training model.

[0312] In some embodiments, the joint training device 111 of the model further includes:

[0313] An optimization module, configured to optimize the first encrypted gradient based on a preset step size when the number of trained rounds does not meet the synchronization condition, so as to obtain an optimized first encrypted gradient;

[0314] A second update module, configured to update its own training model based on the optimized first encrypted gradient to obtain an updated training model.

[0315] In some embodiments, the joint training device 111 of the model further includes:

[0316] A second determination module, configured to determine whether the number of trained rounds meets the synchronization condition based on a first preset threshold;

[0317] The second determination module is further configured to determine that the number of trained rounds meets the synchronization condition when the number of trained rounds is divisible by the first preset threshold;

[0318] The second determination module is further configured to determine that the number of trained rounds does not meet the synchronization condition when the number of trained rounds is not divisible by the first preset threshold.

[0319] In some embodiments, the second acquisition module 113 is further configured to:

[0320] Receive the public key sent by the coordinator;

[0321] Obtain the sample data for joint training and its own training model;

[0322] Input the sample data into the training model for training to obtain a first training result;

[0323] Encrypt the first training result based on the public key to obtain a first encrypted training result.

[0324] In some embodiments, the second acquisition module 113 is further configured to:

[0325] Obtain the identifier of the current round of training samples from the initiator;

[0326] Based on the identifier, screen out the data corresponding to the identifier from its own feature data;

[0327] Determine the data corresponding to the identifier as the sample data for joint training.

[0328] In some embodiments, the first determination module 114 is further configured to:

[0329] Perform regression analysis on the first encrypted training result and the second encrypted training result to obtain an encrypted residual value;

[0330] Determine a first encrypted gradient based on the data corresponding to the identifier and the encrypted residual value.

[0331] In some embodiments, the joint training device 111 of the model may further include:

[0332] A third receiving module, configured to receive the convergence information sent by the coordinator;

[0333] A third determining module, configured to determine the updated training model as the trained target model when the convergence information is converged or the training is completed.

[0334] In some embodiments, the convergence information is determined by the coordinator based on the first decrypted gradient and a preset gradient norm threshold; or, the convergence information is determined by the coordinator based on the number of trained rounds and a preset round threshold; or, the convergence information is determined by the coordinator based on the first loss value sent by itself and the second loss value sent by the initiator;

[0335] When the convergence information is determined by the coordinator based on the first loss value sent by itself and the second loss value sent by the initiator, the joint training device 111 of the model may further include:

[0336] A fourth determining module, configured to determine the first loss value of this round of training based on the public key and the first training result;

[0337] A fourth sending module, configured to send the first loss value to the coordinator, so that the coordinator determines the loss value of this round of training based on the first loss value and the second loss value sent by the initiator, and determines the convergence information of this round of training based on the loss value.

[0338] Based on the above embodiments, an embodiment of the present application further provides a joint training device for a model, which is applied to a collaborator in vertical federated learning. At this time, the software modules in the joint training device of the model may include:

[0339] A generating module, configured to generate a public key for encryption and a private key for decryption;

[0340] A second sending module, configured to send the public key to the passive party and the initiator for model training respectively, so that the passive party and the initiator determine the first encrypted gradient and the second encrypted gradient based on the public key respectively;

[0341] A second receiving module, configured to receive the first encrypted gradient sent by the passive party and the second encrypted gradient sent by the initiator;

[0342] A decryption module, configured to decrypt the first encrypted gradient and the second encrypted gradient respectively based on the private key to obtain a first decrypted gradient and a second decrypted gradient;

[0343] A third sending module, configured to send the first decrypted gradient and the second decrypted gradient to the passive party and the active party respectively, so that the passive party and the active party update their respective training models based on the first decrypted gradient and the second decrypted gradient respectively.

[0344] In some embodiments, the joint training device of the model may further include:

[0345] A third obtaining module, configured to obtain convergence information;

[0346] A fifth sending module, configured to send the convergence information to the passive party and the active party.

[0347] In some embodiments, the third obtaining module is further configured to:

[0348] Determine the difference between the sum of the decrypted gradients and a preset gradient norm threshold; or determine the loss value of this round of training based on the first loss value sent by the passive party and the second loss value sent by the active party; or obtain the number of rounds of training completed by the passive party; the sum of the decrypted gradients is the sum of the first decrypted gradient and the second decrypted gradient;

[0349] When the difference is less than a second preset threshold, or when the loss value is less than a preset loss threshold, determine the convergence information as converged; or when the number of rounds of training completed is greater than a preset round threshold, determine the convergence information as training completed.

[0350] It should be noted here that: the description of the embodiments of the joint training device of the above model is similar to the above method description and has the same beneficial effects as the method embodiments. For the technical details not disclosed in the embodiments of the joint training device of the model in this application, those skilled in the art can refer to the description of the method embodiments of this application for understanding.

[0351] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the joint training method of the model in the above embodiments of the present application.

[0352] An embodiment of the present application provides a storage medium storing executable instructions, where the executable instructions, when executed by a processor, cause the processor to execute the method provided by the embodiment of the present application. For example, as Figures 3 to 7 the method shown.

[0353] In some embodiments, the storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; it may also be various devices including one or any combination of the above memories.

[0354] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0355] As an example, the executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file storing other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (e.g., files storing one or more modules, subroutines, or code portions).

[0356] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.

[0357] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A method for joint training of a model, characterized in that, applied to the passive party in vertical federated learning, where the passive party and the active party in vertical federated learning respectively use their own feature data for model training, and the method includes: Obtain the second ciphertext training result sent by the active party; Obtain the first ciphertext training result and the number of trained rounds; Determine the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result, where determining the first encrypted gradient based on the first ciphertext training result and the second ciphertext training result includes: Perform regression analysis on the first ciphertext training result and the second ciphertext training result to obtain an encrypted residual value; Determine the first encrypted gradient based on the data corresponding to the identifier and the encrypted residual value; When the number of trained rounds meets the synchronization condition, send the first encrypted gradient to the coordinator so that the coordinator determines the first decrypted gradient based on the first encrypted gradient; Receive the first decrypted gradient sent by the coordinator and update its own training model based on the first decrypted gradient to obtain an updated training model; The method further includes: Determine whether the number of trained rounds meets the synchronization condition based on a first preset threshold, where the first preset threshold is any integer value; When the number of trained rounds is divisible by the first preset threshold, determine that the number of trained rounds meets the synchronization condition; When the number of trained rounds is not divisible by the first preset threshold, determine that the number of trained rounds does not meet the synchronization condition; Among them, obtaining the first ciphertext training result includes: Receive the public key sent by the coordinator; Obtain the sample data for joint training and its own training model; Input the sample data into the training model for training to obtain a first training result; Encrypt the first training result based on the public key to obtain the first ciphertext training result; Among them, obtaining the sample data for joint training includes: Obtain the identifier of the current round of training samples from the active party; Based on the identifier, screen out the data corresponding to the identifier from its own feature data; Determine the data corresponding to the identifier as the sample data for joint training.

2. The method according to claim 1, characterized in that, The method further includes: When the number of trained rounds does not meet the synchronization condition, optimize the first encrypted gradient based on a preset step size to obtain an optimized first encrypted gradient; Update its own training model based on the optimized first encrypted gradient to obtain an updated training model.

3. The method according to claim 1, characterized in that, The method further includes: Receive the convergence information sent by the coordinator; When the convergence information is converged or training is completed, determine the updated training model as the trained target model.

4. The method according to claim 3, characterized in that, The convergence information is determined by the coordinator based on the first decrypted gradient and a preset gradient norm threshold; or, the convergence information is determined by the coordinator based on the number of trained rounds and a preset round threshold; or, the convergence information is determined by the coordinator based on the first loss value sent by itself and the second loss value sent by the initiator; When the convergence information is determined by the coordinator based on the first loss value sent by itself and the second loss value sent by the initiator, the method further includes: Determining a first loss value for the current round of training based on the public key and the first training result; Sending the first loss value to the coordinator, so that the coordinator determines the loss value of the current round of training based on the first loss value and the second loss value sent by the initiator, and determines the convergence information of the current round of training based on the loss value.

5. A method for joint training of a model, Characterized in that, Applied to the coordinator according to any one of claims 1-4 in vertical federated learning, the method includes: Generating a public key for encryption and a private key for decryption; Sending the public key to the passive party and the initiator for model training respectively, so that the passive party and the initiator determine a first encrypted gradient and a second encrypted gradient respectively based on the public key; Receiving the first encrypted gradient sent by the passive party and the second encrypted gradient sent by the initiator; Decrypting the first encrypted gradient and the second encrypted gradient respectively based on the private key to obtain a first decrypted gradient and a second decrypted gradient; Sending the first decrypted gradient and the second decrypted gradient to the passive party and the initiator respectively, so that the passive party and the initiator update their respective training models based on the first decrypted gradient and the second decrypted gradient respectively.

6. The method according to claim 5, Characterized in that, The method further includes: Obtaining convergence information; Sending the convergence information to the passive party and the initiator.

7. The method according to claim 6, Characterized in that, The obtaining of the convergence information includes: Determining the difference between the sum of the decrypted gradients and a preset gradient norm threshold; or determining the loss value of the current round of training based on the first loss value sent by the passive party and the second loss value sent by the initiator; or obtaining the number of trained rounds of the passive party; the sum of the decrypted gradients is the sum of the first decrypted gradient and the second decrypted gradient; When the difference is less than a second preset threshold, or when the loss value is less than a preset loss threshold, determining that the convergence information has converged; or when the number of trained rounds is greater than a preset round threshold, determining that the convergence information has completed training.

8. A device for joint training of a model, Characterized in that, Applied to the passive party in vertical federated learning, the passive party and the initiator in vertical federated learning respectively use their own feature data for model training, and the device includes: A first acquisition module, configured to acquire a second ciphertext training result sent by the initiator; A second acquisition module, configured to acquire the first ciphertext training result and the number of trained rounds. Specifically, the second acquisition module is configured to receive the public key sent by the coordinator; acquire the sample data for joint training and its own training model; input the sample data into the training model for training to obtain a first training result; encrypt the first training result based on the public key to obtain a first ciphertext training result. Wherein, the second acquisition module is further specifically configured to acquire the identifier of the training samples in this round from the initiator; filter out the data corresponding to the identifier from its own feature data; and determine the data corresponding to the identifier as the sample data for joint training. A first determination module, configured to determine a first encrypted gradient based on the first ciphertext training result and the second ciphertext training result. Specifically, the first determination module is configured to perform regression analysis on the first ciphertext training result and the second ciphertext training result to obtain an encrypted residual value; and determine the first encrypted gradient based on the data corresponding to the identifier and the encrypted residual value. A first sending module, configured to send the first encrypted gradient to the coordinator when the number of trained rounds meets the synchronization condition, so that the coordinator determines a first decrypted gradient based on the first encrypted gradient. A first receiving module, configured to receive the first decrypted gradient sent by the coordinator. A first update module, configured to update its own training model based on the first decrypted gradient to obtain an updated training model. The device further includes: determining whether the number of trained rounds meets the synchronization condition based on a first preset threshold, where the first preset threshold is any integer value. When the number of trained rounds is divisible by the first preset threshold, it is determined that the number of trained rounds meets the synchronization condition. When the number of trained rounds is not divisible by the first preset threshold, it is determined that the number of trained rounds does not meet the synchronization condition.

9. A device for joint training of a model Characterized in that Applied to the coordinator according to any one of claims 1-4 in vertical federated learning, the device includes: A generation module, configured to generate a public key for encryption and a private key for decryption. A second sending module, configured to send the public key to the passive party and the initiator for model training respectively, so that the passive party and the initiator determine a first encrypted gradient and a second encrypted gradient respectively based on the public key. A second receiving module, configured to receive the first encrypted gradient sent by the passive party and the second encrypted gradient sent by the initiator. A decryption module, configured to decrypt the first encrypted gradient and the second encrypted gradient respectively based on the private key to obtain a first decrypted gradient and a second decrypted gradient. A third sending module, configured to send the first decrypted gradient and the second decrypted gradient to the passive party and the initiator respectively, so that the passive party and the initiator update their respective training models based on the first decrypted gradient and the second decrypted gradient respectively.

10. A device for joint training of a model Characterized in that The device includes: A memory, configured to store executable instructions. A processor, when executing the executable instructions stored in the memory, implements the method according to any one of claims 1 to 4 or claims 5 to 7.

11. A computer-readable storage medium, characterized in that executable instructions are stored on the computer-readable storage medium, and when the processor executes the executable instructions, the method according to any one of claims 1 to 4 or claims 5 to 7 is implemented.

12. A computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the method according to any one of claims 1 to 4 or claims 5 to 7 is implemented.

Citation Information

Patent Citations

  • Longitudinal federated learning model training optimization method and device, equipment and medium

    CN111242316A

  • Model parameter updating method, device and equipment, and readable storage medium

    CN111368196A