Method and apparatus for jointly updating a model

The joint model updating method using asymmetric encryption and randomization algorithms securely combines data from different parties, enhancing model accuracy and privacy protection in machine learning applications.

CN115358387BActive Publication Date: 2025-07-15ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210988137.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-17
Publication Date
2025-07-15
Estimated Expiration
2042-08-17

AI Technical Summary

Technical Problem

When multi-party joint training of deep neural network models, how to improve model accuracy while protecting data privacy, especially in joint data training between different enterprises or institutions, how to effectively protect data privacy and commercial secrets of all parties while improving model performance.

Method used

A combination of asymmetric key encryption and out-of-order algorithms is used to generate public-private key pairs and out-of-order matrix, and only the last layer of the training sample is encrypted and calculated, and other layers are plaintext calculations are performed, and homomorphic encryption and sparse gradient data are used to ensure that data privacy is not leaked.

Benefits of technology

It realizes that while protecting data privacy, accelerates the computing process, ensures model accuracy, prevents data value leakage, and effectively protects the data privacy and commercial secrets of all parties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115358387B_ABST
    Figure CN115358387B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a method and device for two-party joint model update. Under the vertical federated learning architecture, the label holder generates a homomorphic encryption key pair, where the first key can be publicly disclosed to the other party, and the second key is held locally. The non-label holder can generate a scrambling algorithm. The label holder provides the local processing result of the local features to the second party in ciphertext form, and the non-label holder performs ciphertext scrambling fusion on the local processing result of the local features and the local processing result provided by the label holder. The non-label holder provides the ciphertext scrambling fusion result to the label holder, and the label holder completes prediction and determines the model loss in plaintext state based on the decryption of the second key. Based on the gradient backpropagation of the model loss, the label holder obtains the gradient of the scrambling fusion result, and this gradient can be restored to the original order and sparsified by the non-label holder for updating the undetermined parameters of the local models of all parties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a method and device for jointly updating a model. Background Art

[0002] With the development of computer technology, machine learning has been increasingly widely applied in various business scenarios. With the main progress of artificial intelligence technology, deep neural networks (DNNs) have gradually been applied in fields such as risk assessment, speech recognition, face recognition, and natural language processing. However, the DNN network structure under different application scenarios is relatively fixed. To achieve better model performance, more training data is required. In fields such as healthcare and finance, different enterprises or institutions have different data samples. Once these data are jointly trained, the model accuracy will be greatly improved, bringing huge economic benefits to the enterprises. Federated learning based on multi-party secure computing is a method for jointly modeling while protecting private data. However, these original training data contain a large amount of user privacy and business secrets. Once the information is leaked, it will cause irreparable negative impacts. Therefore, while multi-party joint training solves the data silo problem, protecting data privacy has become a key technical issue in related technical fields in recent years. Summary of the Invention

[0003] One or more embodiments of this specification describe a method and device for jointly updating a model to solve one or more problems mentioned in the background art.

[0004] According to a first aspect, a method for jointly updating a model is provided for a first party and a second party to jointly train a model. Among them, the first party holds the first feature of the training sample and the label data, and the second party holds the second feature of the training sample. The model includes a first local model MA and a third model ML provided in the first party, and a second local model MB provided in the second party. The first local model MA and the second local model MB are respectively used to process the first feature and the second feature to obtain tensors of a predetermined dimension as corresponding intermediate results. In the current update cycle, the first party and the second party have n training samples with the same arrangement order in the current batch. The method is executed by the first party and includes: using the first local model MA to process the first feature XA corresponding to the n training samples to obtain an intermediate tensor HA composed of n tensors of a predetermined dimension; encrypting the intermediate tensor HA with the first key Pk in the locally held asymmetric key pair, and obtaining the first ciphertext tensor <ha>Provided to a second party for the second party to feedback the encrypted ciphertext scrambled fusion tensor based on the first key Pk <hs>, wherein the first secret key Pk is publicly disclosed by the first party to the second party, and the encrypted shuffled fusion tensor <hs>Via the first ciphertext tensor <ha>It is obtained by fusing the intermediate tensor HB in a superimposed and fused manner in the encrypted state of the training samples in a scrambled order according to the scrambling algorithm S and encrypted by the first secret key Pk. The intermediate tensor HB is obtained by the second party processing the second feature XB corresponding to the n training samples using the second local model MB; using the third model ML to process the encrypted scrambled fusion tensor decrypted by the second secret key Sk in the asymmetric key pair <hs>Obtain the shuffled fusion tensor Hs, and obtain n predicted tensors Ys after shuffling the n training samples by the model; compare the n predicted tensors Ys with the shuffled labels Ys' corresponding to the n training samples respectively to determine the model loss, where the shuffled labels Ys' are obtained through secure shuffling by the second party based on the first key Pk and the shuffling algorithm S; update the undetermined parameters in the first local model MA and the third model ML with the goal of reducing the model loss, where the undetermined parameters in the third model ML are updated by the gradient determined based on the model loss, and the undetermined parameters in the first local model MA are updated by the gradient data determined based on the secure computation of homomorphic encryption with the second party.

[0005] In one embodiment, the n training samples are determined in advance by the first party and the second party through private set intersection according to the sample identifier that can uniquely describe the training samples.

[0006] In one embodiment, the shuffling algorithm S is a shuffling matrix generated for the current update cycle, and this shuffling matrix is used to disrupt the data arrangement order of different training samples.

[0007] In one embodiment, the superimposed fusion method is one of addition, mean calculation, and weighted average.

[0008] In one embodiment, the shuffled label Ys' is determined in the following way: encrypt the n label data corresponding to the n training samples through the first key Pk to obtain the ciphertext labels <y>; providing the ciphertext tag to a second party <y>for a second party to feed back ciphertext tags processed by the scrambling algorithm S <y>Obtain the ciphertext scrambled tag <Ys'>; use the second key Sk to decrypt the ciphertext scrambled tag <Ys'> to obtain the scrambled tag Ys'.

[0009] In one embodiment, the undetermined parameters in the first local model MA are updated in the following manner: Provide the second party with the gradient data Gs of the model loss with respect to the scrambled fusion tensor Hs, so that the second party can perform the inverse operation S of the scrambled algorithm S -1 Restore the gradient data Gs to the gradient data G according to the training sample restoration order, where the gradient data G describes the gradient of the model loss with respect to the fusion tensor H; Based on the Jacobi matrices of the first local model MA with respect to the undetermined parameters of the last hidden layer and the input data respectively, perform a secure matrix multiplication under the homomorphic encryption of the first key Pk with the sparsified matrix Gp of the gradient data G to obtain the ciphertext gradient data of the undetermined parameters of the last hidden layer and the ciphertext gradient data of the input data of the last hidden layer at the second party; Receive the ciphertext gradient data of the undetermined parameters of the last hidden layer and the ciphertext gradient data of the input data of the last hidden layer from the second party; Use the second key Sk to decrypt the ciphertext gradient data of the undetermined parameters of the last hidden layer for updating the undetermined parameters of the last hidden layer; Use the second key Sk to decrypt the ciphertext gradient data of the input data of the last hidden layer to determine the gradients of the undetermined parameters of the other hidden layers except the last hidden layer in the first local model MA based on gradient backpropagation, so as to update the undetermined parameters of the other hidden layers.

[0010] In one embodiment, each of the first feature and the second feature includes at least one service feature extracted from local service data.

[0011] According to a second aspect, there is provided a method for jointly updating a model by two parties, which is used for the first party and the second party to jointly train a model. Among them, the first party holds the first feature of the training sample and the label data, and the second party holds the second feature of the training sample. The model includes a first local model MA and a third model ML provided at the first party, and a second local model MB provided at the second party. The first local model MA and the second local model MB are respectively used to process the first feature and the second feature to obtain tensors of a predetermined dimension as corresponding intermediate results; In the current update cycle, the first party and the second party correspond to n training samples with the same arrangement order in the current batch. The method is executed by the second party and includes: using the second local model MB to process the second feature XB corresponding to the n training samples to obtain an intermediate tensor HB composed of n tensors of a predetermined dimension; Based on the first key Pk received from the first party and the first ciphertext tensor <ha>, determine the encrypted disordered fusion tensor <hs>and fed back to the first party, wherein the encrypted state scrambled fusion tensor <hs>Via the first ciphertext tensor <ha>It is obtained by fusing the intermediate tensor HB with the shuffled algorithm S in a shuffled manner according to the training samples and in the encrypted form of the first key Pk in a superimposed and fused manner. The first key Pk and the second key Sk held by the first party are an asymmetric key pair; receiving from the first party the gradient data Gs of the model loss with respect to the shuffled fusion tensor Hs, where the shuffled fusion tensor Hs decrypts the encrypted shuffled fusion tensor based on the second key Sk <hs>It is obtained that the model loss is determined by comparing the prediction result Ys of the shuffled fusion tensor Hs processed by the third model ML by the first party with the shuffled labels Ys' of n training samples; based on the inverse operation S of the shuffling algorithm S -1 Restore the gradient data Gs to the gradient data G in the order of the training samples. The gradient data G describes the gradient of the model loss with respect to the fusion tensor H; use the gradient data G to determine the gradients corresponding to the respective undetermined parameters in the second local model MB, so as to update the respective undetermined parameters in the second local model MB using the respective gradients.

[0012] In one embodiment, the shuffling algorithm S is a shuffling matrix generated for the current update period, and this shuffling matrix is used to disrupt the data arrangement order of different training samples.

[0013] In one embodiment, the superposition fusion method is one of summation, averaging, and weighted averaging.

[0014] In one embodiment, the method further includes determining the shuffled label Ys' for the first party in the following manner: receiving, from the first party, the ciphertext labels obtained by encrypting the n label data corresponding to the n training samples with the first key Pk <y>; the ciphertext label is processed by the scrambling algorithm S <y>The obtained ciphertext shuffled tag <Ys'> is fed back to the first party, so that the first party can decrypt the ciphertext shuffled tag <Ys'> using the second secret key Sk to obtain the shuffled tag Ys'.

[0015] In one embodiment, the process of using the gradient data G to determine the gradients corresponding to the respective undetermined parameters in the second local model MB includes: performing random sparsification with a sparsification rate p on the gradient data G to obtain a sparsified gradient Gp; using the sparsified gradient Gp to determine the gradient data of the model loss with respect to the intermediate tensor HB; and performing backpropagation of the gradient data of the model loss with respect to the intermediate tensor HB in the second local model to determine the gradients of the respective undetermined parameters in the second local model, thereby updating the respective undetermined parameters in the second local model.

[0016] In one embodiment, the method further includes: using the sparsified gradient Gp to determine the gradient data of the model loss with respect to the intermediate tensor HA; performing a homomorphic-encryption-based secure multiplication with the first party to obtain the encrypted gradient data of the model loss with respect to the undetermined parameters of the last hidden layer of the first local model MA and the encrypted gradient data of the input data of the last hidden layer; and providing the encrypted gradient data of the model loss with respect to the undetermined parameters of the last hidden layer of the first local model MA and the encrypted gradient data of the input data of the last hidden layer to the first party, so that after the first party decrypts the encrypted gradient data related to the last hidden layer, it can be used to update the respective undetermined parameters in the last hidden layer of the first local model MA, and after decrypting the encrypted gradient data of the input data of the last hidden layer, it can be used to determine the gradients of the undetermined parameters of the other hidden layers except the last hidden layer and update the corresponding other undetermined parameters.

[0017] According to a third aspect, there is provided an apparatus for jointly updating a model by two parties, which is disposed in the first party of the two parties jointly training the model. The first party holds the first feature and label data of the training samples, and the second party of the two parties jointly training the model holds the second feature of the training samples. The model includes a first local model MA and a third model ML disposed in the first party, and a second local model MB disposed in the second party. The first local model MA and the second local model MB are respectively configured to process the first feature and the second feature to obtain tensors of a predetermined dimension as corresponding intermediate results. In the current update cycle, the first party and the second party have n training samples with the same arrangement order in the current batch. The apparatus includes:

[0018] A processing unit configured to use the first local model MA to process the first feature XA corresponding to the n training samples to obtain an intermediate tensor HA composed of n tensors of a predetermined dimension;

[0019] An encryption unit configured to encrypt the intermediate tensor HA using the first key Pk in the locally held asymmetric key pair, and obtain the first encrypted tensor <ha>Provided to a second party for the second party to feedback a ciphertext shuffled fusion tensor encrypted based on the first key Pk <hs>, wherein the first key Pk is publicly disclosed by the first party to the second party, and the encrypted disordered fusion tensor <hs>Via the first ciphertext tensor <ha>It is obtained by fusing the intermediate tensor HB in a superimposed and fused manner in the encrypted state of being shuffled according to the training samples by the shuffling algorithm S and encrypted by the first secret key Pk. The intermediate tensor HB is obtained by the second party processing the second feature XB corresponding to the n training samples using the second local model MB;

[0020] The prediction unit is configured to process the encrypted shuffled fusion tensor by decrypting it with the second secret key Sk in the pair of asymmetric secret keys using the third model ML <hs>Obtain the shuffled fusion tensor Hs, and obtain n predicted tensors Ys after shuffling the n training samples by the model;

[0021] A loss determination unit, configured to compare the n predicted tensors Ys with the shuffled labels Ys' corresponding to the n training samples respectively, so as to determine the model loss, wherein the shuffled labels Ys' are obtained through secure shuffling by the second party based on the first key Pk and the shuffling algorithm S;

[0022] An adjustment unit, configured to update the undetermined parameters in the first local model MA and the third model ML with the goal of reducing the model loss, wherein the undetermined parameters in the third model ML are updated by the gradient determined based on the model loss, and the update gradient of the undetermined parameters in the first local model MA is determined by performing secure calculation with the second party based on the gradient data of the model loss provided to the second party for the shuffled fusion tensor Hs.

[0023] According to a fourth aspect, there is provided a device for jointly updating a model by two parties, which is used for the first party and the second party to jointly train a model. The first party holds the first feature and label data of the training samples, and the second party holds the second feature of the training samples. The model includes a first local model MA and a third model ML provided in the first party, and a second local model MB provided in the second party. The first local model MA and the second local model MB are respectively used to process the first feature and the second feature to obtain tensors of a predetermined dimension as corresponding intermediate results; in the current update cycle, the first party and the second party have n training samples with the same arrangement order in the current batch. The device is provided in the second party and includes:

[0024] A processing unit, configured to process the second feature XB corresponding to the n training samples by using the second local model MB to obtain an intermediate tensor HB composed of n tensors of a predetermined dimension;

[0025] A fusion unit, configured to be based on the first key Pk received from the first party and the first ciphertext tensor <ha>, determine the encrypted state disordered fusion tensor <hs>and fed back to the first party, wherein the encrypted state shuffled fusion tensor <hs>Via the first ciphertext tensor <ha>It is obtained by fusing the intermediate tensor HB with the scrambled algorithm S in a scrambled form according to training samples and in a superimposed and fused manner in the encrypted form of the first key Pk. The first key Pk and the second key Sk held by the first party are an asymmetric key pair;

[0026] A receiving unit, configured to receive, from the first party, gradient data Gs of the model loss with respect to the scrambled fusion tensor Hs, where the scrambled fusion tensor Hs decrypts the encrypted scrambled fusion tensor based on the second key Sk <hs>It is obtained that the model loss is determined by comparing the predicted result Ys of the third model ML processing the shuffled fusion tensor Hs by the first party with the shuffled labels Ys' of n training samples;

[0027] A shuffled recovery unit, configured to restore the gradient data Gs to the gradient data G based on the inverse operation S of the shuffled algorithm S, where the gradient data G describes the gradient of the model loss with respect to the fusion tensor H; -1

[0028] An adjustment unit, configured to use the gradient data G to determine the respective gradients corresponding to the respective undetermined parameters in the second local model MB, and thereby update the respective undetermined parameters in the second local model MB using the respective gradients.

[0029] According to a fifth aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect or the second aspect above.

[0030] According to a sixth aspect, there is provided a computing device, including a memory and a processor, characterized in that the memory stores executable code, and when the processor executes the executable code, the method of the first aspect or the second aspect above is implemented.

[0031] Through the method and device provided in the embodiments of this specification, in the joint split learning process carried out by two training members based on multi-party secure computing, based on the characteristic that the data is vertically split, the training member holding the label (label training member) generates a public-private key pair and discloses the encryption key to other training members, and other training members (ordinary training members) hold the shuffled algorithm. By combining the shuffled algorithm, only the last layer of the embedding model is encrypted for calculation, and the remaining layer models are calculated in plaintext. The ordinary training member fuses the intermediate processing results of the two parties (such as intermediate tensors HA and HB) in the encrypted state, and the label training member restores the plaintext data of the intermediate tensor for the shuffled fusion of the two parties (such as the shuffled fusion tensor Hs), and performs plaintext calculation during the prediction process, which can accelerate the calculation process, save calculation time, and ensure the accuracy of the model. Moreover, the label training member cannot obtain the plaintext output result of the local model of the ordinary training member (such as HB), and the ordinary member cannot obtain the plaintext label of the label member, which can effectively protect the data value of the ordinary training member and the label privacy data of the label training member. And the random sparsification method is used to encrypt the gradients of the backpropagation to prevent the label training member from obtaining the shuffled algorithm and ensure that the data value is not leaked. Description of the Drawings

[0032] ​To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0033] Figure 1 It is a schematic diagram of data interaction under a multi-party secure computing architecture;

[0034] Figure 2 It shows a schematic flowchart of a joint update model under the technical concept of this specification;

[0035] Figure 3 It shows a two-party interaction timing diagram in the process of a joint training model in an embodiment of this specification;

[0036] Figure 4 It shows a two-party interaction timing diagram for disorderly processing of label data in the process of a joint training model in an embodiment of this specification;

[0037] Figure 5 It shows a schematic block diagram of a device of a joint training model provided for a training member holding a label in an embodiment of this specification;

[0038] Figure 6 It shows a schematic block diagram of a device of a joint training model provided for other training members in an embodiment of this specification. Detailed implementation manners

[0039] The solutions provided in this specification will be described below in conjunction with the accompanying drawings.

[0040] Federated Learning, also known as Federated Machine Learning, Joint Learning, Alliance Learning, etc. Federated Machine Learning is a machine learning framework that can effectively help multiple institutions to perform data usage and machine learning modeling while meeting the requirements of user privacy protection, data security, and government regulations.

[0041] Specifically, assume that enterprise A and enterprise B each establish a task model. A single task can be classification or prediction, and these tasks have also been recognized by their respective users when obtaining data. However, due to incomplete data, for example, enterprise A lacks labeled data and enterprise B lacks user feature data, or the data is insufficient and the sample size is not enough to build a good model, then the models at each end may not be established or the effects may not be ideal. The problem that federated learning aims to solve is how to build a high-quality machine learning model at each end of A and B. The training of this model uses the data of each enterprise such as A and B, and the proprietary data of each enterprise is not known to other parties. That is, without violating data privacy regulations, a common model is established. This common model is like the optimal model built by aggregating data from all parties. In this way, the established model only serves the own goals in the regions of each party.

[0042] Each institution participating in federated learning or each training member providing data for federated learning can be called a training member. As a distributed machine learning solution, federated learning can include two implementation architectures. One is to use a trusted or semi-trusted third party as a server to fuse the data of each training member of federated learning, and the other is to perform multi-party secure computation among multiple training members and protect data privacy through methods such as homomorphic encryption and secret sharing, so as to complete joint modeling. Figure 1 Shows a joint learning implementation architecture based on multi-party secure computation. This implementation architecture is a decentralized implementation architecture.

[0043] During the joint learning process, each training member can respectively hold different business data and can each participate in the joint training of the model through devices, computers, servers, etc. Here, the business data can be various data such as characters, pictures, voices, animations, videos, etc. Generally, the business data held by each training member is relevant, and the business parties corresponding to each training member can also be relevant. For example, among multiple business parties involved in the financial business, business party 1 is a bank that provides services such as savings and loans for users and can hold data such as users' income and expenditure flows, loan amounts, and deposit amounts. Business party 2 is an investment management platform that can hold data such as users' borrowing records, investment records, and repayment timeliness. Business party 3 is an e-commerce website that holds data such as users' shopping habits, payment habits, and payment accounts. Another example is that among multiple business parties involved in the medical business, each business party can be various hospitals, physical examination institutions, etc. For example, business party 1 is hospital A, and the local business data includes medical records such as user symptoms, diagnosis results, treatment plans, and treatment results. Business party 2 can be physical examination institution B, and the physical examination record data includes user symptoms, physical examination conclusions, etc. A single training member can hold the business data of one or more business parties that it trusts.

[0044] The model here can be used to process business data and obtain corresponding business processing results. Therefore, it can also be called a business model. Specifically, what kind of business data is processed and what kind of business processing results are obtained depend on actual needs. For example, the business data can be data related to user finance, and the business processing result is the financial credit assessment result of the user. Another example is that the business data can be the customer service conversation data of the user, and the business processing result is the recommended result of the customer service answer, and so on.

[0045] This specification provides corresponding federated learning solutions based on the implementation architecture of multi-party secure computing. It can be understood that according to the relationship between the data held by each training member, the samples of federated learning can be horizontally distributed, corresponding to horizontal federated learning (the samples are aligned horizontally, which can also be called horizontal distribution), or vertically distributed (the features are aligned vertically, which can also be called vertical distribution). More specifically: in the case of horizontal distribution, a single training member can hold the complete data of multiple training samples, that is, all feature data and label data, and the samples among the training members are different; in the case of vertical distribution, the sample spaces of each training member are the same, and a single training member can hold at most partial feature data of each training sample, and some training members also hold label data.

[0046] For vertical distribution, conventional techniques have proposed a split learning method to solve the joint training problem of vertically sliced data. Taking serverless vertical split learning as an example, in a specific split learning training task, each training member has a hidden layer generation model (such as an Embedding model), and the training member with labels (label training member) also has a loss determination model (such as a Loss-computing model). The loss determination model is used to obtain the prediction result and determine the model loss during the model training process. In the loss determination model, the part used to determine the prediction result needs to be used both during model training and when using the trained model for task prediction, while the part used to determine the model loss is only used during the model training process and is not used during the task prediction process using the trained model.

[0047] During the conventional training process, all training members use local private data for the forward propagation of the Embedding model, calculate the corresponding hidden layer output results (such as embedding tensors), and transmit the local hidden layer output results to the label training member. The label training member receives the hidden layer output results of the ordinary training members and fuses them, and then continues the forward propagation to the Loss-computing model to calculate the model loss. The label training member performs backpropagation of the gradient based on the model loss, updates the local model (such as the part of the loss determination model used to obtain the prediction output results and the local hidden layer generation model), and feeds back the gradients of the corresponding hidden layer output results to the ordinary training members, so that each ordinary training member performs backpropagation of the gradient of the local hidden layer generation model to update the model parameters (undetermined parameters in the model) of the local hidden layer generation model.

[0048] During the split learning training and inference process of a certain task, if the training member with the label stores the hidden layer output results of the ordinary training members locally as data for other training tasks, it is not conducive to the protection of the privacy data of the training members and may cause the loss of the data information value of the ordinary training members.

[0049] In view of this, this specification proposes a federated learning method based on the combination of homomorphic encryption and data scrambling. In the case of two training members performing federated learning through multi-party secure computation, a technical solution for effectively maintaining data privacy and protecting the data information value is provided. The applicable scenario of this technical concept is as follows: two training members have the same multiple training samples, one training member holds part of the feature data of these training samples, and hereinafter this training member will be referred to as the second party, and the other training member holds other feature data and label data of these training samples, and hereinafter this training member will be referred to as the first party. Refer to Figure 2 As shown, it is a specific implementation architecture of the technical concept of this specification. Figure 2 In it, Alice is the first party and Bob is the second party.

[0050] Specifically, the first party can generate an asymmetric key pair (Pk, Sk), one key (hereinafter referred to as the first key, such as the public key Pk) can be made public to the second party, and the other key (hereinafter referred to as the second key, such as the private key Sk) is held by the first party itself. The second party can generate a scrambling algorithm S (such as Figure 2 The shuffled matrix S) is used to scramble the data arrangement order of different training samples in the tensor, so as to scramble the order of the corresponding tensor in the sample dimension while maintaining the data alignment of a single training sample. During the federated learning process, the two parties determine a batch of samples through the private intersection method and jointly update the model in one update cycle. Each of the two parties obtains the hidden layer output results of the local feature data (which can also be called the local processing results, such as H_A, H_B) through the local hidden layer generation model (such as Model A, Model B), and encrypts them using the aforementioned first key Pk. The encrypted hidden layer output results can be aggregated to the second party, and the second party fuses the encrypted hidden layer output results of both parties through the encrypted shuffled and superimposed method, and sends the fusion result (such as Figure 2 in <H_s> = <S * (H_A + H_B)>) to the first party. Among them, the purpose of fusing the encrypted hidden layer output results in a superimposed manner is to prevent the first party from inferring the shuffling method based on the simple shuffling of its own hidden layer output results and mining more data secrets. The first party decrypts the encrypted fusion result to obtain the plaintext fusion result (such as Figure 2 in H_s = S * (H_A + H_B)) and uses the forward propagation of the prediction model (such as Figure 2 in Model_L) to determine the prediction results of the model for each training sample (such as Figure 2 in y_s).

[0051] Since the prediction results are shuffled in the training sample dimension, the label data needs to be shuffled in the same way to align the prediction results and the sample labels according to the training samples, so as to determine the model loss. Considering that the labels are usually described by relatively abstract numerical values. For example, in a binary classification model, there may be only two labels, such as represented by numerical values 0 and 1, and the label tensor is composed of the numerical values 0 and 1. Therefore, shuffling the labels usually does not disclose data privacy. In this way, the first party can send the label tensor encrypted by the first key to the second party. The second party performs the same shuffling process as the hidden layer output results on it and then feeds it back to the first party. The first party decrypts it with the second key to obtain the label data aligned with the prediction result samples (such as Figure 2 in S * Y). In this way, the first party can determine the model loss (Loss).

[0052] During the backpropagation of the gradients of each undetermined parameter based on the model loss, the gradient propagation of the prediction model part can be performed locally by the first party until the gradient of the model loss with respect to the fusion result H_s is obtained, denoted as G_s for convenience of description. The first party can provide G_s to the second party, and the second party can use the inverse operation of shuffling to restore the correct order for it to obtain the gradient G of the model loss with respect to the fusion result. In this way, for the second party, based on the fusion method of the hidden layer output results of the two parties obtained above, the gradients of the local hidden layer output result H_B and the first party's hidden layer output result H_A can be determined respectively. Through the gradient of the model loss with respect to the local hidden layer output result H_B, the second party can determine the gradients of each undetermined parameter in the local hidden layer generation model and update them. In order to determine the gradients of each undetermined parameter in the first party's hidden layer generation model and keep the shuffling method from being leaked to the first party, the second party can sparsify the gradient of the model parameter with respect to the fusion result G, and use the sparsified result sparse G to perform a homomorphic encryption-based secure multiplication with the Jacobian matrix of the undetermined parameters (i.e., the input data) of the last hidden layer (the hidden layer that outputs H_A) of the first party's hidden layer generation model, so as to securely determine the gradients of each undetermined parameter of the last hidden layer and the gradient of the input data of the last hidden layer in the first party's local model. These gradients exist in ciphertext form in the second party. After the second party provides them to the first party for decryption, the first party can use them to determine the gradients of other undetermined parameters for updating.

[0053] Under the above technical concept: By using the method of combining row / column shuffling of training samples with homomorphic encryption to process the hidden layer output results and labels, the smooth progress of model training can be ensured. The label training members cannot obtain the plaintext hidden layer output results of ordinary training members, and ordinary members cannot obtain the plaintext labels of label training members, thus protecting the data value of the hidden layer output results of ordinary training members and the label privacy data of label training members; The label training members restore the plaintext of the fused hidden layer output results and perform plaintext calculations on the loss computing model, which can save computing time and ensure the accuracy of the model; By using the method of random sparsification to encrypt the gradients of backpropagation, it is possible to prevent the label training members from obtaining the transformation matrix and ensure that the data information value is not leaked; Only the last layer of the hidden layer generation model is encrypted for calculation, and the rest of the layer models are calculated in plaintext to accelerate the calculation process.

[0054] The following will Figure 3 describe the technical concept of this specification in detail in combination with the

[0055] Please refer to Figure 3 As shown, a flowchart of the joint update model of an embodiment is given. This flowchart is a timing diagram involving the forward processing and the gradient backward propagation process. In this process, there are two training members: the first party that holds part of the feature data and the label data, and the second party that holds other feature data. Among them, for the convenience of description, it is assumed that the label data held by the first party is Y (for a single training sample, the label can be a single value or a vector), and the part of the features held is denoted as the first feature XA, and the other features held by the second party are denoted as XB. Among them, for a single training sample, it can be described by the first feature XA, the second feature XB, and the label Y. According to different business scenarios, the specific feature items represented by the first feature and the second feature can also be different, and each of them can correspond to one or more business feature items. For example, if the first party is a lending platform, the first feature XA may include at least one feature item such as the borrowing amount, borrowing frequency, repayment frequency, etc., and the label data of whether there has been a default. If the second party is a bank, the second feature may include at least one feature item such as the deposit frequency, deposit amount, total deposit amount, average annual income, etc.

[0056] The first party and the second party can perform sample alignment based on the unique sample identifier of the training sample, that is, perform private set intersection (PSI), or determine the training sample and its sorting. The sample alignment process can be completed uniformly before starting training, or can be carried out simultaneously with the training process. At this time, before the end of a model update cycle, at least one batch (battle) of training samples can be determined for the next model update cycle.

[0057] Figure 2 The process shown describes the update process of the undetermined parameters of the model in a model update cycle. Among them, the jointly trained model here can include a first local model MA set in the first party (for example, an embedding model, hereinafter referred to as the first local model, which is used to process part of the feature data held by the first party), a second local model MB set in the second party (for example, an embedding model, hereinafter referred to as the second local model, which is used to process part of the feature data held by the second party), and a third model ML held by the first party (for example, the loss calculation model or the prediction module in the loss calculation model mentioned above, hereinafter referred to as the third model, which is used to make predictions based on the fusion result of the processing results of the first local model and the second local model). Among them, the first local model MA can process the first feature XA, and obtain a tensor of a predetermined dimension (such as m) for each training sample. The second local model MB can process the second feature XB, and also obtain a tensor HB of this predetermined dimension (such as m) for each training sample. Here, the output tensor HA of the first local model MA and the output tensor HB of the second local model MB have the same dimension, and the purpose is to facilitate the fusion of HA and HB in a superimposed fusion form. Optionally, the third model can also include a loss function calculation module, which is used to determine the model loss only during the model training process by comparing the label data with the prediction result of the prediction module.

[0058] Based on the above basic background, the following combines Figure 2 The timing diagram of the process shown describes the operations performed by the first party and the second party respectively, and the interaction process in a model update cycle of the current batch of sample data. Among them, the current batch of sample data is determined by the private intersection of the first party and the second party.

[0059] First, in step 3110, the first party uses the first local model MA to process the first feature XA of the current batch of training samples, and obtains an intermediate tensor HA. The first local model MA can be various reasonable machine learning models suitable for processing the first feature XA. For example, it can include at least one of a fully connected neural network, a recurrent neural network, etc. The first local model MA can output a tensor of a predetermined dimension (such as m, where m is an integer greater than or equal to 1) for a single training sample. This predetermined dimension tensor can be a row vector or a two-dimensional tensor. For the n (n is an integer greater than or equal to 1) training samples in the current batch, n predetermined dimension tensors are output, and these vectors are arranged side by side according to the training sample dimension, forming a two-dimensional tensor or a three-dimensional tensor. For example, n m-dimensional row vectors form an n×m intermediate tensor HA, and n two-dimensional tensors form a three-dimensional tensor with n values in the training sample dimension as the intermediate tensor HA.

[0060] Then, in step 3120, the first party encrypts the intermediate tensor HA with the first key Pk in the public-private key pair (Sk, Pk) to obtain the first ciphertext tensor of the intermediate tensor HA <ha>。

[0061] It can be understood that the public-private key pair (Sk, Pk) can be determined based on an asymmetric encryption algorithm. For a public-private key pair determined based on asymmetric encryption, one that is publicly disclosed for encrypting data is the public key, and one that is held by oneself for decrypting is the private key. Generally, the public key is generated based on the private key, and the private key can be used to decrypt the data encrypted via the public key, and the public key cannot be used to reverse-engineer the private key. In some implementations, the public key and the private key described above can be swapped. In this specification, for convenience, the key in the public-private key pair that is used to encrypt data and can be publicly disclosed to other parties is referred to as the first key, such as Pk in (Sk, Pk), and the other key that is used to decrypt data and is held locally is referred to as the second key, such as Sk in (Sk, Pk). The public-private key pair (Sk, Pk) in the current update cycle can be pre-generated or generated in the current step 2120. Additionally, the public-private key pairs in each update cycle can be the same or different, and this specification does not make any limitations in this regard.

[0062] Thus, in order to prevent the local data HA from being leaked to other training members, the first party can use the first key Pk to encrypt the intermediate tensor HA to obtain the first ciphertext tensor <ha>。The encryption process is a conventional technology in this field and will not be elaborated here.

[0063] After that, through step 3130, the first party can obtain the first ciphertext tensor <ha>Provided to the second party.

[0064] Similarly to the first party, the second party can, via step 3210, process the second feature XB of the current batch of training samples using the second local model MB, thereby obtaining the intermediate tensor HB. Similarly, the second local model MB can be various reasonable machine learning models suitable for processing the second feature XB, and can include, for example, at least one of a fully connected neural network, a recurrent neural network, etc. The number of elements of the intermediate tensor HB is the same as that of the intermediate tensor HA, for example, it is also an n×m two-dimensional tensor formed by arranging n m-dimensional row vectors for n training samples together as the intermediate tensor.

[0065] Upon receiving the first ciphertext tensor provided by the first party <ha>In the case of, the second party can use the scrambling algorithm S to determine the first ciphertext tensor through step 3220 <ha>The out-of-order fusion result of the intermediate tensor HB in the encrypted state <hs>。

[0066] Among them, the second party can encrypt the intermediate tensor HB using the first key Pk in the key pair generated by the first party to obtain the second ciphertext tensor <hb>Among them, the first secret key Pk can be provided by the first party to the second party in advance, or the first ciphertext tensor can be provided in step 3130 <ha>while being provided to a second party at the same time. Therefore, the process of the second party encrypting the intermediate tensor HB using the first key Pk can be carried out after obtaining HB in step 3210. This process can occur after receiving the first ciphertext tensor <ha>Prior to that, it may also occur upon receiving the first ciphertext tensor <ha>After that, this specification does not limit this. In the case of encryption using the asymmetric encryption method, even if the second party obtains the encryption key Pk and the first ciphertext tensor <ha>, nor can the initial data HA be inferred.

[0067] The purpose of the scrambling algorithm S is to scramble the data arrangement order of different training samples. For example, scramble the middle tensors HA, HB, or their ciphertexts in the training sample dimension to prevent the second party from leaking the data of the middle tensor HB to the first party. Assume that both the middle tensors HA and HB are two-dimensional tensors, and each row describes a training sample. Then the scrambling algorithm S can be an algorithm for scrambling rows. For example, the scrambling algorithm S can be implemented through a scrambling matrix. Through the multiplication operation of the scrambling matrix and the middle tensor HA or HB, the purpose of row scrambling is achieved. In this way, since one row of data represents one training sample, after row scrambling, each row still corresponds to each training sample, the data within the row remains unchanged, and the data arrangement order of different rows is swapped. The scrambling algorithm can also be applied to multi-dimensional tensors. For example, for a three-dimensional tensor, the data of one training sample may be represented by a matrix. Then, after being processed by the scrambling matrix S, the arrangement order of each matrix corresponding to each training sample is swapped, while the data inside the matrix remains unchanged..

[0068] It can be understood that both the encryption and scrambling of the middle tensor HB are performed by the second party. Therefore, the second party can first scramble the middle tensor HB using the scrambling algorithm S and then encrypt the scrambled result using Pk, or first encrypt HB using Pk and then scramble it through the scrambling algorithm S. This specification does not make any restrictions on this. For the sake of description, regardless of whether the second party encrypts or scrambles the middle tensor HB first, the result is collectively referred to as the first ciphertext tensor <hb>The scrambled result. For the intermediate tensor HA, the first ciphertext tensor can be processed through the scrambling algorithm <ha>Perform shuffling.

[0069] If the second party directly provides the shuffled data of the first ciphertext tensor and the second ciphertext tensor to the first party, the shuffling algorithm S and the data of the intermediate tensor HB may be leaked. Therefore, the second party can perform operations on the first ciphertext tensor <ha>and the second ciphertext tensor <hb>Fusion is performed on the shuffled data in a superposition manner (the present specification does not support methods such as splicing that may expose the shuffled data of the first ciphertext tensor and the second ciphertext tensor). Among them, the superposition fusion method may include, but is not limited to, one of summation, averaging, weighted averaging, etc. Among them, summation and averaging can also be regarded as special weighted averaging. Similarly, for the first ciphertext tensor <ha>and the second ciphertext tensor <hb>The shuffling and fusion are both executed by the second party. Therefore, either the fusion or the shuffling can be performed first, and this specification does not make any restrictions on this. Among them, in the case of prior fusion, generally, the HB needs to be encrypted first to obtain the second ciphertext tensor <hb>, then perform fusion and scrambling.

[0070] The result of fusion through the superposition fusion method and then scrambling can be called the scrambled fusion result, denoted as Hs. Since the fusion result is in ciphertext at this time, it can be denoted as <hs>。The second party may, in step 3230, use the encrypted shuffled fusion result <hs>Provided to the first party. Due to this scrambled fusion result <hs>After shuffling and the first ciphertext tensor <ha>and the second ciphertext tensor <hb>The fusion of the superimposing method, the first party cannot determine based on the scrambled fusion result <hs>The guessed scrambled algorithm S or the second ciphertext tensor of the second party <hb>。

[0071] Thus, in step 3140, the first party can process the decrypted <hs>Obtain the scrambled fusion result Hs of the plaintext and the predicted result Ys based on the scrambled training samples.

[0072] Since the order of the training samples is scrambled in the predicted result Ys, when determining the model loss, it is necessary to compare with the correct training sample labels, so it is necessary to use the scrambled label data. Suppose the normal label data of n training samples is denoted as Y, then the scrambled label data can be denoted as Ys'. The scrambling algorithm S is held by the second party and cannot be leaked, and this scrambling process needs to be implemented in cooperation with the second party. Figure 4 Shows the process of the first party and the second party cooperating to scramble the label data Y.

[0073] As Figure 4 shown, the scrambling process of the label data Y may include: Step 410, the first party encrypts the label data Y with the first key Pk to obtain the label encrypted data <y>; Step 420, the first party provides the label encrypted data to the second party in the first direction <y>; Step 430, the second party uses the scrambling algorithm S to determine the labeled encrypted data <y>The shuffled ciphertext <S*Y> = S* <y>, where "S * a" represents the result of the action of the scrambling algorithm S on the data a; Step 440, the second party sends the scrambled ciphertext <S * Y> to the first party; Step 450, the first party decrypts <S * Y> with the second key Sk to obtain the scrambled label data Ys' = S * Y.

[0074] It should be noted that in the case where the second party has generated the scrambling algorithm S, Figure 4 The process shown can be executed in parallel with Figure 3 each step that occurs before Step 3150 shown, or executed one after another. This specification does not make any limitation on this. Among them, in an optional embodiment, the second party can generate different scrambling algorithms S for each model update cycle to prevent the first party from inferring the scrambling algorithm S based on the change of the label data after multiple model update cycles.

[0075] In this way, through Step 3150, the first party can compare the prediction result Ys scrambled based on the scrambling algorithm S with the label data Ys' scrambled based on the scrambling algorithm S, so as to determine the model loss of the current update cycle. Since the arrangement order of the training samples remains consistent after being scrambled by the same scrambling algorithm S, and the data processing process of the prediction model ML remains in plaintext, the computational complexity is greatly reduced. Among them, the loss function can be determined based on the variance, cross-entropy, vector similarity, etc. between the prediction result Ys and the label data Ys', which will not be elaborated here.

[0076] Furthermore, according to Step 3160, the first party determines the gradients of each undetermined parameter in the third model ML and the gradient of the scrambled fusion result Hs through the reverse propagation of the gradient.

[0077] It can be understood that since the third model ML is set in the first party, therefore, the gradients corresponding to each undetermined parameter can be determined by the model loss. According to each gradient, the gradient descent method and other parameter update methods related to the gradient can be used to update each undetermined parameter in the third model ML. This update process can be directly performed after determining the corresponding gradient, or can be executed together with the update process of MA in the subsequent process, which is not limited here.

[0078] The shuffled fusion result Hs is the input data of the third model ML and also the shuffled fusion result of the intermediate tensors HA and HB output by the first local model MA and the second local model MB. During the process of backpropagation of gradients, to determine the gradients of the intermediate tensors HA and HB, the gradient of the shuffled fusion result Hs can be determined first. As the input data of the model ML, the shuffled fusion result Hs can be used by the first party to determine the corresponding gradient matrix through the backpropagation process of the model ML, denoted as Gs. This Gs is the gradient matrix based on the processing result of the shuffling algorithm S. Since both the fusion process and the shuffling algorithm S are determined by the second party, in step 3170, the first party can send the gradient matrix Gs of the fusion result Hs to the second party.

[0079] Next, in step 3230, the second party can process the gradient matrix Gs with the inverse operation of the shuffling algorithm S to restore the gradient matrix G before the data arrangement order of the training samples is shuffled.

[0080] Among them, the inverse operation of the shuffling algorithm S is denoted as S -1 , and the processing process of this inverse operation on the gradient matrix Gs is denoted as S -1 Gs = G. In the case where the shuffling algorithm S is implemented via a shuffling matrix, this inverse operation S -1 can be the inverse matrix of the shuffling matrix. This gradient matrix G is the gradient of the fusion result after restoring the training sample order for HA and HB.

[0081] Next, in step 3240, the second party can sparsify the gradient matrix G to obtain the sparsified matrix Gp and determine the gradient data of HA and HB.

[0082] According to the principle described above, considering that directly feeding back the gradient matrix G to the first party will disclose the inverse operation S -1 , which also means disclosing the shuffling algorithm S, thus possibly disclosing the intermediate tensor HB, which is contrary to the technical problem to be solved by the technical solution of this specification. Therefore, this specification uses the sparsified matrix Gp of the gradient matrix G as the basis for backpropagation of backward gradients. The sparsification of a matrix is usually a process of setting a part of the elements to 0. Assuming the coefficient ratio is p (for example, a value between the preset 0.5 and 1), then p portions of the elements in the gradient matrix G can be retained, and the other elements are set to 0 to obtain the sparsified gradient matrix Gp. It can be understood that the matrix sparsification method can include, for example, but is not limited to at least one of random sparsification, top K (retaining the K elements with the largest values) sparsification, etc. To protect data privacy, the random sparsification method can be used to sparsify the gradient matrix G.

[0083] It can be understood that during the backward propagation of gradients, the gradient propagation process between intermediate tensors is usually based on gradient operations. For example, in a fully connected neural network, the gradient of a to-be-determined parameter W in a certain layer is the product of the Jacobian matrix of the output of this layer with respect to the to-be-determined parameter W and the gradient of the model loss with respect to the output result of this layer. Therefore, by setting some elements in the gradient matrix of the model loss with respect to the hidden layer output to 0, and through matrix multiplication operations, gradient data with controllable errors for other to-be-determined parameters can be obtained through backward propagation. The controllable error here can be adjusted via the sparsification ratio p. Generally, if the sparsification ratio is too small, such as p = 0.1, only 0.1 portion of the gradient data in the gradient matrix G is retained as valid data, and the model convergence speed will be greatly reduced. For this reason, the value of the sparsification ratio p can be reasonably determined according to the specific situation of gradient propagation. For example, it can be ensured that there is at least one 0 in each column. The value of the sparsification ratio p can be determined empirically, such as a value between 0.5 and 1.

[0084] In this way, the second party can use the sparsified matrix Gp to determine the gradient data of the model loss with respect to HA and HB respectively according to the fusion method of HA and HB. Since the gradients of the fusion result H (the fusion result after restoring the disordered order of Hs) with respect to HA and HB in the superposition mode are related to the superposition fusion coefficients, the gradients of the model loss with respect to HA and HB can be determined based on the gradient matrix G of the model loss with respect to the fusion result H. For example, in the case where the superposition fusion of HA and HB is a direct sum, the gradient G serves as the gradient of the model loss with respect to both HA and HB at the same time.

[0085] Furthermore, on the one hand, through step 3250, the second party uses the gradient of the model loss with respect to HB to backpropagate the gradient data in the model MB, thereby determining the gradients of each to-be-determined parameter in the model MB, and then updating each to-be-determined parameter in MB.

[0086] On the other hand, the first party can determine the Jacobian matrix of each layer of the neural network in MA according to the structure of the model MA. And in step 3300, the first party performs homomorphic encryption secure multiplication based on the gradient of HA and the Jacobian matrices of the last hidden layer in the model MA with respect to the to-be-determined parameters and the input data respectively, to determine the gradients of each to-be-determined parameter in the last hidden layer of the model MA, and the gradient of the input data of the last hidden layer. Among them, the input data of the last hidden layer is the output result of the penultimate hidden layer. For example, the output result of the last hidden layer is y q = W q x q + c, where x q = y q-1 . Among them, W q and c are the to-be-determined parameters of the last hidden layer, x q is the input data of the last hidden layer, y q-1 This is the output result of the penultimate hidden layer. Therefore, the gradient of the input data of the last hidden layer is also the gradient data of the output result of the penultimate hidden layer. The gradients of the various undetermined parameters in the last hidden layer are, for example, the gradient of the model loss with respect to W q and the gradient of the input data of the last hidden layer is, for example, the gradient of the model loss with respect to x q .

[0087] It can be understood that the Jacobian matrix is a matrix formed by arranging the first-order partial derivatives of a function in a certain way. Therefore, the Jacobian matrix is a gradient matrix. For example, for a fully connected neural network layer, the product between the Jacobian matrix for the undetermined parameters (such as ) and the output matrix of the fully connected neural network (such as y q ) can form the gradient matrix of the undetermined parameters. Therefore, according to the progressive rule in the gradient transmission process, the product between the Jacobian matrix of a certain hidden layer in MA for the undetermined parameters and the gradient matrix of the model loss with respect to the output matrix of this layer is the gradient of the model loss with respect to the undetermined parameters of this layer, and the product between the Jacobian matrix of a certain hidden layer in MA for the input parameters (such as ) and the gradient matrix of the model loss with respect to the output matrix of this layer (such as y q ) is the gradient of the model loss with respect to the input parameters of this layer (which can correspond to the output matrix of the previous layer).

[0088] Since, although sparsification has been performed, directly transmitting Gp or the gradient for HA determined based on Gp to the first party may still disclose the scrambling algorithm S. Therefore, the first party and the second party perform a homomorphic encryption-based secure multiplication on the gradient of HA and the Jacobian matrices of the last layer in the model MA with respect to the undetermined parameters and input data respectively. The first party encrypts each Jacobian matrix with the first key Pk and sends it to the second party. The second party obtains the encrypted gradient data of the various undetermined parameters in the last hidden layer and the encrypted gradient data of the input data of the last hidden layer through matrix multiplication in the ciphertext state. The second party sends the two encrypted gradient data to the first party, and the first party decrypts them to obtain the gradient data of the various undetermined parameters in the last hidden layer and the gradient data of the input data of the last hidden layer (i.e., the gradient data of the output result of the penultimate hidden layer).

[0089] Furthermore, in step 3180, the first party can determine the gradient data of the other undetermined parameters of the first local model MA based on the backpropagation of the gradient data of the input data of the last hidden layer of the first local model MA in MA, so as to update the various undetermined parameters in the first local model MA. It should be noted that Figure 3 In the operation process shown, some steps are described by dashed lines. The order of appearance of these steps in the whole process is usually optional, and they can generally be executed at any feasible moment after the required data is prepared and before using the corresponding operation results. This specification does not make any restrictions in this regard. For example, the operation "encrypt HB with Pk to obtain <hb>", it can be received by the second party after receiving Pk and obtaining HB by executing step 3210, and can be performed at any time before obtaining the shuffled fusion result of HA and HB in step 3220, such as being included in step 3220.

[0090] Reviewing the above process, in the joint split learning process carried out by two training members based on multi-party secure computing, based on the characteristic that the data is vertically split, the training member holding the label generates a public-private key pair and discloses the encryption key to other training members, and other training members hold the shuffling algorithm. By combining the shuffling algorithm, only the output result of the last layer (hidden layer output result) of the local model is encrypted for calculation, and the rest of the layer models are calculated in plaintext. The ordinary member fuses the hidden layer output results of the two parties in the ciphertext state, and the label training member recovers the plaintext data of the embedding tensor for the two-party shuffled fusion. Plaintext calculation is performed during the prediction process, which can accelerate the calculation process, save calculation time, and ensure the accuracy of the model. The label training member cannot obtain the plaintext output result of the local model of the ordinary training member, and the ordinary member cannot obtain the plaintext label of the label member, which can effectively protect the data value of the ordinary training member and the label privacy data of the label training member. And the gradient of the backpropagation is encrypted by using the method of random sparsification to prevent the label training member from obtaining the shuffling algorithm and ensure that the data value is not leaked.

[0091] According to another aspect, this specification also provides corresponding devices for each training member in the process of jointly updating the model by two parties to complete the corresponding process of jointly training the model. Figure 5 Shows the device 500 of the first party provided in Figure 3 According to an embodiment. As Figure 5 shown, the device 500 includes:

[0092] A processing unit 51, configured to process the first feature XA corresponding to n training samples by using the first local model MA to obtain an intermediate tensor HA composed of n tensors with a predetermined dimension;

[0093] An encryption unit 52, configured to encrypt the intermediate tensor HA by using the first key Pk in the locally held asymmetric key pair, and obtain the first ciphertext tensor <ha>Provided to a second party for the second party to feedback the encrypted ciphertext shuffled fusion tensor based on the first key Pk <hs>, wherein a first secret key Pk is publicly disclosed by a first party to a second party, and a ciphertext shuffled fusion tensor <hs>Via the first ciphertext tensor <ha>It is obtained by fusing the intermediate tensor HB in a superimposed and fused manner in the encrypted form of the training samples in disorder according to the disorder algorithm S, and the intermediate tensor HB is obtained by the second party processing the second feature XB corresponding to n training samples using the second local model MB;

[0094] The prediction unit 53 is configured to process the encrypted disorder-fused tensor by using the second key Sk in the asymmetric key pair to decrypt it by using the third model ML <hs>The obtained shuffled fusion tensor Hs, and n predicted tensors Ys after shuffling the model for n training samples are obtained;

[0095] The loss determination unit 54 is configured to compare the n predicted tensors Ys with the shuffled labels Ys' corresponding to n training samples respectively, so as to determine the model loss, wherein the shuffled labels Ys' are obtained through secure shuffling with the second party based on the first secret key Pk and the shuffling algorithm S;

[0096] The adjustment unit 55 is configured to update the undetermined parameters in the first local model MA and the third model ML with the goal of reducing the model loss, wherein the undetermined parameters in the third model ML are updated by the gradient determined based on the model loss, and the undetermined parameters in the first local model MA are updated by the gradient data determined based on the secure computation of homomorphic encryption with the second party.

[0097] Correspondingly, Figure 6 shows the device 600 of the second party provided in Figure 3 according to an embodiment. As Figure 6 shown, the device 600 includes:

[0098] The processing unit 61 is configured to process the second feature XB corresponding to n training samples by using the second local model MB to obtain an intermediate tensor HB composed of n tensors with a predetermined dimension;

[0099] The fusion unit 62 is configured to be based on the first secret key Pk received from the first party and the first ciphertext tensor <ha>, determine the encrypted state disordered fusion tensor <hs>and feed it back to the first party, where the encrypted disordered fusion tensor <hs>Via the first ciphertext tensor <ha>It is obtained by fusing the intermediate tensor HB in a superimposed and fused manner in the encrypted state of the training samples in a scrambled order according to the scrambling algorithm S, and the first key Pk and the second key Sk held by the first party are an asymmetric key pair;

[0100] The receiving unit 63 is configured to receive, from the first party, the gradient data Gs of the model loss with respect to the scrambled fusion tensor Hs, where the scrambled fusion tensor Hs decrypts the encrypted scrambled fusion tensor based on the second key Sk <hs>It is obtained that the model loss is determined by comparing the prediction result Ys of the third model ML processing the shuffled fusion tensor Hs by the first party with the shuffled labels Ys' of n training samples;

[0101] The shuffled recovery unit 64 is configured to restore the gradient data Gs to the gradient data G based on the inverse operation S of the shuffling algorithm S, and the gradient data G describes the gradient of the model loss with respect to the fusion tensor H; -1 The adjustment unit 65 is configured to use the gradient data G to determine the respective gradients corresponding to the respective undetermined parameters in the second local model MB, and thus update the respective undetermined parameters in the second local model MB using the respective gradients.

[0102] It should be noted that

[0103] The apparatuses 500 and 600 shown Figure 5 and Figure 6 correspond to the first party and the second party in the method embodiments shown, and thus cooperate with each other to complete Figure 3 and Figure 4 the process of the two-party joint update of the model in. Therefore, Figure 3 and Figure 4 the relevant descriptions of the first party and the second party in can also be respectively applied to Figure 3 the apparatus 500 shown Figure 4 and the apparatus 600 shown, which will not be elaborated here. Figure 5 Figure 6

[0104] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method described in connection with Figure 3 or Figure 4 etc. regarding the first party or the second party.

[0105] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor, where the memory stores executable code, and when the processor executes the executable code, the method described in connection with Figure 3 or Figure 4 etc. regarding the first party or the second party is implemented.

[0106] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0107] ​​The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the technical concept of this specification. It should be understood that the above description is only the specific embodiments of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of this specification shall be included within the protection scope of the technical concept of this specification.< / hs> < / ha> < / hs> < / hs> < / ha> < / hs> < / ha> < / hs> < / hs> < / ha> < / hb> < / y> < / y> < / y> < / y> < / hs> < / hb> < / hs> < / hb> < / ha> < / hs> < / hs> < / hs> < / hb> < / hb> < / ha> < / hb> < / ha> < / ha> < / hb> < / ha> < / ha> < / ha> < / ha> < / hb> < / hs> < / ha> < / ha> < / ha> < / ha> < / ha> < / hs> < / ha> < / hs> < / hs> < / ha> < / hs> < / ha> < / hs> < / hs> < / ha> < / y> < / y> < / hs> < / ha> < / hs> < / hs> < / ha> < / y> < / y> < / y> < / hs> < / ha> < / hs> < / hs> < / ha>

Claims

1. A method for jointly updating a model, which is used for the first party and the second party to jointly train a model, wherein, The first party holds the first feature and label data of the training samples, and the second party holds the second feature of the training samples. The model includes a first local model MA and a third model ML located at the first party, and a second local model MB located at the second party. The first local model MA and the second local model MB are respectively used to process the first feature and the second feature to obtain tensors of a predetermined dimension as corresponding intermediate results. In the current update cycle, the first party and the second party have n training samples with the same arrangement order in the current batch. The method is executed by the first party and includes: Using the first local model MA to process the first feature XA corresponding to the n training samples, and obtaining an intermediate tensor HA composed of n tensors of a predetermined dimension; Encrypt the intermediate tensor HA with the first key Pk in the locally held asymmetric key pair, and obtain the resulting first ciphertext tensor <ha>Provided to a second party for the second party to feedback a ciphertext shuffled fusion tensor encrypted based on the first key Pk <hs>, wherein the first secret key Pk is publicly disclosed by the first party to the second party, and the encrypted scrambled fusion tensor <hs>Via the first ciphertext tensor <ha>It is obtained by fusing with the intermediate tensor HB in a superimposed fusion manner in the encrypted state of the training samples in disorder and encrypted by the first secret key Pk. The intermediate tensor HB is obtained by the second party using the second local model MB to process the second feature XB corresponding to the n training samples;< / ha> < / hs> < / hs> < / ha> Process the decrypted ciphertext shuffled fusion tensor using a third model ML via a second key Sk in the asymmetric key pair <hs>Obtaining the disordered fusion tensor Hs, and obtaining n prediction tensors Ys after disordering of the model for the n training samples;< / hs> Comparing the n prediction tensors Ys with the disordered labels Ys' corresponding to the n training samples respectively to determine the model loss. Among them, the disordered labels Ys' are obtained through secure disordering by the first party and the second party based on the first secret key Pk and the disorder algorithm S; Updating the undetermined parameters in the first local model MA and the third model ML with the goal of reducing the model loss. Among them, the undetermined parameters in the third model ML are updated by the gradient determined based on the model loss, and the undetermined parameters in the first local model MA are updated based on the gradient data determined by the secure calculation of homomorphic encryption with the second party; 2. The method according to claim 1, wherein The n training samples are determined in advance by the first party and the second party through private set intersection via the sample identifier that can uniquely describe the training samples; 3. The method according to claim 1, wherein The disorder algorithm S is a disorder matrix generated for the current update cycle, and this disorder matrix is used to disrupt the data arrangement order of different training samples; 4. The method according to claim 1, wherein, The superimposed fusion method is one of addition, averaging, and weighted average; 5. The method according to claim 1, wherein The disordered label Ys' is determined in the following way: Encrypt the n pieces of label data corresponding to n training samples with the first key Pk to obtain ciphertext labels <y> ;< / y> Provide the ciphertext label to the second party <y>for a second party to feed back ciphertext tags processed by the scrambling algorithm S <y>Obtaining the ciphertext disordered label <Ys'>;< / y> < / y> Using the second secret key Sk to decrypt the ciphertext disordered label <Ys'> to obtain the disordered label Ys'; 6. The method according to claim 1, wherein, The undetermined parameters in the first local model MA are updated in the following way: Provide the gradient data Gs of the model loss with respect to the shuffled fusion tensor Hs to a second party for the second party to perform the inverse operation S of the shuffled algorithm S -1 Restore the gradient data Gs to gradient data G in the order of training samples, where the gradient data G describes the gradient of the model loss with respect to the fusion tensor H; Based on the Jacobian matrices of the first local model MA with respect to the undetermined parameters of the last hidden layer and the input data respectively, performing a secure matrix multiplication under homomorphic encryption based on the first secret key Pk with the sparsified matrix Gp of the gradient data G, and obtaining the ciphertext gradient data of the undetermined parameters of the last hidden layer and the ciphertext gradient data of the input data of the last hidden layer at the second party; Receiving the ciphertext gradient data of the undetermined parameters of the last hidden layer and the ciphertext gradient data of the input data of the last hidden layer from the second party; Using the second secret key Sk to decrypt the ciphertext gradient data of the undetermined parameters of the last hidden layer for updating the undetermined parameters of the last hidden layer; Decrypt the encrypted gradient data of the input data of the last hidden layer using the second key Sk, and based on gradient backpropagation, determine the gradients of the undetermined parameters of the other hidden layers except the last hidden layer in the first local model MA, so as to update the undetermined parameters of the other hidden layers.

7. The method according to claim 1, wherein The first feature and the second feature each include at least one service feature extracted from local service data.

8. A method for jointly updating a model, which is used for the first party and the second party to jointly train a model, wherein, The first party holds the first feature of the training sample and the label data, and the second party holds the second feature of the training sample. The model includes a first local model MA and a third model ML provided at the first party, and a second local model MB provided at the second party. The first local model MA and the second local model MB are respectively used to process the first feature and the second feature to obtain tensors of a predetermined dimension as corresponding intermediate results. In the current update cycle, the first party and the second party have n training samples with the same arrangement order in the current batch. The method is executed by the second party and includes: Process the second feature XB corresponding to the n training samples using the second local model MB to obtain an intermediate tensor HB composed of n tensors of a predetermined dimension. Based on a first secret key Pk received from a first party and a first ciphertext tensor <ha>, determine the encrypted state disordered fusion tensor <hs>and feed it back to the first party, where the encrypted disordered fusion tensor <hs>Via the first ciphertext tensor <ha>Is obtained by fusing with the intermediate tensor HB in a scrambled and encrypted form under the first key Pk in a superimposed fusion manner according to the scrambling algorithm S for the training samples, and the first key Pk and the second key Sk held by the first party are an asymmetric key pair. < / ha> < / hs> < / hs> < / ha> Receive gradient data Gs of the model loss with respect to the shuffled fusion tensor Hs from a first party, where the shuffled fusion tensor Hs decrypts the encrypted shuffled fusion tensor based on a second key Sk <hs>Obtained, and the model loss is determined by the first party by comparing the prediction result Ys of the scrambled fusion tensor Hs processed by the third model ML with the scrambled labels Ys' of the n training samples. < / hs> Inverse operation S of the scrambling algorithm S -1 Restore the gradient data Gs to the gradient data G according to the training sample restoration order, where the gradient data G describes the gradient of the model loss with respect to the fusion tensor H; Use the gradient data G to determine the respective gradients corresponding to the respective undetermined parameters in the second local model MB, so as to update the respective undetermined parameters in the second local model MB using the respective gradients.

9. The method according to claim 8, wherein, The scrambling algorithm S is a scrambling matrix generated for the current update cycle, and this scrambling matrix is used to disrupt the data arrangement order of different training samples.

10. The method according to claim 8, wherein, The superimposed fusion method is one of addition, averaging, and weighted averaging.

11. The method according to claim 8, wherein The method further includes determining the scrambled label Ys' for the first party in the following manner: Receiving, from a first party, a ciphertext label obtained by encrypting n pieces of label data corresponding to n training samples using the first key Pk <y> ;< / y> The ciphertext label is processed by the scrambling algorithm S <y>The obtained ciphertext scrambled label <Ys'> is fed back to the first party for the first party to decrypt the ciphertext scrambled label <Ys'> using the second key Sk to obtain the scrambled label Ys'. < / y> 12. The method according to claim 8, wherein, The using the gradient data G to determine the respective gradients corresponding to the respective undetermined parameters in the second local model MB includes: Perform random sparsification with a sparsification rate p on the gradient data G to obtain a sparsified gradient Gp. Use the sparsified gradient Gp to determine the gradient data of the model loss with respect to the intermediate tensor HB. Use the backpropagation of the gradient data of the model loss with respect to the intermediate tensor HB in the second local model to determine the gradients of the respective undetermined parameters in the second local model, so as to update the respective undetermined parameters in the second local model.

13. The method according to claim 12, wherein, The method further includes: Use the sparsified gradient Gp to determine the gradient data of the model loss with respect to the intermediate tensor HA. Perform homomorphic encryption-based secure multiplication with the first party to obtain the encrypted gradient data of the model loss with respect to the undetermined parameters of the last hidden layer of the first local model MA, and the encrypted gradient data of the input data of the last hidden layer. Provide the encrypted gradient data of the model loss with respect to the undetermined parameters of the last hidden layer of the first local model MA and the encrypted gradient data of the input data of the last hidden layer to the first party, so that after the first party decrypts the encrypted gradient data related to the last hidden layer, it can be used to update each undetermined parameter in the last hidden layer of the first local model MA, and after decrypting the encrypted gradient data of the input data of the last hidden layer, it can be used to determine the gradients of the undetermined parameters of other hidden layers except the last hidden layer and update the corresponding other undetermined parameters.

14. An apparatus for two-party joint model update, which is disposed in the first party of the two parties for jointly training a model, wherein, The first party holds the first features and label data of the training samples, and the second party in the two parties of the joint training model holds the second features of the training samples. The model includes the first local model MA and the third model ML located in the first party, and the second local model MB located in the second party. The first local model MA and the second local model MB are respectively used to process the first features and the second features to obtain tensors of a predetermined dimension as the corresponding intermediate results. In the current update cycle, the first party and the second party correspond to n training samples with the same order of arrangement in the current batch. The device includes: A processing unit configured to process the first features XA corresponding to the n training samples by using the first local model MA to obtain an intermediate tensor HA composed of n tensors of a predetermined dimension. An encryption unit, configured to encrypt an intermediate tensor HA with a first key Pk in an asymmetric key pair held locally and obtain a first ciphertext tensor <ha>Provided to a second party for the second party to feedback a ciphertext shuffled fusion tensor encrypted based on the first key Pk <hs>, wherein the first key Pk is publicly disclosed by the first party to the second party, and the encrypted shuffled fusion tensor <hs>Via the first ciphertext tensor <ha>It is obtained by fusing with the intermediate tensor HB in an encrypted state encrypted by the first key Pk in a shuffled manner according to the training samples and in a superposition fusion manner. The intermediate tensor HB is obtained by the second party processing the second features XB corresponding to the n training samples by using the second local model MB. < / ha> < / hs> < / hs> < / ha> A prediction unit configured to process the encrypted shuffled fusion tensor decrypted by a second key Sk in the asymmetric key pair by using a third model ML <hs>The obtained shuffled fusion tensor Hs is used to obtain n predicted tensors Ys after shuffling the n training samples of the model. < / hs> A loss determination unit configured to compare the n predicted tensors Ys with the shuffled labels Ys' corresponding to the n training samples respectively to determine the model loss, where the shuffled labels Ys' are obtained by performing secure shuffling with the second party based on the first key Pk and the shuffling algorithm S. An adjustment unit configured to update the undetermined parameters in the first local model MA and the third model ML with the goal of reducing the model loss. Among them, the undetermined parameters in the third model ML are updated by the gradients determined based on the model loss, and the undetermined parameters in the first local model MA are updated based on the gradient data determined by the secure calculation of homomorphic encryption with the second party.

15. An apparatus for two-party joint model update. The two parties for jointly training the model are the first party and the second party. The first party holds the first features of the training samples and the label data, and the second party holds the second features of the training samples. The model includes a first local model MA and a third model ML disposed at the first party, and a second local model MB disposed at the second party. The first local model MA and the second local model MB are respectively used to process the first features and the second features to obtain tensors of a predetermined dimension as corresponding intermediate results. In the current update cycle, the first party and the second party have n training samples with the same current batch permutation order. The apparatus is disposed at the second party and includes: A processing unit configured to process the second features XB corresponding to the n training samples by using the second local model MB to obtain an intermediate tensor HB composed of n tensors of a predetermined dimension; A fusion unit, configured to be based on a first secret key Pk received from a first party and a first ciphertext tensor <ha>, determine the encrypted state disordered fusion tensor <hs>and feed it back to the first party, where the encrypted state shuffled fusion tensor <hs>Via the first ciphertext tensor <ha>Is obtained by fusing with the intermediate tensor HB in a superimposed fusion manner in a ciphertext form encrypted by the first secret key Pk according to the training sample scrambling by the scrambling algorithm S. The first secret key Pk and the second secret key Sk held by the first party are an asymmetric key pair; < / ha> < / hs> < / hs> < / ha> A receiving unit, configured to receive, from a first party, gradient data Gs of a model loss with respect to a shuffled fusion tensor Hs, where the shuffled fusion tensor Hs is obtained by decrypting a ciphertext shuffled fusion tensor based on a second key Sk <hs>Is obtained. The model loss is determined by comparing the predicted result Ys of the scrambled fusion tensor Hs processed by the third model ML by the first party with the scrambled labels Ys' of the n training samples; < / hs> An out-of-order recovery unit, configured to recover the gradient data Gs to gradient data G based on the inverse operation S of the out-of-order algorithm S, where the gradient data G describes the gradient of the model loss with respect to the fused tensor H; -1 ​ An adjustment unit configured to use the gradient data G to determine the respective gradients corresponding to the respective undetermined parameters in the second local model MB, so as to update the respective undetermined parameters in the second local model MB by using the respective gradients.

16. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-13.

17. A computing device, comprising a memory and a processor, characterized in that, Executable code is stored in the memory. When the processor executes the executable code, the method according to any one of claims 1-13 is implemented.

Citation Information

Patent Citations

  • Method and device for jointly updating model

    CN113657611A

  • Longitudinal federal learning method and device for business model

    CN114912624A