Method and apparatus for jointly updating a model based on multi-party secure computation

By constructing a bilinear triangular matrix space mapping structure and an oracle-secure mapping, the problem of excessive communication volume in multi-party secure computation is solved, achieving efficient joint machine learning, especially significantly improving training efficiency in deep neural networks.

CN115526309BActive Publication Date: 2025-10-21ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211138452.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2025-10-21
Estimated Expiration
2042-09-19

AI Technical Summary

Technical Problem

In the joint machine learning of multi-party secure computing, how to reduce the communication volume of data processing to alleviate communication pressure and avoid communication congestion, and improve model training efficiency.

Method used

A matrix space mapping structure based on bilinear triangles is constructed. Through bilinear mappings from the first matrix space, the second matrix space to the third matrix space, and from the third matrix space, the second matrix space to the first matrix space, combined with the secure mapping of oracles, secure transformation of the output matrix and calculation of the gradient matrix are achieved, reducing the amount of communication.

Benefits of technology

By reducing communication volume, the efficiency of joint machine learning for secure multi-party computation is improved, and communication complexity is reduced, especially significantly improving training efficiency in deep neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115526309B_ABST
    Figure CN115526309B_ABST
Patent Text Reader

Abstract

The embodiment of the present specification provides a method and device for jointly updating a model based on multi-party secure calculation. In the process of jointly updating the model based on multi-party secure calculation, for a full connection layer included in the model, a bilinear triangle is constructed based on the dimension characteristics of various matrices representing the nature of the business data according to the forward business processing process and the backward gradient determination process. In the bilinear triangle, the mapping of each two matrix spaces to another matrix space is a predetermined bilinear mapping. In this way, after the forward calculation is performed based on the bilinear mapping of the matrix space where the input matrix is located and the space where the parameter matrix is located to obtain the output matrix for the full connection layer, the other matrix in the same matrix space as the output matrix can be obtained through a secure conversion, and the subsequent calculation is performed with less communication amount by associating the other matrix with the input matrix and the parameter matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the field of computer technology, and in particular, to a method and apparatus for a joint update model based on multi-party secure computing. Background Art

[0002] Advances in computer technology have enabled the increasingly widespread application of machine learning in a wide variety of business scenarios. Federated machine learning is a method for joint modeling while protecting private data. For example, when businesses need to collaborate on secure modeling, federated machine learning can be used. This allows for collaborative training of data processing models using each party's data, while fully protecting the privacy of enterprise data. This allows for more accurate and efficient processing of business data. In a federated machine learning scenario, for example, each party can agree on a machine learning model structure (or agreed-upon model), then train locally using private data. The model parameters are aggregated using secure and reliable methods, and each party improves the local model based on the aggregated model parameters. Federated machine learning effectively breaks down data silos and enables multi-party joint modeling while maintaining privacy.

[0003] Federated machine learning can be categorized as centralized and decentralized. Multi-party secure computation (MPC) can be used as a form of decentralized federated machine learning. Multi-party secure computation, also known as secure multi-party computation (MPC), involves multiple parties jointly computing the result of a function without disclosing the input data of each party. The result is publicly available to one or more of the parties involved. Since the efficiency of MPC is related to the amount of communication, reducing data processing traffic, alleviating communication pressure, and avoiding communication congestion during the joint model training process is a key issue in improving model training efficiency. Summary of the Invention

[0004] One or more embodiments of this specification describe a method and apparatus for a joint update model based on multi-party secure computing to solve one or more problems mentioned in the background art.

[0005] According to a first aspect, a method for jointly updating a model based on multi-party secure computing is provided, wherein the model is jointly updated by multiple participants using local business data, the model includes a first fully connected layer, the first fully connected layer corresponds to the input matrix of the first matrix space, the parameter matrix of the second matrix space, and the output matrix of the third matrix space, the first matrix space, the second matrix space, and the third matrix space constitute a bilinear triangle based on a predetermined bilinear mapping, wherein the bilinear mapping from the first matrix space and the second matrix space to the third matrix space is a first mapping, the bilinear mapping from the third matrix space and the second matrix space to the first matrix space is a second mapping, and the bilinear mapping from the third matrix space and the first matrix space to the second matrix space is a third mapping; the method is executed by a first party among the multiple participants, the first party holds first slices w0 and x0 corresponding to the parameter matrix w and the input matrix x, respectively, and the method includes: using the first slice w0 , x0, determine with other parties the output matrix y of the parameter matrix w processing the input matrix x through a bilinear mapping protocol based on the first mapping, and obtain the first slice y0 of the output matrix y; according to the first slice y0 of the output matrix y, complete the forward calculation of the model with other parties, and determine the overall gradient matrix z of the model loss for the output matrix y during the back propagation process, thereby obtaining the first slice z0 of the overall gradient matrix z; according to z0, determine with other parties the first gradient matrix q of the overall gradient matrix z for the feature matrix x through a bilinear mapping protocol based on the second mapping, and obtain the first slice q0 of the first gradient matrix q; and / or: determine the second gradient matrix s of the overall gradient matrix z for the parameter matrix w based on the bilinear mapping protocol based on the third mapping, and obtain the first slice s0 of the second gradient matrix s; perform a secure update of the model according to the first slice q0 of the first gradient matrix q and / or the first slice s0 of the second gradient matrix s.

[0006] In one embodiment, the method further includes: obtaining first perturbation matrices generated by a third party for the first matrix space, the second matrix space, and the third matrix space respectively The second perturbation matrix The third perturbation matrix and balance items The first shard corresponding to each Among them, each balance term is used to eliminate each predetermined bilinear mapping based on the first perturbation matrix The second perturbation matrix The third perturbation matrix Introduced bias.

[0007] In a further embodiment, the first slice w0, x0 is used to determine the output matrix y of the input matrix x processed by the parameter matrix w through a bilinear mapping protocol based on the first mapping with other parties, and obtaining the first slice y0 of the output matrix y includes: using x0, Determine the first slice δx0 of the first perturbation result matrix δx perturbed by the first perturbation matrix for the feature matrix x; using w0, Determine the first slice δw0 of the second perturbation result matrix δw perturbed by the second perturbation matrix for the parameter matrix w; determine the first perturbation result matrix δx and the second perturbation result matrix δw by disclosing the corresponding slices with other participants; locally use the first mapping to process the first perturbation result matrix δx and the first slice w0 of the parameter matrix w, the first perturbation matrix The first shard and the second perturbation result matrix δw, to obtain the first slice of the first mapping matrix and the first slice of the second mapping matrix in the third matrix space; the balance term The first slice of is fused with the first slice of the first mapping matrix and the first slice of the second mapping matrix to obtain the first slice y0 of the output matrix y.

[0008] In a further embodiment, the balancing term The first perturbation matrix is ​​transformed by the first mapping The second perturbation matrix The first slice δx0 of the first perturbation result matrix δx is determined by x0 and The first slice δw0 of the second perturbation result matrix δw is determined by w0 and The difference is determined; the balance item The first slice of the first mapping matrix and the first slice of the second mapping matrix are merged, including: The first slice of the first mapping matrix is ​​superimposed with the first slice of the first mapping matrix and the first slice of the second mapping matrix.

[0009] In another further embodiment, according to z0, the first gradient matrix q of the overall gradient matrix z to the feature matrix x is determined by other parties via a bilinear mapping protocol based on the second mapping, and the first slice q0 of the first gradient matrix q is obtained: using the third perturbation matrix The first shard Perturb the first slice z0 of the overall gradient matrix z to obtain the first slice δz0 of the third perturbation result δz corresponding to the overall gradient matrix z, thereby restoring the third perturbation result δz based on other slices of the third perturbation result δz obtained by other parties; use the second mapping to process the second perturbation result matrix δw and the first slice z0 of the overall gradient matrix z, the second perturbation matrix The first shard and the third perturbation result matrix δz, to obtain the first slice of the third mapping matrix and the first slice of the fourth mapping matrix of the first matrix space; the balance term The first shard The first slice q0 of the first gradient matrix q is obtained by fusing it with the first slice of the third mapping matrix and the first slice of the fourth mapping matrix.

[0010] In yet another further embodiment, when the second gradient matrix s of the overall gradient matrix z to the parameter matrix w is determined according to z0 based on the bilinear mapping protocol of the third mapping, and the first slice s0 of the second gradient matrix s is obtained: using the third perturbation matrix The first shard The first slice z0 of the overall gradient matrix z is perturbed to obtain the first slice δz0 of the third perturbation result δz corresponding to the overall gradient matrix z, thereby restoring the third perturbation result δz based on other slices of the third perturbation result δz obtained by other parties; the third perturbation result matrix δz is processed with the first slice x0 of the feature matrix x and the third perturbation matrix δz respectively by the third mapping. The first shard and the first perturbation result matrix δx, to obtain the first slice of the fifth mapping matrix and the first slice of the sixth mapping matrix of the second matrix space; the balance term The first shard The first slice s0 of the second gradient matrix s is obtained by fusing the first slice s0 of the fifth mapping matrix and the first slice of the sixth mapping matrix.

[0011] In one embodiment, the overall gradient matrix z of the output matrix y is a matrix in the third matrix space, and the overall gradient matrix z is the processing result of the output matrix y based on the secure mapping of the oracle.

[0012] According to a second aspect, a device for jointly updating a model based on multi-party secure computing is provided, wherein the model is jointly updated by multiple participants using local business data, the model including a first fully connected layer, the first fully connected layer corresponding to an input matrix of a first matrix space, a parameter matrix of a second matrix space, and an output matrix of a third matrix space, the first matrix space, the second matrix space, and the third matrix space forming a bilinear triangle based on a predetermined bilinear mapping, wherein a bilinear mapping from the first matrix space and the second matrix space to the third matrix space is a first mapping, a bilinear mapping from the third matrix space and the second matrix space to the first matrix space is a second mapping, and a bilinear mapping from the third matrix space and the first matrix space to the second matrix space is a third mapping;

[0013] The apparatus is provided at a first party among the multiple participants, the first party holding first slices w0 and x0 corresponding to the parameter matrix w and the input matrix x of the first fully connected layer, respectively, and the apparatus includes:

[0014] The service processing unit is configured to use the first slices w0 and x0 to determine, with another party, an output matrix y obtained by processing the input matrix x by the parameter matrix w through a bilinear mapping protocol based on the first mapping, and obtain a first slice y0 of the output matrix y;

[0015] a secure conversion unit configured to complete the forward calculation of the model with other parties based on the first slice y0 of the output matrix y, and determine the overall gradient matrix z of the model loss with respect to the output matrix y during the backward propagation process, thereby obtaining the first slice z0 of the overall gradient matrix z;

[0016] a gradient determination unit configured to determine, based on z0, a first gradient matrix q of the overall gradient matrix z with respect to the feature matrix x via a bilinear mapping protocol based on the second mapping with another party, to obtain a first slice q0 of the first gradient matrix q; and / or: determine, based on the bilinear mapping protocol based on the third mapping, a second gradient matrix s of the overall gradient matrix z with respect to the parameter matrix w, to obtain a first slice s0 of the second gradient matrix s;

[0017] The updating unit is configured to perform a security update of the model according to the first slice q0 of the first gradient matrix q and / or the first slice s0 of the second gradient matrix s.

[0018] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.

[0019] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, the method of the first aspect is implemented.

[0020] Through the method and device provided by the embodiments of this specification, in the joint machine learning process based on multi-party secure computing, for the fully connected layer, according to its forward business processing process and reverse gradient determination process, a bilinear triangle is constructed based on the dimensional characteristics of various matrices that characterize the nature of the business data, wherein the mapping of each matrix space in the bilinear triangle to another matrix space is a predetermined bilinear mapping. In this way, after the output matrix is ​​obtained by forward calculation based on the bilinear mapping of the matrix space where the input matrix is ​​located and the space where the parameter matrix is ​​located for the fully connected layer, other matrices in the same matrix space as the output matrix can be obtained through secure conversion, and the other matrices are used to associate with the input matrix and the parameter matrix, and subsequent calculations are performed with less communication volume. Therefore, in the process of the joint update model of multi-party secure computing, this technical solution constructs a bilinear triangle through the secure conversion of the same matrix space, thereby reducing the communication volume and improving the efficiency of the joint machine learning of multi-party secure computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 This is a schematic diagram of the system architecture of multi-party secure computing;

[0023] Figure 2 A schematic diagram of a specific implementation architecture of a bilinear map is shown;

[0024] Figure 3 A visual diagram showing a bilinear triangle formed by three matrix spaces;

[0025] Figure 4 A schematic diagram illustrating an interactive process of two parties performing calculations on a bilinear triangle according to an embodiment of the present specification;

[0026] Figure 5 A schematic diagram illustrating a process of a joint training model based on multi-party secure computing according to an embodiment of the present specification is provided;

[0027] Figure 6 A schematic block diagram of a device based on a multi-party secure computing joint training model according to an embodiment of the present specification is shown. DETAILED DESCRIPTION

[0028] The solution provided in this specification is described below in conjunction with the accompanying drawings.

[0029] With the advancement of secure computing technology, multiple data holders can now jointly train business models using their local data (e.g., federated learning) in an increasing number of scenarios. Specifically, suppose that Enterprise A and Enterprise B each establish a task model. Individual tasks could be classification or prediction, and these tasks have already been approved by their respective users when the data was acquired. However, due to data incompleteness—for example, Enterprise A lacks labeled data, Enterprise B lacks user profile data, or the data is insufficient, with insufficient sample size to build a good model—then the model on each end may fail to build or perform poorly. Federated machine learning addresses the problem of building high-quality models on each end. This model is trained using data from both enterprises A and B, while keeping each enterprise's proprietary data private. In other words, a shared model is established without violating data privacy regulations. This shared model can be a superior model created by aggregating data from all parties. In this way, the built model serves only its own objectives within each party's region.

[0030] In the scenario of multi-party joint training of business models, a single participant can hold the private data of one or more trusted data holders. This private data can be, for example, business data in various forms, such as characters, images, voice, animation, and video. Typically, the business data held by each participant is related. For example, multiple banks involved in financial services act as business parties. Each business party can independently provide users with services such as savings and loans, and thus hold data such as users' income and expenditure flow, loan limits, and deposit limits. For another example, multiple hospitals involved in medical services act as business parties. Each business party can use the corresponding medical records of users' symptoms, diagnosis results, treatment plans, treatment results, and so on as local business data.

[0031] The multi-party joint training business model can adopt the implementation architecture of multi-party secure computing, such as Figure 1 As shown in the figure, in a multi-party secure computation architecture, each participant can perform computations using methods such as homomorphic encryption, secret sharing, and obfuscated circuits. During the computation process, each participant obtains a shard of the intermediate results or final business processing results. Because each participant cannot disclose local private data and each holds its own shard of the intermediate results, all data exchanged during the computation process is performed securely. This may generate additional data traffic.

[0032] In the process of jointly training business models, it is usually necessary to calculate functions on variables, such as z = f(x, y). In the two-party secure computation process, x and y are data defined on Abelian groups A and B, respectively, and are constructed and shared between the two participants (e.g., denoted as P0 and P1). In this case, when x and y are both in vector form, f can also be called a bilinear map, that is, a function that generates elements of a third vector space (e.g., z) from elements of two vector spaces (e.g., x and y). Here, z is an element of the third vector space, such as defined on the Abelian group C.

[0033] Figure 2 The following figure shows a schematic diagram of a secure computation protocol for a bilinear map. The computation principle of a bilinear map is to add a perturbation to the vector to be processed, calculate the mapping of the perturbed vector, and eliminate the result deviation caused by the perturbation while ensuring data security. For example:

[0034] f(x,y)=f(xu,y)+f(u,yv)+f(u,v)=f(dx,yO)+f(dx,y1)+f(uO,dy)+f(u1,dy)+f(u,v);

[0035] Among them, u and v can be regarded as perturbations to x and y respectively, and dx = xu, dy = yv, and the suffixes 0 and 1 represent the corresponding shards. The mapping result does not depend on the size of the random perturbations u and v. For the convenience of description, the shards on the participant P0 can be represented by the suffix 0, and the shards on the participant P1 can be represented by the suffix 1. It can be understood that

[0036] like Figure 2 As shown, let b = f(u, v) = b0 + b1, u = u0 + u1, v = v0 + v1, then b0, b1, u0, u1, v0, v1 satisfy the constraint b1 + b0 = f(u0 + u1, v0 + v1). In this way, a third party can assist the calculation party in generating 5 of the random numbers and determine another random number based on the constraint. The third party can be a semi-trusted third party, that is, it cannot know the data of the participants, but can provide auxiliary parameters and shared shard splitting for each participant. Therefore, the third party can also be called a pseudo-random number server.

[0037] It is worth noting that when x and y are vectors, u, v, and b are also vectors of corresponding dimensions. A third party can send shards of the corresponding auxiliary parameters to participants P0 and P1. Assuming P0 obtains u0, v0, and b0, and P1 obtains u1, v1, and b1, then P0 can calculate dx0 = x0 - u0, dy0 = y0 - v0, while P1 can calculate dx1 = x1 - u1, dy1 = y1 - v1. In this way, when P0 and P1 disclose dx0, dy0, dx1, and dy1 to each other, at most the plaintext form of dx and dy can be recovered, without revealing the locally held and shared shards of x and y. Afterwards, P0 can locally calculate z0 = f(dx, y0) + f(u0, dy) + b0, and P1 can locally calculate z1 = f(dx, y1) + f(u1, dy) + b1. Then we have: z = z0 + z1 = f(dx, y0) + f(u0, dy) + b0 + f(dx, y1) + f(u1, dy) + b1 = f(dx, y0) + f(u0, dy) + f(dx, y1) + f(u1, dy) + f(u, v) = f(x, y). In other words, the two slices z0 and z1 form a shared form of the function z = f(x, y).

[0038] Figure 2 In the process shown, since u corresponds to x (defined on the Abelian group A), v corresponds to y (defined on the Abelian group B), and b corresponds to z (defined on the Abelian group C), the total communication volume is: In the offline case, the third party sends b0, b1, u0, u1, v0, and v1 to the first and second parties, generating an offline communication volume of 2(log2|A|+log2|B|+log2|C|), and in the online case, P0 and P1 disclose dx0, dy0, dx1, and dy1 to each other, generating an online communication volume of 2(log2|A|+log2|B|). Where |A|, |B|, and |C| are the number of elements in the Abelian groups A, B, and C, respectively. For example, the number of elements in A is 2 64 , then a single element can be represented by a 64-bit binary number, whose size is log2|2 64 =|, which is 64 bits. Since the shared slices are still elements of the corresponding Abelian group, the communication volume of the shared slices is the communication volume of a single element in the Abelian group. When a single element in A is a vector, the size of the single element is the product of the data size of a single dimension in the vector (e.g., 64 bits) and the number of dimensions.

[0039] In the joint machine learning process, there are usually forward propagation and back propagation (gradient determination) processes. For the fully connected neural network layer (Dence), forward propagation is usually the image obtained under the mapping (such as f) for the input feature matrix (hereinafter denoted as x). For example, f is the mapping of the set of matrices of b×n dimensions defined on the Abelian group A to the set of matrices of n×n' dimensions defined on the Abelian group A, then: y=wx+c, where c is the offset matrix. The definition here on the Abelian group A means that each element in the matrix is ​​a number defined on the Abelian group A. In the reverse gradient process, the following properties exist: the gradient of the model loss with respect to the feature matrix x is the gradient of the model loss with respect to y, multiplied by the gradient of y with respect to x (the transpose of the parameter matrix w), where y is the processing result of the parameter matrix w on the feature matrix x, and is also the mapping result of x and w to the output matrix space by the mapping f. Assuming that the model loss is loss, the following relationship is satisfied between the matrices y, x, w and loss: y=x×w,

[0040] It can be seen that in the forward and backward propagation processes, a single fully connected layer involves a set of matrices in three dimensional spaces. The feature matrix x corresponds to a b×n dimensional matrix, for example, representing the n-dimensional feature vectors of each of the b training samples, the parameter matrix w corresponds to an n×n' dimensional matrix, and the output matrix y corresponds to a b×n' dimensional matrix. Assuming that the elements in each matrix are data defined on the Abelian group A, the b×n dimensional matrix, the n×n' dimensional matrix, and the b×n' dimensional matrix can be regarded as being defined in the matrix space M respectively. b , n (A), M n , n' (A), M b , n' (A), where (A) means that each element in the matrix is ​​an element of the Abelian group A. The three matrix spaces can also be viewed as Abelian groups whose elements are matrices.

[0041] In this way, we can b , n (A), M n , n' (A), M b , n' (A) The above definition is as follows Figure 3 The bilinear maps that interact with each other are shown in Figure 1. Since these bilinear maps form a triangular shape, they can also be called a bilinear triangle. This bilinear triangle can correspond to the following three bilinear maps:

[0042] From the matrix space M b , n (A), M n ,n' (A) To the matrix space M b , n' (A) Mapping f:M b , n (A)×M n , n' (A)→M b , n' (A), for example, the matrix space M b , n The matrix x and M in (A) n , n' The matrix w in (A) is mapped to the matrix space M via mapping f b , n' The element y in (A) can be: y = f(x, w) = xw. Similarly, from M n , n' (A), M b , n' (A) Towards M b , n (A) Mapping g:M n , n '(A)×M b , n '(A)→M b , n (A), for example, the matrix space M n , n The matrix w and M in '(A) b , n The matrix in '(A) Mapped to M b , n The matrix in (A) From M b , n '(A), M b , n (A) Towards M n , n Mapping h of '(A):M b , n' (A)×M b , n (A)→M n , n' (A), for example, the matrix space M b , n' The matrix in (A) With the matrix space M b , n The matrix x in (A) is mapped to M n , n The matrix in '(A)

[0043] For the convenience of description, the matrix space M b , n (A) is denoted as X, where any matrix element is denoted as x, and the matrix space M n , n '(A) is denoted as W, any matrix element in it is denoted as w, and the matrix space M b , n '(A) is denoted as Y, and any matrix element in it is denoted as y. Given the matrices x and w stored in a shared form with two participants P0 and P1, in order to calculate through each bilinear mapping f: X×W→Y, g: W×Y→X, h: Y×X→W, considering that the output matrix of the fully connected network is the matrix space M b , n '(A), and in the back propagation process, the gradient matrix of the model loss for the output matrix is ​​consistent with the output matrix dimension, which is also the matrix space M b , n The elements in '(A), such as Here, z and y belong to the same matrix space M b , n '(A). Therefore, we can introduce the oracle r:Y→Y, which is the matrix space M b , n The elements in '(A) are safely mapped to the matrix space M b , n Elements in '(A) (such as mapping y to z).

[0044] Then we have:

[0045] y=f(x,w); z=r(y); q=g(w,z), s=h(z,x).

[0046] Among them, in the propagation process of the fully connected layer, let The oracle r can ensure the safe calculation of the mapping z = r (y). In this manual, since the forward propagation and back propagation of a single fully connected layer are discussed, the calculations before and after the fully connected layer are not discussed here. The mapping z = r (y) provided by the oracle r can be understood as the process of determining the model loss from the output of the current layer of the fully connected network and determining the gradient of the model loss with respect to the output matrix, such as This process may involve processing of other fully connected layers, activation layers, etc., which will not be discussed here. It is understood that the communication volume of the mapping z = r (y) can be the communication volume of the specific function involved in its calculation process. In the discussion of communication volume in this specification, only the communication volume of the fully connected network at the current layer is discussed, and the communication volume of the mapping z = r (y) is not involved.

[0047] If you use Figure 2 The bilinear mapping protocol shown in the figure is used for calculation. First, y is calculated using the bilinear mapping protocol, then the oracle is called to calculate z. Then, in the reverse process, q and s are calculated using the bilinear mapping protocol. The oracle's communication traffic can be ignored, and the bilinear mapping protocol is executed three times. These three bilinear mappings can share data. Therefore, the total communication complexity is the number of bytes of data generated by the offline third party in the corresponding spaces X, W, and Y: log2|X|+log2|W|+log2|Y|. The three online mutually public data (which overlaps and can be merged) is 2(log2|X|+log2|W|+log2|Y|). Compared with traditional methods, the online communication traffic is reduced by half, and the total online and offline communication traffic is reduced by two-fifths.

[0048] In order to reduce the amount of communication during the joint machine learning process, this specification provides a Figure 3 A secure computation protocol for bilinear triangles. The bilinear triangle computation protocol is based on the principle of the bilinear mapping protocol. The principle of the bilinear mapping protocol is:

[0049]

[0050]

[0051]

[0052] in, They are equivalent to the perturbations of x, w, and z (perturbation matrices), and δx, δw, and δz are the corresponding perturbation results, for example In the case of introducing a disturbance, the mapping result balance term is provided to eliminate the mapping disturbance. This principle is only for illustration. In practice, other methods can also be used to add disturbances. For example, the disturbance results of adding a disturbance matrix to x and w are Then the linear mapping protocol can be etc.

[0053] The following describes the principle of the bilinear triangle calculation protocol in this specification.

[0054] refer to Figure 4 FIG. 1 is a flow chart of a calculation protocol for a bilinear triangle according to an embodiment of the present specification. For ease of description, Figure 4The interactive process involving two participants in a computation is shown. In practice, the number of participants in this interactive process can be expanded to multiple. Initially, the first matrix x in the first matrix space and the second matrix w in the second matrix space can be stored in a shared form on each participant. The first matrix x and the second matrix w are mapped to the matrix y in the third matrix space via a first mapping f. The first matrix space, the second matrix space, and the third matrix space form a bilinear triangle through the first mapping, the second mapping, and the third mapping.

[0055] The following combination Figure 4 Describe the relevant process. Take the number of participants as an example. For the sake of convenience, assume that the first party is recorded as P0 and the second party is recorded as P1. In the result of share form, the shard held by the first party is represented by the suffix 0, and the shard held by the second party is represented by the suffix 1. Any data a (such as And so on) the shard a0 held by the first party (corresponding to suffix 0) and the shard a1 held by the second party (corresponding to suffix 1) are stored in a shared form on both parties P0 and P1.

[0056] First, in step 4000, the semi-trusted third party can generate the perturbation matrices corresponding to the three sets X, Y, and W respectively by pseudo-random number generation. Therefore, according to the corresponding mapping, the various balance terms that eliminate the corresponding mapping disturbance can be calculated: Furthermore, each disturbance can be and various balance items Each is split into two and shared shards (share), and distributed to the first party and the second party. For example, the split is sent to the first party P0, and Sent to the second party P1.

[0057] This step 4000 can be performed offline. The so-called offline execution can be understood as an independent execution process that does not rely on various intermediate results in the online business processing process. The offline execution operation can be pre-executed by a third party, so that P0 and P1 receive the data in advance. The corresponding shards are generated and saved locally. This offline operation can also be performed by a third party when bilinear mapping calculation is required during the business processing of P0 and P1. The third party can be any computer, device, or server with a certain data processing capability, which can realize the processing of generating random numbers, calculating balance items, splitting and sharing shards. The communication volume generated by this step is log2|X|+log2|W|+log2|Y| under the PRF mechanism. Among them, under the PRF mechanism, the third party and the participating parties can generate the corresponding random numbers according to the agreed random number generation method, so that the third party does not need to pass the generated random numbers to the participating parties. As for the last data shard under the dependency constraint (such as the mutual constraints between various disturbance items and balance items), it can be determined by the third party based on other random numbers and provided to the corresponding participating parties, thereby only generating the communication volume of the last data shard.

[0058] It can be understood that, combined with the safe calculation principle of bilinear mapping, in order to calculate the mapping f(x, w), P0 and P1 can safely calculate the first matrix x in the perturbation Since a single participant holds the first matrix x and the perturbation Therefore, the perturbation amount can be added or subtracted on a single slice of the first matrix x , thereby obtaining a single shard of the corresponding perturbation result.

[0059] Thus, in step 4101, party P0 can add a perturbation matrix to the slice x0 of the first matrix x. Shard Disturbed shards after disturbance The perturbation amount is added to the slice w0 of the second matrix w Shard Disturbed shards after disturbance Correspondingly, in step 4102, party P1 can calculate

[0060] Next, in step 4200, δx0, δw0, δx1, and δw1 may be disclosed by both parties P0 and P1. Specifically, the disclosure can be mutually disclosed, i.e., P0 discloses δx0 and δw0 and P1 discloses δx1 and δw1, or one party may disclose the corresponding local shard, such as P0 discloses δx0 and δw0, or P1 discloses δx1 and δw1.

[0061] Thus, parties P0 and P1 can reconstruct δx = δx0 + δx1 and δw = δw0 + δw1 without leaking the values ​​of x0, w0, x1, and w1. If the two parties mutually disclose δx0, δw0, δx1, and δw1, parties P0 and P1 can reconstruct δx and δw, respectively. If one party discloses its local shard, the other party can reconstruct and disclose δx and δw. In both cases, the communication volume is the same, 2(log2|X|+log2|W|).

[0062] Furthermore, in step 4301, party P0 can locally calculate a slice of the f mapping, such as y0=[f(x, Similarly, in step 4302, P1 can calculate another slice of the f map, such as Among them, y0 and y1 are the shared forms of the output matrix y on P0 and P1 sides.

[0063] At this point, a bilinear mapping is completed through steps 4101, 4102, 4200, 4301, and 4302. In a fully connected neural network scenario, this process can be a fully connected neural network calculation process of forward propagation. Based on the above bilinear triangle mapping principle, during the backpropagation process, there is no need to completely re-perform a bilinear mapping. Instead, the relevant bilinear mapping results can be calculated based on partial data in each space of the bilinear triangle during the forward propagation process (such as each perturbation slice, etc.).

[0064] In step 4400, P0 and P1 can jointly call the oracle to securely calculate the mapping of matrix y to matrix z in the third matrix space, and obtain two shared slices z0 and z1 of matrix z respectively. In the training process of the model containing the fully connected layer, the matrix z can be the gradient matrix of the model loss with respect to the output matrix y in the back propagation. For example, P0 and P1 securely calculate the slice r(y0+y1)=z0+z1. Furthermore, in step 4501, P0 further calculates a slice of the perturbation result of the matrix z. Correspondingly, in step 4502, party P1 calculates another slice of the perturbation result of z

[0065] Next, in step 4600, P0 and P1 publicly disclose δz0 and δz1 to reconstruct δz = δz0 + δz1, without revealing the privacy of the z0 and z1 shards. The disclosure and reconstruction methods are identical to those in step 4200 and will not be repeated here. This step generates a communication traffic of 2log2|Y|.

[0066] Furthermore, through steps 4701 and 4702, P0 and P1 can respectively calculate the mapping g and h based on δy, so that each party can obtain the local fragments of q = g(w, z) and s = h(z, x). Based on the above bilinear triangle mapping calculation principle, P0 and P1 can each use the local x and w fragments and the public δz to calculate the corresponding fragments of q and s. For example: P0 calculates the fragment of the mapping result q Sharding of the mapping result s P1 calculates the shards of the mapping result q Sharding of the mapping result s Thus, q0 and q1 constitute the sum and sharing of q at P0 and P1, and s0 and s1 constitute the sum and sharing of s at P0 and P1. No communication may occur in step 4700.

[0067] This completes the calculation of the bilinear triangle. The above process uses an oracle to implement matrix mapping in three-dimensional space. The communication volume is: log2|X|+log2|W|+log2|Y| for offline communication, and 2(log2|X|+log2|W|+log2|Y|) for online communication. Figure 3 The bilinear triangle shown in the figure is used for calculation. Compared with the traditional cubic bilinear mapping, the online communication volume is halved and the total online and offline communication volume is reduced by 2 / 5. By extension, in the joint secure computation process involving multiple participants, the use of Figure 4 The communication overhead of the bilinear triangle computation protocol is shown as: log2|X|+log2|W|+log2|Y| for offline, and 2(log2|X|+log2|W|+log2|Y|) for online. The online communication overhead coefficient decreases from n squared to n squared, which is 3n and 6n when X, Y, and W are all n bits.

[0068] In the process of joint training of the business model, it is assumed that the input of the single fully connected layer contained in the model is the feature matrix (including the business features of multiple training samples) x∈X b,n , the parameter matrix is ​​w∈W n,n ', then the feature matrix x, parameter matrix w, and output matrix y∈Y of the single fully connected layer b,n ' can be formed as follows Figure 3 The bilinear triangle shown. The mapping f maps x and w to the output matrix p of the single fully connected layer, and the oracle mapping r securely maps the output matrix y to the overall gradient matrix z∈Y for the fully connected layer. b,n ', mapping g maps z and w to the gradient matrix q∈W of z to x n,n', mapping h maps z, x to the gradient matrix s∈X of z to w b,n .

[0069] Thus, the bilinear triangle method provided in this specification can be effectively used to complete the model update in the joint learning scenario. When the business model of joint learning includes a multi-layer fully connected neural network, the bilinear triangles corresponding to each fully connected layer can be used for relevant calculations. In a deep neural network, when there are at least one of the following conditions: a large number of participants, a large number of fully connected layers, a large number of feature dimensions, and a large number of training samples, the reduction in communication volume is particularly effective in improving the efficiency of joint learning. The "larger" mentioned here can usually be understood as greater than the corresponding threshold.

[0070] According to the above bilinear triangle construction and safety calculation principles, Figure 5 The figure shows a process for a joint update model based on multi-party secure computation, executed by the first of multiple participants. The multiple participants in the model jointly update the model using their local business data. Furthermore, the model includes at least one fully connected layer, wherein any fully connected layer is denoted as a first fully connected layer. The first fully connected layer corresponds to the input matrix of the first matrix space, the parameter matrix of the second matrix space, and the output matrix of the third matrix space. The first matrix space, the second matrix space, and the third matrix space form a bilinear triangle based on predetermined bilinear mappings. Specifically, the bilinear mapping from the first matrix space and the second matrix space to the third matrix space is the first mapping, the bilinear mapping from the third matrix space and the second matrix space to the first matrix space is the second mapping, and the bilinear mapping from the third matrix space and the first matrix space to the second matrix space is the third mapping. Initially, the first party holds the first slices w0 and x0 corresponding to the parameter matrix w and the input matrix x, respectively. The first slice w0 corresponding to the parameter matrix w forms a shared sum with the other slices held by the other participants, and the first slice x0 corresponding to the input matrix x forms a shared sum with the other slices held by the other participants.

[0071] like Figure 5 As shown, the process of the multi-party secure computation joint update model performed by the first party may include the following steps:

[0072] Step 501: Using the first slices w0 and x0, the system determines the output matrix y obtained by processing the input matrix x using the parameter matrix w through a bilinear mapping protocol based on the first mapping. This results in the first slice y0 of the output matrix y. Simultaneously, other slices of the output matrix y are obtained from other participating parties. These slices constitute the sum and shared form of the output matrix y across all participating parties.

[0073] The bilinear mapping protocol processing process based on the first mapping can refer to Figure 2In the bilinear mapping protocol processing process based on the first mapping, the first party can also use x0, Determine the first slice δx0 of the first perturbation result matrix δx perturbed by the first perturbation matrix for the feature matrix x, and use w0, Determine the first slice δw0 of the second perturbation result matrix δw perturbed by the second perturbation matrix for the parameter matrix w, and determine the first perturbation result matrix δx and the second perturbation result matrix δw with other participants by disclosing the corresponding slices.

[0074] Step 502, based on the first slice y0 of the output matrix y, complete the forward calculation of the model with other parties, and determine the overall gradient matrix z of the model loss for the output matrix y during the back propagation process, thereby obtaining the first slice z0 of the overall gradient matrix z.

[0075] The overall gradient matrix mentioned here can be understood as a matrix that describes the gradient data of the model loss with respect to the entire current fully connected layer. z and y are consistent matrices in the third matrix space. Since this specification uses a single fully connected layer as an example to discuss the reduction of communication traffic, the calculation process from y to z here can be regarded as a mapping between matrices within the third matrix space and does not consume communication traffic. For example, this can be achieved through an oracle r:Y→Y, or it can be expressed as a secure calculation z=r(y) via a secure mapping r.

[0076] Step 503: Based on z0, determine with other parties a first gradient matrix q of the overall gradient matrix z with respect to the feature matrix x via a bilinear mapping protocol based on the second mapping, and obtain a first slice q0 of the first gradient matrix q; and / or: determine a second gradient matrix s of the overall gradient matrix z with respect to the parameter matrix w based on a bilinear mapping protocol based on the third mapping, and obtain a first slice s0 of the second gradient matrix s.

[0077] It can be understood that during the model training process, it is usually necessary to determine the gradient of the model loss with respect to the parameters to be determined (such as the various parameters in the parameter matrix w). However, it is not ruled out that in some scenarios, in order to generate adversarial samples or to mine the importance of features, it is necessary to determine the gradient of the model loss with respect to the feature data (such as the various features in the parameter matrix w). Therefore, in this step 503, according to actual business needs, at least one of the first gradient matrix q of the overall gradient matrix z with respect to the feature matrix x and the second gradient matrix s of the overall gradient matrix z with respect to the parameter matrix w can be calculated. In multi-party secure computing, the first party can obtain the corresponding shards q0 and / or s0.

[0078] Among them, according to the principle of the bilinear triangle protocol mentioned above, during the calculation process, the perturbation result matrices δx and / or δw of the matrices involved in the bilinear mapping protocol under the perturbation matrix can be obtained by the intermediate results in the bilinear mapping protocol processing process based on the first mapping in step 501 without the need for recalculation, thereby saving communication volume.

[0079] Step 504 : Perform a security update of the model according to the first slice q0 of the first gradient matrix q and / or the first slice s0 of the second gradient matrix s.

[0080] Among them, the first gradient matrix q represents the gradient of the model loss with respect to the parameter matrix w. When the model is updated according to the first gradient matrix q, the various model parameters in the parameter matrix w can be adjusted by gradient-based parameter adjustment methods such as gradient descent and Newton's method, thereby achieving model update. For example, under the gradient descent method, the parameter matrix w can be updated to w = w-λq, where λ is a predetermined step size. For the first party, the first slice w0 of the local parameter matrix w can be updated through the first slice q0 of the local first gradient matrix q, and w0 = w0-λq0 can be obtained. In this way, the parameter matrix slices updated by each participating party in the same way constitute the updated parameter matrix and shared form.

[0081] The second gradient matrix s can be used in situations such as generating adversarial samples or mining feature importance. For example, the second gradient matrix s can be used to modify the eigenvalues ​​under the condition that the model loss increases, generate adversarial samples, and use the adversarial samples to update the model. In this case, the feature matrix modification process for generating adversarial samples is similar to the parameter matrix modification process described above and will not be repeated here. For another example, each participant can jointly mine feature importance based on the second gradient matrix s, and jointly determine the features with greater importance to retain, thereby updating the model using the retained features.

[0082] In other business scenarios, the model can also be updated in other ways, which will not be described here.

[0083] According to another embodiment, there is also provided a device for jointly updating a model based on multi-party secure computation. The model here is jointly updated by multiple participants using their local business data. The model may include a first fully connected layer, the first fully connected layer corresponds to the input matrix of the first matrix space, the parameter matrix of the second matrix space, and the output matrix of the third matrix space, and the first matrix space, the second matrix space, and the third matrix space constitute a bilinear triangle based on a predetermined bilinear mapping. Among them, the bilinear mapping from the first matrix space and the second matrix space to the third matrix space is the first mapping, the bilinear mapping from the third matrix space and the second matrix space to the first matrix space is the second mapping, and the bilinear mapping from the third matrix space and the first matrix space to the second matrix space is the third mapping. The first party can hold the first slice w0 corresponding to the parameter matrix w of the first fully connected layer. Before performing the calculation of the fully connected layer, the first party can also obtain the first slice x0 corresponding to its input matrix x.

[0084] Figure 6 FIG. 6 shows an apparatus 600 for jointly updating a model provided at a first party. Figure 6 As shown, the apparatus 600 includes:

[0085] The service processing unit 601 is configured to use the first slices w0 and x0 to determine, with another party, an output matrix y obtained by processing the input matrix x with the parameter matrix w via a bilinear mapping protocol based on the first mapping, and obtain a first slice y0 of the output matrix y;

[0086] The security conversion unit 602 is configured to complete the forward calculation of the model with other parties based on the first slice y0 of the output matrix y, and determine the overall gradient matrix z of the model loss with respect to the output matrix y during the backward propagation process, thereby obtaining the first slice z0 of the overall gradient matrix z;

[0087] The gradient determination unit 603 is configured to determine, based on z0, a first gradient matrix q of the overall gradient matrix z with respect to the feature matrix x via a bilinear mapping protocol based on the second mapping with another party, to obtain a first slice q0 of the first gradient matrix q; and / or to determine a second gradient matrix s of the overall gradient matrix z with respect to the parameter matrix w based on a bilinear mapping protocol based on the third mapping, to obtain a first slice s0 of the second gradient matrix s;

[0088] The updating unit 604 is configured to perform a security update of the model according to the first slice q0 of the first gradient matrix q and the first slice s0 of the second gradient matrix s.

[0089] It is worth mentioning that Figure 6 The device 600 is shown with Figure 5 The method embodiment shown corresponds to the embodiment shown, therefore, Figure 5 The relevant description in can also be applied to Figure 6 The device 600 shown is not described in detail here.

[0090] According to another embodiment, there is also provided a computer readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute a combination of Figure 5 The method described by et al.

[0091] According to another embodiment, a computing device is provided, including a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, the system realizes the combination of Figure 5 The method described by et al.

[0092] Those skilled in the art will appreciate that, in one or more of the above examples, the functions described in the embodiments of this specification may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0093] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the technical concept of this specification. It should be understood that the above is only the specific implementation method of the technical concept of this specification and is not intended to limit the scope of protection of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of this specification should be included in the scope of protection of the technical concept of this specification.

Claims

1. A method for jointly updating a model based on multi-party secure computation, wherein the model is jointly updated by multiple participants using their local business data. The model includes a first fully connected layer, wherein the first fully connected layer corresponds to an input matrix in a first matrix space, a parameter matrix in a second matrix space, and an output matrix in a third matrix space. The first matrix space, the second matrix space, and the third matrix space form a bilinear triangle based on a predetermined bilinear mapping, wherein: The bilinear mapping from the first matrix space and the second matrix space to the third matrix space is a first mapping, the bilinear mapping from the third matrix space and the second matrix space to the first matrix space is a second mapping, and the bilinear mapping from the third matrix space and the first matrix space to the second matrix space is a third mapping; The method is performed by a first party among the multiple participants, the first party holding first slices w0 and x0 corresponding to the parameter matrix w and the input matrix x, respectively, and the method includes: Using the first slices w0 and x0, with other parties, through a bilinear mapping protocol based on the first mapping, determine an output matrix y obtained by processing the input matrix x with the parameter matrix w, and obtain a first slice y0 of the output matrix y; According to the first slice y0 of the output matrix y, complete the forward calculation of the model with other parties, and determine the overall gradient matrix z of the model loss with respect to the output matrix y during the back-propagation process, thereby obtaining the first slice z0 of the overall gradient matrix z; According to z0, determine, with other parties, a first gradient matrix q of the overall gradient matrix z with respect to the feature matrix x via a bilinear mapping protocol based on the second mapping, and obtain a first slice q0 of the first gradient matrix q; and / or: determine a second gradient matrix s of the overall gradient matrix z with respect to the parameter matrix w based on a bilinear mapping protocol based on the third mapping, and obtain a first slice s0 of the second gradient matrix s; The model is securely updated according to the first slice q0 of the first gradient matrix q and / or the first slice s0 of the second gradient matrix s.

2. The method according to claim 1, wherein The method further comprises: Obtain the first perturbation matrix generated by the third party for the first matrix space, the second matrix space, and the third matrix space respectively The second perturbation matrix The third perturbation matrix and balance items The first shard corresponding to each Among them, each balance term is used to eliminate each predetermined bilinear mapping based on the first perturbation matrix The second perturbation matrix The third perturbation matrix Introduced bias.

3. The method according to claim 2, wherein: The method of using the first slices w0 and x0 and determining the output matrix y by processing the input matrix x with the parameter matrix w through a bilinear mapping protocol based on the first mapping with other parties, and obtaining the first slice y0 of the output matrix y includes: Using x0, Determine the first perturbation matrix The first slice δx0 of the first perturbation result matrix δx perturbed for the feature matrix x; Using w0, Determine the second perturbation matrix The first slice δw0 of the second perturbation result matrix δw perturbed for the parameter matrix w; Determine the first perturbation result matrix δx and the second perturbation result matrix δw with other participants by disclosing the corresponding shards; The first perturbation result matrix δx and the first slice w0 of the parameter matrix w, the first perturbation matrix δx are processed locally using the first mapping. The first shard and the second perturbation result matrix δw, obtaining the first slice of the first mapping matrix and the first slice of the second mapping matrix in the third matrix space; The balance term The first slice of is fused with the first slice of the first mapping matrix and the first slice of the second mapping matrix to obtain the first slice y0 of the output matrix y.

4. The method according to claim 3, wherein: Balance Item The first perturbation matrix is ​​transformed by the first mapping The second perturbation matrix The first slice δx0 of the first perturbation result matrix δx is determined by x0 and The first slice δw0 of the second perturbation result matrix δw is determined by w0 and The difference is determined; The balancing item The fusing of the first slice of the first mapping matrix with the first slice of the first mapping matrix and the first slice of the second mapping matrix includes: The balance term The first slice of the first mapping matrix is ​​superimposed with the first slice of the first mapping matrix and the first slice of the second mapping matrix.

5. The method according to claim 1, wherein The overall gradient matrix z of the output matrix y is a matrix in the third matrix space, and the overall gradient matrix z is the processing result of the output matrix y based on the oracle-based secure mapping processing.

6. The method according to claim 3, wherein: According to z0, with other parties through the bilinear mapping protocol based on the second mapping, the first gradient matrix q of the overall gradient matrix z to the feature matrix x is determined, and the first slice q0 of the first gradient matrix q is obtained: Using the third perturbation matrix The first shard Perturbing the first slice z0 of the overall gradient matrix z to obtain the first slice δz0 of the third perturbation result δz corresponding to the overall gradient matrix z, thereby restoring the third perturbation result δz based on other slices of the third perturbation result δz obtained by other parties; The second perturbation result matrix δw and the first slice z0 of the overall gradient matrix z and the second perturbation matrix z are processed respectively by the second mapping. The first shard and the third perturbation result matrix δz, obtaining the first slice of the third mapping matrix and the first slice of the fourth mapping matrix in the first matrix space; The balance term The first shard The first slice q0 of the first gradient matrix q is obtained by fusing it with the first slice of the third mapping matrix and the first slice of the fourth mapping matrix.

7. The method of claim 3, wherein: According to z0, based on the bilinear mapping protocol of the third mapping, the second gradient matrix s of the overall gradient matrix z to the parameter matrix w is determined, and the first slice s0 of the second gradient matrix s is obtained: Using the third perturbation matrix The first shard Perturbing the first slice z0 of the overall gradient matrix z to obtain the first slice δz0 of the third perturbation result δz corresponding to the overall gradient matrix z, thereby restoring the third perturbation result δz based on other slices of the third perturbation result δz obtained by other parties; The third perturbation result matrix δz and the first slice x0 of the feature matrix x, the third perturbation matrix The first shard and the first perturbation result matrix δx to obtain the first slice of the fifth mapping matrix and the first slice of the sixth mapping matrix in the second matrix space; The balance term The first shard The first slice s0 of the second gradient matrix s is obtained by fusing the first slice s0 of the fifth mapping matrix and the first slice of the sixth mapping matrix.

8. A device for jointly updating a model based on multi-party secure computation, wherein the model is jointly updated by multiple participants using their local business data, the model comprising a first fully connected layer corresponding to an input matrix in a first matrix space, a parameter matrix in a second matrix space, and an output matrix in a third matrix space, wherein the first matrix space, the second matrix space, and the third matrix space form a bilinear triangle based on a predetermined bilinear mapping, wherein: The bilinear mapping from the first matrix space and the second matrix space to the third matrix space is a first mapping, the bilinear mapping from the third matrix space and the second matrix space to the first matrix space is a second mapping, and the bilinear mapping from the third matrix space and the first matrix space to the second matrix space is a third mapping; The apparatus is provided at a first party among the multiple participants, the first party holding first slices w0 and x0 corresponding to the parameter matrix w and the input matrix x, respectively, and the apparatus includes: The service processing unit is configured to use the first slices w0 and x0 to determine, with another party, an output matrix y obtained by processing the input matrix x by the parameter matrix w through a bilinear mapping protocol based on the first mapping, and obtain a first slice y0 of the output matrix y; a secure conversion unit configured to complete the forward calculation of the model with other parties based on the first slice y0 of the output matrix y, and determine the overall gradient matrix z of the model loss with respect to the output matrix y during the backward propagation process, thereby obtaining the first slice z0 of the overall gradient matrix z; a gradient determination unit configured to determine, based on z0, a first gradient matrix q of the overall gradient matrix z with respect to the feature matrix x via a bilinear mapping protocol based on the second mapping with another party, to obtain a first slice q0 of the first gradient matrix q; and / or: determine, based on the bilinear mapping protocol based on the third mapping, a second gradient matrix s of the overall gradient matrix z with respect to the parameter matrix w, to obtain a first slice s0 of the second gradient matrix s; The updating unit is configured to perform a security update of the model according to the first slice q0 of the first gradient matrix q and / or the first slice s0 of the second gradient matrix s.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 7.

10. A computing device comprising a memory and a processor, characterized in that: The memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Joint training method and device for business model

    CN111738361A

  • Selection problem processing method based on multi-party security computing

    CN113626841A