Matrix restoration method and matrix fusion method

Through the probability matrix decomposition method, the processing node obtains and updates the matrix to find the average value until the target matrix is reached, solving the problem of complex and poor security of multi-party matrix repair and fusion calculations, and achieving efficient and safe matrix restoration and fusion.

CN114329332BActive Publication Date: 2025-07-11HANGZHOU QULIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111669526.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-07-11
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The prior art has complex calculations and poor security in multi-party matrix repair and fusion, especially in semi-homomorphic processing mode, which leads to insufficient security.

Method used

The probability matrix decomposition method is used to obtain the matrix of multiple participating nodes by processing nodes to find the average value, update and encrypt the gradient matrix until the probability decomposition matrix of the target matrix is reached, and the probability matrix decomposition is performed based on the identification matrix when the matrix is fused to obtain the target matrix.

Benefits of technology

The calculation process is simplified, suitable for multiple scenarios, protects the security of multi-party local data, and realizes the efficiency and security of matrix restoration and fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114329332B_ABST
    Figure CN114329332B_ABST
Patent Text Reader

Abstract

The present application relates to a matrix restoration method and a matrix fusion method. The method includes summarizing matrices held by multiple participating nodes in a processing node in a manner of probabilistic matrix factorization to obtain a target matrix containing all target data. The above matrix restoration method and matrix fusion method adopt the method of probabilistic matrix factorization for matrix fusion and matrix restoration in multi-party interaction, which can be applied to various scenarios and does not require sending local data, protecting the security of local data of multiple parties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of matrix probability decomposition, and in particular to a matrix restoration method and a matrix fusion method. Background Art

[0002] Probabilistic Matrix Factorization (PMF) is an important method in recommendation systems. It aims to learn the uneven information in the matrix and provide help for subsequent recommendations. It can also repair the matrix. The secure probabilistic matrix factorization is a matrix factorization and recommendation that protects the information security of multiple parties when the matrix information of multiple parties is stored in different users and joint probabilistic matrix factorization is required.

[0003] Currently, when performing matrix repair and matrix fusion of multiple parties, a semi-homomorphic method is generally used for processing. Different decomposition methods are required for different scenarios. The calculation is relatively complicated, and there are many interactions between multiple parties, resulting in poor security. Summary of the invention

[0004] Based on this, it is necessary to provide a matrix restoration method and a matrix fusion method to address the above technical problems.

[0005] In a first aspect, the present application provides a matrix restoration method, which is applied to multi-party interaction of multiple participating nodes, each participating node holds a first matrix and a second matrix, and the first matrix and the second matrix have the same latent feature dimension. The method comprises:

[0006] The processing node obtains first matrices of multiple participating nodes, calculates an average value of the multiple first matrices to obtain a third matrix, and sends the third matrix to the multiple participating nodes respectively;

[0007] Each of the participating nodes updates the first matrix based on the third matrix;

[0008] Each of the participating nodes obtains an updated gradient matrix of the first matrix based on a first gradient descent coefficient and sends the updated gradient matrix to the processing node. Each of the participating nodes also updates the second matrix based on a second gradient descent coefficient, where the first gradient descent coefficient and the second gradient descent coefficient are obtained based on a probability matrix decomposition of a target matrix.

[0009] The processing node obtains a fourth matrix based on the third matrix and the gradient matrices sent by the plurality of participating nodes, and sends the fourth matrix to the plurality of participating nodes respectively;

[0010] Each of the participating nodes updates the first matrix based on the fourth matrix, and determines whether the updated first matrix and the second matrix are the probability decomposition matrices of the target matrix. If not, it proceeds to the step where each of the participating nodes obtains the gradient value of the updated first matrix based on the first gradient descent coefficient and sends it to the processing node. Each of the participating nodes also updates the second matrix based on the second gradient descent coefficient until the updated first matrix and the second matrix are the probability decomposition matrices of the target matrix.

[0011] In one embodiment, the participating nodes include the processing node.

[0012] In one embodiment, the step where each of the participating nodes obtains the gradient matrix of the updated first matrix based on the first gradient descent coefficient and sends it to the processing node further includes:

[0013] Each of the participating nodes encrypts the gradient matrix and sends it to the processing node.

[0014] In one embodiment, the step where each of the participating nodes encrypts the gradient matrix includes:

[0015] Each of the participating nodes performs PSA encryption on the gradient matrix.

[0016] In one embodiment, the method further includes:

[0017] Each of the participating nodes performs business recommendation for the target user based on the updated first matrix and the second matrix.

[0018] In one embodiment, the step where the processing node obtains the fourth matrix based on the third matrix and the gradient matrices sent by multiple participating nodes further includes:

[0019] The processing node sums the gradient matrices sent by multiple participating nodes, and obtains the fourth matrix based on the third matrix and the summed gradient matrices.

[0020] In a second aspect, the present application also provides a matrix fusion method, which is applied to multi-party interaction of multiple participating nodes. Each participating node holds a decomposition matrix, and the user space and feature space of the decomposition matrix of each participating node are the same. The method includes:

[0021] Each of the participating nodes obtains an identification matrix based on whether there is data at each position in the decomposition matrix, and sends the identification matrix to the processing node;

[0022] The processing node obtains an indication matrix based on the identification matrices of multiple participating nodes and the number of participating nodes, and sends the indication matrix to the multiple participating nodes. The indication matrix represents whether there is data at corresponding positions in the decomposition matrices of the multiple participating nodes;

[0023] Each participating node performs probabilistic matrix factorization on the decomposition matrix based on the indication matrix to obtain a first gradient matrix and a second gradient matrix, and sends the first gradient matrix and the second gradient matrix to the processing node;

[0024] The processing node obtains a first target matrix and a second target matrix based on the first gradient matrices and the second gradient matrices of the multiple participating nodes, and sends the first target matrix and the second target matrix to the multiple participating nodes.

[0025] In one embodiment, each participating node obtaining an identification matrix based on whether there is data at each position in the decomposition matrix further includes:

[0026] Each participating node fills the positions with data in the decomposition matrix with 1 and fills the positions without data with 0 to obtain the identification matrix.

[0027] In one embodiment, each participating node sending the identification matrix to the processing node further includes:

[0028] Each participating node encrypts the identification matrix and sends the encrypted identification matrix to the processing node.

[0029] In one embodiment, the processing node obtaining an indication matrix based on the identification matrices of the multiple participating nodes and the number of participating nodes further includes:

[0030] The processing node performs matrix summation on the identification matrices of the multiple participating nodes and performs modulo operation on the summed matrix based on the number of participating nodes to obtain the indication matrix.

[0031] The above matrix restoration method and matrix fusion method obtain the first matrices of multiple participating nodes through processing nodes, solve the average value of the multiple first matrices to obtain a third matrix, and send the third matrix to the multiple participating nodes respectively; each participating node updates the first matrix based on the third matrix; each participating node obtains the gradient matrix of the updated first matrix based on the first gradient descent coefficient and sends it to the processing node, and each participating node also updates the second matrix based on the second gradient descent coefficient, and the first gradient descent coefficient and the second gradient descent coefficient are obtained by probability matrix factorization of the target matrix; the processing node obtains a fourth matrix based on the third matrix and the gradient matrices sent by the multiple participating nodes, and sends the fourth matrix to the multiple participating nodes respectively; each participating node updates the first matrix based on the fourth matrix and determines whether the updated first matrix and the second matrix are the probability decomposition matrices of the target matrix. If not, it goes to the step where each participating node obtains the gradient value of the updated first matrix based on the first gradient descent coefficient and sends it to the processing node, and each participating node also updates the second matrix based on the second gradient descent coefficient until the updated first matrix and the second matrix are the probability decomposition matrices of the target matrix, and the way that each participating node obtains an identification matrix based on whether there is data at each position in the decomposition matrix and sends the identification matrix to the processing node; the processing node obtains an indication matrix based on the identification matrices of the multiple participating nodes and the number of participating nodes and sends the indication matrix to the multiple participating nodes, and the indication matrix represents whether there is data at the corresponding position in the decomposition matrices of the multiple participating nodes; each participating node performs probability matrix factorization on the decomposition matrix based on the indication matrix to obtain a first gradient matrix and a second gradient matrix, and sends the first gradient matrix and the second gradient matrix to the processing node; the processing node obtains a first target matrix and a second target matrix based on the first gradient matrices and the second gradient matrices of the multiple participating nodes and sends the first target matrix and the second target matrix to the multiple participating nodes. In multi-party interaction, the method of using probability matrix factorization for matrix fusion and matrix restoration can be applied to various scenarios and does not require sending local data, protecting the security of local data of multiple parties. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic flowchart of the matrix restoration method in an embodiment of the present invention;

[0033] Figure 2 It is a schematic diagram of the application scenario of the matrix restoration method in an embodiment of the present invention;

[0034] Figure 3 Schematic diagram of the application scenario of the matrix restoration method in another embodiment of the present invention;

[0035] Figure 4 Schematic flow chart of the matrix fusion method in an embodiment of the present invention;

[0036] Figure 5 Schematic diagram of the application scenario of the matrix fusion method in another embodiment of the present invention;

[0037] Figure 6 Schematic diagram of the application scenario of the matrix fusion method in another embodiment of the present invention;

[0038] Figure 7 Schematic diagram of the application scenario of the matrix fusion method in another embodiment of the present invention. Detailed implementation manners

[0039] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0040] In one embodiment, as Figure 1 shown, a matrix restoration method is provided, which is applied to multi-party interactions of multiple participating nodes. Each participating node holds a first matrix and a second matrix, and the latent feature dimensions of the first matrix and the second matrix are the same. The latent feature dimension is the dimension formed when a target matrix is decomposed into multiple sub-matrices through probabilistic matrix decomposition. It can be understood that the first matrix and the second matrix respectively contain partial information of the target matrix. Exemplarily, if the target matrix contains complete user identities and user features, the first matrix may only contain user identities, the second matrix only contains user features, or the first matrix only contains user features, and the second matrix only contains user identities.

[0041] Please refer to Figure 2 , Figure 2 Schematic diagram of the application scenario of the matrix restoration method in an embodiment of the present invention, Figure 2 The A matrix and the B matrix shown are respectively held by two participating nodes, have the same feature space and different user spaces. Both the A matrix and the B matrix have four feature spaces I1, I2, I3, and I4. The A matrix has four user spaces U1, U2, U3, and U4, and the B matrix has four user spaces U5, U6, U7, and U8. The dotted filled part is the data owned by the matrix. Please refer to Figure 3 , Figure 3 Schematic diagram of the application scenario of the matrix restoration method in another embodiment of the present invention,Figure 3 The C matrix and the D matrix shown are respectively held by two participating nodes, having the same user space and different feature spaces. Both the C matrix and the D matrix have four user spaces: U1, U2, U3, and U4. The C matrix has four feature spaces: I1, I2, I3, and I4. The D matrix has four feature spaces: I5, I6, I7, and I8. The dotted-filled part is the data held by the matrix. Exemplarily, Figure 2 and Figure 3 the matrices in are the desired matrices that each participating node expects to obtain. The purpose of each participating node is to restore the first matrix and the second matrix to the probability decomposition matrices of the desired matrices of each participating node. It can be understood that the case of three or more parties is similar to the case of two parties and will not be elaborated here.

[0042] In this embodiment, the matrix restoration method includes the following steps:

[0043] Step S101, the processing node obtains the first matrices of multiple participating nodes, calculates the average value of the multiple first matrices to obtain a third matrix, and sends the third matrix to the multiple participating nodes.

[0044] Exemplarily, the first matrix and the second matrix of each participating node are randomly generated by the node, and the data in the matrix is filled in disorder, as long as the difference relationship between the number of rows and the number of columns is satisfied. In other embodiments, the first matrix and the second matrix can be obtained by other means, which will not be specifically limited here.

[0045] In this embodiment, each participating node sends the initialized first matrix to the processing node, and the processing node obtains the first matrix set {V1,..., V N}, and calculates the average value AVG{V1,..., V N} of the first matrix set to obtain a third matrix V, and broadcasts the third matrix V to all participating nodes {1, 2,..., N}.

[0046] Step S102, each participating node updates the first matrix based on the third matrix.

[0047] In this embodiment, each participating node replaces the first matrix with the received third matrix.

[0048] Step S103, each participating node obtains the gradient matrix of the updated first matrix based on the first gradient descent coefficient and sends it to the processing node. Each participating node also updates the second matrix based on the second gradient descent coefficient. The first gradient descent coefficient and the second gradient descent coefficient are obtained based on the probability matrix decomposition of the target matrix.

[0049] In this embodiment, each of the participating nodes obtains the gradient matrix of the third matrix based on the first gradient descent coefficient and sends it to the processing node. The processing node receives the set of gradient matrices {G1,..., G N}. Additionally, the participating nodes also locally update the second matrix based on the second gradient descent coefficient.

[0050] It can be understood that the first gradient descent coefficient and the second gradient descent coefficient are parameters obtained during the process of probabilistic matrix factorization of the target matrix and are pre-stored in each participating node, indicating the directions of the probabilistic factorization matrices of the first matrix and the second matrix approaching the target matrix.

[0051] Step S104: The processing node obtains the fourth matrix from the third matrix and the gradient matrices sent by multiple participating nodes, and sends the fourth matrix to multiple participating nodes.

[0052] In this embodiment, the processing node updates the third matrix using the received set of gradient matrices {G1,..., G N} to obtain the fourth matrix. This update process is a process of making the third matrix approach the probabilistic factorization matrix of the target matrix. It can be understood that after the update is completed, the fourth matrix is sent to multiple participating nodes to update the matrices of multiple participating nodes.

[0053] Step S105: Each participating node updates the first matrix based on the fourth matrix and determines whether the updated first matrix and the second matrix are the probabilistic factorization matrices of the target matrix. If not, it proceeds to the step where each participating node obtains the gradient value of the updated first matrix based on the first gradient descent coefficient and sends it to the processing node, and each participating node also updates the second matrix based on the second gradient descent coefficient, until the updated first matrix and the second matrix are the probabilistic factorization matrices of the target matrix.

[0054] Exemplarily, the participating node replaces the first matrix with the fourth matrix and determines whether the updated first matrix and the second matrix are the probabilistic factorization matrices of the target matrix, that is, whether the data in the first matrix and the second matrix are the target data, and whether the target matrix can be obtained after the first matrix and the second matrix are fused.

[0055] It can be understood that if the updated first matrix and the second matrix are not the probabilistic factorization matrices of the target matrix, it returns to step S102 and continues to execute step S102 and step S103, iterating cyclically to make the first matrix and the second matrix continuously approach the target until the target is reached and the iteration stops.

[0056] In this embodiment, to determine whether the updated first matrix and the second matrix are the probability decomposition matrices of the target matrix, the loss values of the first matrix and the second matrix can be calculated according to a pre-stored loss function. If the loss values converge, the iteration stops. In other embodiments, other determination methods can be adopted, which are not specifically limited herein, as long as it can be determined whether the updated first matrix and the second matrix are the probability decomposition matrices of the target matrix.

[0057] For the above matrix restoration method, the processing node obtains the first matrices of multiple participating nodes, calculates the average value of the multiple first matrices to obtain a third matrix, and sends the third matrix to the multiple participating nodes; the participating nodes update the first matrix based on the third matrix, obtain the gradient matrix of the updated first matrix based on the first gradient descent coefficient, and send it to the processing node. The participating nodes also update the second matrix based on the second gradient descent coefficient, and the first gradient descent coefficient and the second gradient descent coefficient are obtained based on the probability matrix decomposition of the target matrix; the processing node obtains a fourth matrix based on the third matrix and the gradient matrices sent by the multiple participating nodes, and sends the fourth matrix to the multiple participating nodes; the participating nodes update the first matrix based on the fourth matrix, and determine whether the updated first matrix and the second matrix are the probability decomposition matrices of the target matrix. If not, it goes to the step of obtaining the gradient value of the updated first matrix based on the first gradient descent coefficient and sending it to the processing node. The participating nodes also update the second matrix based on the second gradient descent coefficient until the updated first matrix and the second matrix are the probability decomposition matrices of the target matrix. In the multi-party interaction, the probability matrix decomposition method is used for matrix restoration, which has simple calculation, can be applied to various scenarios, has strong applicability, and does not need to send local data, protecting the security of the local data of multiple parties.

[0058] In another embodiment, the participating nodes include the processing node.

[0059] It can be understood that the processing node can also be a participating node. Exemplarily, when there are 2 or more participating nodes, at this time, all nodes are honest nodes. One of the participating nodes can be selected as the processing node to receive the matrices of other participating nodes and jointly process them with the matrices held by itself. When sending the updated matrix, its own matrix is also updated.

[0060] In other embodiments, the processing node can be a third-party node, there are 2 or more participating nodes and all of them are honest nodes, and the processing node is an external semi-honest node. The data of the processing node itself does not participate in the processing or update, but only receives the matrices of other participating nodes and provides the functions of calculation and broadcasting.

[0061] In another embodiment, each of the participating nodes obtaining the gradient matrix of the updated first matrix based on the first gradient descent coefficient and sending it to the processing node further includes the following steps:

[0062] Each of the participating nodes encrypts the gradient matrix and sends it to the processing node.

[0063] It can be understood that the participating nodes encrypt the gradient matrix before sending it, and the processing node solves it and then processes it, which can ensure the security of the local data of the participating nodes.

[0064] In another embodiment, each of the participating nodes encrypting the gradient matrix includes the following steps:

[0065] Each of the participating nodes performs PSA encryption on the gradient matrix.

[0066] Exemplarily, the PSA encryption method can be used to encrypt the gradient matrix. In other embodiments, other encryption methods can be used, which are not specifically limited herein.

[0067] In another embodiment, the method further includes the following steps:

[0068] Each of the participating nodes performs business recommendation for the target user based on the updated first matrix and the second matrix.

[0069] It can be understood that the probability decomposition matrix of the target matrix will be continuously approximated in the updated first matrix and the second matrix, and the data will gradually approach the data in the target matrix, reflecting the actual identity and characteristics of the user. The participating nodes can restore the complete matrix containing the original data according to the first matrix and the second matrix, obtain the user identity and characteristics from the complete matrix, and perform corresponding business recommendations for the target user based on the user identity and characteristics. Exemplarily, the specific business types can be set according to the data of the target matrix and the user requirements, which are not specifically limited herein.

[0070] In another embodiment, the processing node obtaining the fourth matrix based on the third matrix and the gradient matrices sent by multiple participating nodes further includes the following steps:

[0071] The processing node sums the gradient matrices sent by multiple participating nodes, and obtains the fourth matrix based on the third matrix and the summed gradient matrices.

[0072] Specifically, after receiving the gradient matrices sent by multiple participating nodes, the processing node sums them to obtain the comprehensive gradient matrix G = Sum(CG1,...,CG N), and obtain the updated fourth matrix through the comprehensive gradient matrix G and the third matrix V, the fourth matrix V New = V - G

[0073] Specifically, the process of updating the first matrix is as follows:

[0074] In the t-th iteration, the first matrices {V1,..., V N} of each participating node need to calculate the gradients {G1,..., G N} in the local participating nodes. For example, for participating node n, the calculation process of the gradient of the first matrix is as follows:

[0075]

[0076] Among them,

[0077] Exemplarily, t is the number of iterations, is the first gradient descent coefficient, S is the sum of non-zero values at (i, j), the inner product is the predicted value when the matrix decomposition is restored, and the corresponding non-negative γ is the constraint term of the corresponding decomposition matrix, r i,j represents the value at position (i, j).

[0078] Subsequently, update the third matrix in the processing node according to the gradient matrix. Specifically, V t = V t-1 - G. Among them, t is the number of iterations.

[0079] Specifically, the process of updating the second matrix is as follows:

[0080] In the t-th iteration, {U1,..., U N} of all parties are updated locally. For example, for participating node n, the update method of the second matrix is:

[0081]

[0082] Among them,

[0083] Exemplarily, t is the number of iterations, is the first gradient descent coefficient, S is the sum of non-zero values at (i, j), the inner product is the predicted value when the matrix decomposition is restored, and the corresponding non-negative γ is the constraint term of the corresponding decomposition matrix, r i,j represents the value at position (i, j).

[0084] Such as Figure 4As shown, the present invention also provides a matrix fusion method, which is applied to multi-party interactions of multiple participating nodes. Each participating node holds a decomposed matrix, and the user space and feature space of the decomposed matrix of each participating node are the same. It can be understood that each decomposed matrix respectively contains partial information of the target matrix. Exemplarily, the user space and feature space of each decomposed matrix are the same, but each only has partial data of the target matrix. Please refer to Figure 5 , Figure 5 which is a schematic diagram of the application scenario of the matrix fusion method according to another embodiment of the present invention. Figure 5 The E matrix and F matrix shown are respectively held by two participating nodes, both having four user spaces U1, U2, U3, U4 and four feature spaces I1, I2, I3, I4, but each matrix only has partial data. It can be understood that in the Figure 5 embodiment shown, the E matrix and F matrix do not intersect, and there are no identical values at any position. There is no data overlap in the fused G3 matrix either, which may be caused by the overly sparse data collected by each party.

[0085] Please refer to Figure 6 , Figure 6 which is a schematic diagram of the application scenario of the matrix fusion method according to another embodiment of the present invention. Figure 6 The H matrix and I matrix shown are respectively held by two participating nodes, both having four user spaces U1, U2, U3, U4 and four feature spaces I1, I2, I3, I4, but each matrix only has partial data. It can be understood that in the Figure 6 embodiment shown, the H matrix and I matrix intersect, and the values at the same position are the same. For the overlapping positions in the fused G4 matrix, the data of either matrix can be used. For example, both a hospital and a bank have statistically obtained the label data of a person's age.

[0086] Please refer to Figure 7 , Figure 7 which is a schematic diagram of the application scenario of the matrix fusion method according to another embodiment of the present invention. Figure 7 The J matrix and K matrix shown are respectively held by two participating nodes, both having four user spaces U1, U2, U3, U4 and four feature spaces I1, I2, I3, I4, but each matrix only has partial data. It can be understood that in the Figure 7 embodiment shown, the J matrix and K matrix intersect, and the values at the same position are different. For the overlapping positions in the fused G4 matrix, weighted summation is performed using the intersection method for data fusion. For example, the same user has different views on the same movie on the Douban platform and the Zhihu platform at different times, so the criteria are also different.

[0087] In this embodiment, the matrix fusion method includes:

[0088] Step S401: Each of the participating nodes obtains an identification matrix based on whether there is data at each position in the decomposition matrix, and sends the identification matrix to the processing node.

[0089] Exemplarily, the identification matrix is used to represent whether there is data at each position of the decomposition matrix.

[0090] Step S402: The processing node obtains an indication matrix based on the identification matrices of the multiple participating nodes and the number of the participating nodes, and sends the indication matrix to the multiple participating nodes. The indication matrix represents whether there is data at the corresponding positions of the decomposition matrices of the multiple participating nodes.

[0091] In this embodiment, the indication matrix is used to represent that there is data at this position for all the multiple participating nodes, or only some of the multiple participating nodes have data at this position, or there is no data at this position for all the multiple participating nodes.

[0092] Step S403: Each of the participating nodes performs probabilistic matrix factorization on the decomposition matrix based on the indication matrix to obtain a first gradient matrix and a second gradient matrix, and sends the first gradient matrix and the second gradient matrix to the processing node.

[0093] Exemplarily, probabilistic matrix factorization is performed on the decomposition matrix owned by each participating node, and according to the identification at each position in the indication matrix, a first gradient matrix and a second gradient matrix are obtained.

[0094] Step S404: The processing node obtains a first target matrix and a second target matrix based on the first gradient matrices and the second gradient matrices of the multiple participating nodes, and sends the first target matrix and the second target matrix to the multiple participating nodes.

[0095] It can be understood that the processing node receives the first gradient matrices and the second gradient matrices of the multiple participating nodes, performs gradient aggregation to obtain a first target matrix and a second target matrix, and sends the first target matrix and the second target matrix to the multiple participating nodes. Specifically, the first target matrix and the second target matrix are two sub-matrices obtained after probabilistic matrix factorization of a complete matrix containing the original data. Each participating node can perform matrix fusion on the first target matrix and the second target matrix to obtain a complete matrix containing the original data, and the complete matrix includes the data of the decomposition matrices of the multiple participating nodes.

[0096] The above matrix fusion method obtains an identification matrix by each of the participating nodes based on whether there is data at each position in the decomposition matrix, and sends the identification matrix to the processing node; the processing node obtains an indication matrix based on the identification matrices of the multiple participating nodes and the number of the participating nodes, and sends the indication matrix to the multiple participating nodes, where the indication matrix represents whether there is data at the corresponding positions of the decomposition matrices of the multiple participating nodes; each of the participating nodes performs probabilistic matrix factorization on the decomposition matrix based on the indication matrix to obtain a first gradient matrix and a second gradient matrix, and sends the first gradient matrix and the second gradient matrix to the processing node; the processing node obtains a first target matrix and a second target matrix based on the first gradient matrices and the second gradient matrices of the multiple participating nodes, and sends the first target matrix and the second target matrix to the multiple participating nodes. In this way, probabilistic matrix factorization is used for matrix fusion in multi-party interaction, which can be applied to multiple scenarios and does not require sending local data, protecting the security of multi-party local data.

[0097] In another embodiment, the step of each participating node obtaining an identification matrix based on whether there is data at each position in the decomposition matrix further includes the following steps:

[0098] Each of the participating nodes fills the positions with data in the decomposition matrix with 1 and fills the positions without data with 0 to obtain an identification matrix.

[0099] Specifically, multiple participating nodes perform 0-1 conversion on their own decomposition matrices A, B,..., N to generate new 0-1 matrices A′, B′,..., N′. If the data is missing at a position, it is replaced with 0, and if there is data at a position, it is replaced with 1.

[0100] In another embodiment, the step of each participating node sending the identification matrix to the processing node further includes the following steps:

[0101] Each of the participating nodes encrypts the identification matrix and sends the encrypted identification matrix to the processing node.

[0102] Exemplarily, each participating node encrypts its own identification matrix A′, B′,..., N′ through PSA preprocessing to obtain CA′, CB′,..., CN′.

[0103] In another embodiment, the step of the processing node obtaining an indication matrix based on the identification matrices of the multiple participating nodes and the number of the participating nodes further includes the following steps:

[0104] The processing node performs matrix summation based on the identity matrices of multiple participating nodes, and modulo operation on the summed matrix based on the number of participating nodes to obtain the indication matrix.

[0105] In this embodiment, at the trusted third party, i.e., the processing node, matrix summation and modulo operation are performed on the identity matrices of multiple participating nodes. Modulo operation means finding the remainder of the data at each position with respect to N, and rounding up the result to obtain the indicator matrix P. Here, N is the number of participating nodes, and p i,j = 0 indicates that data exists at this position for all parties. If p i,j = 1, it means that data exists at this position for only some nodes among all parties.

[0106] Exemplarily, the steps for each participating node to perform probabilistic matrix factorization on the decomposition matrix based on the indication matrix to obtain the first gradient matrix and the second gradient matrix are as follows:

[0107] The formula for calculating the first gradient matrix is:

[0108] Exemplarily, S is the sum of non-zero values at (i, j), the inner product is the predicted value when the matrix is restored after matrix factorization, the corresponding non-negative γ is the constraint term of the corresponding decomposition matrix, and r i,j represents the value at position (i, j).

[0109] The formula for calculating the second gradient matrix is:

[0110] Exemplarily, S is the sum of non-zero values at (i, j), the inner product is the predicted value when the matrix is restored after matrix factorization, the corresponding non-negative γ is the constraint term of the corresponding decomposition matrix, and r i,j represents the value at position (i, j).

[0111] It can be understood that after each participating node obtains the first gradient matrix and the second gradient matrix, it will and perform PSA preprocessing, that is, execute the operations of PSAEnc and PSAEnc to obtain and and send the encrypted result to the processing node. Additionally, each participating node performs PSA preprocessing on the first gradient matrix and the second gradient matrix, and sends the encrypted first gradient matrix and second gradient matrix to the processing node.

[0112] In another embodiment, the process by which the processing node obtains the first target matrix and the second target matrix based on the first gradient matrix and the second gradient matrix of the plurality of participating nodes is as follows:

[0113] Exemplarily, the parameter Num is used to indicate whether the data at this position is meaningful and whether it needs to be used when calculating the global matrix. The value of this parameter depends on p i,j and r i,j When r = 0, the calculation method of Num is as follows:

[0114]

[0115] Num i = Num i ;

[0116] Num j = Num j ;

[0117] 1) If p i,j = 0

[0118] Num A = 1

[0119] 2) If p i,j = 1 && r i,j ≠ 0

[0120] Num A = 1

[0121] 3) If p i,j = 1 && r i,j = 0

[0122] Num A = 0

[0123] where r i,j represents the value at position (i, j).

[0124] In this embodiment, the first target matrix and the second target matrix are global matrices. After the first target matrix and the second target matrix are aggregated, a complete matrix containing the complete original data can be obtained. The calculation methods of the first target matrix and the second target matrix are as follows:

[0125]

[0126]

[0127] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0128] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0129] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0130] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A matrix restoration method, applied to multi-party interaction of multiple participating nodes, characterized in that Each participating node holds a first matrix and a second matrix. The latent feature dimensions of the first matrix and the second matrix are the same. The first matrix and the second matrix respectively include the obtained user features and user identities. The method includes: The processing node obtains the first matrices of multiple participating nodes, calculates the average value of the multiple first matrices to obtain a third matrix, and sends the third matrix to the multiple participating nodes respectively; Each of the participating nodes updates the first matrix based on the third matrix; Each of the participating nodes obtains the gradient matrix of the updated first matrix based on the first gradient descent coefficient and sends it to the processing node, including: wherein, each of the participating nodes encrypts the gradient matrix and sends it to the processing node; Each of the participating nodes also updates the second matrix based on the second gradient descent coefficient. The first gradient descent coefficient and the second gradient descent coefficient are obtained based on the probabilistic matrix factorization of the target matrix; The processing node obtains a fourth matrix based on the third matrix and the gradient matrices sent by the multiple participating nodes, and sends the fourth matrix to the multiple participating nodes respectively; Each of the participating nodes updates the first matrix based on the fourth matrix, and determines whether the updated first matrix and the second matrix are the probabilistic decomposition matrices of the target matrix. If not, it goes to the step where each of the participating nodes obtains the gradient value of the updated first matrix based on the first gradient descent coefficient and sends it to the processing node, and each of the participating nodes also updates the second matrix based on the second gradient descent coefficient, until the updated first matrix and the second matrix are the probabilistic decomposition matrices of the target matrix; Each of the participating nodes performs business recommendation for the target user based on the updated first matrix and the second matrix.

2. The matrix restoration method according to claim 1, wherein The participating nodes include the processing node.

3. The matrix restoration method according to claim 1, wherein The step where each of the participating nodes encrypts the gradient matrix includes: Each of the participating nodes performs PSA encryption on the gradient matrix.

4. The matrix restoration method according to claim 1, wherein The step where the processing node obtains the fourth matrix based on the third matrix and the gradient matrices sent by the multiple participating nodes further includes: The processing node sums the gradient matrices sent by the multiple participating nodes, and obtains the fourth matrix based on the third matrix and the summed gradient matrices.

5. A matrix fusion method, applied to multi-party interactions of multiple participating nodes, characterized in that, Each participating node holds a decomposition matrix. The user spaces and feature spaces of the decomposition matrices of each of the participating nodes are the same. The method includes: Each of the participating nodes obtains an identification matrix based on whether there is data at each position in the decomposition matrix, and sends the identification matrix to the processing node; The processing node obtains an indication matrix based on the identification matrices of the multiple participating nodes and the number of the participating nodes, and sends the indication matrix to the multiple participating nodes. The indication matrix represents whether there is data at the corresponding positions of the decomposition matrices of the multiple participating nodes; Each of the participating nodes performs probabilistic matrix factorization on the decomposition matrix based on the indication matrix to obtain a first gradient matrix and a second gradient matrix, and sends the first gradient matrix and the second gradient matrix to the processing node; The processing node obtains a first target matrix and a second target matrix based on the first gradient matrix and the second gradient matrix of the multiple participating nodes, and sends the first target matrix and the second target matrix to the multiple participating nodes.

6. The matrix fusion method according to claim 5, wherein Each participating node obtaining an identification matrix based on whether there is data at each position in the decomposition matrix further includes: Each of the participating nodes fills the positions with data in the decomposition matrix with 1 and the positions without data with 0 to obtain an identification matrix.

7. The matrix fusion method according to claim 5, characterized in that Each of the participating nodes sending the identification matrix to the processing node further includes: Each of the participating nodes encrypts the identification matrix and sends the encrypted identification matrix to the processing node.

8. The matrix fusion method according to claim 5, wherein The processing node obtaining the indication matrix based on the identification matrices of the multiple participating nodes and the number of the participating nodes further includes: The processing node performs matrix summation on the identification matrices of the multiple participating nodes, and performs modulo operation on the summed matrix based on the number of the participating nodes to obtain the indication matrix.

Citation Information

Patent Citations

  • Mining method of computer data for recommendation systems

    CN102750360A

  • A tourism product recommendation method based on probability matrix decomposition and feature fusion

    CN109670909A