A Dynamic Association Prediction Method and Device for Protecting Privacy under Vertical Data Partitioning

By adopting vertical data segmentation and multi-party security computing protocols in dynamic correlation analysis, combined with differential privacy technology, the global user matrix is ​​gradually updated, and the problems of insufficient numerical data privacy protection and insufficient federated learning security in the existing technology are solved, and more efficient privacy protection and communication optimization are achieved.

CN116011002BActive Publication Date: 2025-06-17CHINA ACADEMY OF ELECTRONICS AND INFORMATION TECHNOLOGY OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211337038.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-06-17
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

The existing dynamic correlation analysis scheme based on differential privacy has problems such as incomplete information and limited privacy protection when protecting numerical data privacy. Federated learning may leak data information during the use of training models, and the system communication overhead is relatively large under vertical data segmentation.

Method used

Using vertical data segmentation method, at least three data square matrices are obtained, one of which contains the situation where the user uses data items, and the other part of the matrix contains the user's trust relationship. Through multi-party security computing protocols and differential privacy technology, the global user matrix is ​​gradually updated and iterated to the preset number of times to obtain the prediction result.

Benefits of technology

Effectively protect the privacy of related information between numerical users and data, enhance the security and privacy of federated learning, and optimize communication overhead during data fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116011002B_ABST
    Figure CN116011002B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for privacy-preserving dynamic association prediction under vertical data partitioning. The method includes: obtaining at least three data party matrices; performing initialization; respectively determining a first update value, a second update value, and a third update value for the global user matrix; determining the total update value of the global user matrix for the current iteration, and updating the global user matrix according to the total update value of the current iteration; iterating until a preset global model iteration number is reached; and determining a prediction result by using the first data party matrix based on the global user matrix obtained after the iteration is completed. Compared with the prior art, the present invention has at least the following advantages: it protects the privacy of the association information between numerical users and data, enhances the security and privacy in the process of using federated learning, and reduces the communication overhead when fusing data from different data parties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of differential privacy dynamic association, and in particular, to a method and device for privacy-preserving dynamic association prediction under vertical data partitioning. Background Art

[0002] With the development of social informatization and networking, data has become an important basic strategic resource and key production factor in the information age. Dynamic association analysis refers to predicting the data that a specific user is interested in based on the user's past data usage information and the social network information among users. On the one hand, dynamic association analysis can provide recommendation services for users, thereby helping users solve the problem of information overload caused by the explosive growth of data. On the other hand, it can also help the system detect abnormal read and write requirements of users for data, thereby discovering malicious or compromised user accounts by attackers, and thus has a wide range of application scenarios.

[0003] For dynamic association analysis solutions, when more sufficient user data is available, the prediction results are usually more accurate. However, different types of data of the same user may be distributed among different data parties, and at the same time, due to reasons of interest and regulations, each data party cannot directly exchange the original user data it owns. Then, how can data parties improve the performance of dynamic analysis algorithms through effective collaborative sharing while protecting the privacy of their respective user data? Differential privacy adds noise data to the data set for minor perturbation, so as to achieve that changing a single data instance does not affect the usability of the data set for a specific problem. As a theoretical framework for privacy protection, differential privacy can provide strict privacy protection proofs and support quantitative combined analysis of the privacy protection level of the system. Therefore, it has become one of the most promising privacy protection approaches for machine learning and is also an important technology for privacy protection in dynamic association analysis algorithms.

[0004] There is relatively little research work on existing dynamic association analysis schemes based on differential privacy, and several existing methods also have certain limitations, and can only be applied to binary user-data relationships or have limited privacy protection for users. Specifically, the methods proposed in the paper "A privacy-preserving framework for personalized, social recommendations" published by Jorgensen et al. in the International Conference on Extending Database Technology in 2014 and the paper "Differentially private graph-link analysis based social recommendation" published by Guo et al. in the international journal "Information Sciences" in 2018 are only applicable to binary relationships between users and data, that is, whether a user has recently used or not used a certain data, but cannot reflect the frequency of user data usage (i.e., the relationship between users and data is numerical). Therefore, the papers "PPTrustCF: A new privacy protection algorithm for trust collaborative filtering" published by Xian et al. in the International Conference on High Performance Computing and Communications in 2017 and "Personalized privacy-preserving social recommendation" published by Meng et al. in the AAAI Conference on Artificial Intelligence in 2018 respectively proposed differential privacy-based social recommendation algorithms for numerical data. Although the privacy of numerical user-data relationships is protected, whether a user has used a certain data can be inferred from the prediction results, so the privacy protection strength is limited.

[0005] To improve the accuracy and success rate of dynamic association analysis predictions, it is sometimes necessary to fuse data owned by different data parties (such as different websites). According to the situation of data fusion, it can be mainly divided into horizontal data splitting and vertical data splitting. The former refers to different data parties having non-overlapping user evaluations of the same data items (or the corresponding relationships between users and data items such as the frequency of user data usage). The latter usually refers to different data parties having the evaluation data of the same users for non-overlapping data item sets (or the corresponding relationships between users and data items such as the frequency of user data usage). Vertical data splitting is more common in dynamic association analysis. For example, in 2020, Shmueli et al. published the paper "Mediated secure multi-party protocols for collaborative filtering" in the international journal "ACM Transactions on Intelligent Systems and Technology". In addition, in 2019, Yang et al. published the paper "Federated machine learning: Concept and applications" in the international journal "ACM Transactions on Intelligent Systems and Technology", which introduced another concept of vertical data splitting, that is, one data party has the situation of user data usage, while the other data party has user characteristic information about these users, and proposed a federated recommendation algorithm based on matrix factorization technology.

[0006] Federated learning is a distributed machine learning technology for protecting data privacy. The data does not need to leave the local area, but a global shared model is jointly established through parameter exchange. When applying federated learning to dynamic association analysis to achieve privacy protection during data fusion, there are two main problems: First, as a privacy protection technology, federated learning still has defects. It mainly protects the security during the process of collecting data by user terminals, but some information of the training samples can still be recovered during the training model and its usage process. Second, when using federated learning for dynamic association analysis, it often leads to degraded recommendation results and low operating efficiency. Summary of the Invention

[0007] The technical problem to be solved by the present invention is:

[0008] 1) Differential privacy is an effective privacy protection mechanism. However, when applied to dynamic association analysis, problems such as incomplete protected information and limited supported data types will occur. That is, it is applicable to the privacy protection of binary data; when used for numerical data, it can protect the privacy of numerical values, but it cannot protect the sensitive information of whether the user has used the data.

[0009] 2) When dynamic association analysis uses federated learning, its privacy protection of data (i.e., training samples) means that local data does not need to be uploaded to the parameter server, but there is still a possibility that some information of the data may be leaked during the training model and its usage process.

[0010] 3) When it is necessary to fuse data between different data parties, it will lead to an increase in the communication overhead between different data parties in the system, and it is necessary to optimize the communication overhead of the system under the conditions of ensuring system security and accurate recommendation.

[0011] In view of this, the present invention provides a privacy-protected dynamic association prediction method and device under vertical data splitting.

[0012] The technical solution adopted by the present invention is that the privacy-protected dynamic association prediction method under vertical data splitting includes:

[0013] Obtain at least three data party matrices, wherein the situation of user's use of data items is pre-configured in the first data party matrix and the second data party matrix, and the trust relationship of users is pre-configured in the third data party matrix;

[0014] Initialize to obtain a global user matrix and a first data item matrix locally saved by the first data party matrix;

[0015] Use the global user matrix and the first data item matrix to determine a first update value for the global user matrix;

[0016] Use the second data party matrix to determine a second update value for the global user matrix;

[0017] Use the third data party matrix to determine a third update value for the global user matrix;

[0018] Determine the total update value of the current iteration of the global user matrix, and update the global user matrix according to the total update value of the current iteration, wherein the total update value of the current iteration is determined based on the first update value, the second update value, and the third update value;

[0019] Iterate until a preset global model iteration number is reached;

[0020] Based on the global user matrix obtained after the iteration is completed, use the first data party matrix to determine the prediction result.

[0021] In one embodiment, when the number of data party matrices is greater than three, after the first data party matrix has been determined, if a user usage data item is pre-configured in a data party matrix, it is processed as the second data party matrix, and if a user trust relationship is pre-configured in a data party matrix, it is processed as the third data party matrix.

[0022] In one embodiment, in the step of initializing to obtain the global user matrix and the first data item matrix locally saved by the first data party matrix, a random value is used for initialization to obtain the global user matrix.

[0023] In one embodiment, the determining of the first update value for the global user matrix by using the global user matrix and the first data item matrix includes:

[0024] Configure the first data party matrix as the global user matrix;

[0025] Randomly select at least two pieces of data item data, and use a pre-configured first algorithm to determine the gradients of the data item data with respect to the first data party matrix and the first data item matrix, namely the first data party gradient and the first data item gradient;

[0026] Use a pre-configured third algorithm to determine the average value of the first data party gradient and the first data item gradient;

[0027] Use a pre-configured fourth algorithm to determine the noise-added gradient average value corresponding to the average value;

[0028] Use a pre-configured learning rate to update the first data party matrix and the first data item matrix based on the noise-added gradient average value to obtain the updated first data party matrix;

[0029] Determine the difference between the updated first data party matrix and the global user matrix as the update amount of the global user matrix in this iteration;

[0030] Perform iterative update until the preset number of iterations of the first data party matrix is reached to obtain the first update value.

[0031] In one embodiment, the determining of the second update value for the global user matrix by using the second data party matrix includes:

[0032] Configure the second data party matrix as the global user matrix;

[0033] Using a pre-configured first algorithm, determine the gradients corresponding to the second data party matrix and the second data item matrix for the obtained data items, i.e., the second data party gradient and the second data item gradient;

[0034] Locally update the second data party matrix and the second data item matrix respectively using the second data party gradient and the second data item gradient to obtain an updated second data party matrix;

[0035] Determine the difference between the updated second data party matrix and the global user matrix as the update amount of the global user matrix in this iteration;

[0036] Perform iterative updates until a preset number of iterations of the second data party matrix is reached to obtain the second update value.

[0037] In one embodiment, the determining the third update value for the global user matrix using the third data party matrix includes:

[0038] Configure the third data party matrix as the global user matrix;

[0039] Using a pre-configured second algorithm, determine the gradient corresponding to the third data party matrix, i.e., the third data party gradient;

[0040] Locally update the third data party matrix using the third data party gradient to obtain an updated third data party matrix;

[0041] Determine the difference between the updated third data party matrix and the global user matrix as the update amount of the global user matrix in this iteration;

[0042] Perform iterative updates until a preset number of iterations of the third data party matrix is reached to obtain the third update value.

[0043] In one embodiment, based on a multi-party secure computing protocol, use the first update value, the second update value, and the third update value to determine the total update value of this iteration.

[0044] Another aspect of the present invention also provides a privacy-protected dynamic association prediction device under vertical data splitting, including:

[0045] An acquisition module configured to acquire at least three data party matrices, wherein the first data party matrix and the second data party matrix are pre-configured with the situation of user usage data items, and the third data party matrix is pre-configured with user trust relationships;

[0046] An initialization module configured to initialize to obtain a global user matrix and a first data item matrix locally saved by the first data party matrix;

[0047] The first data party matrix module is configured to determine a first update value for the global user matrix by using the global user matrix and the first data item matrix;

[0048] The second data party matrix module is configured to determine a second update value for the global user matrix by using the second data party matrix;

[0049] The third data party matrix module is configured to determine a third update value for the global user matrix by using the third data party matrix;

[0050] The weighting module is configured to calculate the total update value of the current iteration of the global user matrix by using the first data party matrix, and update the global user matrix according to the total update value of the current iteration, where the total update value of the current iteration is determined based on the first update value, the second update value, and the third update value;

[0051] The iteration module is configured to iterate until a preset global model iteration number is reached;

[0052] The prediction module is configured to determine a prediction result by using the first data party matrix based on the global user matrix obtained after the iteration is completed.

[0053] Another aspect of the present invention further provides an electronic device, where the electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps of the privacy-protected dynamic association prediction method under data vertical segmentation as described in any one of the above are implemented.

[0054] Another aspect of the present invention further provides a computer storage medium, where a computer program is stored on the computer storage medium, and when the computer program is executed by a processor, the steps of the privacy-protected dynamic association prediction method under data vertical segmentation as described in any one of the above are implemented.

[0055] By adopting the above technical solutions, the present invention has at least the following advantages:

[0056] The privacy-protected dynamic association prediction method under data vertical segmentation according to the present invention protects the privacy of the association information between numerical users and data (including information on whether a user has used the data); further enhances the security and privacy in the process of using federated learning; and optimizes the communication overhead when fusing data from different data parties. Description of the Drawings

[0057] Figure 1Flowchart of the privacy - protected dynamic association prediction method for vertical data segmentation according to an embodiment of the present invention;

[0058] Figure 2 Schematic diagram of the composition structure of the privacy - protected dynamic association prediction device for vertical data segmentation according to an embodiment of the present invention;

[0059] Figure 3 Schematic diagram of an electronic device according to an embodiment of the present invention;

[0060] Figure 4 Schematic diagram of the algorithm result when the number of global model training iterations is 5 according to an application example of the present invention;

[0061] Figure 5 Schematic diagram of the algorithm result when the number of global model training iterations is 20 according to an application example of the present invention;

[0062] Figure 6 Schematic diagram of the algorithm result when the number of global model training iterations is 50 according to an application example of the present invention;

[0063] Figure 7 Schematic diagram of the result comparison of the algorithm in the Douban dataset when all data owners use differential privacy according to an application example of the present invention. Detailed implementation manners

[0064] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the present invention will be described in detail as follows in combination with the accompanying drawings and preferred embodiments.

[0065] In the accompanying drawings, for the sake of clarity, the thickness, dimensions, and shape of the objects have been slightly exaggerated. The drawings are only for illustration and are not drawn to an exact scale.

[0066] It should also be understood that the terms "comprises", "comprising", "has", "including", and / or "containing", when used in this specification, denote the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. Further, when an expression such as "at least one of..." appears after a list of listed features, it modifies the entire list of listed features rather than individual elements in the list. In addition, when describing the embodiments of the present application, "may" is used to represent "one or more embodiments of the present application". And the term "exemplary" is intended to refer to an example or illustration.

[0067] As used herein, the terms "substantially", "about" and similar terms are used as terms of approximation and not of degree, and are intended to account for the inherent deviations in measured or calculated values that would be recognized by a person of ordinary skill in the art.

[0068] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by a person of ordinary skill in the art to which this application belongs. It should also be understood that terms (such as those defined in a common dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0069] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will describe this application in detail with reference to the drawings and in combination with the embodiments.

[0070] In the description of the method flow in the specification of the present invention and the steps in the flowcharts in the drawings of the specification of the present invention, it is not necessary to strictly execute according to the step numbers. The execution order of the method steps can be changed. Moreover, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0071] In the first embodiment of the present invention, a method for dynamically associating and predicting while protecting privacy under vertical data segmentation, as Figure 1 shown, includes the following specific steps:

[0072] Step S1, obtain at least three data party matrices, wherein the first data party matrix and the second data party matrix are pre-configured with the situation of users using data items, and the third data party matrix is pre-configured with the trust relationships of users;

[0073] Step S2, initialize to obtain a global user matrix and a first data item matrix locally saved by the first data party matrix;

[0074] Step S3, use the global user matrix and the first data item matrix to determine a first update value for the global user matrix;

[0075] Step S4, use the second data party matrix to determine a second update value for the global user matrix;

[0076] Step S5, use the third data party matrix to determine a third update value for the global user matrix;

[0077] Step S6, determine the total update value of the global user matrix for this iteration, and update the global user matrix according to the total update value of this iteration, where the total update value of this iteration is determined based on the first update value, the second update value, and the third update value;

[0078] Step S7, iterate until the preset number of global model iterations is reached;

[0079] Step S8, based on the global user matrix obtained after the iteration is completed, use the first data party matrix to determine the prediction result.

[0080] In this embodiment, three types of entities are mainly involved: users, data parties, and recommendation servers. The goal of the solution is: data party A (the first data party matrix) that has user usage data information (such as user evaluations of data, user usage frequencies of data, etc.), based on its own data, and the user usage data information and user trust relationships owned by the other collaborating data parties B (the second data party matrix) and C (the third data party matrix), to predict the data items that future users are interested in.

[0081] In this embodiment, the attackers include the users of the system. Therefore, the attackers can create multiple user accounts to participate in the system, and then observe the prediction results obtained by these multiple user accounts to obtain the privacy information of other system users. The attackers can also collude to obtain the information of the usage data of all other users except a specific target user of the attack. The goal of the attackers is to infer the usage situation of a specific target user of the attack through the observed prediction results and other auxiliary information, and to determine whether the user has used specific data.

[0082] For the convenience of understanding, a brief description of the multi-party secure computation protocol involved in this embodiment is now given.

[0083] The multi-party secure computation protocol used in this embodiment was proposed by et al. in the article "Multiparty computation from somewhat homomorphic encryption" published in CRYPTO in 2012. This protocol is divided into two stages: the offline stage and the online stage. The offline stage has nothing to do with the function to be specifically computed by the protocol and the input values.

[0084] For the convenience of further understanding, a brief description of the privacy-preserving dynamic association analysis algorithm (d-DPSoreg algorithm) involved in the present invention is now given.

[0085] In 2011, Ma et al. proposed a social recommendation algorithm, namely the Soreg algorithm, in the paper "Recommender systems with social regularization" published in the International Conference on Web Search and Data Mining of ACM. This algorithm incorporates trust relationships into the regularization term during matrix factorization, such that R = U T V, where U is called the user matrix and V is called the data item matrix.

[0086] In this embodiment, the gradient decomposition of the user matrix and the data matrix in the Soreg algorithm is decomposed into two parts. One is the influence of user information on the gradient, and the other is the influence of trust relationships on the gradient, which are represented by equations (1) and (2) respectively:

[0087]

[0088]

[0089] where I ij is the indicator function, which takes the value 1 when user i has used data item j, and 0 otherwise; m is the number of users, n is the number of data items; r ij represents the element value of the i-th row and j-th column of matrix R; U i represents the i-th column vector of the user matrix U; V j represents the j-th column vector of the data item matrix V.

[0090] To ensure that the output result of the designed prediction algorithm meets the requirements of differential privacy, in each round of iterative training, K training data are randomly selected first, and the corresponding update gradients of matrices U and V are calculated using equation (1), denoted as g U and g V respectively. Then, the following gradient clipping is performed on these gradients:

[0091]

[0092] where, ||·||2 represents the l2 norm value, and C represents the set threshold of the sum of the gradients of the user matrix and the data item matrix. Then, the average value of the clipped gradients obtained from the K training data in this round of iteration is calculated, and Gaussian noise is added to the average value, that is, calculate:

[0093]

[0094] where, N(0, σ 2 C 2 ) represents the probability density function of a Gaussian distribution with a mean of 0 and a variance of σ 2 C 2 .

[0095] For data parties with trust relationships, K users are randomly selected in each round of iterative training, and then the user matrix gradient is calculated using formula (2) for the trust relationships of the selected users.

[0096] For the convenience of subsequent description, formulas (1), (2), (3), and (4) are respectively referred to as the first algorithm, the second algorithm, the third algorithm, and the fourth algorithm in the subsequent text. It can be understood that the corresponding algorithms are not unique, and any formula changes based on the same idea as this embodiment and within a reasonable adjustment range should be within the scope of protection of this embodiment.

[0097] It is understandable that the number of iterations mentioned below can be pre-configured according to actual application conditions, and this document will not limit this.

[0098] The input of this embodiment may be: the first to third data square matrices, the first data square matrix A and the second data square matrix B have the user's use of data items, the third data square matrix C has the user's trust relationship, the preset global model training iteration number I, and the local iteration number corresponding to each data square matrix is ​​T A , T B and T C , learning rate γ;

[0099] The output can be: Use differential privacy to obtain prediction results that meet privacy protection requirements.

[0100] The method provided in this embodiment will be described in detail step by step below.

[0101] Step S1, obtaining at least three data square matrices, wherein the first data square matrix and the second data square matrix are pre-configured with situations where users use data items, and the third data square matrix is ​​pre-configured with trust relationships of users.

[0102] In this embodiment, when the number of data square matrices is greater than three, the first data square matrix may be determined first. When the first data square matrix has been determined, if a data square matrix is ​​pre-configured with user usage data items, it is processed as the second data square matrix, and if a data square matrix is ​​pre-configured with user trust relationships, it is processed as the third data square matrix.

[0103] Step S2, initialization to obtain a global user matrix and a first data item matrix of the first data square matrix stored locally.

[0104] In this embodiment, random values ​​are used to initialize the global user matrix U and the data item matrix V stored locally in the first data square matrix. A .

[0105] Step S3: Determine the first update value for the global user matrix by using the global user matrix and the first data item matrix.

[0106] In this embodiment, step S3 specifically includes:

[0107] Configure the first data party matrix as the global user matrix, that is: let U A = U;

[0108] Randomly select at least two (for example, K, and its value can be specifically configured according to the actual application scenario) data item data, and use the pre-configured first algorithm (the above formula (1)) to determine the gradients corresponding to the first data party matrix and the first data item matrix for the data item data, that is, the first data party gradient and the first data item gradient

[0109] Use the pre-configured third algorithm (the above formula (3)) to determine the average value of the first data party gradient and the first data item gradient:

[0110] Use the pre-configured fourth algorithm to determine the noise-added gradient average value corresponding to the average value:

[0111] Use the pre-configured learning rate γ to update the first data party matrix and the first data item matrix based on the noise-added gradient average value to obtain the updated first data party matrix. Specifically, it includes:

[0112] Determine the difference between the updated first data party matrix and the global user matrix as the update amount of the global user matrix in this iteration:

[0113] Perform iterative update until the preset iteration times T of the first data party matrix A is reached to obtain the first update value.

[0114] Step S4: Determine the second update value for the global user matrix by using the second data party matrix;

[0115] In this embodiment, step S4 specifically includes:

[0116] Configure the second data party matrix as the global user matrix: that is, U B = U.

[0117] Use the pre-configured first algorithm to determine the gradients corresponding to the second data party matrix and the second data item matrix for the obtained data item data, that is, the second data party gradient and the second data item gradient

[0118] Locally update the second data party matrix U using the second data party gradient and the second data item gradient respectively B and the second data item matrix V B .

[0119] Determine the difference between the updated second data party matrix and the global user matrix as the update amount of the global user matrix in this iteration:

[0120] Perform iterative updates until the preset number of iterations T of the second data party matrix is reached B to obtain the second update value.

[0121] Step S5, use the third data party matrix to determine the third update value for the global user matrix;

[0122] In this embodiment, step S5 specifically includes:

[0123] Configure the third data party matrix as the global user matrix; i.e., U C = U.

[0124] Use the pre-configured second algorithm to determine the gradient corresponding to the third data party matrix, i.e., the third data party gradient

[0125] Locally update the third data party matrix using the third data party gradient to obtain the updated third data party matrix U C ;

[0126] Determine the difference between the updated third data party matrix and the global user matrix as the update amount of the global user matrix in this iteration:

[0127] Perform iterative updates until the preset number of iterations T of the third data party matrix is reached C to obtain the third update value.

[0128] Step S6, determine the total update value of the global user matrix for this iteration, and update the global user matrix according to the total update value of this iteration, where the total update value of this iteration is determined based on the first update value, the second update value, and the third update value.

[0129] In this embodiment, it is calculated according to the multi-party secure computing protocol used by each party and then calculate to obtain the current global user matrix U = U + G sum .

[0130] Step S7, iterate until the preset number of iterations of the global model is reached.

[0131] Step S8, based on the global user matrix obtained after the iteration is completed, use the first data party matrix to determine the prediction result: Specifically, the prediction result is: U T V A .

[0132] In the second embodiment of the present invention, corresponding to the first embodiment, this embodiment introduces a dynamic association prediction device for protecting privacy under vertical data splitting, as Figure 2 shown, including the following components:

[0133] An acquisition module, configured to acquire at least three data party matrices, wherein the first data party matrix and the second data party matrix are pre-configured with the situation of user usage of data items, and the third data party matrix is pre-configured with the trust relationship of users;

[0134] An initialization module, configured to initialize to obtain a global user matrix and a first data item matrix locally saved by the first data party matrix;

[0135] A first data party matrix module, configured to use the global user matrix and the first data item matrix to determine a first update value for the global user matrix;

[0136] A second data party matrix module, configured to use the second data party matrix to determine a second update value for the global user matrix;

[0137] A third data party matrix module, configured to use the third data party matrix to determine a third update value for the global user matrix;

[0138] A weighting module, configured to determine the total update value of the global user matrix for this iteration, and update the global user matrix according to the total update value of this iteration, wherein the total update value of this iteration is determined based on the first update value, the second update value, and the third update value;

[0139] An iteration module, configured to iterate until a preset global model iteration number is reached;

[0140] A prediction module, configured to use the first data party matrix to determine a prediction result based on the global user matrix obtained after the iteration is completed.

[0141] It can be understood that the device provided in this embodiment can be used to implement the method described in the first embodiment. Its device composition and structure are based on the same concept as the first embodiment, and will not be elaborated herein.

[0142] In the third embodiment of the present invention, an electronic device, as Figure 3As shown, it can be understood as an entity device, including a processor and a memory storing instructions executable by the processor. When the instructions are executed by the processor, the following operations are performed:

[0143] Step S1, obtain at least three data party matrices, wherein the first data party matrix and the second data party matrix are pre-configured with the situation of user usage of data items, and the third data party matrix is pre-configured with the user's trust relationship;

[0144] Step S2, initialize to obtain a global user matrix and a first data item matrix locally saved by the first data party matrix;

[0145] Step S3, use the global user matrix and the first data item matrix to determine a first update value for the global user matrix;

[0146] Step S4, use the second data party matrix to determine a second update value for the global user matrix;

[0147] Step S5, use the third data party matrix to determine a third update value for the global user matrix;

[0148] Step S6, determine the total update value of the current iteration of the global user matrix, and update the global user matrix according to the total update value of the current iteration, wherein the total update value of the current iteration is determined based on the first update value, the second update value, and the third update value;

[0149] Step S7, iterate until a preset global model iteration number is reached;

[0150] Step S8, based on the global user matrix obtained after the iteration is completed, use the first data party matrix to determine the prediction result.

[0151] It can be understood that the software device described in the second embodiment can be configured in any possible electronic device provided in this embodiment to complete the method provided in the first embodiment, which will not be elaborated herein.

[0152] In the fourth embodiment of the present invention, this embodiment is based on the above first embodiment and combines the attached Figures 4 to 7 Introduce an application example of the present invention.

[0153] In this embodiment, the Douban dataset is used as the experimental dataset. This dataset contains 19,971 evaluation records of 2,822 movies by 357 users on the Douban website. The user ratings are integer values between 1 and 5. In addition, the dataset also contains 839 trust relationships established among users. The rating dataset is roughly evenly divided into four evaluation datasets and one trust relationship set, which are used as the evaluation data owned by Data Party A to Data Party D respectively. Data Party E owns the trust relationships among users, as shown in Table 1 specifically.

[0154] Table 1 Division of the experimental dataset

[0155]

[0156] For the evaluation of the accuracy of this embodiment, the performances at the global model training iteration times of 5, 20, and 50 are tested respectively, including the RMSE value and MAE value of rating prediction (the smaller these two values are, the more accurate the prediction result is). Since the existing recommendation algorithms in the case of vertical data segmentation have not considered the privacy issue of rating data for the time being, the algorithm designed by the present invention is compared with the following recommendation algorithms.

[0157] (1) Soreg algorithm: This algorithm is one of the most commonly used social recommendation algorithms, but it does not consider privacy issues.

[0158] (2) d-Soreg algorithm: This algorithm is the Soreg algorithm under vertical data segmentation implemented by each data party calculated according to Equation (1) and Equation (2) proposed in this paper, and it also does not consider privacy issues.

[0159] (3) DPGD algorithm: In 2018, Shin et al. published the paper "Privacy Enhanced matrix factorization for recommendation with local differential privacy" in the international journal "IEEE Transactions on Knowledge and Data Engineering", which proposed a differential privacy recommendation algorithm based on matrix factorization technology. This algorithm is similar to the algorithm in the present invention and can protect the privacy of data values more comprehensively. However, this algorithm does not consider the case of vertical data segmentation and does not consider the trust relationships among users.

[0160] Figures 4 to 6 The performances of the designed recommendation algorithm on the Douban dataset are shown respectively. The abscissa represents the number of data parties participating in the recommendation system, and the corresponding relationship between the number of data parties and the participating data parties is shown in Table 2. Figure 7It shows the results of the algorithm using differential privacy on the Douban dataset by all parties owning the rating data. It can be seen from these figures that the prediction algorithm can better protect privacy while retaining the accuracy of prediction (the smaller the RMSE and MAE values, the more accurate the rating prediction result).

[0161] Table 2 Correspondence Table between the Number of Data Parties and the Participating Data Parties

[0162]

[0163] In the fifth embodiment of the present invention, the process of the method for dynamically associated prediction while protecting privacy under vertical data splitting in this embodiment is the same as that in the first, second, or third embodiment. The difference is that in engineering implementation, this embodiment can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the method of the present invention can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a device to execute the method described in the embodiments of the present invention.

[0164] In summary, compared with the prior art, the present invention has at least the following advantages:

[0165] 1) The present invention protects the privacy of numerical user-data association information (including information on whether the user has used the data).

[0166] 2) The present invention enhances the security and privacy in the process of using federated learning.

[0167] 3) The present invention reduces the communication overhead when fusing data from different data parties.

[0168] Through the description of the specific implementation manners, it should be possible to understand more deeply and specifically the technical means and effects adopted by the present invention to achieve the predetermined purpose. However, the accompanying drawings are only for reference and illustration, and are not used to limit the present invention.

Claims

1. A method for dynamic association prediction that protects privacy under vertical data partitioning, characterized in that, Including: Obtain at least three data - party matrices, where the usage situation of user - used data items is pre - configured in the first data - party matrix and the second data - party matrix, and the trust relationship of users is pre - configured in the third data - party matrix; Initialize to obtain a global user matrix and a first data - item matrix locally saved by the first data - party matrix; Use the global user matrix and the first data - item matrix to determine a first update value for the global user matrix; Use the second data - party matrix to determine a second update value for the global user matrix; Use the third data - party matrix to determine a third update value for the global user matrix; Determine the total update value of the current iteration of the global user matrix, and update the global user matrix according to the total update value of the current iteration, where the total update value of the current iteration is determined based on the first update value, the second update value, and the third update value; Iterate until a preset number of global model iterations is reached; Based on the global user matrix obtained after the iteration is completed, use the first data - party matrix to determine a prediction result; The step of using the global user matrix and the first data - item matrix to determine a first update value for the global user matrix includes: Configure the first data - party matrix as the global user matrix; Randomly select at least two data - item data, and use a pre - configured first algorithm to determine the gradients of the data - item data with respect to the first data - party matrix and the first data - item matrix, namely the first data - party gradient and the first data - item gradient; Use a pre - configured third algorithm to determine the average value of the first data - party gradient and the first data - item gradient; Use a pre - configured fourth algorithm to determine the average value of the noise - added gradients corresponding to the average value; Use a pre - configured learning rate to update the first data - party matrix and the first data - item matrix based on the average value of the noise - added gradients to obtain the updated first data - party matrix; Determine the difference between the updated first data - party matrix and the global user matrix as the update amount of the global user matrix in the current iteration; Perform iterative updates until a preset number of first data - party matrix iterations is reached to obtain the first update value.

2. The method for dynamic association prediction that protects privacy under vertical data partitioning according to claim 1, characterized in that: When the number of data - party matrices is greater than three, after the first data - party matrix has been determined, if the usage situation of user - used data items is pre - configured in a data - party matrix, it is processed as the second data - party matrix, and if the trust relationship of users is pre - configured in a data - party matrix, it is processed as the third data - party matrix.

3. The method for dynamic association prediction that protects privacy under vertical data partitioning according to claim 1, characterized in that, In the step of initializing to obtain a global user matrix and a first data - item matrix locally saved by the first data - party matrix, a random value is used for initialization to obtain the global user matrix.

4. The method for dynamic association prediction that protects privacy under vertical data partitioning according to claim 1, characterized in that, The step of using the second data - party matrix to determine a second update value for the global user matrix includes: Configure the second data - party matrix as the global user matrix; Using a pre-configured first algorithm, determine the gradients corresponding to the second data party matrix and the second data item matrix for the acquired data items, namely the second data party gradient and the second data item gradient; Locally update the second data party matrix and the second data item matrix respectively using the second data party gradient and the second data item gradient to obtain the updated second data party matrix; Determine the update amount of the global user matrix in this iteration as the difference between the updated second data party matrix and the global user matrix; Perform iterative updates until the preset number of iterations of the second data party matrix is reached to obtain the second update value.

5. The method for dynamic association prediction that protects privacy under vertical data partitioning according to claim 1, characterized in that, The determining of the third update value for the global user matrix using the third data party matrix includes: Configure the third data party matrix as the global user matrix; Using a pre-configured second algorithm, determine the gradient corresponding to the third data party matrix, namely the third data party gradient; Locally update the third data party matrix using the third data party gradient to obtain the updated third data party matrix; Determine the update amount of the global user matrix in this iteration as the difference between the updated third data party matrix and the global user matrix; Perform iterative updates until the preset number of iterations of the third data party matrix is reached to obtain the third update value.

6. The method for protecting privacy in dynamic association prediction under vertical data segmentation according to claim 1, wherein: Based on the multi-party secure computing protocol, use the first update value, the second update value, and the third update value to determine the total update value for this iteration.

7. A device for protecting privacy in dynamic association prediction under vertical data segmentation, characterized in that, Including: An acquisition module configured to acquire at least three data party matrices, where the first data party matrix and the second data party matrix are pre-configured with the usage situation of user data items, and the third data party matrix is pre-configured with the trust relationship of users; An initialization module configured to initialize to obtain a global user matrix and a first data item matrix locally saved by the first data party matrix; A first data party matrix module configured to use the global user matrix and the first data item matrix to determine the first update value for the global user matrix; A second data party matrix module configured to use the second data party matrix to determine the second update value for the global user matrix; A third data party matrix module configured to use the third data party matrix to determine the third update value for the global user matrix; A weighting module configured to determine the total update value for this iteration of the global user matrix and update the global user matrix according to the total update value for this iteration, where the total update value for this iteration is determined based on the first update value, the second update value, and the third update value; An iteration module configured to iterate until the preset number of iterations of the global model is reached; A prediction module configured to determine a prediction result using the first data party matrix based on the global user matrix obtained after the iteration is completed; The first data party matrix module is further configured to: Configure the first data party matrix as the global user matrix; Randomly select at least two pieces of data item data, and use a pre-configured first algorithm to determine the gradients corresponding to the first data party matrix and the first data item matrix for the data item data, namely the first data party gradient and the first data item gradient; Use a pre-configured third algorithm to determine the average value of the first data party gradient and the first data item gradient; Use a pre-configured fourth algorithm to determine the noisy gradient average value corresponding to the average value; Use a pre-configured learning rate to update the first data party matrix and the first data item matrix based on the noisy gradient average value to obtain the updated first data party matrix; Determine the difference between the updated first data party matrix and the global user matrix as the update amount of the global user matrix in this iteration; Perform iterative updates until the preset number of iterations of the first data party matrix is reached to obtain the first update value.

8. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the privacy-protected dynamic association prediction method under data vertical segmentation as described in any one of claims 1 to 6.

9. A computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for protecting privacy in dynamic association prediction under vertical data segmentation according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Matrix decomposition recommendation method based on localized differential privacy

    CN113849772A

  • Federal learning method meeting personalized local differential privacy requirements

    CN114841364A