Privacy protection federal learning method for personalized recommendation system
By introducing a processing method of noise reduction matrix and global noise gradient in the federal recommendation system, the problem of taking into account both privacy protection and model accuracy in the prior art is solved, and high-precision personalized recommendation list generation is realized, while protecting user privacy.
Patent Information
- Application Number
- CN202510094806.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing local differential privacy federal recommendation algorithm cannot take into account the problem of privacy protection and model accuracy.
By initializing the user vector and project matrix based on personalized product information, updating the user vector and calculating the local gradient, and requesting the noise reduction matrix, adding noise and aggregating the local gradients, obtaining a global noise gradient, removing noise based on the noise reduction matrix and updating the project matrix, iteratively terminates and outputting the recommended list.
This method can achieve the same model accuracy as in plain text scenarios while protecting user privacy, solve the problem of taking into account privacy protection and model accuracy, and overcome the impact of low model accuracy caused by noise accumulation in local differential privacy federated learning by introducing trusted third-party designs.
Smart Images

Figure CN120012878A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of privacy protection technology, and in particular to a privacy protection federated learning method for a personalized recommendation system. Background Art
[0002] A recommendation system can be defined as an application whose core is a recommendation algorithm. The main purpose of a recommendation system is to explore user preferences in a specific scenario, find information that users may be interested in from a large amount of data, and generate a recommendation list for the user. As an effective technology to deal with the problem of "information overload", recommendation systems are widely used in e-commerce, advertising, music, video, news, job hunting, social networking, tourism, food delivery, finance, education, medical care, industrial production and other fields. According to a McKinsey report, 35% of the content purchased by customers on Amazon and 75% of the content watched by users on Netflix comes from recommendation systems.
[0003] Recommendation system technology brings convenience to our lives, but also brings new challenges. Recommendation systems use the historical interaction data between users and items and their inherent content attribute characteristics, and integrate a variety of additional information (such as social information, text information, image information, etc.) for personalized modeling to achieve the function of accurately predicting items that users may be interested in in the future. Since there is inevitably user personal data and sensitive information in massive information, the contradiction between the platform's need to collect more training data to improve recommendation performance and users' need to share as little personal data as possible to protect privacy has gradually become prominent. Whether it is an individual, an enterprise, or an organization, no one wants their private data to be leaked, but existing technologies cannot provide good data protection capabilities. In order to strengthen the protection of data privacy and security, countries have begun to introduce various laws and regulations to regulate and protect data security from a legal perspective.
[0004] In traditional centralized learning, the server needs to collect the original rating data of users and complete the model training, and then feed back the recommendation results to the user. This training method easily leads to the leakage of user behavior information to untrusted servers. Federated learning is a distributed machine learning technology, which is characterized by: user behavior data is stored locally, and the server and the user train the model by co-optimizing the intermediate parameters of the model. Therefore, applying the federated learning framework to the recommendation system can fundamentally solve the privacy protection problem of user behavior information. However, the latest research (Wang Z, Song M, Zhang Z, Song Y, Wang Q, Qi H. Beyond inferring class representatives: User-level privacy leakage from federated learning. In: Proceedings of the IEEE Conference on Computer Communications, 2019, 16: 2512-2520) shows that the gradient information of the transmission model may still be subject to reverse attacks and leak user privacy and model structure information.
[0005] The current privacy protection methods for federated recommendation systems mainly include homomorphic encryption and differential privacy. The former requires high communication and computing costs and is not suitable for terminal devices with limited computing resources. The latter inevitably affects the accuracy of the model due to strict mathematical assumptions. Therefore, a lightweight, high-precision, and privacy-preserving federated learning method for personalized recommendation systems is still in urgent need of research. Summary of the invention
[0006] The purpose of the present invention is to provide a privacy-preserving federated learning method for a personalized recommendation system, aiming to solve the problem that the existing local differential privacy federated recommendation algorithm cannot take into account both the degree of privacy protection and the accuracy of the model.
[0007] To achieve the above object, the present invention provides a privacy-preserving federated learning method for a personalized recommendation system, comprising the following steps:
[0008] Initialize user vector and item matrix based on personalized product information;
[0009] Update the user vector and calculate the local gradient, and request the denoising matrix;
[0010] Noising and aggregating the local gradients to obtain a global noise gradient;
[0011] The noise in the global noise gradient is removed based on the noise reduction matrix, and the item matrix is updated, and the iteration is terminated and a recommendation list is output.
[0012] The specific method of initializing the user vector and the item matrix based on the personalized product information is as follows:
[0013] Randomly initialize user feature vectors locally based on personalized product information;
[0014] The trusted third party randomly initializes the item feature matrix and sends it to the users participating in this round of training.
[0015] The specific method of updating the user vector and calculating the local gradient and requesting the denoising matrix is as follows:
[0016] The users participating in this round of training respectively update the user feature vector and calculate the local gradient locally;
[0017] Request the noise matrix from a trusted third party to randomly generate the Laplace noise matrix and weight parameters.
[0018] The specific method of adding noise to the local gradient and aggregating to obtain the global noise gradient is:
[0019] The users participating in this round of training respectively add noise to the local gradient locally, and send the local noisy gradient to the federated server;
[0020] The federated server aggregates the local noise gradients to obtain the global noise gradient and sends it to the users participating in this round of training.
[0021] The specific method of removing the noise in the global noise gradient based on the noise reduction matrix, updating the item matrix, iterating and terminating and outputting the recommendation list is as follows:
[0022] removing noise from the global noise gradient based on the noise reduction matrix and updating the project matrix;
[0023] When the training reaches the model accuracy or the iteration upper limit threshold, the iteration is terminated, and the user calculates the rating vector locally to obtain the recommendation list.
[0024] Among them, the update formula of the user vector is:
[0025]
[0026] The local gradient is calculated as:
[0027]
[0028] Where i is the index of the user, j is the index of the item, and u i is the user feature vector of the ith user, v jis the item feature vector of the jth item, V is the item matrix, G i is the local gradient calculated by the ith user, γ and λ 1 is a hyperparameter.
[0029] The users have the same list of items and a rating vector formed by the user's ratings on the items. The rating vector is a sparse vector containing a large number of missing values, and the rating vectors of all users form a rating matrix.
[0030] The present invention provides a privacy-preserving federated learning method for a personalized recommendation system, which initializes a user vector and an item matrix based on personalized product information; updates the user vector and calculates a local gradient, and requests a denoising matrix; adds noise to the local gradient, aggregates it, and obtains a global noise gradient; removes noise in the global noise gradient based on the denoising matrix, and updates the item matrix, terminates the iteration, and outputs a recommendation list, which can prevent user ratings, local gradients, and recommendation results from being leaked to a server, while achieving the same model accuracy as in a plaintext scenario. The local gradient of user-server communication satisfies local differential privacy, and the global gradient satisfies joint differential privacy. Both local data and gradient privacy are protected, and the personalized recommendation list finally generated is generated locally and can only be accessed by the user himself, thus protecting the privacy of the user's personalized recommendation results. By introducing a trusted third party, a mechanism is designed for users to remove global gradient noise locally, overcoming the impact of low model accuracy caused by noise accumulation in local differential privacy federated learning, so that the proposed method has strong privacy protection capabilities while maintaining model accuracy. This method does not require all users to participate in each round of training. By selecting the user set participating in each round of training through the server, it can effectively deal with the "straggler effect" of federated learning, making federated learning efficient and robust, and solving the problem that the existing local differential privacy federated recommendation algorithm cannot take into account both the degree of privacy protection and the model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1 It is a schematic diagram of the mathematical principle of the matrix decomposition recommendation system.
[0033] Figure 2 This is a schematic diagram of privacy-preserving federated learning for a personalized recommendation system provided by the present invention.
[0034] Figure 3 It is a flowchart of a privacy-preserving federated learning method for a personalized recommendation system provided by the present invention.
[0035] Figure 4 This is a flowchart of a privacy-preserving federated learning method for a personalized recommendation system provided by the present invention.
[0036] Figure 5 It is a flowchart of the specific way to initialize the user vector and item matrix.
[0037] Figure 6 It is a flowchart of a specific method of updating the user vector, calculating the local gradient, and requesting the denoising matrix.
[0038] Figure 7 It is a flowchart of a specific method of adding noise to the local gradient and aggregating it to obtain the global noise gradient.
[0039] Figure 8 It is a flowchart of a specific method of removing noise in the global noise gradient based on the noise reduction matrix, updating the item matrix, iteratively terminating and outputting a recommendation list. DETAILED DESCRIPTION
[0040] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0041] See also Figures 1 to 8 The present invention provides a privacy-preserving federated learning method for a personalized recommendation system, comprising the following steps:
[0042] S1 initializes the user vector and item matrix based on personalized product information;
[0043] In the embodiment of the present invention, the user randomly initializes the user feature vector locally, and the trusted third party randomly initializes the project feature matrix, and sends it to the n users participating in this round of training. Each user has the same project list and a rating vector composed of the user's rating of the project. Usually, the rating vector is a sparse vector with a large number of missing values, and the rating vectors of all users constitute a rating matrix.
[0044] The updating formula of the user vector is:
[0045]
[0046] The local gradient is calculated as:
[0047]
[0048] Where i is the index of the user, j is the index of the item, and u i is the user feature vector of the ith user, v j is the item feature vector of the jth item, V is the item matrix, G i is the local gradient calculated by the ith user, γ and λ 1 is a hyperparameter.
[0049] Specific method:
[0050] S11 randomly initializes the user feature vector locally based on the personalized product information;
[0051] S12 The trusted third party randomly initializes the project feature matrix and sends it to the users participating in this round of training.
[0052] S2 updates the user vector and calculates the local gradient, and requests a denoising matrix;
[0053] Specific method:
[0054] S21 The users participating in this round of training respectively update the user feature vector locally and calculate the local gradient;
[0055] In the embodiment of the present invention, the n users participating in this round of training respectively update the user feature vectors and calculate the local gradients locally.
[0056] S22 requests the noise matrix from the trusted third party and randomly generates the Laplace noise matrix and weight parameters.
[0057] In the embodiment of the present invention, the noise matrix is requested from the trusted third party, and the trusted third party randomly generates the Laplace noise matrix N and the weight parameter w i (in, ), w i Multiply by N to get n noise matrices N i (i=1,2,L,n), and n copies (N,N i ) are randomly distributed to n users participating in this round of training.
[0058] S3 adds noise to the local gradient and aggregates it to obtain a global noise gradient;
[0059] Specific method:
[0060] S31 The users participating in this round of training respectively add noise to the local gradient locally, and send the local noise gradient to the federal server;
[0061] In this embodiment of the present invention, the n users participating in this round of training respectively add noise N to the local gradient locally.i , and send the local noise gradient to the federated server.
[0062] Local gradient noise formula: G i ′=G i +N i ;
[0063] Federated server (aggregation) formula: G′=∑G i ′.
[0064] S32 The federal server aggregates the local noise gradients to obtain the global noise gradient, and sends it to the users participating in this round of training.
[0065] In the embodiment of the present invention, the federation server aggregates the local noise gradients of the users participating in the current round of training to obtain the global noise gradient, and sends it to the users participating in the current round of training. The server selects users in good condition from the user set as the next round of training users.
[0066] S4 removes the noise in the global noise gradient based on the noise reduction matrix, updates the item matrix, terminates the iteration and outputs the recommendation list.
[0067] Specific method:
[0068] S41 removes the noise in the global noise gradient based on the noise reduction matrix, and updates the project matrix;
[0069] In an embodiment of the present invention, the n users participating in this round of training respectively use N locally to remove noise in the global noise gradient to obtain a true global gradient, and use it to update the project matrix.
[0070] Global gradient denoising formula: G = G′-N;
[0071] Project matrix update formula: V t+1 =V t -γ(λ 2 V t -G t ).
[0072] S42 When the training reaches the model accuracy or the iteration upper limit threshold, the iteration is terminated, and the user calculates the rating vector locally to obtain a recommendation list.
[0073] In order to better understand the present technical solution, the following embodiments are provided for further explanation:
[0074] The example contains three types of entities, namely users, servers (CS) and trusted third parties (TPA). The server is assumed to be honest but curious. The method of the present invention aims to train a high-precision recommendation model through three-party collaboration and ensure that the user's original ratings, gradients, and recommendation lists and other private information will not be stolen by untrusted servers. Three small, medium, and large movie datasets that are widely used in the evaluation of recommendation algorithms are selected in the example. The specific information is shown in Table 1.
[0075] Table 1 Recommendation algorithm evaluation dataset
[0076]
[0077] according to Figure 3 The flowchart of the privacy-preserving federated learning method for personalized recommendation system is shown in the figure. The specific implementation steps are as follows:
[0078] Initialization phase: The server selects online users in good condition to participate in this round of training. The trusted third party sets the model's hyperparameter λ 1 ,λ 2 ,γ, and randomly initialize the item matrix And (V,λ 1 ,λ 2 ,γ) is sent to all users, and all users randomly initialize user vectors locally Among them, k is the latent feature dimension, λ 1 ,λ 2 for is the regularization parameter of the matrix decomposition model, and γ is the learning rate of the gradient descent algorithm. The specific parameter settings are shown in Table 2.
[0079] Table 2 Hyperparameter settings
[0080]
[0081] Model training phase: The n users participating in this round of training request the trusted third party to send noise, and use the user vector update formula to locally update the user vector u i , the local gradient calculation formula is used to calculate the local gradient G′. After receiving the noise request from n users, the trusted third party randomly generates the Laplace noise matrix Randomly generate n weight parameters w i (i=1,2,L,n), so that Then N and w i Multiply to get n noise matrices The trusted third party will send n copies (N,N i) is randomly sent to n users. After receiving the noise sent by the trusted third party, the user adds noise to the local gradient according to the local gradient noise formula to obtain the disturbed local gradient G i ′, and sends it to the server. After the server receives the perturbed gradients sent by n users, it performs the aggregation operation of the local gradient aggregation formula to obtain the perturbed global gradient G′, and sends G′ to the n users participating in this round of training. The user receives the perturbed global gradient, and performs the denoising operation of the global gradient denoising formula locally to obtain the true global gradient G, and then uses the project matrix update formula to complete the update of the project matrix. After this round of training is completed, the server selects the user set participating in the next round of training, and iterates until the termination condition is met, and the training phase ends. Among them, each round of training uses 5-fold cross validation to evaluate the prediction effect of the model on the test set. The evaluation indicators include: root mean square error (RMSE) and mean absolute error (MAE), which are defined as:
[0082]
[0083] Where N is the number of scoring samples in the test set; r ij , are the true rating and predicted rating of item j by user i, respectively.
[0084] Training completion phase: After the training phase is completed, the optimal model parameters are output V * , users perform calculations locally: Get the user's predicted scores for all items, and use the top K items with higher predicted scores as the user's recommendation list, so that the recommendation task is completed.
[0085] Finally, the method of the present invention (Ours) is compared with the training method in plain text (FLMF) and the federated learning method based on local differential privacy (LDPFLMF). The cross-validation comparison results are shown in Table 3. From the comparison results, it can be seen that the prediction accuracy of the method of the present invention is comparable to that of the method trained in plain text, and is significantly better than the federated learning method based on local differential privacy.
[0086] Table 3 Comparison of model prediction accuracy
[0087]
[0088] What is disclosed above is only a preferred embodiment of the privacy-preserving federated learning method for a personalized recommendation system of the present invention. Of course, this cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiments and equivalent changes made according to the claims of the present invention still fall within the scope of the invention.
Claims
1. A privacy-preserving federated learning method for personalized recommendation systems, characterized in that: The following steps are involved: Initialize user vector and item matrix based on personalized product information; Update the user vector and calculate the local gradient, and request the denoising matrix; Noising and aggregating the local gradients to obtain a global noise gradient; The noise in the global noise gradient is removed based on the noise reduction matrix, and the item matrix is updated, and the iteration is terminated and a recommendation list is output.
2. The privacy-preserving federated learning method for a personalized recommendation system according to claim 1, It is characterized by: The specific method of initializing the user vector and item matrix based on personalized product information is as follows: Randomly initialize user feature vectors locally based on personalized product information; The trusted third party randomly initializes the item feature matrix and sends it to the users participating in this round of training.
3. The privacy-preserving federated learning method for personalized recommendation systems according to claim 1, characterized in that: The specific method of updating the user vector and calculating the local gradient and requesting the denoising matrix is: The users participating in this round of training respectively update the user feature vector and calculate the local gradient locally; Request the noise matrix from a trusted third party to randomly generate the Laplace noise matrix and weight parameters.
4. The privacy-preserving federated learning method for personalized recommendation systems according to claim 1, characterized in that: The specific method of adding noise to the local gradient and aggregating to obtain the global noise gradient is: The users participating in this round of training respectively add noise to the local gradient locally, and send the local noisy gradient to the federated server; The federated server aggregates the local noise gradients to obtain the global noise gradient and sends it to the users participating in this round of training.
5. The privacy-preserving federated learning method for a personalized recommendation system according to claim 1, It is characterized by: The specific method of removing the noise in the global noise gradient based on the noise reduction matrix, updating the item matrix, iterating and terminating and outputting the recommendation list is as follows: removing noise from the global noise gradient based on the noise reduction matrix and updating the project matrix; When the training reaches the model accuracy or the iteration upper limit threshold, the iteration is terminated, and the user calculates the rating vector locally to obtain the recommendation list.
6. The privacy-preserving federated learning method for personalized recommendation systems according to claim 1, characterized in that ; The updating formula of the user vector is: The local gradient is calculated as: Where i is the index of the user, j is the index of the item, and u i is the user feature vector of the i-th user, v j is the item feature vector of the jth item, V is the item matrix, G i is the local gradient calculated by the ith user, and γ and λ1 are hyperparameters.
7. The privacy-preserving federated learning method for personalized recommendation systems according to claim 2, characterized in that ; The users have the same list of items and a rating vector formed by the user's ratings on the items. The rating vector is a sparse vector containing a large number of missing values. The rating vectors of all users form a rating matrix.
Citation Information
Patent Citations
Matrix decomposition recommendation method based on homomorphic encryption
CN110209994A
Method and system for protecting federal learning privacy data, medium and electronic equipment
CN116861486A
Federal matrix decomposition recommendation method with efficient communication
CN118332165A
Training Method, Apparatus, and Device for Federated Neural Network Model, Computer Program Product, and Computer-Readable Storage Medium
US20230023520A1