Federal privacy protection sequence recommendation method based on knowledge increment collaboration
By adopting a knowledge-increment collaborative federated privacy-preserving sequence recommendation method, the problems of privacy exposure and low recommendation accuracy are solved, thereby improving recommendation accuracy and model performance while protecting user privacy.
Patent Information
- Application Number
- CN202511132251.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-18
AI Technical Summary
Existing federal privacy-preserving sequence recommendation methods suffer from privacy exposure and low recommendation accuracy.
We adopt a knowledge-increment collaborative federated privacy-preserving sequence recommendation method. We obtain the scores of items in the candidate pool through local optimization on the client side, and use knowledge increments for filtering and uploading. Combined with global embedding matrix update, we introduce user-item consistency and global item similarity regularization constraints to avoid directly uploading high-entropy personalized information.
It achieves improved accuracy of recommendation models while ensuring privacy and security, avoids leakage of user privacy, and maintains good learning ability of personalized preferences and global patterns, thereby improving recommendation accuracy.
Smart Images

Figure CN120974540A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data privacy protection, in particular to a federated privacy protection sequence recommendation method based on knowledge increment collaboration. BACKGROUND
[0002] Sequence recommendation is an important recommendation system method, which aims to predict and recommend items or content that users may be interested in according to the user's historical interaction sequence, and provide personalized and time-sequential recommendation results by modeling the time sequence interaction relationship between users and each interaction item in the sequence and the correlation between items. As an important supporting technology for personalized services, federated sequence recommendation (FSR) has become a research hotspot in recent years under the background of increasing demand for data privacy protection. Federated sequence recommendation technology optimizes the personalized recommendation model by decentralizing the model training process to the user device locally, without directly transmitting the original behavior data, thereby to a certain extent, avoiding the risk of user privacy leakage in the centralized system.
[0003] With the rapid growth of the number of users in the system and the complexity of user behavior data, the existing federated recommendation framework faces many challenges in privacy protection, model personalization ability and communication efficiency. On the one hand, the user's behavior sequence often has a high degree of structure and uniqueness, and contains a large number of high-entropy personal features, which may leak user privacy through uploaded model gradients even without uploading original data. Especially under the assumption of "honest but curious" server, attackers can reconstruct the user's behavior sequence by means of gradient reverse reasoning technology, causing serious privacy exposure. On the other hand, the existing federated learning architecture generally relies on global model aggregation strategy, which is easy to drown out the long-tail interests and heterogeneous behavior patterns of a small number of users in the aggregation process, resulting in insufficient personalized modeling ability and low recommendation accuracy. SUMMARY
[0004] The present application aims to solve the problems of privacy exposure and low recommendation accuracy in the existing federated privacy protection sequence recommendation method, and proposes a federated privacy protection sequence recommendation method based on knowledge increment collaboration.
[0005] A federated privacy protection sequence recommendation method based on knowledge increment collaboration, specifically comprising:
[0006] Obtaining user historical interaction items, inputting the user historical interaction items into an optimized client, the optimized client obtaining the score of each item in the candidate pool, and forming a recommendation list by the L items with the highest scores and outputting the recommendation list;
[0007] The client obtains the score of each item in the candidate pool, specifically:
[0008]
[0009] wherein, is a client rating of an item, is a preference vector of the client is a global embedding parameter of the item is a personalized bias vector of the item for the client u;
[0010] The items in the candidate pool include private item pools of clients and items in a system global item library.
[0011] Further, the optimized client is obtained by the following way:
[0012] Step one, initialize the client, the server and the iteration round respectively, specifically:
[0013] initialize the global embedding vector to a matrix of , the element value in the initialized global embedding vector is obtained by random sampling in a normal distribution with a mean of 0 and a preset standard deviation;
[0014] wherein, is the total number of items, is a preset embedding dimension;
[0015] The i-th row of the global embedding matrix represents the global embedding parameter of item i maintained by the server;
[0016] The client u initializes the client preference vector and the item personalized bias matrix to a random vector;
[0017] wherein, the personalized bias matrix stores the personalized bias vector of item i for the client u;
[0018] Initialize the iteration round t=0;
[0019] Step two, the server obtains the global item similarity graph using the global embedding matrix , then the server selects the participating client group of this round from the available client pool, finally the server broadcasts to all clients in and ;
[0020] Step 3: Join the client group The client in the middle receives the global embedding matrix Similarity map with global items Then, using , Combine with client Preference vector And item personalization bias matrix Perform local model optimization to obtain the updated client. Preference vector and the updated item personalization bias matrix Then, knowledge increments are obtained using the item personalization bias matrix. ;
[0021] Step 4: The client filters the knowledge increments extracted in Step 3 and uploads the filtered knowledge increments to the server.
[0022] Step 5: The server aggregates the received knowledge increments, updates the global embedding matrix using the aggregated knowledge increments, and sends the updated global embedding matrix to the client.
[0023] Step Six: Obtain the average knowledge increment norm of round t. Compare the average knowledge increment norm of round t with a preset norm threshold. If the average knowledge increment norm of round t is lower than the preset norm threshold or t=T, then obtain the optimized server-side and optimized client-side; otherwise, let t=t+1. , , And return to step two;
[0024] The optimized client stores the updated client preference vector and the updated item personalization bias matrix;
[0025] Where T is the preset maximum number of iterations.
[0026] Furthermore, the server in step two utilizes a global embedding matrix. Get global item similarity map Specifically:
[0027]
[0028] in, It is the L2 norm. yes The element in the i-th row and j-th column.
[0029] Furthermore, the server selects the participating client group for this round from the pool of available clients. Specifically:
[0030] First, obtain the pool of available clients based on preset client selection criteria. ;
[0031] Then, the server uses an unbiased random sampling method to sample from the client pool. Mid-sampling to obtain client groups .
[0032] Furthermore, the participating client group in step three The client in the middle receives the global embedding matrix Similarity map with global items Then, using , Combine with client Preference vector And item personalization bias matrix Perform local model optimization to obtain the updated client. Preference vector and the updated item personalization bias matrix Then, knowledge increments are obtained using the item personalization bias matrix. Specifically:
[0033] Step 31: The client obtains the personalized item embedding vector using the global embedding parameters, specifically:
[0034]
[0035] in, It is an item Personalized item embedding vectors, It is the client u to items Personalized bias vector, It is the global embedding parameter of item i;
[0036] Step 3.2: Construct a training set using personalized item embedding vectors;
[0037] The positive samples in the training set are Negative samples are ;
[0038] in, It is the client identifier. It is a client Interacted items, It is a client Items that have not been interacted with are labeled 1 for positive samples and 0 for negative samples.
[0039] Step three, the client inputs the training set into the local model, and calculates the overall loss function of the local model on the training set;
[0040] The working process of the local model is specifically:
[0041]
[0042] Wherein, is the matching score of the client and the item , is the scoring function, is the preference vector of the client ;
[0043] Step three, the client inputs the training set into the local model, and calculates the overall loss function of the local model on the training set; Step four, the client updates the item personalized bias matrix and the preference vector of the client based on the overall loss function of the local model on the training set;
[0044] Step five, the knowledge increment is obtained by using the item personalized bias matrix of the item i before the update and the item personalized bias matrix of the item i after the update.
[0045] Further, the overall loss function of the local model in step three on the training set is specifically:
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052] Wherein, is the core recommendation ranking loss, is the user-item consistency regularization term, is the item-item similarity regularization term, , is the weight coefficient, is the item sample set in the training data, is the item that the user has interacted with, is the item that the user has not interacted with, is the Sigmoid function, is the client and the item matching score, is the client and the item matching score, is the recommendation ranking loss of the user u, is the global embedding parameter of the item , is the personalized bias vector of the item , is the set of items interacted by the client u, is the aggregation function, is the cosine similarity of and , is the cosine similarity of and .
[0053] Further, the client in step three four updates the item personalized bias matrix and the preference vector of the client based on the loss function of the trained local model, specifically:
[0054]
[0055]
[0056] wherein, is the local learning rate of the client, is the gradient value of , is the total loss function, is the preference vector of the client in the t-th round of training, is the preference vector of the client in the t+1-th round of training.
[0057] Further, the step three five uses the item personalized bias matrix of the item i before updating and the item personalized bias matrix of the item i after updating to obtain the knowledge increment, specifically:
[0058] .
[0059] Further, the step four filters the knowledge increment extracted in step three by the client, and uploads the filtered knowledge increment to the server side, specifically:
[0060] obtains the L2 norm of each row vector in , and sorts each row vector in descending order of L2 norm, and obtains The first K vectors are used as the filtered knowledge increments, and the filtered knowledge increments are uploaded to the server.
[0061] Furthermore, in step five, the server aggregates the received knowledge increments and updates the global embedding matrix using the aggregated knowledge increments, specifically as follows:
[0062] First, the server aggregates the received knowledge increments to obtain the aggregated knowledge increments, specifically:
[0063]
[0064] in, It is a federal average algorithm. This is the first knowledge increment uploaded by the client u. row vectors It is a collection of all clients that have uploaded information about item i. This represents the number of samples in the training set of client u;
[0065] Then, utilize the aggregated knowledge increment Update the global embedding parameters as follows:
[0066]
[0067] in, These are the global embedding parameters of item i in the t-th iteration. These are the global embedding parameters of item i in the (t+1)th iteration. This represents the learning rate of the global aggregation.
[0068] The beneficial effects of this invention are as follows:
[0069] This invention achieves both privacy and recommendation model performance by employing an innovative "knowledge increment collaboration" mechanism. The information uploaded by the client does not contain the original gradients; instead, it is uploaded as knowledge increments represented by state differences, supplemented by Top-K subset filtering to ensure concise and low-sensitivity content. This ensures that the user's preference vector and other high-entropy features remain locally on the client, fundamentally cutting off the channels for sensitive information leakage through gradient paths and preventing user privacy exposure. By leveraging user-item consistency and global item similarity as structural regularization constraints, this invention maintains a strong ability to learn personalized preferences and global patterns while protecting privacy. This avoids the model accuracy degradation caused by relying solely on noise perturbations and the inability to effectively utilize global knowledge, thus improving recommendation accuracy. Attached Figure Description
[0070] Figure 1A schematic diagram of the architecture of the present application;
[0071] Figure 2 A flowchart of client-server collaborative training. DETAILED DESCRIPTION
[0072] Embodiment one: the specific process of the federated privacy protection sequence recommendation method based on knowledge increment collaboration in this embodiment is as follows:
[0073] The user historical interaction items are obtained, and the user historical interaction items are input into the optimized client. The optimized client obtains the score of each item in the candidate pool, and the L items with the highest scores are combined to form a recommendation list and output the recommendation list.
[0074] The client uses the score of each item in the candidate pool, specifically:
[0075]
[0076] wherein, is the score of the client to the item , is the preference vector of the client , is the global embedding parameter of the item , is the personalized bias vector of the client u to the item ;
[0077] The items in the candidate pool include the private item pool of the client and the global system total item library of the global end;
[0078] The private item pool of the client stores the items that the client has interacted with;
[0079] The and are stored in the optimized client;
[0080] The is a parameter in the global embedding matrix;
[0081] In this step, if the item i is not an item that the client u has interacted with, then it is a random matrix.
[0082] As shown in Figure 2 , the optimized client is obtained by the following way:
[0083] Step one, initialize the client and server respectively:
[0084] Initialize the global embedding matrix to a matrix of initialized global embedding matrix The element value in the matrix is randomly sampled in a normal distribution with a mean of 0 and a preset standard deviation (for example, σ = 0.01);
[0085] wherein, is the total number of items, is a preset embedding dimension;
[0086] The i-th row in the global embedding matrix represents the global embedding parameter of item i maintained by the server ;
[0087] The client u initializes the client preference vector and the item personalized bias matrix to random values or zero vectors;
[0088] wherein, the personalized bias matrix stores the personalized bias vector of item i of the client u ;
[0089] Initialize the iteration round t = 0;
[0090] As shown in Figure 1 , the present application proposes a FeudalGCN framework, which includes a central server and a plurality of client user devices. In each client, there is a local sequence recommendation model parameter related to the user, including the preference vector of the client and the personalized item embedding vector.
[0091] Step two, at the beginning of the t-th round of federated training, the server first calculates the updated global item similarity graph using the global embedding matrix it currently holds, the latest version of the global embedding matrix , then the server selects the participating client group from the available client pool in this round, and finally the server broadcasts the public knowledge in the t-th round, i.e. and , to all clients in ;
[0092] The global item similarity graph in the t-th round stores the cosine similarity between the global embedding parameter of item i and the global embedding parameter of item in the t-th round, specifically:
[0093]
[0094] wherein, is the L2 norm, is an element in the i-th row and j-th column of the matrix;
[0095] the server selects a group of participating clients in this round from the available client pool by the following method:
[0096] First, obtain the available client pool based on preset client selection conditions ;
[0097] The preset client selection conditions are: the device has accessed the WIFI network, is charging and is in an inactive use state;
[0098] Then, the server uses an unbiased random sampling method to sample the client group from the client pool ;
[0099] The unbiased random sampling method is specifically:
[0100] The server randomly selects a fixed number or a fixed proportion (for example, 10%) of clients from with uniform probability to form .
[0101] Step three, the clients in the client group receive the global embedding matrix and the global item similarity graph , and then use , combined with the preference vector and the item personalized bias matrix of the client held locally by the client to perform local model optimization, specifically:
[0102] Step three one, the client uses the global embedding parameter to obtain the personalized item embedding vector, specifically:
[0103]
[0104] Wherein, is the personalized item embedding vector of the item , and is the personalized bias vector of the client u to the item ;
[0105] The is only saved locally on the client and never uploaded;
[0106] In this step, the client decouples the embedding representation of each item i into "the item's personalized bias vector + the item's global embedding". Each user also has a corresponding preference vector on the client. This vector, representing the overall interest, is also updated only locally. Through this decoupled embedding structure, high-entropy, personalized information in the user sequence is confined to the local device and does not appear directly in the data sent to the server, reducing the risk of sensitive information leakage at the architectural level.
[0107] Step 32: Construct the training set;
[0108] The positive samples in the training set are Negative samples are ;
[0109] in, It is the client identifier. It is a client Interacted items, It is a client Items that have not been interacted with; the labels of the samples in the training set are "self-supervised" labels, with positive samples labeled as 1 and negative samples labeled as 0;
[0110] Step 3: The client inputs the training set into the local model (the local model updated in the last iteration) and calculates the overall loss function of the local model on the training set. ;
[0111] The working process of the local model is as follows:
[0112]
[0113] in, It is a client and items Match score, It is a scoring function. It is a client The preference vector;
[0114] Steps 3 and 4: The client updates the overall loss function based on the trained local model using gradient descent. and Specifically:
[0115]
[0116]
[0117] in, This is the client-side local learning rate, a preset scalar hyperparameter. It controls the step size of the client's private parameter updates during gradient descent optimization. yes gradient value, It is the overall loss function. It is the client for the t-th round of training. The preference vector, It is the client for the (t+1)th round of training. The preference vector;
[0118] Step 3.5: Using the item personalization bias matrix of item i before the update and the item personalization bias matrix of item i after the update, obtain the knowledge increment, specifically:
[0119]
[0120] In this step, each client uses its local user behavior sequence data to perform several gradient update training iterations on the FeudalGCN model parameters (FeudalGCN includes the client's local model and the server-side parameter maintenance and aggregation mechanism). To balance personalized recommendation performance and model robustness, the client uses an objective function that includes a structure regularization term for optimization. It is a matrix of the same size as the total number of items, where each row... These represent the changes in the global embedding of item i during the client's local training process. These difference vectors condense the new information learned by the client from local data, but do not directly contain explicit traces of a specific interaction. This collaborative approach, which replaces gradients with state differences, fundamentally disrupts the highly sensitive correspondence between the original gradients and training data, increasing the difficulty of defending against attacks.
[0121] Step 4: The client filters the extracted knowledge increments and uploads the filtered knowledge increments to the server. Specifically:
[0122] Get Find the L2 norm of each row vector in the dataset, and sort the row vectors in descending order of their L2 norm to obtain the L2 norm. The first K vectors are used as the filtered knowledge increments, and the filtered knowledge increments are uploaded to the server.
[0123] In this step, the client performs Top-K filtering on the extracted knowledge increments and uploads them. Since the parameter update magnitudes of different items may vary significantly, to reduce noise interference and the amount of uploaded information, this invention requires each client to only retrieve... The K largest changes are packaged and uploaded, and the rest of the increments are discarded locally and do not participate in subsequent aggregation. Here, "K" is much smaller than the total number of items interacted by a single user , meaning that most of the user's behavior-related parameter updates are not visible to the server in a single communication, achieving a similar "K-anonymity" effect. This strategy ensures that only the part of the update that best represents general knowledge is retained in the uploaded content, thereby reducing the likelihood of sensitive information leakage while reducing communication overhead. After Top-K truncation, each client packages a filtered knowledge increment set and sends it to the server. Since the client uploads a number of optimized parameter differences in the present invention, rather than a sequence of iterative gradients, it is difficult for an attacker to easily infer the impact of a single behavior record on the model at a certain step of training, even if they have obtained a sparse knowledge increment set. This constitutes a "computational confusion" mechanism, increasing the difficulty of reverse reasoning.
[0124] Step five, the server aggregates the received knowledge increments, updates the global embedding matrix using the aggregated knowledge increments, and distributes the updated global embedding matrix to the clients, specifically:
[0125] First, the server aggregates the received knowledge increments to obtain the aggregated knowledge increments, specifically:
[0126]
[0127] wherein, is the federated averaging algorithm, is the row vector of the knowledge increment uploaded by the client u, is the set of all clients that uploaded information about item i, denotes the number of samples in the training set of client u;
[0128] Then, the aggregated knowledge increments are used to update the global embedding parameters, specifically:
[0129]
[0130] wherein, is the global embedding parameter of item i in the tth iteration, is the global embedding parameter of item i in the (t+1)th iteration, denotes the global aggregation learning rate, which is used to control the step size of the server's update of the global embedding after aggregating the knowledge increments uploaded by each client.
[0131] In this step, since each client u only uploads its Top-K sparse knowledge increments, the sets of items corresponding to the increments uploaded by different clients are usually different. The server's aggregation process is an independent aggregation process on an item-by-item basis. Specifically, for any item i in the system, the server first identifies all the knowledge increments about item i that were actually uploaded in this round of training. A subset of clients, which we denote as Then, the server only uses this subset. The information in the data is used to calculate the aggregate increment of item i via FedAvg(). If for a certain item i, no client uploads its increment in this round (i.e. If it is an empty set, then its aggregation increment is... This is a zero vector. By repeating this process for all items i, the server eventually constructs a complete, but potentially highly sparse, aggregated increment matrix. .
[0132] In this step, the server aggregates the knowledge increments from each client and updates the global model. After receiving Top-K increments from multiple clients in each round, the server aggregates the increments of the same item from different users using an element-wise weighted average method to obtain the updated global parameters, which are then applied to the global model. Unlike traditional solutions, this invention additionally considers the aforementioned dual-structure regularization constraint when the server aggregates and updates. The server adds the sum of all client regularization terms to the global loss and ensures through iterative optimization that the updated global model simultaneously minimizes these regularization terms. Regarding user-item consistency, if the global update causes a deviation between a user's preference vector and their historical behavior, the regularization term will penalize this, thereby prompting the update algorithm to adjust in the direction of maintaining user preference consistency. Regarding item similarity, if the update of some item embeddings violates the global similarity prior, the regularization term will pull the parameters in the opposite direction. This global constraint based on structural knowledge makes knowledge fusion more robust and efficient.
[0133] Step Six: Obtain the average knowledge increment norm of round t. Compare the average knowledge increment norm of round t with a preset norm threshold. If the average knowledge increment norm of round t is lower than the preset norm threshold or t=T, then obtain the optimized server-side global parameters (referred to as the optimized server) and the optimized local models of each client (referred to as the optimized client); otherwise, let t=t+1. , , And return to step two;
[0134] Where T is the preset maximum number of iterations, and in this invention, T is set to 200.
[0135] In this step, if the average knowledge increment norm of the t-th round is lower than the preset norm threshold, it is considered that the client has no obvious update, and the system tends to converge.
[0136] The optimized client stores its final and optimal local model parameters (mainly including the client preference vector and the personalized bias embedding matrix ).
[0137] In this step, the server obtains the updated global embedding parameters, and then distributes them to each client. The client receives and merges these updates, replaces the corresponding global embedding part in the local model with the new parameters distributed by the server, and prepares for the next round of training. The user preference vector and the personalized bias vector learned by the item i can continue to accumulate optimization between rounds, but the global embedding part will be reset to the new value distributed by the server every round to play the role of global "anchor". This setting ensures that the knowledge increment calculation of each round is always based on a unified global reference, making the collaborative process more stable and controllable. As multiple iterations proceed, the FeudalGCN framework gradually completes the training of the sequence recommendation model. Throughout the process, most of the user's behavior information does not directly appear in the communication content, but is reflected through a small amount of high-order knowledge increment; even if faced with a curious server, each user only exposes a very small amount of "shadow" information about their own data each round, and this information is also mixed with global prior and multi-step optimization results, greatly increasing the difficulty of inference for attackers. Therefore, the present application can ensure the practical performance of the model while achieving high-strength privacy protection for user sequence data. Finally, when the training is completed, the server obtains a recommendation model that integrates global knowledge while maintaining the personalized features of each user, and each client holds a local model component that matches its own preferences, thereby achieving the goal of providing high-quality personalized sequence recommendations while protecting user privacy.
[0138] Specific implementation method two: the total loss function of the local model in step three on the training set is:
[0139]
[0140]
[0141]
[0142]
[0143]
[0144]
[0145] in, It is the core recommendation ranking loss. It is a user-item consistency regularization term. It is an item-to-item similarity regularization term. , These are weighting coefficients. It is the set of item samples in the training data. User Interacted items, User Items that have not been interacted with It is the Sigmoid function. It is a client and items Match score, It is a client and items Match score, It is the recommendation ranking loss for user u. It is an item Global embedding parameters, It is an item Personalized bias vector, It is a collection of items that the client has interacted with. yes, It is an aggregate function. yes and cosine similarity, yes and The cosine similarity.
[0146] In this step, Defined using Bayesian personalized ranking criteria. The scoring method of this invention ensures that the final representation of items includes both global common knowledge and... Furthermore, it incorporates user u's personalized learning preferences for local learning. The two regularization components in this step aim to leverage structural information to improve the model's personalization and generalization capabilities: user-item consistency regularization. To ensure consistency between the user representation and its historical item representations, a regular expression is used. Requires user preference vectors It can be accurately reconstructed or fitted by the complete set of embedded representations of the items it has interacted with, thus ensuring that the model maintains consistency in its abstract representation of user preferences at the structural level. This means that even with the application of new privacy mechanisms, the model will not deviate from its portrayal of the user's true interests. Item-item similarity regularization. The global item similarity information issued by the server is used to constrain the local item representation, so that the model is prevented from overfitting on sparse local data, thereby improving the generalization performance of recommendation. By minimizing the difference between global and local similarity, the model is guided to align with global knowledge when updating local item embedding, avoiding pushing item representation to an extreme only according to a small amount of behavior, thereby explicitly balancing "global commonality" and "local individuality". The size of the adjustment parameter β can control the influence of the regular term, and this design enables the present application to theoretically explain the adjustment of the individualization degree of the model under different privacy protection strengths, providing a guidance basis for balancing privacy and individuality according to the needs in practical applications. The client can obtain the updated local model parameters after completing the local model training .
[0147] The present application provides a user privacy protection mechanism in a FeudalGCN framework, aiming at the problem that gradient update in the existing federated learning process may leak user sequence preferences, a scheme of replacing traditional gradient upload with "knowledge increment cooperation" is proposed. From the architecture, by decoupling the model embedding structure and screening the uploaded information, the transmission of local high-entropy individualized information to the server is blocked, thereby protecting the user sequence data privacy. At the same time, a double structure regularization strategy is introduced to ensure that the model still has good individualized expression ability and global generalization ability under the premise of protecting privacy, overcoming the defects in the prior art that only adding noise leads to model precision decline or inability to effectively utilize global knowledge. The federated sequence recommendation model constructed by the present application has low sensitivity and controllable upload information entropy, even in the face of "honest but curious" attackers trying to reconstruct user sequence data, it is difficult for the attacker to successfully restore the original behavior sequence of the user according to the obtained information, thereby significantly improving the privacy security of user data in the federated learning environment. The present application not only avoids the privacy risks brought by uploading high-entropy individualized behavior information, but also improves the balance between model recommendation individualized expression and global generalization ability through the knowledge increment screening and structure regularization cooperation mechanism, thereby achieving better recommendation effect and system efficiency while protecting user privacy.
Claims
1. A federated privacy-preserving sequence recommendation method based on knowledge increment collaboration, characterized in that: The specific process of the method is as follows: Retrieve the user's historical interaction items, input the user's historical interaction items into the optimized client, the optimized client obtains the score of each item in the candidate pool, and combines the L items with the highest scores into a recommendation list and outputs the recommendation list; The client obtains a score for each item in the candidate pool, specifically as follows: in, It is a client For items The rating, It is a client The preference vector, It is an item Global embedding parameters, Client u for items Personalized bias vector; The items in the candidate pool include items from the client's private item pool and items from the system's full item library.
2. The federated privacy-preserving sequence recommendation method based on knowledge increment collaboration according to claim 1, characterized in that: The optimized client is obtained through the following method: Step 1: Initialize the client, server, and iteration rounds respectively, specifically as follows: global embedding vector Initialize to The matrix, the initialized global embedding vector The element values are obtained by random sampling from a normal distribution with a mean of 0 and a preset standard deviation; in, It is the total number of items. It is a preset embedding dimension; The i-th row of the global embedding matrix represents the global embedding parameters of item i maintained by the server. ; Client u will have client preference vectors And item personalization bias matrix Initialize as a random vector; Among them, personalized bias matrix Personalized bias vector of client u for item i in storage ; Initialize the iteration round t=0; Step 2: The server utilizes the global embedding matrix. Get global item similarity map Subsequently, the server selects the participating client group from the pool of available clients for this round. Finally, the server sent... All client broadcasts and ; Step 3: Participate in the client group The client in the middle receives the global embedding matrix Similarity map with global items Then, using , Combine with client Preference vector And item personalization bias matrix Perform local model optimization to obtain the updated client. Preference vector and the updated item personalization bias matrix Then, knowledge increments are obtained using the item personalization bias matrix. ; Step 4: The client filters the knowledge increments extracted in Step 3 and uploads the filtered knowledge increments to the server. Step 5: The server aggregates the received knowledge increments, updates the global embedding matrix using the aggregated knowledge increments, and sends the updated global embedding matrix to the client. Step Six: Obtain the average knowledge increment norm of round t. Compare the average knowledge increment norm of round t with a preset norm threshold. If the average knowledge increment norm of round t is lower than the preset norm threshold or t=T, then obtain the optimized server-side and optimized client-side; otherwise, let t=t+1. , , And return to step two; The optimized client stores the updated client preference vector and the updated item personalization bias matrix; Where T is the preset maximum number of iterations.
3. The federated privacy-preserving sequence recommendation method based on knowledge increment collaboration according to claim 2, characterized in that: The server in step two utilizes a global embedding matrix. Get global item similarity map Specifically: in, It is the L2 norm. yes The element in the i-th row and j-th column.
4. The federated privacy-preserving sequence recommendation method based on knowledge increment collaboration according to claim 3, characterized in that: The server selects the participating client group for this round from the pool of available clients. Specifically: First, obtain the pool of available clients based on preset client selection criteria. ; Then, the server uses an unbiased random sampling method to sample from the client pool. Mid-sampling to obtain client groups .
5. The federated privacy-preserving sequence recommendation method based on knowledge increment collaboration according to claim 4, characterized in that: The participating client group in step three The client in the middle receives the global embedding matrix Similarity map with global items Then, using , Combine with client Preference vector And item personalization bias matrix Perform local model optimization to obtain the updated client. Preference vector and the updated item personalization bias matrix Then, knowledge increments are obtained using the item personalization bias matrix. Specifically: Step 31: The client obtains the personalized item embedding vector using the global embedding parameters, specifically: in, It is an item Personalized item embedding vectors, It is the client u to items Personalized bias vector, It is the global embedding parameter of item i; Step 3.2: Construct a training set using personalized item embedding vectors; The positive samples in the training set are Negative samples are ; in, It is the client identifier. It is a client Interacted items, It is a client Items that have not been interacted with are labeled 1 for positive samples and 0 for negative samples. Step 3: The client inputs the training set into the local model and calculates the overall loss function of the local model on the training set. The working process of the local model is as follows: in, It is a client and items Match score, It is a scoring function. It is a client The preference vector; Steps three and four: The client updates the item-specific bias matrix based on the overall loss function of the local model on the training set, and the client... The preference vector; Step 3.5: Use the item personalization bias matrix of item i before the update and the item personalization bias matrix of item i after the update to obtain knowledge increment.
6. The federated privacy-preserving sequence recommendation method based on knowledge increment collaboration according to claim 5, characterized in that: The overall loss function of the local model on the training set in step 33 is as follows: in, It is the core recommendation ranking loss. It is a user-item consistency regularization term. It is an item-to-item similarity regularization term. , These are weighting coefficients. It is the set of item samples in the training data. User Interacted items, User Items that have not been interacted with It is the Sigmoid function. It is a client and items Match score, It is a client and items Match score, It is the recommendation ranking loss for user u. It is an item Global embedding parameters, It is an item Personalized bias vector, It is a collection of items that the client has interacted with. It is an aggregate function. yes and cosine similarity, yes and The cosine similarity.
7. The federated privacy-preserving sequence recommendation method based on knowledge increment collaboration according to claim 6, characterized in that: In steps three and four, the client updates the item-personalized bias matrix based on the loss function of the trained local model and the client... The preference vector is as follows: in, It is the client's local learning rate. yes gradient value, It is the overall loss function. It is the client for the t-th round of training. The preference vector, It is the client for the (t+1)th round of training. The preference vector.
8. The federated privacy-preserving sequence recommendation method based on knowledge increment collaboration according to claim 7, characterized in that: The knowledge increment obtained in step three-five, using the item personalization bias matrix of item i before and after the update, is as follows: 。 9. A federated privacy-preserving sequence recommendation method based on knowledge increment collaboration according to claim 8, characterized in that: In step four, the client filters the knowledge increments extracted in step three and uploads the filtered knowledge increments to the server. Specifically: Get Find the L2 norm of each row vector in the dataset, and sort the row vectors in descending order of their L2 norm to obtain the L2 norm. The first K vectors are used as the filtered knowledge increments, and the filtered knowledge increments are uploaded to the server.
10. A federated privacy-preserving sequence recommendation method based on knowledge increment collaboration according to claim 9, characterized in that: In step five, the server aggregates the received knowledge increments and updates the global embedding matrix using the aggregated knowledge increments. Specifically: First, the server aggregates the received knowledge increments to obtain the aggregated knowledge increments, specifically: in, It is a federal average algorithm. This is the first knowledge increment uploaded by the client u. row vectors It is a collection of all clients that have uploaded information about item i. This represents the number of samples in the training set of client u; Then, utilize the aggregated knowledge increment Update the global embedding parameters as follows: in, These are the global embedding parameters of item i in the t-th iteration. These are the global embedding parameters of item i in the (t+1)th iteration. This represents the learning rate of the global aggregation.
Citation Information
Cited By
Efficient lightweight federal recommendation method
CN121598334A