Personalized cross-domain recommendation method and system based on federal learning

Through the two-stage training method of federated learning and model parameter decomposition technology, the problems of privacy protection and personalized recommendation in cross-domain recommendation are solved, and personalized models are trained on users' local devices are realized, "cold start" is alleviated, and the accuracy and efficiency of recommendation services are improved.

CN120492743APending Publication Date: 2025-08-15NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510450184.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, there are problems such as risk of privacy data leakage, differences in user potential feature mapping relationships, differences in user-item interaction data distribution, and insufficient performance of personalized recommendation services. Especially in cross-domain recommendations, how to achieve personalized recommendation services and alleviate the problem of "cold start" while protecting user privacy.

Method used

Using a two-stage training method based on federated learning, the single-domain score prediction model is first trained in a single domain by neural collaborative filtering, and then the mapping relationship of user potential feature representation is captured through the migration module of the multi-layer neural network, and each layer of the network of the cross-domain recommendation model is decomposed into base vectors and personalized vectors. The server aggregates the base vectors, and the user retains the personalized vectors to achieve personalized recommendations.

Benefits of technology

While protecting user privacy, it improves the personalization of recommendation services, reduces the computational complexity and communication costs, effectively alleviates the problem of "cold start", and realizes high-precision personalized cross-domain recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492743A_ABST
    Figure CN120492743A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized cross-domain recommendation method and system based on federal learning, and the method provides a personalized cross-domain recommendation service for a user on the premise that original data of interaction between the user and an article and user parameters are kept locally. The method comprises two stages of federated training: stage 1, intra-domain users cooperatively train a single-domain score prediction model by using a neural collaborative filtering method; in the second stage, overlapping users of the two domains cooperatively train a migration module based on a multi-layer neural network to capture a mapping relation represented by potential user features between the two domains; besides, each layer of network of the cross-domain recommendation model is decomposed into a base vector and a personalized vector which respectively represent common knowledge among different users and unique knowledge of the users, and a local model obtained by final training can provide personalized recommendation services for registered users in a target domain, so that the recommendation efficiency is improved, and the user experience is improved. Meanwhile, the global model obtained through training provides effective initial recommendation for new users in the target domain, and the cold start problem in a recommendation algorithm is effectively relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of the intersection of federated learning and recommendation systems, and specifically relates to a personalized cross-domain recommendation method and system based on federated learning. Background Art

[0002] In the age of information explosion, recommender systems have been used to support various applications in our daily lives to alleviate information overload, such as e-commerce, social media, and online news. Recommender systems intelligently filter out low-value information based on user preferences and recommend items that meet their needs. The key to personalized recommendation systems is modeling item preferences based on users' past interactions with items, a process known as collaborative filtering. Matrix factorization, a mainstream approach to collaborative filtering, projects users and items into a shared latent space, using a latent feature vector to represent each user or item. User-item interactions are modeled as the inner product of their latent vectors. However, its performance is limited by its linear representation (the inner product of two vectors). Consequently, researchers have begun exploring the use of deep neural networks to learn the mapping between user and item interactions, achieving significant success. This is the neural network collaborative filtering approach.

[0003] Cross-domain recommendation is a branch of recommendation algorithms. Its core idea is to transfer user knowledge from the source domain (data-rich domain) to the target domain (data-free domain) to provide initial recommendations for users in the target domain. It was proposed to address the "cold start" problem in recommendation algorithms. Specifically, any application constantly adds new users. However, since these new users have no historical interaction records with items in the domain, traditional recommendation algorithms are unable to provide effective recommendations for these users. Cross-domain recommendation, on the other hand, leverages users' knowledge in the source domain to provide recommendations in the target domain. While cross-domain recommendation can effectively alleviate the "cold start" problem, traditional cross-domain recommendation requires centralized collection of user-item interaction data from multiple platforms, which seriously threatens user privacy. Federated learning provides a decentralized model training method that leverages the collective strengths of multiple users without the need for centralized data aggregation, effectively protecting user privacy. However, due to the differences in individual user preferences for items and the heterogeneity of user data distribution, simple federated averaging methods lead to over-smoothing of user local features, making personalized cross-domain recommendation services ineffective. Therefore, to achieve personalized cross-domain recommendation services that protect user privacy, the following issues need to be addressed:

[0004] 1) Due to privacy concerns, users’ original data and user-related parameters should be retained on their local devices. The server needs to train the cross-domain recommendation model without obtaining user-related information.

[0005] 2) There is a mapping relationship between the user potential feature vectors of the same user in different domains, but it is necessary to transfer the user's knowledge in the source domain to the target domain through reasonable and effective algorithms and modeling.

[0006] 3) Due to differences in user item preferences, the distribution of user-item interaction data, and the amount of user-item interaction data, local model training can vary significantly between users. If the server simply averages and aggregates all local models, the performance of the final global model will be reduced, and personalized recommendation services for each user will not be achieved.

[0007] 4) The actual computing, storage, and communication costs of the user’s local device must be considered. On the one hand, the transmission of all parameters of the cross-domain recommendation model results in high bandwidth consumption. On the other hand, the user’s local device does not support the deployment of large models.

[0008] To address these issues, this paper studies how to use a federated learning framework to provide users with privacy-protected personalized cross-domain recommendation services and deploy them on a large scale. However, this is not an easy task for the following reasons:

[0009] a. The latent spaces of users and items in different domains have different characteristics, and the embedding layer dimensions between different domains may vary. Effectively learning the latent feature vectors of users and items in each domain and then learning the mapping relationship between latent feature vectors of users or items across domains, given the different embedding layer dimensions, to transfer knowledge between domains is a complex problem.

[0010] b. Existing federated learning frameworks assume that user data is independent and uniformly distributed, but user behavior in actual federated recommendation scenarios exhibits significant heterogeneity. An effective method is needed to extract shared knowledge among users while preserving their unique, personalized knowledge, while safeguarding their privacy. This method can be used to collaboratively train distributed user groups under server coordination. Achieving personalized cross-domain recommendations is a challenge that this invention aims to overcome.

[0011] c. The user objects of the "cold start" problem include not only partially overlapping users with sparse data in the target domain, but also new users with no interaction data in the target domain. It is necessary not only to maintain a personalized local cross-domain recommendation model for overlapping users, but also to maintain a general global model to implement initial recommendations for new users in the target domain, which makes the problem more complicated. Summary of the Invention

[0012] Purpose of the invention: To address the problems of privacy data leakage risks, differences in user potential feature mapping relationships, differences in user-item interaction data distribution, and insufficient performance of personalized recommendation services in the above-mentioned prior arts, the present invention provides a personalized cross-domain recommendation method and system based on federated learning. This method realizes effective personalized cross-domain recommendation services while retaining the user's original data and specific user parameters locally.

[0013] Technical solution: A personalized cross-domain recommendation method based on federated learning, which includes two stages of federated training:

[0014] Phase 1: In-domain users collaborate to train a single-domain rating prediction model based on neural collaborative filtering;

[0015] Phase 2: Collaborative training of overlapping users in the two domains using a multi-layer neural network-based transfer module to capture the mapping relationship between the user's latent feature representations in the two domains.

[0016] The method decomposes model parameters in the following manner: each layer of the cross-domain recommendation model is decomposed into basis vectors and personalized vectors. The basis vectors represent the common knowledge between different users and are aggregated by the server. The personalized vectors represent the user's unique knowledge and can be retained locally. The local model finally trained can provide high-precision personalized recommendation services for registered users in the target domain, and the global model trained can implement initial recommendations for new users in the target domain.

[0017] Furthermore, the method is specifically used for training the single-domain rating prediction model in stage 1 as follows:

[0018] (11) Define the user set U in the domain = {u1,u2,…,u M}, item set V = {v1,v2,…,v N}, user-item real rating matrix And the user-item prediction rating matrix generated by the rating prediction model

[0019] From a single-domain perspective, discrete users and items are projected into a shared latent space, and an embedded feature vector is used to replace a user or an item. The mathematical expression is as follows:

[0020]

[0021] in, represents the embedding feature vector of a single user, i.e. user embedding, Represents the embedding feature vector of a single item, i.e., item embedding, represents the item embedding matrix formed by concatenating all items in the domain, K represents the embedding dimension, N represents the number of items, h represents the embedding layer, θ and Represent the embedding layer parameters of user and item sets respectively;

[0022] (12) Define the optimization objective of the rating prediction model to minimize the total prediction loss. Its mathematical expression is as follows:

[0023]

[0024] Among them, Θ f The model parameter f representing the interaction function is defined as a multi-layer neural network, which can be expressed as:

[0025]

[0026] Among them, φ out and φ X Represent the neural collaborative filtering mapping functions of the output layer and the x-th layer respectively, embedding the user Item embeddings for interactive items The concatenation is used as the input of the rating prediction model, and the output is the user's predicted rating for the item;

[0027] (13) According to the mean square error, the loss between the predicted value and the true value is measured. i The mathematical expression of the local model loss function is as follows:

[0028]

[0029] Among them, r ij Represents user u i The real rating data, It is the predicted score data output by the local model;

[0030] (14) In the initial stage of federated learning, the server will send the initialized global item embedding matrix parameters Q and neural network model parameters Θ f For the selected participating users, each user initializes their own user embedding locally In local training, Q,Θ f All of them are used as trainable parameters, and the model is trained based on the loss function in step (13), and the local user embedding parameters are updated through reverse gradient propagation Item embedding parameters and neural network model parameters Updated are kept locally as user-specific parameters, while the updated Upload to the server for aggregation, and the aggregation process is defined as:

[0031]

[0032] In the next round of federated learning, the server will send the aggregated results as the training parameters for the next round to the user. The user uses private data to Q (t+1) ,Θ f (t+1) Conduct local training;

[0033] (15) After each round of federated learning, the performance of the global model is evaluated based on the mean absolute error and root mean square error, and the similarity between the predicted value of the rating prediction model and the true value of the user-item is obtained. Its mathematical expression is as follows:

[0034]

[0035] Among them, D test Represents the test data of each user. The test set size is set by dividing the data set. ij and Represents user u i Item v j The true and predicted ratings of

[0036] (16) Perform multiple rounds of federated learning iterations until the model converges or reaches the predetermined training rounds, and the server saves the trained global item embedding parameter Q A and Q B , and the single domain score prediction model parameters and Used for training the cross-domain recommendation model in stage 2.

[0037] Furthermore, the method is specifically as follows for cross-domain recommendation model training in stage 2:

[0038] (21) The federated cross-domain recommendation framework includes domain A and domain B. The cross-domain recommendation task from source domain A to target domain B is considered as follows: domain A and domain B contain user sets U A and U B , Item Set V A and V B , and the user-item rating matrix R A and R B , the overlapping users of the two domains are defined as U O =U A ∩U B , overlapping user data

[0039] (22) Construct a cross-domain recommendation model to capture the mapping relationship between user embeddings from the source domain to the target domain, using the trained user embeddings of overlapping users in domain A and domain B respectively. and To capture the mapping relationship of user embedding between two domains, its mathematical expression is as follows:

[0040]

[0041] Among them, Φ out and Φ X Represent the output layer and x-th layer model parameters of the neural network respectively, is the overlapping user u i The user embedding trained in domain A is used as the input of the cross-domain recommendation model, and the model outputs the user u i The predicted user embedding in domain B is K and Z represent the embedding layer dimensions of domain A and domain B, respectively;

[0042] (23) Decompose the cross-domain recommendation model parameters in step (22) into a migration module and a personalization module, specifically:

[0043] Step (22) provides a local model g for each user A→B , the model parameters of each layer are i and O represent the input dimension and output dimension of the network parameters of this layer, respectively. The high-order matrix Φ of each layer is decomposed into the outer product of two vectors to separate the public knowledge and personalized knowledge of the local network and reduce the network complexity. Its mathematical expression is as follows:

[0044]

[0045] Among them, a k is the basis vector, representing the common knowledge between clients, b k It is a personalized vector, representing the knowledge unique to each client;

[0046] The local model for the user ends up being:

[0047]

[0048] (24) Based on the global item embedding and single-domain rating prediction model trained in stage 1, the final recommendation result is optimized by adding item information to the optimization objective. Specifically:

[0049] Embed the reconstructed predicted user The item embedding Q of the domain B After splicing, it is input into the rating prediction model f trained in stage 1 BIn the algorithm, the average prediction loss is required to be minimized to further optimize the cross-domain recommendation model parameters. The mathematical expression of the loss function is as follows:

[0050]

[0051] Among them, α and β are adjustable hyperparameters, the first term of the loss function is the reconstruction loss, and the second term is the rating loss of using the reconstructed user embedding for prediction; It means using the user embedding of overlapping users in domain A to reconstruct the user embedding of the user in domain B. Indicates the real user embedding of the user in domain B; represents the user embedding of domain B reconstructed using this user and the items in domain B are embedded in Q B Predict user ratings for items in domain B;

[0052] (25) In the initial stage of federated learning, the server uniformly sends the basis vectors of each layer of the network For the selected participating users, during local training, each local user will and As trainable parameters, the cross-domain recommendation model is trained using the loss function in step (24) and updated by reverse gradient propagation, where As the user's personalized knowledge does not participate in aggregation, the updated Upload to the server for aggregation, and the aggregation process is defined as:

[0053]

[0054] In the next round of federated learning, the server will send the aggregated results as the training parameters for the next round to the user. The user will use private data for local training and update the and Upload updated

[0055] (26) After each round of federated learning, the performance of the cross-domain recommendation model is evaluated, and its mathematical expression is as follows:

[0056]

[0057] in, represents the test data of overlapping users in target domain B, and Represent overlapping users u i For item v in target domain B j The true and predicted ratings of

[0058] (27) Multiple rounds of federated learning iterations are performed until the model converges or reaches the predetermined training rounds, and finally each user uses the trained global model parameters and locally updated The mathematical expression for personalized local cross-domain recommendation model fusion is as follows:

[0059]

[0060] Before the end of federated training, the server collects personalized knowledge of all users in the last round And perform federal averaging, the mathematical expression is as follows:

[0061]

[0062] The server then uses and Perform global model fusion to provide initial recommendation services for new users in the target domain.

[0063] Furthermore, the method is suitable for completing the initial cross-domain recommendation of items to new users in the target domain. The items include products, movies and books, and can provide users with personalized cross-domain recommendation services while retaining the original data of the user-item interaction and user parameters in their local environment.

[0064] The present invention provides a personalized cross-domain recommendation method based on federated learning. Based on the federated learning framework, it is aimed at collaborative training of multi-user models, with the goal of optimizing the performance of individual user models and the performance of general global models. It collaboratively trains multi-layer neural networks to capture the mapping relationship between domains. Based on the model parameter decomposition method, the network parameters are decomposed into common knowledge parts and personalized knowledge parts. An aggregation method is designed to retain local personalized knowledge while optimizing global common knowledge.

[0065] Based on the above method, the present invention constructs a personalized cross-domain recommendation system based on federated learning. The system has the following deployment on the user's client:

[0066] The single-domain rating prediction module performs the single-domain rating prediction model training in stage 1 of the above method;

[0067] The cross-domain recommendation module performs the cross-domain recommendation model training in stage 2 of the above method.

[0068] Furthermore, the cross-domain recommendation module includes a migration module and a personalization module. The migration module is used to realize the migration of user preferences between different domains, and the personalization module is used to provide personalized recommendations for heterogeneous data distribution.

[0069] Beneficial effects: Compared with the prior art, the present invention has substantial progress and significant effects as follows:

[0070] 1) This paper designs a novel personalized cross-domain recommendation model training process architecture based on federated learning. Under the premise of protecting user data privacy and security, users collaborate to train global models and personalized models, and ultimately realize personalized new domain recommendation services, effectively alleviating the "cold start" problem in the recommendation system.

[0071] 2) This paper comprehensively considers the factors affecting cross-domain recommendation under the federated learning framework, including the use of a parameter decomposition-based method to retain user personalized knowledge, thereby improving local model performance while reducing computational complexity and communication costs. In addition, considering the differences in mapping relationships between different domains and the differences in feature dimensions of the latent space, a cross-domain recommendation model based on a neural network is constructed, and user embedding and item embedding are used in synergistic manner to transfer knowledge between domains, thereby achieving effective optimization of the cross-domain recommendation model.

[0072] 3) Throughout the training process of Phase 1 and Phase 2, the user's original data and user-specific parameters are retained on the user's local device. The server is only responsible for aggregating the model parameters of the item embedding matrix, the rating prediction model parameters, and the basis vector part of the cross-domain recommendation model, thereby effectively protecting the user's privacy.

[0073] 4) The present invention verifies the effectiveness of the method by performing three different cross-domain rating prediction tasks on the representative Amazon 5-core dataset, namely "Movies & Books", "Music & Books", and "Music & Movies". It also verifies that the present invention can effectively separate the public knowledge and personalized knowledge in the model, achieving effective personalized recommendations while quickly converging. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 This is a schematic diagram of the training process of the single-domain rating prediction model in stage 1;

[0075] Figure 2 This is a schematic diagram of the cross-domain recommendation model training process in stage 2;

[0076] Figure 3 This is the first stage of the cross-domain recommendation task "Movies & Books" movie domain ( Figure 3 (a)) and the book domain ( Figure 3 (b) Single domain score prediction experimental results;

[0077] Figure 4 This is the first stage of the cross-domain recommendation task "Music & Books" music domain ( Figure 4 (a)) and the book domain ( Figure 4(b) Single domain score prediction experimental results;

[0078] Figure 5 This is the first stage of the cross-domain recommendation task "Music & Movies" music domain ( Figure 5 (a)) and the film domain ( Figure 5 (b) Single domain score prediction experimental results;

[0079] Figure 6 is from the source domain "movie" to the target domain "book" ( Figure 6 (a)) and from the source domain “books” to the target domain “movies” ( Figure 6 (b)) Cross-domain rating prediction experimental results;

[0080] Figure 7 is from the source domain "music" to the target domain "books" ( Figure 7 (a)) and from the source domain “books” to the target domain “music” ( Figure 7 (b)) Cross-domain rating prediction experimental results;

[0081] Figure 8 is from the source domain "music" to the target domain "movie" ( Figure 8 (a)) and from the source domain "movie" to the target domain "music" ( Figure 8 (b)) Cross-domain rating prediction experimental results;

[0082] Figure 9 This is a graph of experimental results of different decomposition ranks when the model parameters are decomposed. DETAILED DESCRIPTION

[0083] In order to illustrate the technical solution disclosed in the present invention in detail, further elaboration is given below in conjunction with the accompanying drawings and specific embodiments.

[0084] Training a cross-domain recommendation system based on a federated learning framework can effectively protect the privacy and security of users. However, differences in user preferences for items, differences in the distribution of user-item interaction data, and differences in the amount of user-item interaction data will lead to huge differences in the training of local models between different users. If the server simply averages and aggregates all local models, it will not only fail to obtain a global model with good generalization performance, but also fail to provide personalized recommendation services for each user. In addition, there are feature differences in the latent spaces of users and items between different domains, and the embedding layer dimensions of the latent spaces of different domains may be different, which also increases the difficulty of training cross-domain recommendation models, especially the initial recommendation and "cold start" problems for new users in the target domain. In response to these problems, the present invention provides a personalized cross-domain recommendation method based on federated learning. Under the premise of retaining the user's original data and specific user parameters in their local area, users collaborate to train global models and personalized models to achieve effective personalized cross-domain recommendations. This method involves two stages of federated training. In the first stage, users in each domain use the federated neural network collaborative filtering method to train a single-domain rating prediction model, and the server is responsible for aggregating the global model and item embeddings. In the second stage, overlapping users in the two domains collaboratively train a migration module based on a multi-layer neural network to capture the mapping relationship between the user's potential feature representations between the two domains.

[0085] After users have implemented cross-domain recommendation model training based on federated learning, the present invention constructs a model parameter decomposition method that decomposes each layer of the cross-domain recommendation model into a basis vector and a personalized vector. The basis vector represents the common knowledge between different users, and the server is responsible for aggregating it. The personalized vector represents the user's own unique knowledge and can be retained locally. The local model finally trained can provide high-precision personalized recommendations for registered users in the target domain, and the global model implements initialization recommendations for new users in the target domain, effectively alleviating the "cold start" problem. The present invention verifies the effectiveness of this method by performing three different cross-domain rating prediction tasks on the representative Amazon 5-core dataset; at the same time, it verifies that the present invention can effectively separate the common knowledge and personalized knowledge in the model, achieving effective personalized recommendations while quickly converging.

[0086] The specific implementation steps of the technical solution provided by the present invention can be described as follows:

[0087] Stage 1: Single-domain rating prediction model training

[0088] (11) The training process of the single-domain rating prediction model based on federated learning is as follows Figure 1 As shown in Figure 2, some independent users and some overlapping users in the two data domains are selected to form the federated learning subjects. From a single-domain perspective, the federated neural network collaborative filtering method is used to train the rating prediction models of the two domains respectively.

[0089] (12) Define the user set U in the domain = {u1,u2,…,u M}, item set V = {v1,v2,…,v N}, user-item real rating matrix And the user-item prediction rating matrix generated by the rating prediction model The user and item are projected into a shared latent space through the embedding layer, and the embedded feature vector is used to replace a user or an item. The mathematical expression is as follows:

[0090]

[0091] in, represents the embedding feature vector of a single user (referred to as user embedding for short), represents the embedding feature vector of a single item (referred to as item embedding), represents the embedding feature matrix of all item sets in the domain, K represents the embedding dimension, h represents the embedding layer, θ and represent the embedding layer parameters of user and item sets respectively.

[0092] In the formula, the original user-item data and the user's embedded feature vector parameters are retained on the local device to protect user privacy. The embedded feature vectors of the items are combined to form the item embedding matrix Q of the entire domain for subsequent training.

[0093] (13) The rating prediction model is trained using the federated neural network collaborative filtering method and the optimization objective is defined. Its mathematical expression is as follows:

[0094]

[0095] Among them, D i is the user's local data, Θ f The model parameter f representing the interaction function is defined as a multi-layer neural network, which can be expressed as:

[0096]

[0097] Among them, φ out and φ X Represent the neural collaborative filtering mapping functions of the output layer and the x-th layer respectively. Item embeddings for interactive items The spliced together are used as the input of the rating prediction model f, and the model outputs the user's predicted rating for the item

[0098] (14) Rating prediction in the recommendation system is a regression problem, that is, predicting the user-item rating matrix. Therefore, the mean square error (MSE) is used to measure the loss between the predicted value and the true value. i The mathematical expression of the local loss function is as follows:

[0099]

[0100] Among them, r ij Represents user u i The real rating data, It is the predicted score data output by the model.

[0101] (15) In the initial stage of federated learning, the server will send the initialized global item embedding matrix parameters Q and neural network model parameters Θ f For the selected participating users, the users generate the initialized user embedding locally In local training, Q,Θ f All of them are used as trainable parameters, and the model is trained using the loss function in step (14). The local user embedding parameters are updated through reverse gradient propagation. Item embedding parameters and neural network model parameters Updated are kept locally as user-specific parameters, while the updated Upload to the server for aggregation, the aggregation process can be defined as:

[0102]

[0103] In the next round of federated learning, the server will send the aggregated results as the training parameters for the next round to the user. The user uses private data to Q (t+1) ,Θ f (t+1) Perform local training and repeat the training process.

[0104] (16) After each round of federated learning, the performance of the single-domain rating prediction model is evaluated. The present invention uses the mean absolute error (MAE) and root mean square error (RMSE) to evaluate the performance of the global model. MAE and RMSE evaluate the similarity between the predicted value of the rating prediction model and the actual value of the user-item. They have been widely used in recommendation systems. Their mathematical expressions are as follows:

[0105]

[0106] Among them, D testRepresents the test data of each user. The test set size can be set by dividing the data set. ij and Represents user u i Item v j The true and predicted ratings.

[0107] The training results are visualized as follows Figure 3 、 4 ,5, the model performance measurement indicators are the above-mentioned MAE and RMSE, and all models converge within the predetermined training rounds. Figure 3 This is the first stage of the cross-domain recommendation task "Movies & Books" movie domain ( Figure 3 (a)) and the book domain ( Figure 3 (b) Single domain score prediction experimental results; Figure 4 This is the first stage of the cross-domain recommendation task "Music & Books" music domain ( Figure 4 (a)) and the book domain ( Figure 4 (b) Single domain score prediction experimental results; Figure 5 This is the first stage of the cross-domain recommendation task "Music & Movies" music domain ( Figure 5 (a)) and the film domain ( Figure 5 (b)) Single domain score prediction experimental results.

[0108] (17) Perform multiple rounds of federated learning iterations until the model converges or reaches the predetermined number of training rounds. The server saves the trained global item embedding parameters Q A and Q B , and the single domain score prediction model parameters and Used for training the second-stage cross-domain recommendation model.

[0109] Phase 2: Cross-domain recommendation model training:

[0110] (21) The federated cross-domain recommendation framework is defined to contain two different domains: domain A and domain B. Domain A and domain B contain user sets U A and U B , Item Set V A and V B , and the user-item rating matrix R A and R B , the overlapping users of the two domains are defined as U O =U A ∩U B , overlapping user data

[0111] (22) The training process of personalized cross-domain recommendation model based on federated learning is as follows Figure 2As shown in Figure 2, the overlapping users who participated in the first stage of training in the two data domains are selected to form the federated learning subjects in the second stage. Assuming that the goal is to transfer the user's knowledge in the source domain A to the target domain B, the method is to use the user embeddings of the overlapping users in domains A and B respectively. and and the global item embedding Q of the target domain B and single domain score prediction model f B To learn the mapping relationship between user embeddings of the two domains, and use the method of model parameter decomposition from a dual-domain perspective to perform personalized cross-domain recommendations.

[0112] (23) Design a migration module based on a multi-layer neural network, using the trained user embeddings of overlapping users in domain A and domain B. and To capture the mapping relationship of user embedding between two domains, its mathematical expression is as follows:

[0113]

[0114] Among them, Φ out and Φ X Represent the output layer and x-th layer model parameters of the neural network respectively, is the overlapping user u i User embedding in domain A is used as model input, and the model outputs overlapping users u i The predicted user embedding in domain B is K and Z represent the embedding layer dimensions of domain A and domain B respectively, and they may be equal or unequal. During model training, the user embeddings of overlapping users are and It is always kept locally, and the server is only responsible for parameter aggregation of the migration module.

[0115] (24) Step (23) provides a local model g for each user, and the model parameters of each layer are I and O represent the input and output dimensions of the network parameters of this layer, respectively. We decompose the high-order matrix Φ of each layer into the outer product of two vectors to separate the public knowledge and personalized knowledge of the local network and reduce network complexity. The mathematical expression is as follows:

[0116]

[0117] Among them, a k is the basis vector, representing the common knowledge between clients, b k is a personalized vector that represents the unique knowledge of each client. The user's local model can finally be rewritten as follows:

[0118]

[0119] (25) Define the optimization goal of the cross-domain recommendation model, requiring the reconstructed predicted user embedding Embedded with real users The average distance between them is the smallest. However, it is difficult to achieve the expected optimization effect by using only user features to learn transfer knowledge, because the scoring task is not only closely related to the user, but also closely related to the domain item embedding. Therefore, based on the one-stage trained global item embedding and single-domain recommendation model, the present invention further optimizes the performance of the final recommendation result by adding item information to the optimization target. Specifically, the reconstructed predicted user embedding The item embedding Q of the domain B After splicing, it is input into the rating prediction model f trained in the first stage B In the above example, the average prediction loss is required to be minimized to achieve further optimization. The mathematical expression of the loss function is as follows:

[0120]

[0121] Among them, α and β are adjustable hyperparameters, the first term of the loss function is the reconstruction loss, and the second term is the rating loss of using the reconstructed user embedding for prediction; It means using the user embedding of overlapping users in domain A to reconstruct the user embedding of the user in domain B. Indicates the real user embedding of the user in domain B; It means using the user embedding of domain B reconstructed by the user and the item embedding of domain B to predict the user's rating of the item in domain B.

[0122] (26) Formulate a federated learning process. In the initial stage of federated learning, the server will uniformly send the basis vectors of each layer of the network. To the selected participating users; in local training, each local user will and As a trainable parameter, the model is trained using the loss function in step (25) and updated by back-gradient propagation. As the user's personalized knowledge does not participate in aggregation, the updated Upload to the server for aggregation, the aggregation process can be defined as:

[0123]

[0124] In the next round of federated learning, the server will send the aggregated results as the training parameters for the next round to the user. The user will use private data for local training and update the Upload updated Repeat this training process.

[0125] (27) After each round of federated learning, the performance of the cross-domain recommendation model is evaluated. Similar to step (26) in phase 1, its mathematical expression is as follows:

[0126]

[0127] in, represents the test data of overlapping users in target domain B, and Represent overlapping users u i For item v in target domain B j The true and predicted ratings.

[0128] The final evaluation results of the three cross-domain tasks are as follows Figure 6 、 7 As shown in Figure 8, the model performance measurement indicator is the MAE mentioned above, and all models converged within the predetermined training rounds. Under the premise that other experimental settings remain unchanged, "BaseModel" in the figure means that the cross-domain recommendation model is not decomposed, and users collaborate to train the complete cross-domain recommendation model; "Agg_a" represents the method of the present invention, that is, the cross-domain recommendation model network is parameter-decomposed, and the server only aggregates the basis vectors to save the user's personalized knowledge; "Agg_a&b" means that the server aggregates both the basis vectors and the personalized vectors; "Agg_a&b" represents the experimental results in which the server only aggregates the personalized vectors. The final experimental results show that a k is the basis vector, b k It is the user's personalized vector. Aggregating only the basis vector, that is, the user's common knowledge, and not aggregating the personalized vector, that is, the user's personalized knowledge, will make the final model perform better.

[0129] Specifically, Figure 6 is from the source domain "movie" to the target domain "book" ( Figure 6 (a)) and from the source domain “books” to the target domain “movies” ( Figure 6 (b)) Cross-domain rating prediction experimental results; Figure 7 is from the source domain "music" to the target domain "books" ( Figure 7 (a)) and from the source domain “books” to the target domain “music” ( Figure 7 (b)) Cross-domain rating prediction experimental results; Figure 8 is from the source domain "music" to the target domain "movie" ( Figure 8 (a)) and from the source domain "movie" to the target domain "music" ( Figure 8 (b)) Cross-domain rating prediction experimental results.

[0130] In addition, we conducted experiments on the size of the network decomposition rank of the cross-domain recommendation model, such as Figure 9 When the decomposition rank is 1, each network layer is decomposed into a basis vector and a personalized vector; when the decomposition rank is 2 or above, each network layer is decomposed into a basis matrix and a personalized matrix. We tested the model performance with decomposition ranks of 1, 2, 5, 10, and 20, and found that the models all converged to the same interval within the predetermined training rounds, and the communication volume between the user and the server was minimized when the decomposition rank was 1.

[0131] (28) Multiple rounds of federated learning iterations are performed until the model converges or reaches the predetermined number of training rounds. Finally, each user uses the trained global model parameters and locally updated The mathematical expression for personalized local cross-domain recommendation model fusion is as follows:

[0132]

[0133] Before the end of federated training, the server collects personalized knowledge of all users in the last round And perform federal averaging, the mathematical expression is as follows:

[0134]

[0135] The server then uses and Perform global model fusion to provide initial recommendation services for new users in the target domain.

[0136] We validated our approach by running three different cross-domain rating prediction tasks on a representative Amazon 5-core dataset. We also verified that our approach effectively separates common and personalized knowledge from the model, achieving effective personalized recommendations while achieving rapid convergence. The experiments show that our proposed approach consistently achieves the best performance across all cross-domain recommendation tasks, while effectively preserving the user's personalized knowledge.

Claims

1. A personalized cross-domain recommendation method based on federated learning, characterized by: The method involves two stages of federated training: Phase 1: In-domain users collaborate to train a single-domain rating prediction model based on neural collaborative filtering; Phase 2: Collaborative training of overlapping users in the two domains using a multi-layer neural network-based transfer module to capture the mapping relationship between the user's latent feature representations in the two domains. The method decomposes model parameters in the following manner: each layer of the cross-domain recommendation model is decomposed into basis vectors and personalized vectors. The basis vectors represent the common knowledge between different users and are aggregated by the server. The personalized vectors represent the user's unique knowledge and can be retained locally. The local model finally trained can provide high-precision personalized recommendation services for registered users in the target domain, and the global model trained can implement initial recommendations for new users in the target domain.

2. The personalized cross-domain recommendation method according to claim 1, characterized in that The method for training the single-domain rating prediction model in stage 1 is as follows: (11) Define the user set U in the domain = {u1,u2,…,u M }, item set V = {v1,v2,…,v N }, user-item real rating matrix And the user-item prediction rating matrix generated by the rating prediction model From a single-domain perspective, discrete users and items are projected into a shared latent space, and embedded feature vectors are used to replace users or items. The mathematical expression is as follows: in, represents the embedding feature vector of a single user, i.e. user embedding, Represents the embedding feature vector of a single item, i.e., item embedding, represents the item embedding matrix formed by concatenating all items in the domain, K represents the embedding dimension, N represents the number of items, h represents the embedding layer, θ and Represent the embedding layer parameters of user and item sets respectively; (12) Define the optimization objective of the rating prediction model to minimize the total prediction loss. Its mathematical expression is as follows: Among them, Θ f The model parameter f representing the interaction function is defined as a multi-layer neural network and is expressed as: Among them, φ out and φ X Represent the neural collaborative filtering mapping functions of the output layer and the x-th layer respectively, embedding the user Item embeddings for interactive items The concatenation is used as the input of the rating prediction model, and the output is the user's predicted rating for the item; (13) According to the mean square error, the loss between the predicted value and the true value is measured. i The mathematical expression of the local model loss function is as follows: Among them, r ij Represents user u i The real rating data, It is the predicted score data output by the local model; (14) In the initial stage of federated learning, the server will send the initialized global item embedding matrix parameters Q and neural network model parameters Θ f For the selected participating users, each user initializes their own user embedding locally During local training, Q,Θ f All of them are used as trainable parameters, and the model is trained based on the loss function in step (13), and the local user embedding parameters are updated through reverse gradient propagation Item embedding parameters and neural network model parameters Updated are kept locally as user-specific parameters, while the updated Upload to the server for aggregation, and the aggregation process is defined as: In the next round of federated learning, the server will send the aggregated results as the training parameters for the next round to the user. The user uses private data to Q (t+1) ,Θ f (t+1) Conduct local training; (15) After each round of federated learning, the performance of the global model is evaluated based on the mean absolute error and root mean square error, and the similarity between the predicted value of the rating prediction model and the true value of the user-item is obtained. Its mathematical expression is as follows: Among them, D test Represents the test data of each user. The test set size is set by dividing the data set. ij and Represents user u i Item v j The true and predicted ratings of (16) Perform multiple rounds of federated learning iterations until the model converges or reaches the predetermined training rounds. The server saves the trained global item embedding parameters Q and the single domain rating prediction model parameters Θ f , used for training the cross-domain recommendation model in stage 2.

3. The personalized cross-domain recommendation method according to claim 1, characterized in that The method for cross-domain recommendation model training in stage 2 is as follows: (21) The federated cross-domain recommendation framework includes domain A and domain B. The cross-domain recommendation task from source domain A to target domain B is considered as follows: domain A and domain B contain user sets U A and U B , Item Set V A and V B , and the user-item rating matrix R A and R B , the overlapping users of the two domains are defined as U O =U A ∩U B , overlapping user data and the global item embedding parameter Q trained in stage 1 A and Q B , and the single domain score prediction model parameters and (22) Construct a cross-domain recommendation model to capture the mapping relationship between user embeddings from the source domain to the target domain, using the trained user embeddings of overlapping users in domain A and domain B respectively. and To capture the mapping relationship of user embedding between two domains, its mathematical expression is as follows: Among them, Φ out and Φ X Represent the output layer and x-th layer model parameters of the neural network respectively, is the overlapping user u i The user embedding trained in domain A is used as the input of the local cross-domain recommendation model, and the model outputs the user u i The predicted user embedding in domain Q is K and Z represent the embedding layer dimensions of domain A and domain B, respectively; (23) Decompose the cross-domain recommendation model parameters in step (22) into a migration module and a personalization module, specifically: Step (22) provides a local model g for each user A→B , the model parameters of each layer are I and O represent the input dimension and output dimension of the network parameters of this layer, respectively. The high-order matrix Φ of each layer is decomposed into the outer product of two vectors to separate the public knowledge and personalized knowledge of the local network and reduce the network complexity. Its mathematical expression is as follows: Among them, α k is the basis vector, representing the common knowledge between clients, b k It is a personalized vector, representing the knowledge unique to each client; The local model for the user ends up being: (24) Based on the global item embedding and single-domain rating prediction model trained in stage 1, the final recommendation result is optimized by adding item information to the optimization objective. Specifically: Embed the reconstructed predicted user The item embedding Q of the domain B After splicing, it is input into the rating prediction model f trained in stage 1 B In the algorithm, the average prediction loss is required to be minimized to further optimize the cross-domain recommendation model parameters. The mathematical expression of the loss function is as follows: Among them, α and β are adjustable hyperparameters, the first term of the loss function is the reconstruction loss, and the second term is the rating loss of using the reconstructed user embedding for prediction; It means using the user embedding of overlapping users in domain A to reconstruct the user embedding of the user in domain B. Indicates the real user embedding of the user in domain B; represents the user embedding of domain B reconstructed using this user and the items in domain B are embedded in Q B Predict user ratings for items in domain B; (25) In the initial stage of federated learning, the server uniformly sends the basis vectors of each layer of the network For the selected participating users, during local training, each local user will and As trainable parameters, the cross-domain recommendation model is trained using the loss function in step (24) and updated by reverse gradient propagation, where As the user's personalized knowledge does not participate in aggregation, the updated Upload to the server for aggregation, and the aggregation process is defined as: In the next round of federated learning, the server will send the aggregated results as the training parameters for the next round to the user. The user will use private data for local training and update the and Upload updated (26) After each round of federated learning, the performance of the cross-domain recommendation model is evaluated, and its mathematical expression is as follows: in, represents the test data of overlapping users in target domain B, and Represent overlapping users u i For item u in target domain B j The true and predicted ratings of (27) Multiple rounds of federated learning iterations are performed until the model converges or reaches the predetermined training rounds, and finally each user uses the trained global model parameters and locally updated The mathematical expression for personalized local cross-domain recommendation model fusion is as follows: Before the end of federated training, the server collects personalized knowledge of all users in the last round And perform federal averaging, the mathematical expression is as follows: The server then uses and Perform global model fusion to achieve initial recommendations for new users in the target domain.

4. The personalized cross-domain recommendation method according to claim 1, characterized in that The method is suitable for completing the initial cross-domain recommendation of items to new users in the target domain. The items include products, movies and books. It can provide users with personalized cross-domain recommendation services while retaining the original data of the user-item interaction and user parameters in their local area.

5. A personalized cross-domain recommendation system based on federated learning, characterized by: The system has the following client deployments for any user: A single-domain rating prediction module, performing the single-domain rating prediction model training of stage 1 in the method according to any one of claims 1 to 4; The cross-domain recommendation module performs the cross-domain recommendation model training of stage 2 in the method according to any one of claims 1 to 4.

6. The personalized cross-domain recommendation system based on federated learning according to claim 5, characterized in that The cross-domain recommendation module includes a migration module and a personalization module. The migration module is used to realize the migration of user preferences between different domains, and the personalization module is used to provide personalized recommendations for heterogeneous data distribution.

Citation Information

Cited By

  • Federal recommendation method based on dual decoupling

    CN121436223A

  • Federal recommendation method for heterogeneous AIoT edge devices

    CN122262415A