Cross-domain recommendation cold start method based on alignment loss
By using an encoder-decoder framework for dimensionality reduction and user alignment in cross-domain recommendations, the problems of few overlapping users and insufficient historical data are solved, thus improving recommendation performance.
Patent Information
- Application Number
- CN202211320186.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-10-26
AI Technical Summary
Existing cross-domain recommendation models face difficulties in embedding high-dimensional mapping functions when there are few overlapping users and insufficient historical data, leading to degraded recommendation performance and insufficient training of mapping functions.
We adopt a cross-domain recommendation method based on alignment loss. By establishing an encoder-decoder framework, we reduce the dimensionality of data in the source and target domains. We then construct a fully connected network using an autoencoder and decoder to calculate the low-dimensional alignment loss of overlapping users, thereby achieving user alignment in the low-dimensional space.
It effectively alleviates the cold start problem across domains, improves recommendation performance, reduces the embedding difficulty of high-dimensional mapping functions, and enhances recommendation results when there are few overlapping users.
Smart Images

Figure CN115600001B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a cold start recommendation method, and more particularly to a cross-domain cold start recommendation method based on alignment loss, belonging to the technical field of recommendation methods. Background Technology
[0002] In the internet age of information overload, effective recommendation systems have provided convenience for production and life in various industries, enabling people to efficiently extract information of interest. Cross-domain recommendations can extract user evaluation records, feedback information, search records, and click information from different domains to extract data of interest to users and then push it to users. When users have no historical data or very little data, the cold start of cross-domain recommendations becomes a difficult problem for recommendation systems to overcome.
[0003] Current cross-domain recommendation models often have excessively high embedding dimensions when mapping cold-start user embeddings to the target domain. The high-dimensional functions used as input and output functions of the mapping function make it difficult to optimize the recommendation model. At the same time, the small number of overlapping users between the source and target domains leads to insufficient training of the mapping function for cold-start cross-domain recommendations, resulting in poor generalization performance of the mapping function.
[0004] Therefore, finding a cross-domain recommendation cold start method based on alignment loss, reducing the dimension of the mapping function and achieving low-dimensional space overlapped user alignment, can alleviate the cold start of cross-domain recommendation when there are not many overlapping users, effectively improve the performance of cross-domain recommendation, and reduce the difficulty of embedding high-dimensional mapping functions. Summary of the Invention
[0005] The main objective of this invention is to address the problems of poor cross-domain recommendation performance and insufficient training of mapping functions caused by the difficulty of embedding high-dimensional mapping functions when there are few overlapping users and insufficient historical data in existing cross-domain recommendation methods. Therefore, this invention provides a cold-start method for cross-domain recommendation based on alignment loss.
[0006] The objective of this invention can be achieved by adopting the following technical solution:
[0007] A cross-domain recommendation cold start method based on alignment loss includes the following steps:
[0008] A1. Establish a mathematical model for the recommendation problem: Create source and target domain datasets, and denote the source domain as... The target domain is denoted as Project sets are created in both the source and target domains. User set and evaluation matrix ,Will Let be the user in the evaluation matrix. For the project The rating, denoted by the source domain's project set, user set, and rating matrix, is denoted as , and The evaluation matrix of the target domain's item set and user set is denoted as , and ,exist and There are overlapping users, and the set of overlapping users is denoted as... In the cold start problem, the target user is not part of the source domain user set. And belong to the target domain user set Users;
[0009] A2. Establish a single-domain latent factor model: In the source domain, the user embedding and the item embedding are denoted as follows: and ,in Let be the size of the embedding, and construct the evaluation matrix. The probability model;
[0010] A3. Calculate the minimum loss function of the model: using the maximum likelihood method. The evaluation is performed, the minimum loss is solved, and this estimation method is used to train independent probabilistic models for the source and target domains respectively.
[0011] A4. Using the trained user embedding function, embed users from the source domain into the target domain, create an encoder-decoder framework, and combine overlapping and non-overlapping data to construct a low-dimensional loss training model: connect the source domain to the first layer autoencoder and the target domain to the second layer autoencoder. The first and second layers of autoencoders are respectively connected to the first and second layers of autodecoders. Construct a fully connected network using the first and second layers of autoencoders, the first and second layers of autodecoders, and the second and third layers of autodecoders. Use the first and second encoders to extract invariant factors in the source and target domains and calculate the low-dimensional alignment loss of overlapping users.
[0012] A5. Use the second decoder in the target domain to map the invariant factors to specific items in the target domain for recommendation.
[0013] As a further embodiment of the present invention, in A2, The probability model is as follows: In the formula: Record as a score The probability, Described as a matrix and The inner product, Let be denoted as variance, where Satisfying the mean is variance is Gaussian distribution, Recorded as The model satisfies a Gaussian distribution function, and the single-domain latent factor recommendation model used is a deep learning-based model. As a fully connected network conversion parameter, the fully connected network conversion parameter is the conversion parameter between the first layer autoencoder and the second layer autodecoder.
[0014] As a further aspect of the present invention, It equals the minimum value of the minimum loss function, the formula for which is:
[0015] ,
[0016] In the formula: This is denoted as the minimum loss value. Let be the user in the evaluation matrix. For the project The impact of the length of the rating set on the value of the minimum loss function. Record as user For the project The true probability of the rating Let be the user in the evaluation matrix. For the project The sum of squares of the differences between the true probability of the rating and the maximum likelihood estimate.
[0017] As a further aspect of the present invention, the loss function formula used for training the embedding mapping function in A4 is as follows:
[0018] ,
[0019] In the formula: Let it be denoted as the embedding mapping function. Let the variable be denoted as . At that time, overlapping users in the source domain The corresponding mapping function value and the corresponding overlapping user in the target domain The difference in values forms the L-2 norm of the matrix.
[0020] As a further aspect of the present invention, the transformation relationship of the first autoencoder in the source domain is as follows:
[0021] ,
[0022] ,
[0023] In the formula: Recorded as Low-dimensional representation under the action of the first autoencoder Recorded as the original source domain user The reconstructed representation;
[0024] The transformation relationship of the second autoencoder in the target domain is as follows:
[0025] ,
[0026] ,
[0027] In the formula: Recorded as Low-dimensional representation under the action of the first autoencoder Recorded as the original source domain user The reconstructed representation.
[0028] As a further aspect of the present invention, the comprehensive loss function after the source and target domains are reconstructed is expressed as a weighted sum of the reconstruction loss of the source domain, the reconstruction loss of the target domain, and the reconstruction loss of overlapping users. The formula for the comprehensive loss function is:
[0029] ,
[0030] In the formula: , , represent the adjustment relative weight hyperparameters for the source domain data reconstruction loss, target domain data reconstruction loss, and overlapping user dimensionality reduction alignment loss, respectively. , , , Let represent the source domain data reconstruction loss, the target domain data reconstruction loss, and the overlapping user dimensionality reduction alignment loss, respectively, calculated from the loss function used to train the embedding function. , , The formula is as follows:
[0031] ;
[0032] ;
[0033] ;
[0034] In the formula: Let the overlapping user set in the source domain be denoted as the user. Dimensionality reduction representation, Let these be users in the overlapping user set in the target domain. Dimensionality reduction representation.
[0035] The beneficial technical effects of this invention are as follows: According to the cross-domain recommendation cold start method based on alignment loss of this invention, by adopting an autoencoder-autodecoder framework to form a fully connected network between the source domain and the target domain, the data in the source domain and the target domain are reduced in dimensionality, and a dimensionality reduction mapping function is created. This avoids the problem of the difficulty of embedding high-dimensional mapping functions when there are few overlapping users and historical data. At the same time, the overlapping user alignment is performed on the loss of the dimensionality reduction source domain and the target domain, which can effectively alleviate the cross-domain cold start when there are few overlapping users and improve the performance of cross-domain recommendation. Attached Figure Description
[0036] Figure 1 A flowchart illustrating the workflow of the cross-domain recommendation cold start method based on alignment loss according to the present invention; Detailed Implementation
[0037] To enable those skilled in the art to understand the technical solution of the present invention more clearly, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0038] like Figure 1 As shown, the cross-domain recommendation cold start method based on alignment loss provided in this embodiment includes the following steps:
[0039] A1. Establish a mathematical model for the recommendation problem: Create source and target domain datasets, and denote the source domain as... The target domain is denoted as Project sets are created in both the source and target domains. User set and evaluation matrix ,Will Let be the user in the evaluation matrix. For the project The rating, denoted by the source domain's project set, user set, and rating matrix, is denoted as , and The evaluation matrix of the target domain's item set and user set is denoted as , and ,exist and There are overlapping users, and the set of overlapping users is denoted as... In the cold start problem, the target user is not part of the source domain user set. And belong to the target domain user set Users;
[0040] A2. Establish a single-domain latent factor model: In the source domain, the user embedding and the item embedding are denoted as follows: and ,in Let be the size of the embedding, and construct the evaluation matrix. The probability model;
[0041] A3. Calculate the minimum loss function of the model: using the maximum likelihood method. The evaluation is performed, the minimum loss is solved, and this estimation method is used to train independent probabilistic models for the source and target domains respectively.
[0042] A4. Using the trained user embedding function, embed users from the source domain into the target domain, create an encoder-decoder framework, and combine overlapping and non-overlapping data to construct a low-dimensional loss training model: connect the source domain to the first-layer autoencoder and the target domain to the second-layer autoencoder. The first-layer autoencoder and the second-layer autoencoder are respectively connected to the first-layer autodecoder and the second-layer autodecoder. Construct a fully connected network using the first-layer autoencoder, the second-layer autoencoder, the first-layer autodecoder, and the second-layer autodecoder. Use the first-layer encoder and the second-layer encoder to extract invariant factors in the source domain and the target domain, and calculate the low-dimensional alignment loss of overlapping users.
[0043] A5. Use the second decoder in the target domain to map the invariant factors to specific items in the target domain for recommendation.
[0044] It equals the minimum value of the minimum loss function, the formula for which is:
[0045] ,
[0046] In the formula: This is denoted as the minimum loss value. Let be the user in the evaluation matrix. For the project The impact of the length of the rating set on the value of the minimum loss function. Record as user For the project The true probability of the rating Let be the user in the evaluation matrix. For the project The sum of squares of the differences between the true probability of the rating and the maximum likelihood estimate.
[0047] The loss function formula used in A4 for training the embedding mapping function is:
[0048] ,
[0049] In the formula: Let it be denoted as the embedding mapping function. Let the variable be denoted as . At that time, overlapping users in the source domain The corresponding mapping function value and the corresponding overlapping user in the target domain The difference in values forms the L-2 norm of the matrix.
[0050] The transformation relationship of the first-layer autoencoder in the source domain is as follows:
[0051] ,
[0052] ,
[0053] In the formula: Recorded as Low-dimensional representation under the action of the first autoencoder Recorded as the original source domain user The reconstructed representation;
[0054] The transformation relationship of the second-level autoencoder in the target domain is as follows:
[0055] ,
[0056] ,
[0057] In the formula: Recorded as Low-dimensional representation under the action of the first-layer autoencoder Recorded as the original source domain user The reconstructed representation.
[0058] The combined loss function after reconstructing the source and target domains is expressed as a weighted sum of the reconstruction loss of the source domain, the reconstruction loss of the target domain, and the reconstruction loss of overlapping users. The formula for the combined loss function is:
[0059] ,
[0060] In the formula: , , represent the adjustment relative weight hyperparameters for the source domain data reconstruction loss, target domain data reconstruction loss, and overlapping user dimensionality reduction alignment loss, respectively. , , , Let represent the source domain data reconstruction loss, the target domain data reconstruction loss, and the overlapping user dimensionality reduction alignment loss, respectively, calculated from the loss function used to train the embedding function. , , The formula is as follows:
[0061] ;
[0062] ;
[0063] ;
[0064] In the formula: Let the overlapping user set in the source domain be denoted as the user. Dimensionality reduction representation, Let these be users in the overlapping user set in the target domain. Dimensionality reduction representation. Example
[0065] First, data from Amazon-5cores was obtained from Amazon1 as a data sample. Two call detail records (CDRs) were selected. The source domain of CDR Scenario 1 was set to Movies and the target domain was set to Music. The source domain of CDR Scenario 2 was set to Books and the target domain was set to Movies.
[0066] Secondly, the model was implemented using PyTorch2 and a cross-domain recommendation cold start method based on alignment loss. The code was run on a GPU (NVIDIA GEFORCE GTX) 1080, and the model was optimized using the Adam optimizer. The average performance of 5 runs was reported as the final result.
[0067] Contrast Model
[0068] For comparison, five representative embedding mapping-based models—CMF, EMCDR, SSCDR, DCDCSR, and LACDR_a—were selected as comparison models. The code was run on a GPU (NVIDIA GEFORCE GTX) at 1080p, and the models were implemented using PyTorch2. The models were optimized using the Adam optimizer, and the average performance of the five runs was reported as the final result.
[0069] Through practical operation, experimental results show that the cross-domain recommendation cold start method based on alignment loss can improve the performance of cross-domain recommendation by combining overlapping and non-overlapping user data. When the amount of training data for the mapping function is small, the cross-domain recommendation cold start method based on alignment loss outperforms the performance of the five baseline models CMF, EMCDR, SSCDR, DCCDCSR, and LACDR_a.
[0070] By employing an autoencoder-autodecoder framework to form a fully connected network between the source and target domains, and then reducing the dimensionality of the data in the source and target domains to create a dimensionality-reduced mapping function, the problem of embedding high-dimensional mapping functions with high difficulty in cases of overlapping users and limited historical data is avoided. At the same time, the loss of the dimensionality-reduced source and target domains is aligned with overlapping users, which can effectively alleviate cross-domain cold start and improve the performance of cross-domain recommendation when there are few overlapping users.
[0071] The above description is merely a further embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope disclosed in the present invention, based on the technical solution and concept of the present invention, shall fall within the scope of protection of the present invention.
Claims
1. A cross-domain recommendation cold start method based on alignment loss, characterized in that, Includes the following steps: A1. Establish a mathematical model for the recommendation problem: Create source and target domain datasets, and denote the source domain as... The target domain is denoted as Project sets are created in both the source and target domains. User set and evaluation matrix ,Will Let be the user in the evaluation matrix. For the project The rating, denoted by the source domain's project set, user set, and rating matrix, is denoted as , and The evaluation matrix of the target domain's item set and user set is denoted as , and ,exist and There are overlapping users, and the set of overlapping users is denoted as... In the cold start problem, the target user is not part of the source domain user set. And belong to the target domain user set Users; A2. Establish a single-domain latent factor model: In the source domain, the user embedding and the item embedding are denoted as follows: and ,in Let be the length of the embedding, and construct the evaluation matrix. The probability model; A3. Calculate the minimum loss function of the model: using the maximum likelihood method. The evaluation is performed, the minimum loss is solved, and this estimation method is used to train independent probabilistic models for the source and target domains respectively. A4. Using the trained user embedding function, embed users from the source domain into the target domain, create an encoder-decoder framework, and combine overlapping and non-overlapping data to construct a low-dimensional loss training model: connect the source domain to the first-layer autoencoder, connect the target domain to the second-layer autoencoder, and connect the first-layer autodecoder and the second-layer autodecoder to the first-layer autodecoder and the second-layer autodecoder, respectively. Construct a fully connected network using the first-layer autoencoder, the second-layer autoencoder, the first-layer autodecoder, and the second-layer autodecoder. Use the first-layer encoder and the second-layer encoder to extract invariant factors in the source and target domains, and calculate the low-dimensional alignment loss of overlapping users. A5. Use the second decoder in the target domain to map the invariant factors to specific items in the target domain for recommendation; In A2, The probability model is as follows: , In the formula: Record as a score The probability, Described as a matrix and The inner product, Let be denoted as variance, where Satisfying the mean is variance is Gaussian distribution, Recorded as The model satisfies a Gaussian distribution function, and the single-domain latent factor recommendation model used is a deep learning-based model. As a fully connected network conversion parameter, the fully connected network conversion parameter is the conversion parameter between the first layer autoencoder and the second layer autodecoder.
2. The cross-domain recommendation cold start method based on alignment loss as described in claim 1, characterized in that, It equals the minimum value of the minimum loss function, the formula for which is: , In the formula: This is denoted as the minimum loss value. Let be the user in the evaluation matrix. For the project The impact of the length of the rating set on the value of the minimum loss function. Record as user For the project The true probability of the rating Let be the user in the evaluation matrix. Item The sum of squares of the differences between the true probability of the rating and the maximum likelihood estimate.
3. The cross-domain recommendation cold start method based on alignment loss as described in claim 1, characterized in that, The loss function formula used in A4 for training the embedding mapping function is: , In the formula: Let it be denoted as the embedding mapping function. Let the variable be denoted as . At that time, overlapping users in the source domain The corresponding mapping function value and the corresponding overlapping user in the target domain The difference in values forms the L-2 norm of the matrix.
4. The cross-domain recommendation cold start method based on alignment loss as described in claim 3, characterized in that, The transformation relationship of the first-layer autoencoder in the source domain is as follows: , , In the formula: Recorded as Low-dimensional representation under the action of the first-layer autoencoder Recorded as the original source domain user The reconstructed representation; The transformation relationship of the second-level autoencoder in the target domain is as follows: , , In the formula: Recorded as Low-dimensional representation under the action of the first autoencoder Recorded as the original source domain user The reconstructed representation.
5. The cross-domain recommendation cold start method based on alignment loss as described in claim 4, characterized in that, The combined loss function after reconstructing the source and target domains is expressed as a weighted sum of the reconstruction loss of the source domain, the reconstruction loss of the target domain, and the reconstruction loss of overlapping users. The formula for the combined loss function is: , In the formula: , , represent the adjustment relative weight hyperparameters for the source domain data reconstruction loss, target domain data reconstruction loss, and overlapping user dimensionality reduction alignment loss, respectively. , , , Let represent the source domain data reconstruction loss, the target domain data reconstruction loss, and the overlapping user dimensionality reduction alignment loss, respectively, calculated from the loss function used to train the embedding function. The formula is as follows: ; ; ; In the formula: Specified as overlapping users in the source domain Dimensionality reduction representation, Specified as overlapping users in the target domain Dimensionality reduction representation.
Citation Information
Patent Citations
Recommendation method and system based on metadata enhancement
CN113449205A
Scalable cross domain recommendation system
EP2860672A2