Multi-interest cross-domain recommendation method, model training method, system and device and medium
By combining the diffusion model and the multi-interest cross-domain recommendation method of the fusion network, the problem of low cold start user recommendation accuracy in the existing technology is solved, and more effective user interest migration and recommendation system performance improvement is achieved.
Patent Information
- Application Number
- CN202510132625.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-06
AI Technical Summary
Existing cross-domain recommendation methods have difficulties in providing accurate recommendations for cold-start users, especially in capturing multiple interests of users and handling distribution differences.
A multi-interest cross-domain recommendation method is proposed to obtain the user's potential intentions and improve the accuracy of recommendations through the combination of diffusion model and converged network. Specific steps include obtaining target domain and source domain data, preprocessing to obtain user embeddings and interest embeddings, migrating interest distributions using diffusion models, and generating predictive scores through the fusion network.
Effectively capture multiple interests of users, solve the problem of cross-domain distribution differences, improve the recommendation accuracy of cold-start users, and enhance the performance of the recommendation system.
Smart Images

Figure CN120216755A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of deep learning, and in particular, to a multi-interest cross-domain recommendation method, a model training method, a system, a device, and a medium. Background Art
[0002] Recommendation systems help users discover relevant products and services. In recent years, improving the performance of recommendation systems has attracted extensive research attention. However, recommendation systems face challenges in providing accurate recommendations for new users lacking historical interactions, and this limitation is known as the cold start problem.
[0003] Cross-domain recommendation improves the recommendation performance in the target domain by migrating knowledge from the source domain with historical data and using the rich historical behaviors in the source domain. Related cross-domain recommendation methods usually migrate user embeddings across domains through personalized preference bridges or more fine-grained bridging functions. However, these methods often ignore the distribution differences and fail to fully capture the multiple interests of users. In addition, cold start users lack behavioral data in the target domain, and most methods are difficult to fully learn the impact of diverse interests on user representations from these interactions, resulting in low accuracy of recommendation results. Summary of the Invention
[0004] The main objective of the embodiments of the present disclosure is to propose a multi-interest cross-domain recommendation method, a model training method, a system, a device, and a medium, which can better capture the potential intentions of users and improve the accuracy of recommendations for cold start users.
[0005] To achieve the above objective, on the one hand, an embodiment of the present application proposes a method for training a multi-interest cross-domain recommendation model, including the following steps:
[0006] Obtain target domain data and source domain data, where the target domain data includes user information, first item information, and first ratings, and the source domain data includes user information, second item information, and user behavior information;
[0007] Preprocess the target domain data to obtain target domain user embeddings and target domain item embeddings, and preprocess the source domain data to obtain user behavior embeddings and source domain interest embeddings;
[0008] Based on a diffusion model, obtain real noise and predicted noise according to the target domain user embeddings, the source domain interest embeddings, and the user behavior embeddings;
[0009] Adjust the parameters of the diffusion model according to the real noise and the predicted noise to obtain the trained diffusion model;
[0010] Use the trained diffusion model to obtain predicted interest embeddings according to the source domain interest embeddings and the user behavior embeddings;
[0011] Based on the fusion network, a predicted score is obtained according to the predicted interest embedding, the user behavior embedding, and the target domain item embedding;
[0012] According to the predicted score and the first score, a loss function of the fusion network is obtained;
[0013] According to the loss function, the parameters of the fusion network are adjusted to obtain the trained fusion network;
[0014] The trained fusion network and the trained diffusion model are combined to obtain a multi-interest cross-domain recommendation model.
[0015] In some embodiments, the preprocessing of the source domain data to obtain the user behavior embedding and the source domain interest embedding includes the following steps:
[0016] According to the user information and the second item information, a first interest embedding is obtained;
[0017] Feature extraction is performed on the user behavior information to obtain the user behavior embedding;
[0018] According to the first interest embedding and the user behavior embedding, a source domain interest embedding is obtained.
[0019] In some embodiments, the obtaining of the first interest embedding according to the user information and the second item information includes the following steps:
[0020] Feature extraction is performed on the second item information to obtain the attribute information of the second item information;
[0021] According to the relationship between the user information, the second item information, and the attribute information, multiple meta-paths are constructed;
[0022] The multiple meta-paths are aggregated to obtain the first interest embedding.
[0023] In some embodiments, the obtaining of the source domain interest embedding according to the first interest embedding and the user behavior embedding includes the following steps:
[0024] Through an attention network, the weight of the first interest embedding is obtained according to the first interest embedding and the user behavior embedding;
[0025] A meta-network is used to obtain the parameters of the bridging function according to the weight of the first interest embedding;
[0026] According to the first interest embedding and the parameters of the bridging function, a source domain interest embedding is obtained.
[0027] In some embodiments, based on the diffusion model, obtaining the real noise and the predicted noise according to the target domain user embedding, the source domain interest embedding, and the user behavior embedding includes the following steps:
[0028] Obtain Gaussian noise;
[0029] Sample the Gaussian noise to obtain the real noise;
[0030] Add the real noise to the target domain user embedding to obtain a second user embedding;
[0031] Add the real noise to the source domain interest embedding to obtain a second interest embedding;
[0032] Through the approximator of the diffusion model, according to the second user embedding, the second interest embedding, and the user behavior embedding, obtain the predicted noise, where the predicted noise is used to denoise the second interest embedding to obtain the predicted interest embedding.
[0033] In some embodiments, based on the fusion network, obtaining the predicted score according to the predicted interest embedding, the user behavior embedding, and the target domain item embedding includes the following steps:
[0034] Obtain the fusion network parameters;
[0035] According to the predicted interest embedding and the fusion network parameters, obtain the attention score of the predicted interest embedding;
[0036] According to the attention score and the predicted interest embedding, obtain a third interest embedding;
[0037] According to the user behavior embedding and the third interest embedding, obtain the weight of the predicted interest embedding;
[0038] According to the weight of the predicted interest embedding and the predicted interest embedding, obtain the target user embedding;
[0039] According to the target user embedding and the target domain item embedding, obtain the predicted score.
[0040] On the other hand, an embodiment of the present invention proposes a multi-interest cross-domain recommendation method, including the following steps:
[0041] Obtain the source domain data of the user and the target domain items;
[0042] Preprocess the source domain data to obtain the source domain interest embedding and the user behavior embedding;
[0043] Extract features from the target domain items to obtain the target domain item embedding;
[0044] Input the source domain interest embedding, the user behavior embedding, and the target domain item embedding into the multi-interest cross-domain recommendation model to obtain the predicted score for the target domain items of the user. The multi-interest cross-domain recommendation model is obtained by the multi-interest cross-domain recommendation model training method described in any one of the previous embodiments.
[0045] On the other hand, an embodiment of the present invention proposes a multi-interest cross-domain recommendation model training system, including:
[0046] A first module for obtaining target domain data and source domain data, where the target domain data includes user information, first item information, and first scores, and the source domain data includes user information, second item information, and user behavior information;
[0047] A second module for preprocessing the target domain data to obtain target domain user embeddings and target domain item embeddings, and preprocessing the source domain data to obtain user behavior embeddings and source domain interest embeddings;
[0048] A third module for obtaining real noise and predicted noise based on a diffusion model according to the target domain user embedding, the source domain interest embedding, and the user behavior embedding;
[0049] A fourth module for adjusting the parameters of the diffusion model according to the real noise and the predicted noise to obtain the trained diffusion model;
[0050] A fifth module for using the trained diffusion model to obtain predicted interest embeddings according to the source domain interest embedding and the user behavior embedding;
[0051] A sixth module for obtaining predicted scores based on a fusion network according to the predicted interest embedding, the user behavior embedding, and the target domain item embedding;
[0052] A seventh module for obtaining the loss function of the fusion network according to the predicted score and the first score;
[0053] An eighth module for adjusting the parameters of the fusion network according to the loss function to obtain the trained fusion network;
[0054] A ninth module for combining the trained fusion network and the trained diffusion model to obtain a multi-interest cross-domain recommendation model.
[0055] On the other hand, an embodiment of the present invention proposes an electronic device, including:
[0056] At least one processor;
[0057] At least one memory for storing at least one program;
[0058] When the at least one program is executed by the at least one processor, such that at least one of the processors implements the multi-interest cross-domain recommendation model training method or the multi-interest cross-domain recommendation method as described in the previous embodiments.
[0059] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the multi-interest cross-domain recommendation model training method or the multi-interest cross-domain recommendation method as described in the previous embodiments.
[0060] At least one of the above technical solutions of the present invention has the following advantages or beneficial effects: By modeling the distribution of user interests using a diffusion model, gradually migrating the interest distribution characteristics from the source domain to the target domain to ensure distribution alignment, while solving the cross-domain distribution difference, it can more effectively migrate the multiple interests of users and reconstruct the user interests in the target domain, thereby effectively capturing the potential intentions of users. The fusion network aggregates different interests, making the multi-interest migration process more adaptable to the target domain information, thereby further improving the performance in the fusion process and enhancing the accuracy of recommendations to users. Description of the Drawings
[0061] Figure 1 is a flowchart of the multi-interest cross-domain recommendation model training method provided by an embodiment of the present application;
[0062] Figure 2 is a schematic diagram of the heterogeneous network structure provided by an embodiment of the present application;
[0063] Figure 3 is a schematic diagram of the influence of multi-interests on cross-domain recommendation provided by an embodiment of the present application;
[0064] Figure 4 is a schematic diagram of the multi-interest cross-domain recommendation model structure provided by an embodiment of the present application;
[0065] Figure 5 is a schematic diagram of meta-path aggregation provided by an embodiment of the present application;
[0066] Figure 6 is a flowchart of the diffusion model operation provided by an embodiment of the present application;
[0067] Figure 7 is a flowchart of the multi-interest cross-domain recommendation method provided by an embodiment of the present application;
[0068] Figure 8 is a flowchart of the multi-interest cross-domain recommendation model operation provided by an embodiment of the present application;
[0069] Figure 9It is the effect diagram of the DICDR generalization experiment based on different neural networks provided by the embodiments of this application;
[0070] Figure 10 It is the schematic diagram of the hardware structure of the electronic device provided by the embodiments of this application. Detailed implementation manners
[0071] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0072] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0074] First, several nouns involved in this application are analyzed:
[0075] Diffusion Models: A class of deep learning methods based on probabilistic generative models that have achieved remarkable results in fields such as image generation, speech synthesis, and text generation in recent years. Diffusion models simulate a process of gradually transforming from a data distribution to a Gaussian noise distribution (forward diffusion), and then learn the inverse process to reconstruct high-quality data samples from the noise (backward diffusion).
[0076] Cross-Domain Recommendation (CDR): A technology that applies transfer learning in recommendation systems, aiming to solve the cold start and data sparsity problems and improve the performance and diversity of recommendation systems. Cross-domain recommendation uses the source domain information from rich data to assist the target domain with sparse data, thereby improving the recommendation effect.
[0077] In recent years, improving the performance of recommendation systems has attracted extensive research attention. However, recommendation systems face challenges in providing accurate recommendations for new users lacking historical interactions, and this limitation is called the cold start problem. Cross-domain recommendation provides a feasible solution to this problem by migrating knowledge from the source domain with historical data to the target domain.
[0078] Mapping functions are the most common method in cross - domain recommendation, which establish a paradigm for migrating user embeddings using a bridging function. However, these mapping - based methods often ignore the diversity of user preferences. Therefore, more and more research efforts are dedicated to improving performance by mining users' diverse interests. Among them, in order to more intuitively represent the relationship between users and items, many studies have tried to capture users' diverse interests by constructing Heterogeneous Information Networks (HINs). As Figure 2 shown Figure 2 shows the heterogeneous network structure. By aggregating multi - dimensional interactions using meta - paths in the HIN, researchers can reveal complex relationships and gain richer insights into user preferences. Interactions involving users, items, and brands can be represented as meta - paths, and these paths reflect different impacts on user interests.
[0079] Existing methods often fail to consider the inherent distribution differences of user interests across domains, which limits their ability to effectively generalize to the target domain. For example, as Figure 3 shown Figure 3 illustrates the impact of different interests on the target result. When existing methods make recommendations for movie fans who are keen on the NBA, they often emphasize the basketball theme, resulting in the recommended music being mainly basketball - related, while ignoring the users' potential preferences for other energetic music genres. Secondly, due to the diversity of user interests, these methods are difficult to accurately capture the user intentions in the target domain, especially for cold - start users. Therefore, the items recommended to users by these methods often lack novelty and diversity.
[0080] Based on this, the embodiments of the present disclosure provide a multi - interest cross - domain recommendation method, a model training method, a system, a device, and a medium, which can provide more accurate recommendations for cold - start users in the target domain, capture users' potential intentions, and make the recommendations more in line with actual interests.
[0081] Referring to Figure 1 shown Figure 1 is an optional flowchart of a multi - interest cross - domain recommendation model training method provided by some embodiments of the present application. A multi - interest cross - domain recommendation model training method according to an embodiment of the present invention includes but is not limited to steps S100 to S900.
[0082] Step S100, obtain target - domain data and source - domain data, where the target - domain data includes user information, first item information, and first ratings, and the source - domain data includes user information, second item information, and user behavior information;
[0083] Step S200: Preprocess the target domain data to obtain target domain user embeddings and target domain item embeddings, and preprocess the source domain data to obtain user behavior embeddings and source domain interest embeddings;
[0084] Step S300: Based on the diffusion model, obtain the real noise and predicted noise according to the target domain user embeddings, source domain interest embeddings, and user behavior embeddings;
[0085] Step S400: Adjust the parameters of the diffusion model according to the real noise and predicted noise to obtain the trained diffusion model;
[0086] Step S500: Use the trained diffusion model to obtain predicted interest embeddings according to the source domain interest embeddings and user behavior embeddings;
[0087] Step S600: Based on the fusion network, obtain the predicted score according to the predicted interest embeddings, user behavior embeddings, and target domain item embeddings;
[0088] Step S700: Obtain the loss function of the fusion network according to the predicted score and the first score;
[0089] Step S800: Adjust the parameters of the fusion network according to the loss function to obtain the trained fusion network;
[0090] Step S900: Combine the trained fusion network and the trained diffusion model to obtain a multi-interest cross-domain recommendation model.
[0091] In step S100 of some embodiments, the data is divided into a source domain and a target domain. The data includes a set of users u ∈ U d , a set of items i ∈ I d , and a set of scores r ui ∈ R d , where d ∈ {s, t} represents the source domain (s) and the target domain (t) respectively. The target domain users can be represented as U t , and the source domain users can be represented as U s . For overlapping users, they are defined as u o ∈ U s ∩ U t . The user information can be represented as the overlapping user u o . The items between the target domain and the source domain usually do not overlap. The target domain items are represented as I t , which is the first item information, and the source domain items are represented as I s , which is the second item information. The score can represent the degree of preference of the user for the item. The scores of the target domain users for the target domain items can be represented as R t , which is the first score; the scores of the source domain users for the source domain items can be represented as R sFor each user, some historical behaviors can be found in the source domain, represented as This historical behavior can be represented as user behavior information, where n represents the number of interactive items in the source domain. In the pre-trained model trained in the target domain, and represent the embeddings of the i-th user and the j-th item respectively, where n represents the dimension of all embeddings.
[0092] In step S200 of some embodiments, the target domain data is preprocessed to obtain target domain user embeddings and target domain item embeddings. The target domain user embedding is a vector, which can be obtained by feature extraction through an encoder or the like according to the users in the target domain and the interaction behaviors of users with items (such as ratings). It can represent the overall preference of users in the target domain and is a summary of users' interests and behaviors. The target domain item embedding is a vector representation of the target domain item, and its feature extraction can be performed through an encoder or the like. For example, the target domain item embedding can characterize the category, brand, label, etc. of the item. The source domain data is preprocessed to obtain user behavior embeddings and source domain interest embeddings. The user behavior embedding is a vector that can represent the historical behavior characteristics of users in the source domain. The historical behavior can be records such as users' purchases, searches, and browsing of items. The source domain interest embedding is a vector that can represent the specific preference of users for a certain item in the source domain.
[0093] In some embodiments, step S200 may include but is not limited to steps S210 to S230:
[0094] Step S210, obtaining a first interest embedding according to user information and second item information;
[0095] Step S220, performing feature extraction on the user behavior information to obtain user behavior embeddings;
[0096] Step S230, obtaining source domain interest embeddings according to the first interest embedding and the user behavior embeddings.
[0097] Please refer to Figure 4 , Figure 4It represents the overall architecture of the multi-interest cross-domain recommendation model. In the figure, the source-domain user interests are represented as the first interest embedding, which is obtained through meta-path aggregation. The historical behaviors of users in the source domain contain rich information, such as operations on different items (purchase, browsing, rating, etc.) and interactions with other users. The behavior encoder can process and extract features from the complex and diverse historical behavior data, convert the historical behaviors into specific vector representations (user behavior embeddings), and extract key features and patterns through encoding operations, so that this information can exist in a form more suitable for subsequent calculations. Based on this, the bridging function uses the output of the behavior encoder to affect the calculation of the attention scores for different items, and then affects the final personalized interest encoding (source-domain interest embedding), which can better reflect the actual interest preferences of users and improve the personalization degree of recommendations.
[0098] In some embodiments, step S210 may include but is not limited to steps S211 to S213:
[0099] Step S211, extract features from the second item information to obtain the attribute information of the second item information;
[0100] Step S212, construct multiple meta-paths according to the relationships among the user information, the second item information, and the attribute information;
[0101] Step S213, aggregate the multiple meta-paths to obtain the first interest embedding.
[0102] In steps S211 to S213 of some embodiments, the second item information includes descriptions of the appearance, content, etc. of the source-domain items. By extracting key information from this information, the attribute information related to the source-domain items can be obtained, including but not limited to the category, brand, etc. of the items. Please refer to Figure 5 , which defines multiple meta-paths to represent the influences of factors such as item categories, brands, and other users. By aggregating information at the node and path levels, the user's interest representation is obtained from multiple dimensions.
[0103] In some embodiments, four types of nodes are defined: user U s , item I s , category C s and brand B s , where each item has a category c i ∈C s and a brand b i ∈B s , and a heterogeneous network is constructed using these nodes. By combining these nodes, multiple meta-paths can be constructed. Specifically, there are relationships between users and items, items and categories, and items and brands, and a set of meta-paths {ρ1, ρ2,..., ρk}, for each user, aggregate different meta - paths from this set, where each meta - path generates a unique embedding that captures a specific aspect of the user's interest.
[0104] To improve efficiency, it is crucial to select meta - path - based neighbors for aggregation. As Figure 5 shown, starting from user , selecting n meta - paths can generate at most n meta - path - based neighbors, and each neighbor belongs to N ρ . During the aggregation process, two types of nodes are included: the user himself and the neighbor user nodes. The attention score of each meta - path is defined as shown in Equation (1).
[0105]
[0106] where f ρ (·, ·) is a linear layer, represents the final user of each meta - path. Starting from a given user, a corresponding set of neighbor nodes is determined according to different meta - paths, and the set of these nodes is N ρ .
[0107] The embedding of each meta - path is obtained by assigning weights to each neighbor node and summing them up, as shown in Equation (2).
[0108]
[0109] where σ represents the activation function.
[0110] By evaluating the importance of each meta - path, integrate all meta - paths related to the current user to generate the final representation of the user's interest, as shown in Equation (3).
[0111]
[0112] where, represents the importance of each path, and its definition is shown in Equation (4) and Equation (5).
[0113]
[0114] where W p is the weight matrix, b is the bias matrix. In the linear layer, W p and b are essentially learnable parameter matrices. W p is used to perform a linear transformation on the input features and add an offset. Specifically, W p is multiplied by the input to map the input to a new space and appropriately add an offset through b.
[0115] q is a semantic-level attention vector and is a learnable parameter vector. Specifically, q performs a dot product with the hidden representation that has undergone the non-linear transformation tanh to calculate the attention score for each input, indicating the importance of each input in the final representation.
[0116] Finally, a set of interests is obtained , where represents the embedding vector of the j-th interest of user u i . Each interest embedding is represented in the space , where k represents the number of interests and n represents the dimension of the embedding. The first interest embedding is obtained from the interest embeddings of multiple users, that is, the first interest embedding contains multiple E i , and the first interest embedding is represented as Figure 4 which is the source domain user interest in
[0117] In some embodiments, step S230 may include but is not limited to steps S231 to S233:
[0118] Step S231, through the attention network, obtain the weight of the first interest embedding according to the first interest embedding and the user behavior embedding;
[0119] Step S232, adopt the meta-network to obtain the parameters of the bridging function according to the weight of the first interest embedding;
[0120] Step S233, obtain the source domain interest embedding according to the first interest embedding and the parameters of the bridging function.
[0121] In step S210, the interest set of the source user is obtained through meta-path aggregation In steps S231 to S233 of some embodiments, it is encoded through the interest personalization bridging function , where j represents the j-th bridging function. Specifically, since different items contribute differently to interests, different items are weighted through the attention mechanism, and the formula is shown in Equation (6).
[0122]
[0123] where represents the personalization weight of the j-th interest of user u i , and the weight of the first interest embedding represents the weights of multiple interests of each user of the first interest embedding, including multiple a l is the attention score of item v l . The attention score a l is obtained from Equation (7).
[0124] a l = Softmax(h(v j ; φ h )), (7)
[0125] where h(·) represents the attention network, and φ h represents the parameters of h(·).
[0126] The obtained is input into a meta-network to obtain the parameters of the bridging function as shown in Equation (8). as shown in Equation (8).
[0127]
[0128] where f e (·) is the meta-network, parameterized by the parameters φ e The meta-network is a two-layer feed-forward network. is a vector whose size depends on the structure of the bridging function.
[0129] To adapt to the size of the bridging parameters, the vector is reshaped into a matrix Finally, the personalized transformed interest embedding is obtained, as shown in Equation (9).
[0130]
[0131] where represents the interest embedding encoded by the personalized bridge. Through the above multiple personalized bridges, the personalized interest of the i-th user can be obtained for input into the diffusion model for interest transfer.
[0132] In step S300 of some embodiments, please refer to Figure 4 , the part within the dashed box represents the diffusion model, and the rest belongs to the fusion network.
[0133] In the diffusion stage, the diffusion model gradually transforms the data into pure noise by gradually adding noise. For the input x0, at each step t of the diffusion process, a new data point is generated by adding noise, and this process can be represented by Equation (10).
[0134]
[0135] where represents a Gaussian distribution with a mean of and β tLet \(I\) denote the variance of the noise and \(I\) denote the identity matrix, indicating that the added noise is isotropic. The noise injection at each step is controlled by a predefined schedule \(\beta\) t where \(t\) represents the current diffusion step.
[0136] Since the diffusion process follows a Markov chain, each state \(x\) t depends only on the previous state \(x\) t-1 such that \(x\) can be directly derived from the initial input \(x_0\) by leveraging the properties of the Markov chain. t This relationship is expressed as shown in Equation (11).
[0137]
[0138] where and \(\alpha\) t \(= 1 - \beta\) t .
[0139] In the reverse process, the original data \(x_0\) is recovered from the pure noise \(x\) t by gradually removing the noise. Specifically, for the current denoising step \(x\) t , the reverse process for the next step \(x\) t-1 can be expressed as shown in Equations (12) and (13).
[0140]
[0141] where \(x_0\) is unknown and needs to be predicted by a neural network.
[0142] The reverse process approximated to the forward process is shown in Equation (14).
[0143]
[0144] where \(\mu\) θ (x t , \(t\)) is the predicted mean and \(\sum\) θ (x t , \(t\)) is the predicted variance, both of which are learned by a neural network with parameters \(\theta\), and the predicted noise is used to denoise the noisy data in the reverse process.
[0145] In some embodiments, step S300 may include but is not limited to steps S310 to S350:
[0146] Step S310, obtain Gaussian noise;
[0147] Step S320, sample the Gaussian noise to obtain the true noise;
[0148] Step S330, add the true noise to the target domain user embedding to obtain the second user embedding;
[0149] Step S340: Add real noise to the source domain interest embedding to obtain a second interest embedding;
[0150] Step S350: Through the approximator of the diffusion model, according to the second user embedding, the second interest embedding, and the user behavior embedding, obtain the predicted noise, where the predicted noise is used to denoise the second interest embedding to obtain the predicted interest embedding.
[0151] In steps S310 to S350 of some embodiments, please refer to Figure 6 , Gaussian noise refers to a type of noise whose probability density function follows a Gaussian distribution (i.e., normal distribution). Through tools such as the numpy.random.normal function in Python, a Gaussian noise sequence is generated according to the set mean and standard deviation. The generated Gaussian noise is a continuous distribution, but the model requires specific samples as real noise at each diffusion time step. At this time, random sampling is performed from the generated Gaussian noise, and the samples obtained each time may be different at different time steps. This randomness ensures the diversity and uncertainty of the diffusion process. For example, at a certain time step, a sample is drawn from the generated Gaussian noise, and this is the real noise at that time step, which will be used to add noise to the target domain user embedding and the source domain interest embedding subsequently.
[0152] For the acquisition of the predicted noise, it mainly depends on the approximator of the diffusion model. The approximator usually consists of a complex neural network structure, such as a multi-layer perceptron (MLP), which receives information such as the second user embedding (the target domain user embedding after adding real noise), the second interest embedding (the source domain interest embedding processed by adding real noise), and the user behavior embedding. Inside the approximator, a series of linear transformations and non-linear activation functions are used to process these inputs, and finally the predicted noise is output. The predicted noise is used to effectively remove the noise added to the source domain interest embedding, and restore more user interest information closer to the real situation, providing a reliable basis for subsequent recommendation tasks.
[0153] In the forward process, Gaussian noise is gradually added to the user embedding in the target domain. Let represent the concatenation of multiple user embeddings of the i-th user in the target domain . According to formula (11), the process of adding noise at each step is shown in formula (15).
[0154]
[0155] where t represents the time step of the diffusion process. To ensure that the reverse process does not start from completely noisy, a smaller α t is used to control the noise scale. Through this process, a user representation U composed of random noise is finally obtained.i .
[0156] In the reverse process, an approximator is trained to gradually remove the noise and recover the original embedding representation. Noise is added to the previously extracted user interest at time step t to simulate the interest reconstruction process.
[0157] This embodiment adopts a classifier-free guidance method to generate samples aligned with the user embeddings in the target domain. This method calculates the difference between the conditional approximator and the unconditional approximator, i.e., ∈ θ (x t |y) and ∈ θ (x t ) to ensure that the final embedding captures the user interest more accurately. The classifier-free guidance score function is defined as shown in Equation (16).
[0158]
[0159] where is the conditional guidance embedding, which is composed of the user's behavior information in the source domain controlled by the attention mechanism.
[0160] Furthermore, use to represent the noise signal obtained by adding noise to according to Equation (11), where encodes the multiple interests of the overlapping i-th user through personalized bridging. Starting from the noise embedding representation , the noise is gradually removed to generate the interest embedding, as shown in Equation (17).
[0161]
[0162] where z t is the noise sampled from the standard normal distribution. In each denoising step, the noise prediction is adjusted through classifier-free guidance to ensure that the generated embedding is aligned with the target user interest. To accelerate the reverse process, DPM-solver is adopted as the sampler, thereby improving the generation efficiency and reducing the model complexity, while eliminating the need for an external classifier.
[0163] In this embodiment, a multi-layer perceptron (MLP) is adopted as the backbone network of the approximator, and two zero linear layers are used in the outermost layer. The time encoder generates sinusoidal position embeddings for each time step, while the conditional encoding layer generates the conditional transition input of the diffusion model. Specifically, the user's behavior information in the source domain is used as the conditional information. To simplify the model, both the time layer and the conditional layer are implemented as linear layers. This combination can accurately predict the noise added during the diffusion process.
[0164] In the diffusion model, after obtaining the personalized encoding, noise is added and sampled to obtain the interest containing the target domain distribution. The fusion network then calculates the weighted fusion of multiple interests to generate the final user representation for prediction.
[0165] In step S400 of some embodiments, to train the diffusion model, a simplified loss function can be defined. This is achieved by minimizing the difference between the true noise and the model-predicted noise at each step. The loss function is shown in Equation (18).
[0166]
[0167] Among them, represents the true noise, and ∈ θ (x t , t) is the noise predicted by the model. The goal is to prompt the model to learn an accurate reverse process to denoise the data.
[0168] In this embodiment, all parameters of the diffusion model are represented by θ. A time step t is randomly selected, where t ∈ [0, T], and the noise schedule β is incorporated into the target domain user embedding U i , and then the approximator is used to predict the noise therein. represents the difference between the true noise and the predicted noise, as defined in Equation (19).
[0169]
[0170] By adjusting all parameters of the diffusion model, the difference between the true noise and the predicted noise is reduced, so that the interest embedding generated by the diffusion model can be aligned with the target user interest, achieving interest consistency.
[0171] In step S500 of some embodiments, the personalized interest of the i-th user is input into the trained diffusion model and the corresponding user behavior can obtain The predicted interest embedding is obtained from the H of multiple users
[0172] In step S600 of some embodiments, for cold start users, the lack of sufficient information in the target domain means that generating the target user embedding still heavily depends on the regulation of the source domain user interest. Therefore, this problem is solved by guiding the target domain interest encoding and source domain interaction.
[0173] In some embodiments, step S600 may include but is not limited to steps S610 to S630:
[0174] Step S610, obtaining the fusion network parameters;
[0175] Step S620: Obtain the attention scores of the predicted interest embedding according to the predicted interest embedding and the fusion network parameters.
[0176] Step S630: Obtain the third interest embedding according to the attention scores and the predicted interest embedding.
[0177] Step S640: Obtain the weights of the predicted interest embedding according to the user behavior embedding and the third interest embedding.
[0178] Step S650: Obtain the target user embedding according to the weights of the predicted interest embedding and the predicted interest embedding.
[0179] Step S660: Obtain the predicted score according to the target user embedding and the target domain item embedding.
[0180] In steps S610 to S650 of some embodiments, obtain from the trained diffusion model where each interest embedding Based on the interest embedding, attention scores can be calculated to determine the weights of each interest.
[0181] By represent different manifestations of the interest in the user embedding, and are composed of to form the third interest embedding, where is defined as shown in Equation (20).
[0182]
[0183] where is a parameter to be learned and belongs to a parameter of the fusion network, is represented as the attention score.
[0184] To calculate the user embedding in the target domain, guidance generated by source domain signal alignment is also required. Therefore, use f h (·;·;φ h ) to calculate the influence of the interaction history of the source domain on the target domain interest as shown in Equation (21).
[0185]
[0186] where represents the weight of each interest under the source domain guidance. For f h use two linear layers with parameters φ h to learn the interest differences between different users. consists of multiple items It can be represented as user behavior information, and the weights for predicting interest embeddings can be represented as being composed of multiple components.
[0187] Finally, a weighted sum is performed on all to obtain the target user embedding, which contains multiple as shown in Equation (22).
[0188]
[0189] Among them, β j is the weight of each transformed interest embedding and can be obtained through training.
[0190] In step S660 of some embodiments, when calculating the predicted score, specific mathematical operations are usually used to combine these two embeddings. The feature information contained in the user embedding and the item embedding can reflect the potential scoring tendency of the user towards the item after specific operations. The predicted score can provide a decision basis for the recommendation system. The target domain items are sorted according to the score, and the items with higher scores are recommended to the user, thereby achieving the purpose of cross-domain recommendation, improving the efficiency of the user to discover interesting items, and improving the performance of the recommendation system.
[0191] In step S700 of some embodiments, to improve the training performance, an optimization method based on mapping and task is adopted, and the performance of the final recommendation task is directly used as the optimization goal. φ represents the parameters of the fusion network. The optimization loss based on mapping is defined as shown in Equation (23).
[0192]
[0193] Among them, is the target user embedding generated by the fusion network. This method ensures that is closer to the target user embedding.
[0194] Furthermore, the focus of this embodiment is on the scoring task, and the task-based loss function is defined as shown in Equation (24).
[0195]
[0196] Among them, represents the interaction of overlapping users in the target domain, r ui is the actual score, and u t (i t ) T is the predicted score.
[0197] Since and Loss task, it will affect the optimization of the fusion network, and the final loss function is shown as in Equation (25).
[0198]
[0199] Where λ is a hyperparameter.
[0200] In step S800 of some embodiments, an optimization algorithm is usually used to adjust the parameters of the fusion network according to the loss function, and the optimization algorithm updates the parameters based on the gradient information of the loss function.
[0201] After multiple rounds of training and parameter adjustment, the loss function is made as small as possible, and the fusion network can better learn the relationship between the source domain and target domain information, and more accurately use the predicted interest embedding, user behavior embedding, and target domain item embedding to generate predicted scores, thereby improving the performance of the entire multi-interest cross-domain recommendation model and enabling it to provide better quality recommendation results even in complex scenarios such as cold-start users.
[0202] In step S900 of some embodiments, combining the trained fusion network and the diffusion model can give full play to the advantages of both and achieve more accurate cross-domain recommendation.
[0203] The trained diffusion model plays a key role in handling the distribution migration of user interests. By receiving the interest embedding encoded by the personalized bridging function, adding Gaussian noise to the target domain user embedding in the forward process, and gradually removing the noise using the source domain interaction information in the reverse process, it realizes the effective migration and distribution alignment of the user's multiple interests, and at the same time outputs the denoised interest embedding to the fusion network. For example, in cross-domain recommendation from the movie domain to the music domain, the diffusion model can capture the user's interests in different types of movies (such as action movies, comedies, etc.) in the movie domain and transform this interest distribution information into the music domain, providing a richer interest basis for subsequent recommendations.
[0204] The trained fusion network focuses on dynamically generating accurate user embeddings and uses source domain interactions to guide the calculation of interest weights. It encodes different interests through a personalized meta-network, fully considering the differences between various user interests, and inputs the interests encoded by the personalized bridging function into the diffusion model. Before generating the target user embedding, receiving the interest embedding output by the diffusion model, the fusion network can accurately assign weights to different interests according to the historical behavior information in the source domain. For example, in a music recommendation scenario, if there is a potential association between the user's preference for a specific type of movie in the source domain movie and certain music styles, the fusion network can use this information to make corresponding adjustments when calculating the user's interest weights for music, thereby generating a target user embedding that better conforms to the user's actual interests.
[0205] The processed and migrated interest embedding information provided by the diffusion model is input into the fusion network, providing a more comprehensive and accurate basis for the fusion network to calculate user interest weights and generate target user embeddings. The fusion network further optimizes the user embedding based on the output of the diffusion model and information such as target domain item embeddings, and finally generates a predicted score to achieve cross-domain recommendation.
[0206] Please refer to Figure 7 , Figure 7 A multi-interest cross-domain recommendation method provided by an embodiment of the present invention includes but is not limited to steps S910 to S940.
[0207] Step S910, obtaining the source domain data of the user and target domain items;
[0208] Step S920, preprocessing the source domain data to obtain source domain interest embeddings and user behavior embeddings;
[0209] Step S930, extracting features from the target domain items to obtain target domain item embeddings;
[0210] Step S940, inputting the source domain interest embeddings, user behavior embeddings, and target domain item embeddings into the multi-interest cross-domain recommendation model to obtain the predicted score for the target domain items of the user. The multi-interest cross-domain recommendation model is obtained through the multi-interest cross-domain recommendation model training method as in the previous embodiment.
[0211] In steps S910 to S940 of some embodiments, by integrating the source domain data and target domain item information, multi-interest cross-domain recommendation is used to generate accurate predicted scores. The predicted scores can accurately reflect the potential preference degree of the user for the target domain items, thereby providing a decision basis for the recommendation system, achieving efficient multi-interest cross-domain recommendation, effectively alleviating the cold start problem, and improving the generality and accuracy of the recommendation system between different domains.
[0212] In some embodiments, as Figure 8 shown, rate a music album. In the source movie domain, the user is interested in genres such as action, family, and love, which can be classified into different movie genres or actors. Different source path information can be interpreted as the influence of movie actors, users, and actor types on the user. After migrating different source path information, it can be clearly seen that movie actors cannot effectively guide the generation of user information in the target music domain.
[0213] In some embodiments, a series of experiments were conducted to evaluate the performance and robustness of the multi-interest cross-domain recommendation model.
[0214] First, obtain the experimental dataset. The Amazon review dataset is one of the most widely used public datasets in e-commerce recommendation systems and is suitable for evaluating various recommendation algorithms. In this experiment, the Amazon-5 core dataset is used, which ensures that each user and item has at least five rating records. This feature guarantees the richness and diversity of the data and helps improve the generalization ability of the recommendation model.
[0215] As shown in Table 1, Table 1 presents the detailed information of the dataset and scenarios. For the cross-domain recommendation experiment, three popular categories are selected: movies and TV (movies_and_tv), music (cds_and_vinyl), and books (books). Based on these categories, three specific cross-domain recommendation scenarios are defined in this experiment: Scenario 1: movies → music, Scenario 2: books → movies, and Scenario 3: books → music. To challenge the cold start problem from different perspectives, different cold start ratios are set for the test users in the three scenarios. Specifically, the test user ratio β is set to 20%, 50%, and 80%, and the remaining users are used for training. As β increases, the recommendation task becomes more challenging.
[0216] Table 1 Cross-domain recommendation scenario statistics
[0217]
[0218] Among them, overlap represents the number of overlapping users, and proportion represents the proportion of overlap in the total number of users.
[0219] The Amazon review dataset contains rating data (0 - 5 stars), and the mean absolute error (MAE) and root mean square error (RMSE) are used to evaluate the performance on the test set, as shown in Equations (26) and (27).
[0220]
[0221] Among them, Φ represents the test set, r i,j and represent the true rating and predicted rating respectively.
[0222] Among them, the best results are shown in bold, and * represents the paired t-test result of DICDR and the best baseline at the 0.05 level.
[0223] TGT is a simple target model that is specifically trained on the target domain data;
[0224] CMF is trained using all overlapping users in the source domain and target domain;
[0225] EMCDR uses matrix factorization to learn embeddings and then migrates the user embeddings from the source domain to the target domain through a network;
[0226] SSCDR proposes a CDR framework based on semi-supervised mapping, with a shared bridging function, which uses overlapping users or items for training;
[0227] DCDCSR is a bridging-based method that calculates the differences between domains from the perspective of users or items;
[0228] LACDR adopts an encoder-decoder structure and aligns data through a low-dimensional latent space;
[0229] PTUPCDR uses a meta-network, taking user feature embeddings as input, and generates personalized bridging functions for each user.
[0230] In this experiment, the multi-interest cross-domain recommendation model is the Diffusion Multi-Interest framework for Cross-Domain Recommendation (DICDR).
[0231] To ensure the fairness of the experiment, the number of rounds of pre-training and formal training is set to 10 for all models in each scenario. The initial learning rate of the Adam optimizer is adjusted within the range of {0.001, 0.005, 0.01, 0.02, 0.1} through grid search.
[0232] During the process of extracting interests, the meta-path set is set to {uiu, uibcu, uiciu}, where u represents users, i represents items, b represents the brand of items, and c represents the category of items. For λ, it is set to 0.1. The personalized bridging functions in PTUPCDR and DICDR use the same meta-network structure, which consists of two linear layers and an activation function, with 2×k hidden units, where k represents the embedding dimension. The embedding dimensions of users, items, and meta-path neighbor nodes are set to k. In the diffusion network, the input and output dimensions of the approximator are set to k, and the MLP consists of three layers, with the size of the linear layer and hidden units being 256.
[0233] In Table 2, DICDR is compared with the above five baseline models to verify its effectiveness. The results of three cross - domain recommendation (CDR) scenarios under different β settings are shown in the experiment, where β represents the proportion of cold - start users. The experimental results clearly show that DICDR has achieved remarkable results in all three scenarios. The TGT model trained only with target - domain data performs the worst. In contrast, CMF utilizes auxiliary source - domain information, thus improving performance. However, CMF fails to distinguish different domains and ignores the domain - transfer problem. EMCDR constructs a cross - domain embedding vector model and makes more effective use of data, further enhancing performance. PTUPCDR achieves personalized transfer by independently constructing bridging functions for different users, but the transfer of single - user embeddings fails to capture users' multiple interests, and aligning inter - domain information through bridging functions is challenging. Compared with the best baseline PTUPCDR, DICDR improves MAE by 13.99% and RMSE by 13.29% in Scenario 1, 10.96% / 9.76% in Scenario 2, and 16.84% / 15.47% in Scenario 3.
[0234] This is attributed to the fact that the DICDR method makes better use of the rich and diverse information between users or items in the source domain and the target domain, especially considering the influence of multiple item attributes and other users. Transferring source - domain information through the diffusion model and achieving distribution alignment alleviates the burden of the bridging function to a certain extent. The fusion network then aggregates different interests, making the multi - interest transfer process more adaptable to target - domain information, thus further improving performance during the fusion process.
[0235] In the experiment, matrix factorization (MF) was initially used for evaluation. However, MF is a non - neural - network model and may be too simple to effectively handle the complexity of large - scale real - world recommendation data. Although matrix - factorization algorithms perform well in recommendation systems, they cannot fully demonstrate the robustness and compatibility of DICDR.
[0236] Therefore, EMCDR, PTUPCDR, and MAFCDR are further applied to two more complex neural - network models: GMF and YouTube DNN. GMF assigns different weights in the dot - product prediction function through a deeper neural network to improve prediction performance. YouTube DNN is a two - tower model. To ensure the reliability of the experiment, all settings are kept consistent. To demonstrate the stronger performance of generalized DICDR in cold - start scenarios, tests are conducted under the condition of β = 80%.
[0237] As Figure 9As shown, generalization experiments on three basic models (a) MF, (b) GMF, and (c) YouTube DNN are presented. The figure shows the average results of five runs. EMCDR, PTUPCDR, and DICDR all improve the cold-start recommendation performance. With the improvement of the basic models, GMF and YouTube DNN have achieved significant improvements compared to MF. The MAE and RMSE results in the figure indicate that generalized DICDR has achieved good results in most cold-start scenarios.
[0238] Table 2 Comparative data on the performance of cross-domain recommendation models under different scenarios
[0239]
[0240] Table 3 Ablation experiment data
[0241]
[0242] As shown in Table 3, the ablation experiment further explores the impact of each component in the DICDR model on performance. Specifically, the following models are evaluated:
[0243] I-CDR: Remove the diffusion model, retain the optimization method of DICDR, and only use the fusion network to generate target domain user embeddings.
[0244] D-CDR: Remove the fusion network and only use the average weighting method to generate target domain user embeddings.
[0245] In I-CDR, it can be regarded as a variant of PTUPCDR. If the interest is replaced by user embeddings, it will degrade to PTUPCDR. Therefore, it can be concluded that introducing more user information improves the model's performance to a certain extent. In D-CDR, only relying on the diffusion model to generate user embeddings, its performance is still better than PTUPCDR, indicating that the introduction of the diffusion model effectively solves the problem of insufficient distribution information in the bridging method. When the fusion network and the diffusion model are combined, DICDR is significantly better than other algorithms, demonstrating the key role of each module.
[0246] The embodiment of the present invention also provides a multi-interest cross-domain recommendation model training system, including:
[0247] The first module is used to obtain target domain data and source domain data, where the target domain data includes user information, first item information, and first ratings, and the source domain data includes user information, second item information, and user behavior information;
[0248] The second module is used to preprocess the target domain data to obtain target domain user embeddings and target domain item embeddings, and preprocess the source domain data to obtain user behavior embeddings and source domain interest embeddings;
[0249] The third module is used to obtain the real noise and predicted noise based on the diffusion model according to the target domain user embedding, source domain interest embedding, and user behavior embedding;
[0250] The fourth module is used to adjust the parameters of the diffusion model according to the real noise and predicted noise to obtain the trained diffusion model;
[0251] The fifth module is used to use the trained diffusion model to obtain the predicted interest embedding according to the source domain interest embedding and user behavior embedding;
[0252] The sixth module is used to obtain the predicted score based on the fusion network according to the predicted interest embedding, user behavior embedding, and target domain item embedding;
[0253] The seventh module is used to obtain the loss function of the fusion network according to the predicted score and the first score;
[0254] The eighth module is used to adjust the parameters of the fusion network according to the loss function to obtain the trained fusion network;
[0255] The ninth module is used to combine the trained fusion network and the trained diffusion model to obtain a multi-interest cross-domain recommendation model.
[0256] It can be understood that the content in the above embodiments of the multi-interest cross-domain recommendation model training method is applicable to the embodiments of this system. The functions specifically implemented by the embodiments of this system are the same as those of the above embodiments of the multi-interest cross-domain recommendation model training method, and the beneficial effects achieved are also the same as those of the above embodiments of the multi-interest cross-domain recommendation model training method.
[0257] Next, in conjunction with Figure 10 The electronic device of the embodiments of this application will be introduced in detail.
[0258] As Figure 10 , Figure 10 schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0259] A processor 1100, which can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this disclosure;
[0260] The memory 1200 can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 1200 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1200 and are called by the processor 1100 to execute the multi-interest cross-domain recommendation model training method or the multi-interest cross-domain recommendation method of the embodiments of the present disclosure;
[0261] The input / output interface 1300 is used to implement information input and output;
[0262] The communication interface 1400 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0263] The bus 1500 transmits information between the various components of the device (such as the processor 1100, the memory 1200, the input / output interface 1300, and the communication interface 1400);
[0264] Among them, the processor 1100, the memory 1200, the input / output interface 1300, and the communication interface 1400 are communicatively connected to each other inside the device through the bus 1500.
[0265] The embodiments of the present disclosure also provide a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the above-mentioned multi-interest cross-domain recommendation model training method or the multi-interest cross-domain recommendation method.
[0266] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can include memories that are remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0267] The embodiments described in the embodiments of the present disclosure are to more clearly illustrate the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are also applicable to similar technical problems.
[0268] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0269] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0270] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0271] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0272] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0273] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0274] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0275] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0276] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0277] The preferred embodiments of the embodiments of the present disclosure have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present disclosure. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present disclosure shall fall within the scope of the rights of the embodiments of the present disclosure.
Claims
1. A multi-interest cross-domain recommendation model training method, characterized in that: The following steps are involved: Acquire target domain data and source domain data, wherein the target domain data includes user information, first item information, and a first score, and the source domain data includes user information, second item information, and user behavior information; Preprocessing the target domain data to obtain target domain user embedding and target domain item embedding, and preprocessing the source domain data to obtain user behavior embedding and source domain interest embedding; Based on the diffusion model, obtaining real noise and predicted noise according to the target domain user embedding, the source domain interest embedding and the user behavior embedding; Adjusting the parameters of the diffusion model according to the real noise and the predicted noise to obtain the trained diffusion model; Using the trained diffusion model, a predicted interest embedding is obtained according to the source domain interest embedding and the user behavior embedding; Based on the fusion network, a prediction score is obtained according to the predicted interest embedding, the user behavior embedding and the target domain item embedding; Obtaining a loss function of the fusion network according to the predicted score and the first score; Adjusting the parameters of the fusion network according to the loss function to obtain the trained fusion network; The trained fusion network and the trained diffusion model are combined to obtain a multi-interest cross-domain recommendation model.
2. The multi-interest cross-domain recommendation model training method according to claim 1, characterized in that: The preprocessing of the source domain data to obtain user behavior embedding and source domain interest embedding includes the following steps: Obtaining a first interest embedding according to the user information and the second item information; Extract features from user behavior information to obtain user behavior embedding; A source domain interest embedding is obtained according to the first interest embedding and the user behavior embedding.
3. The multi-interest cross-domain recommendation model training method according to claim 2 is characterized in that: The step of obtaining a first interest embedding according to the user information and the second item information comprises the following steps: Extracting features of the second item information to obtain attribute information of the second item information; constructing a plurality of meta-paths according to the relationship between the user information, the second item information and the attribute information; Aggregate the multiple meta-paths to obtain the first interest embedding.
4. The multi-interest cross-domain recommendation model training method according to claim 2, characterized in that: The step of obtaining the source domain interest embedding according to the first interest embedding and the user behavior embedding comprises the following steps: Obtaining a weight of the first interest embedding according to the first interest embedding and the user behavior embedding through an attention network; Using a meta-network, according to the weights of the first interest embedding, a parameter of a bridge function is obtained; The source domain interest embedding is obtained according to the first interest embedding and the parameters of the bridging function.
5. The multi-interest cross-domain recommendation model training method according to claim 1, characterized in that: The method of obtaining the real noise and the predicted noise based on the target domain user embedding, the source domain interest embedding and the user behavior embedding based on the diffusion model includes the following steps: Get Gaussian noise; Sampling the Gaussian noise to obtain the real noise; Adding the real noise to the target domain user embedding to obtain a second user embedding; Adding the real noise to the source domain interest embedding to obtain a second interest embedding; Through the approximator of the diffusion model, prediction noise is obtained according to the second user embedding, the second interest embedding and the user behavior embedding, wherein the prediction noise is used to denoise the second interest embedding to obtain the predicted interest embedding.
6. The multi-interest cross-domain recommendation model training method according to claim 1, characterized in that: The method of obtaining a predicted score based on the fusion network according to the predicted interest embedding, the user behavior embedding and the target domain item embedding comprises the following steps: Get fusion network parameters; Obtaining an attention score of the predicted interest embedding according to the predicted interest embedding and the fusion network parameters; Obtaining a third interest embedding according to the attention score and the predicted interest embedding; Obtaining a weight of a predicted interest embedding according to the user behavior embedding and the third interest embedding; Obtaining a target user embedding according to the weight of the predicted interest embedding and the predicted interest embedding; A predicted score is obtained according to the target user embedding and the target domain item embedding.
7. A multi-interest cross-domain recommendation method, characterized in that: The following steps are involved: Obtain the user's source domain data and target domain items; Preprocessing the source domain data to obtain source domain interest embedding and user behavior embedding; Extracting features of the target domain items to obtain target domain item embedding; The source domain interest embedding, the user behavior embedding and the target domain item embedding are input into a multi-interest cross-domain recommendation model to obtain a predicted score for the target domain item of the user, wherein the multi-interest cross-domain recommendation model is obtained by the multi-interest cross-domain recommendation model training method according to any one of claims 1 to 6.
8. A multi-interest cross-domain recommendation model training system, characterized in that: include: The first module is used to obtain target domain data and source domain data, wherein the target domain data includes user information, first item information and a first rating, and the source domain data includes user information, second item information and user behavior information; The second module is used to preprocess the target domain data to obtain target domain user embedding and target domain item embedding, and preprocess the source domain data to obtain user behavior embedding and source domain interest embedding; The third module is used to obtain real noise and predicted noise based on the target domain user embedding, the source domain interest embedding and the user behavior embedding based on a diffusion model; A fourth module is used to adjust the parameters of the diffusion model according to the real noise and the predicted noise to obtain the trained diffusion model; A fifth module is used to obtain a predicted interest embedding according to the source domain interest embedding and the user behavior embedding by using the trained diffusion model; A sixth module is used to obtain a predicted score based on the predicted interest embedding, the user behavior embedding and the target domain item embedding based on a fusion network; A seventh module is used to obtain a loss function of the fusion network according to the predicted score and the first score; An eighth module is used to adjust the parameters of the fusion network according to the loss function to obtain the trained fusion network; The ninth module is used to combine the trained fusion network and the trained diffusion model to obtain a multi-interest cross-domain recommendation model.
9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the multi-interest cross-domain recommendation model training method as described in any one of claims 1 to 6 or the multi-interest cross-domain recommendation method as described in claim 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement the multi-interest cross-domain recommendation model training method as described in any one of claims 1 to 6 or the multi-interest cross-domain recommendation method as described in claim 7 when executed by the processor.
Citation Information
Patent Citations
User interest diffusion recommendation method and system based on knowledge graph
CN117216281A
Sequence recommendation method and system based on prime generalization diffusion model
CN118364376A
Connectome Ensemble Transfer Learning
US20240161017A1
Cross-domain recommendation via contrastive learning of user behaviors in attentive sequence models
US20240161165A1