A self-supervised learning multi-modal recommendation method and system based on a diffusion model
By using the diffusion model to generate interactive data in self-supervised learning, the problem of self-supervised signal failure in the prior art is solved, and the recommendation accuracy of multi-modal recommendation system is improved.
Patent Information
- Application Number
- CN202510412906.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The prior art is difficult to inject modal perceptual signals in self-supervised learning, resulting in the failure of the extracted self-supervised signals and cannot effectively improve the recommendation accuracy of the multimodal recommendation system.
A self-supervised learning method based on diffusion model is adopted, and modal alignment and ID guidance are performed by generating interactive data to inject modal information and ID information, and self-supervised losses are calculated to improve the accuracy of the recommendation system.
By injecting modal-aware collaboration signals, effective self-supervision losses are obtained, which significantly improves the recommendation accuracy of the multimodal recommendation system and can provide users with more accurate and personalized recommendations.
Smart Images

Figure CN119919217B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an item recommendation technology in the field of data processing technology, and in particular to a self-supervised learning multimodal recommendation method and system based on a diffusion model. Background Art
[0002] With the rapid development of the Internet, people's lifestyles have undergone tremendous changes. More and more users choose to conduct shopping, entertainment, learning and other activities through online platforms. This has led to the accumulation of massive amounts of user data and item data on various e-commerce platforms, content platforms, etc. How to accurately recommend items that meet users' interests and needs from these massive data has become a key issue in improving user experience and platform competitiveness. Item recommendation technology has emerged as the times require.
[0003] In recent years, commonly used data generation techniques in the field of recommendation systems include generative adversarial networks, variational autoencoders, and diffusion models. These methods can generate high-fidelity image, text, audio and other data by learning the potential distribution of data through neural networks. In the recommendation system, since there is less historical interaction data between users and items, their potential interests can be simulated through generation technology, thereby improving the accuracy of recommendations. In addition, data generation technology can also be used to generate feature data for long-tail products or unpopular content, making up for the problem of uneven data distribution in the recommendation system and improving the ability to recommend unpopular products. Data generation technology enriches the historical interaction data between users and items, alleviates the data sparsity problem to a certain extent, and effectively improves the performance of the recommendation model. Using generated data to compare with original data to generate self-supervisory signals is a common strategy. However, if the modal perception signal cannot be injected into the self-supervised learning task, it cannot reflect the different preferences of users for items, and the extracted self-supervisory signal will lose its effect. Summary of the invention
[0004] Technical problem to be solved by the present invention: In view of the above-mentioned problems in the prior art, a self-supervised learning multimodal recommendation method and system based on a diffusion model is provided. The present invention aims to use a diffusion model to generate interactive data to inject modal information and ID information, compare the injected data to meet the requirements of modal alignment and ID-guided feature representation, and inject modal-aware collaborative signals into self-supervised learning, thereby obtaining an effective self-supervised loss to improve the recommendation accuracy of the multimodal recommendation system.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] A self-supervised learning multimodal recommendation method based on a diffusion model includes the following steps: embedding the input user original ID into 、Item original ID embed And the modal features of the two modalities of the item image and text A multimodal recommendation model is used to obtain a list of recommended items, wherein the loss function used in the training of the multimodal recommendation model includes a self-supervised loss , and the self-supervised loss The calculation includes: normalizing the interaction matrix of user items Generating user-item collaboration graph using diffusion model , and the user-item collaboration graph Perform ID guidance and extract modal semantic commonality to obtain ID-guided collaboration graph and each mode Guided Collaboration Diagram , respectively calculate the loss of ID-guided modality representation and the loss of enhanced semantic consistency and sum them as the self-supervised loss ,in The value is or , and Represent image modality and text modality respectively.
[0007] Optionally, the normalized interaction matrix of the user item Generating user-item collaboration graph using diffusion model When the normalized interaction matrix The calculation function expression is:
[0008] ,
[0009] in, is the original interaction matrix of user items, is a normalized degree value matrix; the diffusion model is based on the normalized interaction matrix of the input user items , the user-item collaboration graph is generated through multiple iterations using the reverse generation process shown in the following formula :
[0010] ,
[0011] in, is the time step The data distribution, is the time step The data distribution, Depend on The reverse prediction is: represents a normal distribution, is the time step The noise, is the time step The diffusion coefficient, is the cumulative diffusion coefficient at time step , is the cumulative diffusion coefficient at time step .
[0012] Optionally, the user-item collaboration graph is subjected to ID guidance and extraction of modal semantic commonalities to obtain an ID-guided collaboration graph and the collaboration graph guided by each modality has a functional expression as follows:
[0013] , ,
[0014] wherein, is the interaction matrix of user items, is the concatenation operation, and are the modal features of the user and the modal features of the item respectively, and are the original ID embeddings of the user and the original ID embeddings of the item respectively, such that the ID-guided collaboration graph and the collaboration graph guided by each modality are both composed of two embedding parts representing the user and the item.
[0015] Optionally, the function expression for calculating the loss of the ID-guided modal representation and the loss of enhancing semantic consistency and summing them as the self-supervised loss is as follows:
[0016] , ,
[0017] ,
[0018] ,
[0019] ,
[0020] wherein, is the loss of the ID-guided modal representation, is the loss of enhancing semantic consistency; and are intermediate variables, is the set of modalities, is the dataset of this batch, is the similarity function, is the temperature coefficient, The collaborative graphs guided by ID respectively and each modality guided collaborative graph the embedded part of the user in it, The collaborative graphs guided by ID respectively and each modality guided collaborative graph the embedded part of the item in it. Similarly, is the modality guided collaborative graph the embedded part of the user in it, is the modality guided collaborative graph the embedded part of the item in it. The user and the user are the datasets of this batch the users and items and the item are the datasets of this batch the items in it; are respectively the modality the collaborative graphs guided when it is the image modality and the text modality the embedded part of the user in it, is the modality the modality when it is the text modality guided collaborative graph the embedded part of the user in it, are respectively the modality the collaborative graphs guided when it is the image modality and the text modality the embedded part of the item in it, is the modality the modality when it is the text modality guided collaborative graph the embedded part of the item in it.
[0021] Optionally, the loss function adopted by the multi-modal recommendation model during training is:
[0022] ,
[0023] wherein, is the loss function adopted by the multi-modal recommendation model during training, is the Bayesian loss, is the self-supervised loss coefficient, is the self-supervised loss.
[0024] Optionally, embedding the original user ID of the input , embedding the original item ID and the modal features in two modalities of the image and text of the item Using a multi-modal recommendation model to obtain an item recommendation list includes: respectively embedding the original user ID and the original item ID embedding to perform similarity graph enhancement to obtain an updated user ID embedding and the item ID embedding , and using the updated user ID embedding and the item ID embedding to obtain a collaborative view embedding through graph convolution ; projecting the modal features in each modality and respectively calculating user preferences to obtain the modal features of the user , performing modal enhancement to obtain the modal features of the item , and using the modal features of the user and the modal features of the item to obtain the preference view embeddings of each modality through graph convolution ; regularizing the preference view embeddings of each modality and fusing them with the collaborative view embedding to obtain a view fusion embedding , splitting the view fusion embedding to obtain the final user representation and the item representation , taking the inner product of the final user representation and the item representation to obtain the score of each user for each item , and selecting the top K items with high scores for each user as the output of the obtained item recommendation list , and for each user selecting the scores and outputting the top K items as the obtained item recommendation list
[0025] Optionally, the function expression for performing similarity graph enhancement on the original user ID embedding and the original item ID embedding to obtain an updated user ID embedding and the item ID embedding is:
[0026] The original user ID embedding and the original item ID embedding Perform similar graph enhancement to obtain an updated user ID embedding With the item ID embedding The functional expression is:
[0027] ,
[0028] ,
[0029] ,
[0030] ,
[0031] Among them, And Respectively represent the user node similarity graph and the item node similarity graph, Is the normalized interaction matrix of the user and item, and the superscript Represents the transpose operation; The modal features in each modality After projection, calculate the user preference respectively to obtain the user Modal features And modal enhancement to obtain the modal features of the item The functional expression is:
[0032] ,
[0033] ,
[0034] ,
[0035] Among them, Is the neighbor set of the user On the original interaction matrix of the user and item, Is the neighbor set Size of, Is the item Embedded representation of the item combined with similar modal features, Is the modal feature of the item projected into a unified embedding space, Is the modal semantic graph used to establish associations between items with similar modal features, Is the transformation matrix.
[0036] In addition, the present invention also provides a self-supervised learning multi-modal recommendation system based on a diffusion model, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the self-supervised learning multi-modal recommendation method based on the diffusion model.
[0037] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the self-supervised learning multimodal recommendation method based on the diffusion model through a processor.
[0038] In addition, the present invention also provides a computer program product, including a computer program or an instruction, wherein the computer program or the instruction is programmed or configured to execute the diffusion model-based self-supervised learning multimodal recommendation method through a processor.
[0039] Compared with the prior art, the present invention has the following advantages: the loss function used in the training of the multimodal recommendation model of the present invention includes self-supervised loss , and the self-supervised loss The calculation includes: normalizing the interaction matrix of user items Generating user-item collaboration graph using diffusion model , and the user-item collaboration graph Perform ID guidance and extract modal semantic commonality to obtain ID-guided collaboration graph and each mode Guided Collaboration Diagram , respectively calculate the loss of ID-guided modality representation and the loss of enhanced semantic consistency and sum them as the self-supervised loss , the present invention generates interactive data by using a diffusion model to inject modal information and ID information, compares the injected data to meet modal alignment, ID-guided feature representation, and injects modality-aware collaborative signals in self-supervised learning, thereby obtaining an effective self-supervised loss to improve the recommendation accuracy of the multimodal recommendation system. The present invention effectively improves the prediction accuracy of the recommendation system by combining the self-supervised signal generated by the diffusion model with multimodal information. This technology can provide users with more accurate personalized recommendations and is suitable for various multimodal recommendation application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Schematic diagram of the calculation process of the self-supervision loss in an embodiment of the present invention.
[0041] Figure 2 Schematic diagram of the principle of multimodal recommendation model and self-supervisory loss calculation in an embodiment of the present invention.
[0042] Figure 3 The figure is a schematic diagram of the workflow of the multimodal recommendation model in an embodiment of the present invention. DETAILED DESCRIPTION
[0043] To enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.
[0044] As Figure 1 and Figure 2 shown, the self-supervised learning multi-modal recommendation method based on the diffusion model in this embodiment includes the following steps: embedding the input user original ID , item original ID embedding and modal features of the item in two modalities of image and text to obtain an item recommendation list by using a multi-modal recommendation model, and the loss function adopted by the multi-modal recommendation model during training includes a self-supervised loss , and the calculation of the self-supervised loss includes: generating a user-item collaboration graph by using a diffusion model for the normalized interaction matrix of user-items , and performing ID guidance and extracting modal semantic commonalities on the user-item collaboration graph to obtain an ID-guided collaboration graph and collaboration graphs guided by each modality , calculating the loss of the ID-guided modal representation and the loss of enhancing semantic consistency respectively and summing them as the self-supervised loss , where , and , where takes values of or , and represent the image modality and the text modality respectively.
[0045] In this embodiment, when generating the user-item collaboration graph by using a diffusion model for the normalized interaction matrix of user-items , the calculation function expression of the normalized interaction matrix is:
[0046] ,
[0047] where is the original interaction matrix of user-items, and is the normalized degree value matrix.
[0048] In the forward propagation process of the diffusion model, noise is added to the original interaction graph and finally transformed into an approximate Gaussian distribution, which can be expressed as:
[0049] ,
[0050] where Represents the original interaction graph, Indicates adding to the t step noise, where represents Gaussian noise ; is the time step cumulative diffusion coefficient. In this embodiment, the diffusion model generates the user-item collaboration graph through multiple iterations using the reverse generation process shown in the following formula according to the normalized interaction matrix of the input user items : :
[0051] ,
[0052] where is the data distribution at time step , is the data distribution at time step , is reversely predicted by , represents the normal distribution, is the noise at time step , is the diffusion coefficient at time step , is the cumulative diffusion coefficient at time step , is the cumulative diffusion coefficient at time step .
[0053] In this embodiment, the function expression for performing ID guidance and extracting modal semantic commonalities on the user-item collaboration graph to obtain the ID-guided collaboration graph and the collaboration graphs guided by each modality is: :
[0054] , ,
[0055] where is the interaction matrix of user items, is the concatenation operation, and are the modal features of the user and the modal features of the item respectively, and are the original ID embeddings of the user and the original ID embeddings of the item respectively, such that the ID-guided collaboration graph and the collaboration graphs guided by each modality Both consist of two embedding parts representing users and items.
[0056] In this embodiment, the losses of the ID-guided modality representation and the enhanced semantic consistency are calculated separately and summed as the self-supervised loss The functional expression of which is:
[0057] , ,
[0058] , ,
[0059] ,
[0060] where is the loss of the ID-guided modality representation, is the loss of the enhanced semantic consistency; and are intermediate variables, is the modality set, is the dataset of this batch, is the similarity function, is the temperature coefficient, are the embedding parts of the user in the ID-guided collaboration graph and each modality guided collaboration graph respectively, are the embedding parts of the item in the ID-guided collaboration graph and each modality guided collaboration graph respectively. Similarly, is the embedding part of the user in the collaboration graph guided by the modality respectively, is the embedding part of the item in the collaboration graph guided by the modality respectively. The users and the user are the users in the dataset of this batch , and the items and the item are the items in the dataset of this batch respectively; are the embedding parts of the user in the collaboration graph guided by the image modality and the text modality respectively, is a modality When it is a text modality, it is the modality that guides the collaboration graph in which the user embedding part of is respectively the modality When it is an image modality and a text modality, it is the collaboration graph guided in which the item embedding part of is a modality When it is a text modality, it is the modality that guides the collaboration graph in which the item embedding part of
[0061] In this embodiment, the loss function adopted by the multi-modal recommendation model during training is:
[0062] ,
[0063] where is the loss function adopted by the multi-modal recommendation model during training, is the Bayesian loss, is the self-supervised loss coefficient, is the self-supervised loss. Among them, the self-supervised loss coefficient reflects the degree of influence of the corresponding part on the optimization during the model optimization process, and the required value can be selected according to needs. Combining the above loss with the Bayesian loss, using this as the optimized loss function to optimize the model, the trained model can utilize the user's historical interaction behavior and multi-modal information to perform accurate personalized recommendations for users. Among them, the calculation function expression of the Bayesian loss is:
[0064] ,
[0065] where and are the user 's predicted score for the item , and and respectively represent the items that have interacted with the user and the items that have not interacted with the user , is the activation function.
[0066] Figure 3 This is the schematic diagram of the working process of the multi-modal recommendation model in this embodiment. As Figure 3 shown, in this embodiment, the input user's original ID is embedded , the item's original ID is embedded And modal features in two modalities of the image and text of the item Using a multi-modal recommendation model to obtain an item recommendation list includes:
[0067] Multi-modal data preprocessing: Extracting the original user ID embedding And the original item ID embedding And modal features in each modality ; Respectively perform similarity graph enhancement on the original user ID embedding And the original item ID embedding To obtain the updated user ID embedding And the item ID embedding , Project the modal features in each modality And calculate the user preference respectively to obtain the user 's modal features , Modal enhancement to obtain the modal features of the item ;
[0068] Multi-modal information representation: Pass the updated user ID embedding And the item ID embedding Through graph convolution to obtain a collaborative view embedding ; Pass the user 's modal features , The modal features of the item Through graph convolution to obtain the preference view embeddings of each modality ;
[0069] Multi-modal feature fusion: Regularize the preference view embeddings of each modality And fuse them with the collaborative view embedding To obtain a view fusion embedding ;
[0070] Recommendation result prediction: Split the view fusion embedding To obtain the final user representation And the item representation , Take the inner product of the final user representation And the item representation To obtain the score of each user For each item 's score , And for each user Select the score Of the top K items as the obtained item recommendation list for output.
[0071] In this embodiment, for the original user ID embedding And the original item ID embedding Perform similar graph enhancement to obtain the updated user ID embedding and the item ID embedding The functional expression is:
[0072] ,
[0073] ,
[0074] ,
[0075] ,
[0076] where, and represent the user node similarity graph and the item node similarity graph respectively, is the normalized interaction matrix of user-items. The superscript represents the transpose operation; as an alternative implementation, the user's original ID embedding and the item's original ID embedding are extracted from the original dataset containing user and item interactions using the built-in tool class of MMRec. The normalized interaction matrix of user-items is used to represent the interaction relationship between users and items. It is a two-dimensional matrix, where rows represent users, columns represent items, and each element in the matrix represents whether there is an interaction between a user and an item. Each element takes a value of 1 or 0, where 1 means there is an interaction (such as browsing, clicking, etc.) and 0 means there is no interaction; in addition, a more specific quantitative interaction method can also be adopted. For example, a rating from 1 to 5 can be quantified as a value between 0 and 1.
[0077] In this embodiment, the modal features under each modality are projected and then the user preferences are calculated respectively to obtain the modal feature of user u, and the modal enhancement is performed to obtain the modal feature of the item. The functional expression is:
[0078] ,
[0079] ,
[0080] ,
[0081] where, is the neighbor set of user on the original interaction matrix of user-items, is the size of the neighbor set , is the embedding representation of the item after combining the similar modal features, are the modal features of items projected into a unified embedding space, is the modal semantic graph for establishing associations between items with similar modal characteristics, is the transformation matrix. In this embodiment, the modal features in each modality include the modal features in the image modality (visual modality) and the modal features in the text modality , which are respectively extracted from the original dataset based on the pre-trained Transformer model and ResNet model. First, the modal features can be projected into a unified embedding space through the following formula:
[0082] ,
[0083] Then, the modal semantic graph updates the modal embedding through the following formula to obtain the modal features of the item :
[0084] ,
[0085] Finally, the user's modal preference is obtained by aggregating the item modal features through the following formula:
[0086] ,
[0087] so as to obtain the user's modal features . The acquisition of the modal semantic graph is an existing method, specifically by constructing an association graph between items with similar modal features and then weighted summing the association graphs of multiple modalities.
[0088] In this embodiment, the updated user ID embedding and the item ID embedding are used to obtain the collaborative view embedding through graph convolution. The modal features of the user , and the modal features of the item are used to obtain the modal preference view embeddings through graph convolution. The graph convolution used includes
[0089] In this embodiment, when the updated user ID embedding and the item ID embedding are used to obtain the collaborative view embedding through graph convolution, for any The functional expression of the graph convolutional layer for learning high-order information of ID embeddings through graph propagation is as follows:
[0090] ,
[0091] ,
[0092] where and are the user ID embedding and item ID embedding output by the -th graph convolutional layer respectively, and are the user ID embedding and item ID embedding output by the -th graph convolutional layer respectively, and represent the set of one-hop neighbors of user and item on the normalized interaction matrix of user-items. To avoid an increase in scale during graph convolution operations, the symmetric normalization coefficient is adopted, and the initial user ID embedding and item ID embedding are , . After layers of message passing, the representations of each graph convolutional layer are integrated using summation, and then combined to obtain the representation of the collaborative view embedding, which can be expressed as:
[0093] , ,
[0094] ,
[0095] where and are the user ID embedding and item ID embedding respectively, is the row-wise concatenation operation, is the collaborative view embedding obtained by graph convolution of the updated user ID embedding and the item ID embedding .
[0096] In this embodiment, when obtaining the modal preference view embeddings of the modal features of user and the modal features of items through graph convolution, the functional expression of the -th graph convolutional layer for learning high-order information of the modal graph propagation is:
[0097] ,
[0098] ,
[0099] Among them, and are respectively the modality-unified embeddings output by the th graph convolutional layer, and are respectively the modality-unified embeddings output by the th graph convolutional layer, and respectively represent the set of one-hop neighbors of user and item on the normalized interaction matrix of user-items . To avoid the increase in scale during graph convolutional operations, a symmetric normalization coefficient is adopted, and the initial user ID embedding and item ID embedding are , . After layers of message passing, the representations of each graph convolutional layer are integrated using summation, and then combined to obtain the representation of the collaborative view embedding, which can be expressed as:
[0100] , ,
[0101] ,
[0102] Among them, and are respectively the user modality-unified embedding and the item modality-unified embedding, is the modality preference view embedding of modality of user , the modality feature of the item , obtained through graph convolution. In this embodiment, each modality preference view embedding
[0103] is regularized and fused with the collaborative view embedding to obtain the view fusion embedding which can be expressed as:
[0104] ,
[0105] Among them, is the normalization function for regularization, which is used to alleviate the value scale difference between modalities.
[0106] In this embodiment, the view fusion embedding is split to obtain the final user representation and item representation , the final user represents With item representation Perform the inner product to get each user For each item Rating , which can be expressed as:
[0107] ,
[0108] Among them, the superscript represents the transposition operation. Finally, for each user Select Rating The top K items are output as the resulting item recommendation list.
[0109] In order to verify the performance of the self-supervised learning multimodal recommendation method based on the diffusion model in this embodiment, the datasets used in this embodiment are the comment datasets of Amazon_Baby (referred to as the Baby dataset) and Amazon_Sports (referred to as the Sports dataset). The existing recommendation methods compared with the method of this embodiment (SSLDM) include: LATTICE, DualGNN, MICRO, DiffMM and BM3. The evaluation indicators used are recall rate (Recall) and normalized discounted cumulative gain (NDCG), and experiments are carried out for different numbers of items K recommended to users. The results are shown in Tables 1 and 2.
[0110] Table 1: Comparative experimental results on the Baby dataset
[0111]
[0112] Table 2: Comparative experimental results on the Sports dataset
[0113]
[0114] In Tables 1 and 2, R@10 and R@20 are respectively the recall rates (Recall) when the number of items K recommended to the user is 10 and 20, and N@10 and N@20 are respectively the normalized discounted cumulative gain (NDCG) when the number of items K recommended to the user is 10 and 20. As shown in Tables 1 and 2, the method of this embodiment (SSLDM) is superior to the existing recommendation methods in R@10, R@20, N@10, and N@20.
[0115] To verify the effectiveness of each part of the self-supervised learning multi-modal recommendation method based on the diffusion model in this embodiment, the datasets used are the Baby dataset and the Sports dataset. Using R@20 and N@20 as evaluation metrics, the impact of each part on the overall model of the method (SSLDM) in this embodiment was analyzed, and the results are shown in Table 3 below.
[0116] Table 3: Ablation experiment results on the Baby dataset and the Sports dataset
[0117]
[0118] In Table 3, w / o denotes the variant obtained by removing the user node similarity graph and the item node similarity graph ( and ) from the method (SSLDM) in this embodiment. w / o denotes the variant obtained by removing the modal semantic graph from the method (SSLDM) in this embodiment. w / o denotes the variant obtained by removing the self-supervised loss from the method (SSLDM) in this embodiment. As can be seen from Table 3, the node similarity graph, the item semantic graph, and the self-supervised learning part used in the method (SSLDM) in this embodiment all have varying degrees of impact on SSLDM, fully demonstrating the effectiveness of each part of the method (SSLDM) in this embodiment.
[0119] It can be seen that in this embodiment, by using the diffusion model to generate interaction data to inject modal information and ID information, and comparing the injected data to meet modal alignment, ID-guided feature representation, and injecting modal-aware collaborative signals in self-supervised learning, an effective self-supervised loss is obtained to improve the recommendation accuracy of the multi-modal recommendation system. By combining the self-supervised signals generated by the diffusion model with multi-modal information, this embodiment effectively improves the prediction accuracy of the recommendation system. This technology can provide more accurate personalized recommendations for users and is applicable to various application scenarios of multi-modal recommendations.
[0120] In addition, this embodiment also provides a self-supervised learning multi-modal recommendation system based on the diffusion model, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the self-supervised learning multi-modal recommendation method based on the diffusion model.
[0121] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the self-supervised learning multi-modal recommendation method based on the diffusion model through a processor.
[0122] In addition, this embodiment also provides a computer program product, including a computer program or instruction, which is programmed or configured to execute the self-supervised learning multimodal recommendation method based on the diffusion model through a processor.
[0123] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0124] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A self-supervised learning multimodal recommendation method based on a diffusion model, characterized in that: The steps include: embedding the input user's original ID into 、Item original ID embed And the modal features of the two modalities of the item image and text A multimodal recommendation model is used to obtain a list of recommended items. The loss function used in the training of the multimodal recommendation model includes a self-supervised loss. , and the self-supervised loss The calculation includes: normalizing the interaction matrix of user items Generating user-item collaboration graph using diffusion model , and the user-item collaboration graph Perform ID guidance and extract modal semantic commonality to obtain ID-guided collaboration graph and each mode Guided Collaboration Diagram , respectively calculate the loss of ID-guided modality representation and the loss of enhanced semantic consistency and sum them as the self-supervised loss ,in The value is or , and Represent image modality and text modality respectively; The input user's original ID is embedded in 、Item original ID embed And the modal features of the two modalities of the item image and text Using the multimodal recommendation model to obtain the item recommendation list includes: embedding the user's original ID Embedded with the original item ID Perform similarity graph augmentation to obtain updated user ID embeddings Embedded with item ID , embed the updated user ID into Embedded with item ID Collaborative view embedding via graph convolution ; The modal features of each mode After projection, user preferences are calculated to obtain user The modal characteristics of , modal enhancement to obtain the modal features of items , the user The modal characteristics of , modal characteristics of items Obtaining each modality preference view embedding through graph convolution ; Embed each modal preference view into Regularized and collaborative view embedding Fusion gets view fusion embedding , embedding the view fusion Split to get the final user representation With item representation , the final user represents With item representation Perform the inner product to get each user For each item Rating , and for each user Select Rating The top K items are output as the resulting item recommendation list.
2. The self-supervised learning multimodal recommendation method based on the diffusion model according to claim 1, characterized in that: The normalized interaction matrix of the user item Generating user-item collaboration graph using diffusion model When the normalized interaction matrix The calculation function expression is: , in, is the original interaction matrix of user items, is a normalized degree value matrix; the diffusion model is based on the normalized interaction matrix of the input user items , the user-item collaboration graph is generated through multiple iterations using the reverse generation process shown in the following formula : , in, is the time step The data distribution, is the time step The data distribution, Depend on The reverse prediction is: represents a normal distribution, is the time step The noise, is the time step The diffusion coefficient, is the time step The cumulative diffusion coefficient, is the time step The cumulative diffusion coefficient.
3. The self-supervised learning multimodal recommendation method based on the diffusion model according to claim 1, characterized in that: The user-item collaboration graph Perform ID guidance and extract modal semantic commonality to obtain ID-guided collaboration graph and each mode Guided Collaboration Diagram The function expression is: , , in, is the interaction matrix of user items, For splicing operation, and For users The modal features of and the modal features of items, and The original ID of the user and the original ID of the item are embedded respectively, so that the ID-guided collaboration graph and each mode Guided Collaboration Diagram Both consist of two embedding parts representing users and items.
4. The self-supervised learning multimodal recommendation method based on the diffusion model according to claim 1, characterized in that: The loss of ID-guided modality representation and the loss of enhanced semantic consistency are calculated separately and summed as the self-supervised loss The function expression is: , , , , , in, is the loss of ID-guided modality representation, To enhance the loss of semantic consistency; and is the intermediate variable, is the mode set, For this batch of data, is the similarity function, is the temperature coefficient, The collaboration diagrams guided by ID and each mode Guided Collaboration Diagram Medium User The embedded part, The collaboration diagrams guided by ID and each mode Guided Collaboration Diagram Medium Items The embedded part of For modal Guided Collaboration Diagram Medium User The embedded part, For modal Guided Collaboration Diagram Medium Items The embedded part of the user and users The dataset for this batch Users, items in and items The dataset for this batch Items in The modal Collaboration diagram for image mode and text mode Users in The embedded part, For modal Modal when it is text modal Guided Collaboration Diagram Medium User The embedded part, The modal Collaboration diagram for image mode and text mode Medium Items The embedded part, For modal Modal when it is text modal Guided Collaboration Diagram Medium Items The embedded part.
5. The self-supervised learning multimodal recommendation method based on diffusion model according to claim 1, characterized in that: The loss function used in the training of the multimodal recommendation model is: , in, is the loss function used in training the multimodal recommendation model. is the Bayesian loss, is the self-supervised loss coefficient, is the self-supervision loss.
6. The self-supervised learning multimodal recommendation method based on diffusion model according to claim 1, characterized in that: The user's original ID is embedded Embedded with the original item ID Perform similarity graph augmentation to obtain updated user ID embeddings Embedded with item ID The function expression is: , , , , in, and Respectively represent the user node similarity graph and the item node similarity graph, is the normalized interaction matrix of user items, represents the transposition operation; the modal features under each mode are After projection, user preferences are calculated to obtain user The modal characteristics of , modal enhancement to obtain the modal features of items The function expression is: , , , in, For users The neighbor set on the original interaction matrix of user-items, Neighborhood Set The size of For items The embedding representation after combining items with similar modality features, is the modal features of the items projected into a unified embedding space, A modal semantic graph used to establish associations between items with similar modal characteristics. is the transformation matrix.
7. A self-supervised learning multimodal recommendation system based on a diffusion model, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the self-supervised learning multimodal recommendation method based on the diffusion model as described in any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the self-supervised learning multimodal recommendation method based on the diffusion model as described in any one of claims 1 to 6 through a processor.
9. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the self-supervised learning multimodal recommendation method based on the diffusion model as described in any one of claims 1 to 6 through a processor.
Citation Information
Patent Citations
Multi-mode intelligent question answering and recommending system supporting emotional speech output
CN119739840A
Streamlined image to message and action replacement workflow with multi-modality machine-learned large language model
US20240378656A1