Project recommendation method and system with diffusion, equipment and medium
By constructing a project-solid graph and generating a contrasting view using a hierarchical modal perceptual diffusion model, the problem of noise information pollution in a multimodal recommendation system is solved, and more accurate and diversified project recommendations are achieved.
Patent Information
- Application Number
- CN202510183873.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
After the existing multimodal recommendation system performs graph convolution operations on the user-project graph, it is easy to inject noise information, resulting in a degradation of recommended performance.
By building a project-entity graph, modal-specific features and semantic entities are extracted, and a semantic and content-level comparison view is generated using a hierarchical modal-aware diffusion model, and a main view is generated to obtain modal-aware user preferences and project relationships.
Effectively capture modal-aware user preferences, prevent noise from propagating into the embedding, improve the performance of the recommendation system, and solve the noise and data sparsity problems in multimodal recommendation systems.
Smart Images

Figure CN120123583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of project recommendation, and particularly to a project recommendation method, system, device and medium with diffusion. Background Art
[0002] Multimedia-based recommendation (MMRec), as an important part of the modern information retrieval field, its core goal is to provide personalized recommendation services for users by analyzing the complex interactions between users and multimedia content. These systems integrate information from different modalities (such as text, images, audio, etc.) to more comprehensively understand user preferences, thereby improving the accuracy and diversity of recommendations. However, although multimedia-based recommendation has great potential in theory, it faces many challenges in practical applications, especially problems such as data sparsity, insufficient multimodal information fusion, and interference from noise information. These problems seriously restrict the performance improvement of recommendation systems and the optimization of user experience.
[0003] In the current field of multimedia-based recommendation, although graph neural network (GNNs) technology has made significant progress and demonstrated powerful capabilities in capturing high-order relationships in the user-project network and mining multimodal user preference clues, there are still many challenges in practical applications.
[0004] In the prior art in the field of multimedia-based recommendation systems, the multimodal content of projects in the real scenario inevitably contains noise information that has nothing to do with user interests. These information may come from user popularity bias or misclick behavior. After performing graph convolution operations on the user-project graph, all user and project representations will explicitly inject this noise information, resulting in the enhanced contrast view being possibly contaminated by multimodal noise, thereby introducing inaccurate self-supervised signals and leading to a decline in recommendation performance. Summary of the Invention
[0005] The purpose of the present invention is to address the above deficiencies of the prior art, and provide a project recommendation method, system, device and medium with diffusion to solve the problem in the prior art that after performing graph convolution operations on the user-project graph, all user and project representations will explicitly inject this noise information, resulting in the enhanced contrast view being possibly contaminated by multimodal noise, thereby introducing inaccurate self-supervised signals and leading to a decline in recommendation performance.
[0006] The present invention specifically provides the following technical solutions:
[0007] A project recommendation method with diffusion, comprising the following steps:
[0008] Obtain a user-project graph, and construct a project-entity graph through the semantic entities of each project in the user-project graph;
[0009] Extract features from the project-entity graph to obtain modality-specific features, construct a hierarchical modality-aware diffusion model, and input the modality-specific features and semantic entities of different modalities into the hierarchical modality-aware diffusion model for diffusion enhancement processing to generate a content-level contrast view and a semantic-level contrast view respectively;
[0010] Use the content-level contrast view and the semantic-level contrast view to combine and generate a main view, obtain the relationship between the preferences of the modality-aware user and the project through the main view, and recommend projects through the relationship.
[0011] Preferably, the obtaining of the user-project graph and the construction of a project-entity graph through the semantic entities of each project in the user-project graph include:
[0012] Define the user-project graph as G=(U, V, Y); where U={u 1 ,…,u i ,…,u I}, U is a set of users, V={v 1 ,…,v j ,…,v J}, V is a set of projects, the numbers of users and projects are represented as I and J respectively, Y=[y i,j I×J ∈{0,1}, Y is an interaction relationship, where y ij =1 indicates that there is an interaction between user u i and project v j ;
[0013] Extract the semantic entities of each project v to construct a project-entity graph G vo ={(v, r vo , o)}; where o∈O is a multimodal entity, and O is the set of all multimodal entities.
[0014] Preferably, when constructing the hierarchical modality-aware diffusion model, it further includes:
[0015] During the diffusion process, define the user u who interacts with a project set in the user-project graph as where, indicates whether there is an interaction between user u and project v j ;
[0016] By adding noise during the diffusion stage, disrupt the user-project interaction x 0 in the user-project graph G; the specific expression is:
[0017]
[0018] where q(x t |x 0 ) represents the process of obtaining x 0 from x t , x 0 = a u , t ∈ {1, …, T} represents the diffusion step, I represents an identity matrix, N represents a Gaussian distribution, represents the noise level, and β t ∈(0, 1) controls the noise of adding Gaussian noise at each step;
[0019] In the inverse process, the state is iteratively recovered from the Gaussian noise x T , and the specific expression is:
[0020] p θ (x t-1 |x t ) = N(x t-1 ; μ θ (x t , t), Σ θ (x t , t));
[0021] where p θ (x t-1 |x t ) represents the result of recovering the state from the Gaussian noise, x t represents the state at the t-th diffusion step, x t-1 represents the state at the (t - 1)-th diffusion step, μ θ (x t , t) and Σ θ (x t , t) represent the mean and covariance of the Gaussian distribution;
[0022] The mean μ θ is reparameterized to obtain the noise added at the time step; the specific expression is:
[0023]
[0024] The hierarchical modal perception diffusion model parameters are updated using the maximum variational lower bound of the log-likelihood with the time step, and the specific expression is:
[0025]
[0026] where L elbo represents the maximum variational lower bound loss of the log-likelihood; E t~U(1,T) represents the expectation calculation for the time step; It is to perform an expectation calculation on x 0 ; ∈ θ (x t , t) is a deep neural network with parameters θ, which can predict the noise vector ∈ given x T and t; denotes the Euclidean norm, that is, the sum of the squares of the vector elements, which measures the difference between ∈ θ (x t , t) and x 0 .
[0027] Preferably, the modality-specific features and semantic entities of different modalities are respectively input into a hierarchical modality-aware diffusion model for diffusion enhancement processing, where generating semantic-level contrast views includes:
[0028] Using relation-aware GNNs to aggregate the aligned semantic entity embeddings of item v j ; the specific expression is:
[0029]
[0030] where z j ∈R d and z o ∈R d respectively represent the ID embedding and entity embedding of the item associated with an item v j and an entity o e , N j represents the adjacent entities of item v vo through various relationships in the item-entity graph G j , the function Norm represents normalization, a is a learnable weight, represents the semantic entity embedding of item v j ;
[0031] Aggregate the semantic entity embeddings with the predicted user-item interaction probability Then aggregate the item ID embedding z j with the observed user-item interaction x 0 , and obtain the semantic-level mean square error loss between the two aggregated embeddings; the specific expression is:
[0032]
[0033] where L s is the semantic-level mean square error loss;
[0034] After obtaining the reconstructed , use to modify the user-item graph to obtain the reconstructed semantic-level contrast view.
[0035] Preferably, the modality-specific features and semantic entities of different modalities are respectively input into a hierarchical modality-aware diffusion model for diffusion enhancement processing, where generating content-level contrast views includes:
[0036] Aggregate the m modality features of each item v j , and the specific expression is:
[0037]
[0038] where M is the number of modalities, the vector is the feature of the m-th modality of item v j , d m is the dimension of these features, and MLP is the aggregation of a multi-layer perceptron;
[0039] Aggregate the modality feature with the predicted user-item interaction probability , then aggregate the item ID embedding z j with the observed user-item interaction x 0 to obtain the content-level mean squared error loss between the two aggregated embeddings, and the specific expression is:
[0040]
[0041] where L c is the content-level mean squared error loss;
[0042] After obtaining the reconstructed user-item probability , use to modify the user-item graph structure to obtain the reconstructed content-level contrast view
[0043] Preferably, after generating the main view by combining the content-level contrast view and the semantic-level contrast view, it further includes:
[0044] Optimize the L elbo loss together with L s and L c , and the specific expression is:
[0045] L dm = L elbo + θ 0 (L s + L c );
[0046] where θ 0 is a hyperparameter used to control the contributions of modality-specific features and semantic entities;
[0047] Take the embeddings from two multimodal hierarchical contrast views as anchors, and adopt the InfoNCE loss function to maximize the mutual information between the two multimodal hierarchical contrast views, obtaining the first contrastive learning loss function The specific expression is:
[0048]
[0049] where, are the positive pairs of the same node between the two hierarchical contrast views, while are the negative pairs of any two different nodes between the two hierarchical contrast views, and respectively represent the embeddings of user u i from two multimodal hierarchical contrast views, s(·) is the cosine similarity, and τ is the temperature; the item contrast loss is defined in the same way as ;
[0050] Optimize the two loss terms simultaneously, and the specific expression is:
[0051]
[0052] Adopt the user behavior patterns in the two contrast views to guide and enhance the learning of the MMRec task, and adopt InfoNCE to maximize its mutual information with the embeddings of the two hierarchical contrast views with the main view embedding as the anchor, obtaining the second contrastive learning loss, and the specific expression is:
[0053]
[0054] where, and are the positive pairs of the same node between the main view and the two hierarchical contrast views, while and are the negative pairs of any two different nodes between the main view and the two hierarchical contrast views respectively, and the item contrast loss is defined in the same way and the two loss terms are optimized simultaneously, and the specific expression is:
[0055]
[0056] Jointly optimize the contrast loss L mm 、L cl and the main objective function to obtain the final loss function, and the specific expression is:
[0057]
[0058] where, θ 1 and θ2 is a hyperparameter for controlling contributions, Θ is a model parameter, and the contrastive loss L mm or L cl is used as the main objective function for the MMRec task. L r is the main objective function, and its specific expression is:
[0059]
[0060] where y u,i defines the predicted scores of a pair of positive items for user u, while y u,j defines the predicted scores of a pair of negative items for user u;
[0061] The hierarchical modality-aware diffusion model is optimized through the optimized final loss function to obtain the optimized hierarchical modality-aware diffusion model.
[0062] Preferably, when extracting features from the item-entity graph, the features include text features, video features, or audio features.
[0063] The present invention provides an item recommendation system with diffusion, including:
[0064] An acquisition module for obtaining a user-item graph and constructing an item-entity graph through the semantic entities of each item;
[0065] A generation processing module for extracting features from the item-entity graph to obtain modality-specific features, constructing a hierarchical modality-aware diffusion model, and respectively inputting the modality-specific features and semantic entities of different modalities into the hierarchical modality-aware diffusion model for diffusion enhancement processing to generate a content-level contrast view and a semantic-level contrast view respectively;
[0066] A recommendation module for using the content-level contrast view and the semantic-level contrast view in combination to generate a main view, obtaining the relationship between the preferences of the modality-aware user and the items through the main view, and making item recommendations based on the relationship.
[0067] The present invention provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of the above-mentioned item recommendation method with diffusion.
[0068] The present invention provides a storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the steps of the above-mentioned item recommendation method with diffusion.
[0069] Compared with the prior art, the present invention has the following remarkable advantages:
[0070] The present invention constructs an item-entity graph through semantic entities of each item, and after feature extraction, obtains modality-specific features and semantic entities of different modalities. By constructing a semantic-level contrast view and a content-level contrast view through a hierarchical modality-aware diffusion model, it effectively captures modality-aware user preferences and prevents noise from spreading into the embeddings, solving the noise problem and data sparsity problem in multi-modal recommendation systems. Moreover, it uses multi-modal hierarchical contrast views to enhance node semantic information and reduce noise. Finally, by synthesizing the main view, it obtains the relationship between user preferences and items, realizing the enhancement of node semantic information through modality awareness and low noise, and improving the performance of the recommendation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 is the overall flowchart of a method for item recommendation with diffusion according to the present invention;
[0072] Figure 2 is the overall architecture diagram of MHDiff in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The following will clearly and completely describe the technical solutions of the embodiments of the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] The present invention proposes a multi-modal hierarchical graph contrast learning architecture with diffusion enhancement (MHDiff) based on multimedia recommendation. The purpose of the present invention is to eliminate the noise of multi-modal content and enhance the modality-aware user representation learning by constructing two multi-modal hierarchical contrast views (semantic-level contrast view and content-level contrast view).
[0075] To achieve this goal, the present invention designs a hierarchical modality-aware diffusion model, which adopts a hierarchical mode for generating contrast views. By injecting modality-specific features and semantic entities respectively to guide the generation of two multi-modal hierarchical contrast views, the present invention can prevent noise from specific modality features from flowing into the ID embeddings during message propagation, while capturing modality-aware user preferences. In addition, the present invention also uses diffusion enhancement technology to optimize the structure and information of the contrast views, further improving the recommendation performance. In this way, the present invention effectively solves the sparsity and noise problems existing in the prior art in the field of multi-modal recommendation systems.
[0076] The proposed MHDiff model is introduced in detail. This model combines a graph diffusion mechanism with hierarchical modality awareness and a contrast learning framework with hierarchical modality awareness, as Figure 2As shown below. First, a hierarchical modality-aware diffusion model is constructed, aiming to generate semantic-level and content-level contrast views to capture users' modality preferences and suppress the propagation of noise to the ID embedding. Subsequently, leveraging the guiding role of these two contrast views, a main view is generated, which not only enhances the semantic information of the nodes but also possesses modality-aware characteristics and low-noise characteristics.
[0077] As Figure 1 shown below, a project recommendation method with diffusion provided by the present invention will be described, which specifically includes the following steps:
[0078] Step S1: Obtain a user-project graph and construct a project-entity graph through the semantic entities of each project in the user-project graph.
[0079] Define the user-project graph as G = (U, V, Y); where U = {u 1 , …, u i , …, u I}, |U| = I is a set of users, V = {v 1 , …, v j , …, v J}, |V| = J is a set of projects, and the numbers of users and projects are represented as I and J respectively. Y = [y i,j I×J ∈{0, 1}, Y is the interaction relationship, where if y ij = 1, it means there is an interaction between user u i and project v j , otherwise y ij = 0.
[0080] To enhance the user-project interaction graph G with different modalities, first extract the semantic entity v of each project to construct a project-entity graph G vo = {(v, r vo , o)}; where o ∈ O is a multimodal entity, and O is the set of all multimodal entities. If r vo = 1, it means there is an interaction between project v and entity o, otherwise r vo = 0.
[0081] Step S2: Extract features from the project-entity graph to obtain modality-specific features, construct a hierarchical modality-aware diffusion model, and input the modality-specific features and semantic entities of different modalities into the hierarchical modality-aware diffusion model for diffusion enhancement processing to generate content-level contrast views and semantic-level contrast views respectively.
[0082] Extract features from the project-entity graph to obtain modality-specific features and semantic entities of different modalities, including:
[0083] Extract modality-specific features for each item v; the features include text features, video features, or audio features, defined as: M represents the number of modalities, and the vector is defined as the modality m feature of item v, and d m is the dimension of these features; semantic entities of different modalities are extracted for each item v.
[0084] Define the MMRec problem as: Given a graph G=(U, V, Y), first design a hierarchical modality-aware diffusion model by injecting semantic entity embeddings into the diffusion model to guide the generation of semantic-level contrast views where Ψ(·) is defined as a relation-aware GNNs for mapping semantic entities to embedding vectors. Then inject modality-specific feature variables into the diffusion model to guide the generation of content-level contrast views This can capture modality-aware user preferences and prevent noise from propagating into the ID embeddings. Finally, use two contrast views to guide the generation of the main view The MMRec task objective of this method is based on the modality-aware contrast learning paradigm, through the input to obtain the corresponding encoded representation to predict the unobserved user-item interaction relationship.
[0085] The proposed MHDiff framework in the present invention designs a hierarchical modality-aware diffusion model specifically for generating contrast views. Specifically, the MHDiff framework aims to prevent noise from interfering with the ID embeddings in multimodal content and strengthen the modeling of modality-aware user preferences. To this end, semantic entities from different modalities are introduced to guide the generation of semantic-level contrast views; at the same time, diverse modality features derived from semantic entities are injected to guide the generation of content-level contrast views. In this way, modality-specific features and semantic entities are independently propagated during the graph convolution process rather than being fused into an integrated representation. This strategy effectively avoids the influence of noise in modality-specific features on ID embeddings and successfully captures the user's modality-aware preferences.
[0086] Graph diffusion model: The graph diffusion model on the user-item graph consists of two key processes. The diffusion process focuses on destroying the original user-item graph by gradually introducing Gaussian noise, while the reverse process aims to learn and denoise the damaged graph connection structure by gradually refining the damaged graph to restore the original interaction between users and items. Specifically as follows:
[0087] When constructing a hierarchical modality-aware diffusion model, it also includes:
[0088] During the diffusion process, the user u who interacts with an item set in the user-item graph G is defined as If it means that there is an interaction between user u and item v j otherwise,
[0089] Initialize x 0 = a u , and by adding noise during the diffusion stage, the user-item interaction x in the original user-item graph G is disrupted 0 ; The specific expression is:
[0090]
[0091] where q(x t |x 0 ) represents the process of obtaining x 0 from x t , x 0 = a u , t ∈ {1, …, T} represents the diffusion step, I represents an identity matrix, N represents a Gaussian distribution, represents the amount of noise, and β t ∈ (0, 1) controls the noise of adding Gaussian noise at each step. When T → ∞, the state of x T converges to a standard Gaussian distribution.
[0092] In the inverse process, the (user-item interaction) state x is iteratively recovered from the Gaussian noise x T . It is defined that the diffusion model learns a denoising neural network to recover x 0 from x t , and the specific expression is: t-1 p
[0093] θ (x t-1 |x t ) = N(x t-1 ; μ θ (x t , t), Σ θ (x t , t));
[0094]
[0094] where p θ (x t-1 |x t ) represents the result of recovering the state from the Gaussian noise, x t represents the state in the t-th diffusion step, x t-1 represents the state in the (t - 1)-th diffusion step, μ θ (x t , t) and Σ θ (xt , t) represents the mean and covariance of the Gaussian distribution, obtained through a neural network constructed with learnable parameters θ.
[0095] The mean μ θ is reparameterized to obtain the noise added at the time step; the specific expression is:
[0096]
[0097] where μ θ (x t , t) is implemented using a Multi-Layer Perceptron (MLP) to obtain the prediction results for the input x t and t The above formula makes close to x t to update the model.
[0098] To recover the hierarchical modality-aware diffusion model parameters are updated using the maximum log-likelihood variational lower bound ELBO with time steps {1, 2, …, T}, and the specific expression is:
[0099]
[0100] where L elbo represents the maximum log-likelihood variational lower bound loss; E t~U(1,T) represents the expectation calculation over the time step, and t is a uniform distribution on the set {1, 2, …, T}; is the expectation calculation over x 0 ; ∈ θ (x t , t) is a deep neural network with parameters θ that can predict the noise vector ∈ given x T and t; represents the Euclidean norm, that is, the sum of the squares of the vector elements, measuring the difference between ∈ θ (x t , t) and x 0 .
[0101] Semantic-level contrast view: The MHDiff framework aims to prevent noise from spreading to the ID embeddings in multimodal content while learning modality-aware user preferences. To this end, semantic entities of items in different modalities are first injected into the diffusion module to guide the generation of semantic-level contrast views.
[0102] The modality-specific features and semantic entities of different modalities are respectively input into the hierarchical modality-aware diffusion model for diffusion enhancement processing, where generating semantic-level contrast views includes:
[0103] Specifically, first, semantic entities are extracted for each item v to construct an item-entity graph G vo ={(v, r vo , o)}, where o ∈ O defines a multimodal entity and O defines a set of multimodal entities. If r vo = 1, it indicates that there is an interaction between item v and entity o; otherwise, r vo = 0.
[0104] Relational-aware GNNs are adopted to aggregate the aligned semantic entity embeddings of item v j ; the specific expression is:
[0105]
[0106] where z j ∈ R d and z o ∈ R d represent the ID embedding of the item associated with an item v j and the entity embedding of an entity o e respectively. N j represents the adjacent entities of item v through various relationships in the item-entity graph G vo , the function Norm represents normalization, a is a learnable weight, j and represents the semantic entity embedding of item v j .
[0107] To inject semantic entities, first, the semantic entity embeddings are aggregated with the predicted user-item interaction probability Then, the item ID embedding z j is aggregated with the observed user-item interaction x 0 , and the semantic-level mean squared error loss (MSE) between the two aggregated embeddings is obtained and optimized with L elbo ; the specific expression is:
[0108]
[0109] where L s is the semantic-level mean squared error loss.
[0110] In this way, it enriches the diffusion module with semantic entities, and this module can learn modality-aware user preferences. After obtaining the reconstructed , the user-item graph is modified using to obtain a reconstructed semantic-level contrast view. Specifically, users u highly relevant to the task are selectedi and item v j the top-k relationships to adjust the user-item graph structure, thus retaining the reconstructed user-item graph 's information structure.
[0111] Content-level comparison view: To obtain a content-level comparison view, different modal features of semantic entities are injected into the diffusion module.
[0112] The modality-specific features and semantic entities of different modalities are respectively input into the hierarchical modality-aware diffusion model for diffusion enhancement processing, where generating the content-level comparison view includes:
[0113] Aggregate the m modal features of each item v j through an MLP, and the specific expression is:
[0114]
[0115] where M is the number of modalities, the vector is the feature of the m-th modality of item v j , d m is the dimension of these features, and MLP is the aggregation of the multi-layer perceptron.
[0116] To inject modal features, the modal features are aggregated with the predicted user-item interaction probability , and then the item ID embedding z j is aggregated with the observed user-item interaction x 0 to obtain the content-level mean squared error loss (MSE) between the two aggregated embeddings, and optimize it with L elbo , and the specific expression is:
[0117]
[0118] where L c is the content-level mean squared error loss.
[0119] Through the designed hierarchical mode, the modality-specific features and semantic entities are respectively propagated during the graph convolution operation. Therefore, the noise in the modality-specific features can prevent the flow to the ID embedding. After obtaining the reconstructed user-item probability , use to modify the user-item graph structure to obtain the reconstructed content-level comparison view Specifically, users u i highly relevant to the task and item v jThe top-k relationship between to adjust the user-item graph structure, thus preserving the reconstructed user-item graph of all information structures.
[0120] Multi-modal graph learning:
[0121] To generate a modality-aware and less noisy main view Use the graph message passing layer to aggregate two multi-modal hierarchical contrast views, embedding user and item nodes into a d-dimensional latent space. Among them, the graph message passing layer takes user u i and item v j as embedding vectors e i and e j with a dimension of R d . The embedding matrix is defined as: E (u) ∈R I×d and E (v) ∈R I×d , representing the embeddings of users and items respectively. Therefore, a simplified graph embedding propagation layer is constructed to aggregate two multi-modal hierarchical contrast views by removing the feature transformation matrix and the non-linear activation function. The specific expression is:
[0122]
[0123] where A, and are the normalized adjacency matrices of the user-item graph G, the semantic-level contrast view and the content-level contrast view respectively. and have a dimension of R d , representing the aggregated information with hierarchical pattern awareness from neighboring users or items to the central nodes u i and v j . Therefore, less noisy hierarchical modality-aware user preferences can be captured from the two multi-modal hierarchical contrast views.
[0124] Then, use multiple embedding propagation layers to aggregate local neighborhood information to optimize the embeddings of users or items. Use and to represent the embeddings of user u i and v j in the l-th layer of GCN respectively. Therefore, the message passing process from the (l - 1)-th layer to the l-th layer can be obtained by the following formula:
[0125]
[0126] Sum the embeddings of the nodes across all layers to obtain the final embedding. Then, use the inner product between the final embedding of user u i and item v j to predict the preference of user u i as shown in the following equation.
[0127]
[0128] Time complexity analysis: Define L as the number of graph embedding propagation layers, as the number of edges in the main view, as the number of edges in the two contrast views, d as the dimension, M as the set of modalities, I and J as the number of users and items. B represents the number of nodes included in a single batch, and T is the number of diffusion steps. The time complexity of the main view is Hierarchical contrastive learning: O(B×L×(I + J)×d), and the diffusion model for generating the two contrast views is:
[0129] Step S3: Combine the content-level contrast view and the semantic-level contrast view to generate the main view, obtain the relationship between the preferences of the modality-aware user and the item through the main view, and make item recommendations based on the relationship.
[0130] Model optimization: The optimization of the MHDiff model mainly includes the training of the hierarchical modality-aware graph diffusion module and the training of the MMRec task.
[0131] After combining the content-level contrast view and the semantic-level contrast view to generate the main view, it also includes:
[0132] For the hierarchical modality-aware graph diffusion module, optimize L elbo loss together with L s and L c The specific expression is:
[0133] L dm = L elbo + θ 0 (L s + L c );
[0134] where θ 0 is a hyperparameter used to control the contributions of modality-specific features and semantic entities.
[0135] In the multi-modal hierarchical contrast view, two different contrast methods are used to train the MMRec task. Specifically, one method uses two hierarchical contrast views as anchors, while the other method uses the main view as the anchor. Based on the correlation of user behavior patterns under different contrast views, the embeddings from the two multi-modal hierarchical contrast views are used as anchors, and the InfoNCE loss function is adopted to maximize the mutual information between the two multi-modal hierarchical contrast views, obtaining the first contrast learning loss function; the specific expression is:
[0136]
[0137] where, is the positive pair of the same node between the two hierarchical contrast views, while is the negative pair of any two different nodes between the two hierarchical contrast views, and respectively represent the embeddings of user u i from the two multi-modal hierarchical contrast views, which can be obtained through relation-aware GNNs and the input , s(·) is the cosine similarity, and τ is the temperature; the item contrast loss is defined in the same way as .
[0138] Both loss terms are optimized simultaneously, and the specific expression is:
[0139]
[0140] In this way, the contrast loss function maximizes the consistency of the same node in the two hierarchical contrast views and minimizes the consistency between any two different nodes, thereby promoting the learning of modality-aware user preferences and preventing the noise in modality-specific features from spreading to the ID embeddings.
[0141] The user behavior patterns in the two contrast views are used to guide and enhance the learning of the MMRec task, and InfoNCE is adopted to maximize the mutual information between it and the embeddings of the two hierarchical contrast views with the main view embedding as the anchor, obtaining the second contrast learning loss, and the specific expression is:
[0142]
[0143] where, and are the positive pairs of the same node between the main view and the two hierarchical contrast views, while and are the negative pairs of any two different nodes between the main view and the two hierarchical contrast views respectively, and the item contrast loss Define in the same way and optimize the two loss terms simultaneously. The specific expression is:
[0144]
[0145] Jointly optimize the contrastive loss \(L\) mm or \(L\) cl with the main objective function to obtain the final loss function. The specific expression is:
[0146]
[0147] where \(\theta\) 1 and \(\theta\) 2 are hyperparameters used to control the contribution, \(\Theta\) are the model parameters. Use or select the contrastive loss \(L\) mm or \(L\) cl as the main objective function for the MMRec task. \(L\) r is the main objective function. The specific expression is:
[0148]
[0149] where \(y\) u,i defines the predicted scores of a pair of positive items for user \(u\), while \(y\) u,j defines the predicted scores of a pair of negative items for user \(u\).
[0150] Optimize the hierarchical modality-aware diffusion model through the optimized final loss function to obtain the optimized hierarchical modality-aware diffusion model.
[0151] Based on the above method, the present invention provides an item recommendation system with diffusion, including: an acquisition module, a generation processing module, and a recommendation module.
[0152] Among them, the acquisition module is used to obtain the user-item graph and construct an item-entity graph through the semantic entities of each item; the generation processing module is used to extract features from the item-entity graph, obtain modality-specific features, construct a hierarchical modality-aware diffusion model, and input the modality-specific features and semantic entities of different modalities into the hierarchical modality-aware diffusion model for diffusion enhancement processing, respectively generating a content-level contrastive view and a semantic-level contrastive view; the recommendation module is used to combine the content-level contrastive view and the semantic-level contrastive view to generate a main view, obtain the relationship between the preferences of the modality-aware user and the items through the main view, and recommend items through the relationship.
[0153] The present invention also provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of an item recommendation method with diffusion.
[0154] According to the disclosed embodiments, the computer device may communicate with one or more external devices (such as a keyboard, a pointing device, Bluetooth communication, etc.), or communicate with any device (such as a router, a demodulator, etc.) that enables the computing device to communicate with one or more other computing devices.
[0155] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a project recommendation method with diffusion are implemented.
[0156] According to the disclosed embodiments, the storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0157] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A project recommendation method with diffusion, characterized in that: include: Get the user-item graph and build an item-entity graph using the semantic entities of each item in the user-item graph; Extracting features from the item-entity graph to obtain modality-specific features, and constructing a hierarchical modality perception diffusion model, and inputting the modality-specific features and semantic entities of different modalities into the hierarchical modality perception diffusion model for diffusion enhancement processing, respectively generating a content-level comparison view and a semantic-level comparison view; The content-level comparison view and the semantic-level comparison view are combined to generate a main view. The relationship between the modality-aware user's preference and the project is obtained through the main view, and the project is recommended based on the relationship.
2. A method for recommending items with diffusion as claimed in claim 1, characterized in that: The method of obtaining a user-item graph and constructing an item-entity graph through the semantic entity of each item in the user-item graph includes: The user-item graph is defined as G = (U, V, Y); where U = {u1, ..., u i ,…,u I }, U is a group of users, V = {v1,…,v j ,…,v J }, V is a set of projects, the number of users and projects are represented by I and J respectively, Y = [y i,j ] I×J ∈{0,1}, Y is the interaction relationship, where y ij =1 indicates user u i and project v j There is interaction between them; Extract the semantic entities of each item v to construct an item-entity graph G vo ={(v,r vo ,o)}; where o∈O is a multimodal entity and O is the set of all multimodal entities.
3. A method for recommending items with diffusion as claimed in claim 2, characterized in that: When constructing the hierarchical modal perception diffusion model, the method further includes: In the diffusion process, a user u interacting with an item set in the user-item graph is defined as in, Represents user u and item v j Is there any interaction between them? By adding noise in the diffusion stage, the user-item interaction x0 in the user-item graph G is destroyed; the specific expression is: Among them, q(x t |x0) means get x from x0 t process, x0=a u , t∈{1,…,T} represents the diffusion step, I represents an identity matrix, N represents the Gaussian distribution, Indicates the amount of noise, with β t ∈(0,1) controls the noise added to Gaussian noise at each step; In the inverse process, from the Gaussian noise x T The iterative recovery state is expressed as: p θ (x t-1 |x t )=N(x t-1 ;μ θ (x t ,t),Σ θ (x t ,t)); Among them, p θ (x t-1 |x t ) represents the result of recovering the state from Gaussian noise, x t represents the state in the tth diffusion step, x t-1 represents the state in the t-1th diffusion step, μ θ (x t ,t) and Σ θ (x t ,t) represents the mean and covariance of Gaussian distribution; The mean μ θ Reparameterize to obtain the noise added at the time step; the specific expression is: The maximum variation lower bound of log-likelihood with time step is used to update the parameters of the hierarchical modal perception diffusion model. The specific expression is: Among them, L elbo represents the maximum variation lower bound loss of log-likelihood; E t~U(1,T) Indicates the expected calculation of the time step; It is the expected calculation of x0; ∈ θ (x t ,t) is a deep neural network with parameters θ, which can be used to generate T and t, the predicted noise vector ∈; represents the Euclidean norm, which is the sum of the squares of the vector elements, measuring ∈ θ (x t ,t) and the difference between x0.
4. A method for recommending items with diffusion as claimed in claim 3, characterized in that: The modality-specific features and semantic entities of different modalities are respectively input into the hierarchical modality perception diffusion model for diffusion enhancement processing, wherein generating a semantic level contrast view includes: Adopt relation-aware GNNs to aggregate items v j The aligned semantic entity embedding of is: Among them, z j ∈R d and z o ∈R d Respectively represent a project v j and an entity o e The ID embedding and entity embedding of the associated item, N j By Project-Entity Graph G vo The various relationships in the project v j The function Norm represents normalization, and a is a learnable weight. Indicates project v j Semantic entity embeddings; Aggregate semantic entity embedding and the predicted user-item interaction probability Then embed the project ID into z j Aggregate with the observed user-item interaction x0 and obtain the semantic-level mean squared error loss between the two aggregated embeddings; the specific expression is: Among them, L s is the semantic level mean square error loss; After being reconstructed After that, use Modify the user-item graph to obtain a reconstructed semantic-level comparison view.
5. A method for recommending items with diffusion as claimed in claim 4, characterized in that: The modality-specific features and semantic entities of different modalities are respectively input into the hierarchical modality perception diffusion model for diffusion enhancement processing, wherein generating content-level comparison views includes: Aggregate each item v j The m modal features of are expressed as follows: Where M is the number of modes, and the vector It is project v j The characteristic of the mth mode, d m is the dimension of these features, MLP is a multi-layer perceptron aggregation; Modal Features and the predicted user-item interaction probability Aggregate and embed the project ID into z j Aggregate with the observed user-item interaction x0 to obtain the content-level mean squared error loss between the two aggregated embeddings, which is specifically expressed as: Among them, L c is the content level mean square error loss; In obtaining the reconstructed user-item probability Afterwards, use To modify the user-project graph structure, get a reconstructed content-level comparison view 6. A method for recommending items with diffusion as claimed in claim 5, characterized in that: After the main view is generated by combining the content-level comparison view and the semantic-level comparison view, the method further includes: L elbo Loss and L s and L c Optimize together, the specific expression is: L dm =L elbo +θ0(L s +L c ); Among them, θ0 is a hyperparameter used to control the contribution of modality-specific features and semantic entities; The embeddings from two multimodal hierarchical contrastive views are used as anchors, and the InfoNCE loss function is used to maximize the mutual information between the two multimodal hierarchical contrastive views to obtain the first contrastive learning loss function The specific expression is: in, u i ∈U is a positive pair of the same node between two hierarchically contrasted views, and u i ,u i' ∈U,u i ≠u i' is the negative pair of any two different nodes between the two hierarchical comparison views, and Represents user u i The embeddings from two multimodal hierarchical contrastive views, s(·) is the cosine similarity, τ is the temperature; item contrastive loss The definition method and same; Optimize two loss terms at the same time, the specific expression is: The user behavior patterns in the two contrasting views are used to guide and enhance the learning of the MMRec task, and InfoNCE is used to maximize its mutual information with the two hierarchical contrasting view embeddings with the main view embedding as the anchor to obtain the second contrastive learning loss, which is specifically expressed as: in, and (u i ∈U) is the positive pairing of the same node between the main view and the two hierarchical contrast views, and and (u i ,u i' ∈U,u i ≠u i' ) are the negative pairs of any two different nodes between the main view and the two hierarchical comparison views, respectively, and the item comparison loss Defined in the same way, and optimizing two loss terms at the same time, the specific expression is: Jointly optimize the contrast loss L mm , L cl With the main objective function, the final loss function is obtained, and the specific expression is: Among them, θ1 and θ2 are hyperparameters used to control the contribution, Θ is the model parameter, and the contrast loss L is used or selected. mm or L cl As the main objective function of the MMRec task, L r is the main objective function, and its specific expression is: Among them, y u,i defines the predicted score of a pair of positive items for user u, and y u,j Defines the prediction score of a pair of negative items for user u; The hierarchical modal perception diffusion model is optimized through the optimized final loss function to obtain the optimized hierarchical modal perception diffusion model.
7. The method for recommending items with diffusion according to claim 1, characterized in that: When extracting features from the item-entity graph, the features include text features, video features or audio features.
8. A project recommendation system with diffusion, characterized in that: include: The acquisition module is used to obtain the user-item graph and construct an item-entity graph through the semantic entities of each item in the user-item graph; A generation processing module is used to extract features from the item-entity graph to obtain modality-specific features, and to construct a hierarchical modality perception diffusion model, and to input the modality-specific features and semantic entities of different modalities into the hierarchical modality perception diffusion model for diffusion enhancement processing, and to generate a content-level comparison view and a semantic-level comparison view respectively; The recommendation module is used to generate a main view by combining the content-level comparison view and the semantic-level comparison view, obtain the relationship between the modality-aware user's preference and the project through the main view, and recommend the project based on the relationship.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a method for recommending items with diffusion as claimed in any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for recommending items with diffusion according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Recommendation method and system for multi-mode expert network and contrast diffusion based on ID guidance
CN121919407A