Knowledge graph diffusion recommendation method based on information gain guidance

By introducing information gain-guided knowledge graph diffusion model, heterogeneous knowledge aggregation mechanism and comparison learning strategies into the knowledge graph recommendation system, the limitations of traditional recommendation systems in noise and heterogeneous information environments are solved, and efficient and accurate personalized recommendations are achieved.

CN120144880APending Publication Date: 2025-06-13XIDIAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510173085.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional collaborative filtering recommendation systems show limitations when facing noise and heterogeneous information in large-scale data. There is a lot of noise, incompleteness and redundant information in knowledge graph data. How to effectively extract high-quality relationship information and improve recommendation results has become an important technical challenge.

Method used

A knowledge graph diffusion recommendation method based on information gain guidance is proposed. By combining the knowledge graph diffusion model, heterogeneous knowledge aggregation mechanism and comparative learning strategies, the knowledge graph denoising and training of the recommendation model is realized to generate a high-quality recommendation list.

Benefits of technology

The performance of the recommendation system is significantly improved, and the high-quality information related to the recommendation task is retained through the denoising process, the accuracy of the recommendation results and the generalization ability of the model are improved, and the robustness of the recommendation results is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144880A_ABST
    Figure CN120144880A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph diffusion recommendation method based on information gain guidance, and the method comprises the steps: carrying out the preprocessing operation of an original knowledge graph and user-article interaction data, calculating information gain, and obtaining a user feature weight matrix; constructing a knowledge graph denoising process based on a diffusion model based on the user feature weight matrix, and training the diffusion model by using an original knowledge graph and user-article interaction data; training a recommendation model by using the de-noised knowledge spectrogram, heterogeneous knowledge aggregation and contrast reinforcement learning to obtain a trained recommendation model; and inputting a to-be-predicted user ID and user-article interaction data into the trained recommendation model to obtain a related recommendation list of the user. By combining the knowledge graph diffusion model, the heterogeneous knowledge aggregation mechanism and the comparative learning strategy, the performance of the recommendation system is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data mining, and particularly relates to a knowledge graph diffusion recommendation method guided by information gain. Background Art

[0002] With the rapid development of the Internet, the scale of interaction data between users and items has increased exponentially. Here, items refer to objects that users may be interested in or interact with in the recommendation system. For example, products on e-commerce platforms such as clothes, electronic products, etc., or movies in a movie recommendation system, songs in a music recommendation system. However, traditional collaborative filtering recommendation systems gradually show their limitations when facing noise and heterogeneous information in large-scale data. As a data form that can structurally represent entities and their relationships, knowledge graphs provide richer semantic information for recommendation systems. However, in practical applications, knowledge graph data often contains a large amount of noise, incompleteness, and redundant information. How to effectively extract high-quality relationship information and improve the recommendation effect in such a complex environment has become an important technical challenge in the current field.

[0003] In recent years, generative diffusion models (Diffusion Model) have gradually received attention due to their excellent characteristics in graph data processing. Through the process of forward noise addition and backward denoising, this model can extract more discriminative subgraph structures from complex graph structures. However, the current research on applying diffusion models to knowledge graph recommendation is still in its infancy, and only a few attempts have preliminarily verified their potential in this field. In addition, by dynamically evaluating the importance of nodes in combination with the information gain matrix and focusing the diffusion process on nodes highly relevant to user preferences and recommendation goals, the accuracy and robustness of the recommendation system can be further improved while suppressing noise interference. Therefore, how to make full use of the potential of the combination of generative diffusion models and knowledge graphs, guide the diffusion process through information gain, and achieve more efficient and accurate personalized recommendation has become a hot issue in current technical research. Summary of the Invention

[0004] To solve the above problems existing in the prior art, the present invention provides a knowledge graph diffusion recommendation method guided by information gain. The technical problems to be solved by the present invention are achieved through the following technical solutions:

[0005] The present invention provides a knowledge graph diffusion recommendation method guided by information gain, including:

[0006] S1: Perform preprocessing operations on the original knowledge graph and user-item interaction data, calculate the information gain, and obtain the user feature weight matrix;

[0007] S2: Construct a denoising process of the knowledge graph based on the diffusion model based on the user feature weight matrix, and use the original knowledge graph and the user-item interaction data to train the diffusion model;

[0008] S3: Use the denoised knowledge spectrum graph, heterogeneous knowledge aggregation and contrastive reinforcement learning to train the recommendation model to obtain the trained recommendation model;

[0009] S4: Input the user ID to be predicted and the user-item interaction data into the trained recommendation model to obtain the relevant recommendation list of the user.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0011] 1. The present invention proposes a knowledge graph diffusion recommendation method based on information gain guidance. By combining a knowledge graph diffusion model, a heterogeneous knowledge aggregation mechanism, and a contrastive learning strategy, the performance of the recommendation system is significantly improved. The present invention proposes a denoising mechanism based on the knowledge graph. Through the forward diffusion and reverse diffusion processes, the noise in the knowledge graph is effectively removed, while the high-quality information related to the recommendation task is retained. By dynamically adjusting the node noise to protect important nodes and combining information gain to quantify the weights of features, the key relationships between user behaviors and item attributes are accurately retained during the denoising process. Compared with the method of directly using the original knowledge graph, the present invention can generate a more pure and relevant subgraph, greatly improving the accuracy of the recommendation results and the generalization ability of the model.

[0012] 2. The present invention designs a heterogeneous knowledge aggregation mechanism. By using the graph attention mechanism to distinguish different relationships in the knowledge graph, corresponding attention weights are assigned to each relationship to capture diverse semantic information. At the same time, the local graph embedding propagation layer is used to mine high-order collaborative signals to generate more comprehensive and accurate embedding representations for users and items. By introducing a contrastive learning strategy, the present invention enhances the denoised knowledge graph subgraph, effectively reducing the interference of irrelevant information. At the same time, through the optimization of positive and negative sample pairs, the ability of the recommendation model to capture the semantics of the knowledge graph and the robustness of the recommendation results are further improved.

[0013] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings

[0014] Figure 1 is a flowchart of a knowledge graph diffusion recommendation method based on information gain guidance provided by an embodiment of the present invention;

[0015] Figure 2 is a schematic diagram of a knowledge graph diffusion process provided by an embodiment of the present invention. Detailed Embodiments

[0016] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the method for knowledge graph diffusion recommendation guided by information gain proposed based on the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0017] The foregoing and other technical contents, features and effects of the present invention can be clearly presented in the following detailed description in conjunction with the accompanying drawings. Through the description of the specific embodiments, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the intended purpose can be obtained. However, the accompanying drawings are only for reference and illustration, and are not used to limit the technical solution of the present invention.

[0018] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of another identical element in the article or device including the element.

[0019] Example 1

[0020] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for knowledge graph diffusion recommendation guided by information gain provided by an embodiment of the present invention. The knowledge graph diffusion recommendation method includes:

[0021] S1: Perform preprocessing operations on the original knowledge graph and user-item interaction data, calculate the information gain, and obtain the user feature weight matrix.

[0022] This step specifically includes the following steps:

[0023] S1.1: Construct a user-item interaction matrix according to the interaction data (such as clicks, purchases, collections, etc.) between users and items where represents the number of users, represents the number of items. The matrix elements are binary (0 or 1), that is, Y uj = 1 indicates that user u has an interaction behavior with item j, otherwise Y uj = 0.

[0024] S1.2: Represent the knowledge graph in the form of triples as Among them, ε represents the entity set, represents the relationship set, represents the triple set, and the item is mapped to the entity e ∈ ε; based on the triple construct the entity adjacency matrix Among them, e h represents the head entity in the triple, and e t represents the tail entity in the triple, r represents the relationship between e h and e t A h,t = 1 indicates that there is a direct association relationship between the head entity e h and the tail entity e t Otherwise, A h,t = 0, indicating that there is no association relationship between the head entity e h and the tail entity e t has no association relationship.

[0025] S1.3: Based on the original knowledge graph and the interaction data between the user and the item, obtain the user feature weight matrix.

[0026] S1.3 of this embodiment specifically includes:

[0027] S1.3.1: For each entity e ∈ ε in the original knowledge graph, calculate the global information entropy H(e) of each entity e:

[0028]

[0029] Among them, p(r|e) is the conditional probability between entity e and relationship r, indicating the probability between entity e and a certain relationship r given entity e, or rather, it describes the possibility of the association between entity e and relationship r; represents all triples with entity e as the head entity, represents all triples with entity e as the head entity and relationship r. The global information entropy H(e) quantifies the relationship diversity of entity e. A high entropy value (such as the "director" relationship) reflects a more complex association, and a low entropy value (such as the "production company") indicates a more single semantics.

[0030] To further capture the multi-hop semantic association of entities, construct a semantic feature set for each item j:

[0031]

[0032] Among them, represents the semantic feature set of item j, and H(e t ) represents the global information entropy of the tail entity e t , Indicates that item j is mapped to the head entity e h 。

[0033] S1.3.2: Extract the associated entities of the interaction items of each user based on the user-item interaction data, and perform weighted aggregation on the associated entities of the interaction items to generate a user attribute context vector, which comprehensively reflects the user's interest characteristics.

[0034] Specifically, user behavior data reflects the user's potential interest preferences. In order to combine user behavior with the semantics of the knowledge graph, a user attribute context vector is constructed. For the user set and its set of historical interaction items extract the associated entities of the interaction items of user :

[0035]

[0036] where ε u represents the set of entities corresponding to all items with historical interaction records with user u. ∪ represents the union operation, that is, taking the union of the entity sets corresponding to each item j, and finally obtaining the set of all associated entities of the interaction items of user u.

[0037] Subsequently, generate a user attribute context vector by weighted aggregation of the embedding vectors of the user-associated entities:

[0038]

[0039] where C u represents the user attribute context vector of user u. H(e) represents the global information entropy of entity e, H(e′) represents the global information entropy of entity e′, and η represents the entropy sensitivity coefficient, which is used to adjust the weight of the entity (higher entropy entities have higher weights). represents the embedding vector of entity e (with dimension d), and encodes the discrete entity into a dense representation in the continuous vector space through the distributed representation method.

[0040] Specifically, the process of mapping entity e to the embedding vector R e is completed through the Embedding layer module of the PyTorch deep learning framework. In the model initialization stage, a trainable parameter matrix is created for all entities. During the model training process, this embedding matrix participates in the end-to-end optimization as a learnable parameter, and dynamically adjusts the numerical distribution of the embedding vector based on the loss function.

[0041] S1.3.3: Combine user behavior with the semantics of the knowledge graph to construct a positive sample data set and a negative sample data set for each user, providing training data for modeling the user's decision-making behavior.

[0042] In a recommendation system, the decision-making behavior of users is modeled through a positive and negative sample learning model. To this end, based on user behavior and the semantics of the knowledge graph, a decision feature space is constructed. For the set of historical interaction items of user u Define the positive sample dataset (that is, the positive sample dataset is the union of the semantic feature sets of all items j in the set of historical interaction items ) and the negative sample dataset (randomly sampled from the non-interaction items of user u, satisfying ). Combine the positive sample dataset and the negative sample dataset to construct the user feature dataset:

[0043]

[0044] Among them, represents the user feature dataset of user u, f is the feature selected from the positive sample dataset and the negative sample dataset of user u, and each feature f is related to user preferences and usually represents some attribute related to items or user behavior, such as an attribute of an item (such as "Director: James Cameron") or the user's interest in a certain type of item (such as "action movies"); x f represents the feature vector corresponding to feature f, such as "Director: James Cameron", and y f is the label, indicating positive and negative samples. The label of the positive sample is 1, and the label of the negative sample is 0.

[0045] S1.3.4: Use global feature entropy and conditional entropy to quantify the impact of features on user decisions and dynamically adjust feature weights to adapt to different user contexts.

[0046] Specifically, to further measure the impact of features on user decisions, define the information gain of feature f corresponding to user u as IG(f|u):

[0047] IG(f|u) = H(f) - H(f|C u ),

[0048] Among them, H(f) represents the global feature entropy, which reflects the uncertainty of the feature in the overall user behavior, while H(f|C u ) represents the conditional entropy, which is used to evaluate the remaining uncertainty of the feature under specific user context conditions. The global feature entropy H(f) is a measure of the distribution complexity of feature f in the overall user behavior, and the calculation formula is:

[0049] H(f) = -∑ y∈{0,1} p(y)log 2 p(y),

[0050] Among them, p(y) represents the probability that the feature f is a positive or negative sample. High-entropy features usually cover more diverse user behaviors and have a wider influence. For example, the feature related to "Director: James Cameron" shows a high entropy value because it can attract a large number of different types of users. The global feature entropy can identify which features play an important role in the overall user behavior, so that they can be mainly retained during the denoising process.

[0051] On the other hand, the conditional entropy H(f|C u ) measures the uncertainty of the feature f under the given user attribute context vector C u . The specific calculation is as follows:

[0052]

[0053] Among them, p(y|C u ) represents the conditional probability in the conditional entropy. represents the expected value under the user attribute context vector C u .

[0054] The conditional entropy is used to measure the explanatory ability of the feature for the user's decision under the given context. A low conditional entropy value means that a specific user context strengthens the important role of the feature in the decision. For example, the influence of the 'director' feature on science fiction movies is higher than that on romantic comedies. In this way, the diffusion model can dynamically adjust the feature weights to make them more in line with the specific interests of users. To estimate the conditional probability p(y|C u ) in the conditional entropy, a neural network is used for modeling, and the specific expression is:

[0055] p(y = 1|C u ) = σ(WC u + b f ),

[0056] Among them, σ represents the Sigmoid activation function, and W and b f are model parameters to be learned. Through this estimation method, the diffusion model can accurately predict the probability that the feature is a positive sample in a specific user context, and then calculate the conditional entropy.

[0057] S1.3.5: Based on the calculation results of the information gain, assign a weight W u to each user and feature to obtain the user feature weight matrix Among them, the feature weight of user u is expressed as:

[0058]

[0059] Among them, To suppress the popularity feature weight, β = 0.3 is used to balance personalization and debiasing. Specifically, is an indicator function that is 1 when the feature f belongs to the feature set of user u′ and 0 otherwise.

[0060] S2: Construct a knowledge graph denoising process based on the diffusion model based on the user feature weight matrix, and use the original knowledge graph and the user-item interaction data to train the diffusion model.

[0061] Before use, the diffusion model of this embodiment needs to be trained first. During the training process, the original knowledge graph, user-item interaction data, and the user feature weight matrix W obtained in step S1 are input into the diffusion model IG , and the diffusion model adds noise through forward diffusion and denoises through reverse diffusion to generate a denoised subgraph of the original knowledge graph and uses the diffusion loss for optimization.

[0062] S2.1: Gradually add Gaussian noise to the knowledge graph to simulate the process of graph structure damage, and finally make the knowledge graph state approach the standard Gaussian distribution.

[0063] It should be noted that, as Figure 2 shown, the diffusion model usually consists of a forward diffusion process and a reverse diffusion process. Among them, forward diffusion: gradually add noise to the original knowledge graph to make the information in the knowledge graph transition from the original data distribution to the pure Gaussian noise distribution, simulating the gradual damage of the graph structure. Reverse diffusion: gradually denoise from the state completely covered by noise, restore the semantic information of the knowledge graph, and generate a high-quality denoised graph.

[0064] In the forward process, by gradually adding Gaussian noise to the original knowledge graph, the process of damage to the original structure of the knowledge graph is simulated. First, the input of the diffusion model, that is, the initial state z 0 , is initialized as the original entity adjacency matrix A. The forward process usually adds noise according to the following formula:

[0065]

[0066] where z t is the knowledge graph state at time step t obtained using the above formula, z t-1 is the knowledge graph state at time step t - 1, ∈ t is the standard Gaussian noise, α t is the attenuation coefficient used to control the noise level, is the standard Gaussian distribution with a mean of 0 and a variance of 1.

[0067] The user feature weight matrix W calculated based on step S1 IG In the present invention, node noise is dynamically regulated during the diffusion process to protect important nodes, thereby improving the semantic recovery quality of reverse denoising. The update formula is as follows:

[0068]

[0069] where z t is the knowledge graph state at time step t obtained using the above formula. ⊙ represents element-wise multiplication, and W IG reflects the importance weight of the user for the graph nodes, and α t represents the weighting factor of the current time step z t-1 .

[0070] Furthermore, in order to capture the time step features and provide time dynamic information to the diffusion model, the present invention uses sine and cosine functions to generate the time step embedding φ(t). The specific definition is as follows:

[0071]

[0072] where k is the dimension index and d is the dimension of the embedding vector of entity e. By combining multi-dimensional sine and cosine functions to guide the diffusion to distinguish the noise characteristics at different times, φ(t)[2k] represents the sine encoding based on time step t (even positions), and φ(t)[2k + 1] is the cosine encoding based on time step t (odd positions). These embeddings can provide a discriminative time step representation for the diffusion model through the combination of sine and cosine functions.

[0073] As the forward process time progresses, the noise gradually increases, and the relationship between items and entities becomes blurred. The whole process can be regarded as a Markov chain z 1:T :

[0074]

[0075] where t ∈ 1,…,T represents the time step, and α t represents the weighting factor of the current time step z t-1 , usually defined as a decreasing sequence. 1 - α t represents the noise variance at time step t.

[0076] When the time step T approaches positive infinity, the state z T converges to the standard Gaussian distribution. By using the reparameterization trick and the additivity of two independent Gaussian noises, the diffusion process can be recursively expanded, and the state z 0 at any time step can be directly derived from the initial state z t :

[0077]

[0078] Furthermore, to adjust the noise addition in z 1:T , the present invention employs a linear noise scheduler to control the amount of noise added at each step. Its definition is as follows:

[0079]

[0080] where the hyperparameter s ∈ [0, 1] controls the scale of the noise, and n low < n high ∈ (0, 1) are respectively the lower and upper bounds of the set noise.

[0081] S2.2: Gradually denoise the results obtained from the forward process through conditional probability modeling, and dynamically adjust the node weights in combination with user interaction information to obtain the denoised knowledge graph.

[0082] Specifically, the reverse diffusion process uses a denoising network to perform the denoising process through conditional probability modeling, and gradually restores the semantic information of the knowledge graph from the state z T completely covered by noise. The specific formula is:

[0083]

[0084] where, ∈ θ (z t , t, W IG ) is the noise predicted by the denoising network at the current time step.

[0085] In this embodiment, based on the standard reverse denoising formula, at each time step, the coupling relationship between time dynamics and node importance is captured by combining user-item interaction information and time step features to guide the personalized restoration of the denoising process. The user feature weight matrix W JG is mapped to the conditional embedding ψ(W IG ) = WW IG + b through a linear transformation, where W and b are learnable parameters. Then, the noise at the current time step is predicted through a multi-layer perceptron:

[0086] ∈ θ (z t , t, W IG ) = MLP(z t , φ(t), ψ(W IG ),

[0087] After T backtracks (T time steps), denoising is gradually performed at each time step, and finally a state representation closer to the real data is obtained i.e., the denoised data.

[0088] S2.3: The loss function of the diffusion model is constructed by maximizing the evidence lower bound (ELBO) to train and optimize the diffusion model during the denoising process, obtaining the trained diffusion model, and guiding the learning of the denoising distribution through the reconstruction loss function.

[0089] Specifically, to optimize the diffusion model, the method of maximizing the evidence lower bound (ELBO) is adopted to restore the relationships in the original knowledge graph. Specifically, the optimization objective of the probabilistic diffusion process can be summarized as the following formula:

[0090]

[0091] where, is used to calculate the expected value of the random variable. KL(·) represents the Kullback-Leibler divergence, which is used to measure the difference between two probability distributions. This optimization objective enables the diffusion model to accurately learn the correct denoising relationships at each diffusion time step by maximizing the likelihood of the denoising distribution. Since directly optimizing the ELBO is complex, the diffusion model uses the Monte Carlo method to approximate it. At the same time, using the Gaussian transformation hypothesis and Bayes' theorem, the above ELBO loss can be further simplified to the reconstruction loss in the form of mean squared error (MSE):

[0092]

[0093] where, is the denoising loss of the diffusion model, and ∈ represents the noise actually injected during the forward process.

[0094] Furthermore, to alleviate the potential limitations of the diffusion model in generating denoised knowledge graphs containing relationships related to downstream recommendation tasks, this embodiment proposes a collaborative knowledge graph convolution mechanism. This mechanism uses the supervision signal of the recommendation task to guide the process of knowledge graph diffusion training. Specifically, first, the user-item interaction information and the predicted relationship probability in the knowledge graph are aggregated to obtain an updated user-item interaction matrix. Then, this updated user-item interaction matrix is combined with the user embedding to obtain an item embedding that integrates the knowledge graph and user information. Finally, by calculating the mean squared error (MSE) between the aggregated item embedding and the original item embedding, the diffusion model is assisted in optimization:

[0095]

[0096] where, is the interaction information fusion loss, Y represents the user-item interaction matrix, E u and E jThey represent user embeddings and item embeddings respectively. At the initialization of the model, user embeddings and item embeddings are randomly initialized with dimension d. These embedding vectors will be continuously updated during the training process. Through iterative training, the model can gradually optimize the embedding representation, making the value of the loss function gradually decrease and obtaining the best embedding vectors at the end of the training.

[0097] The training of the model of the present invention consists of the training of the recommendation task and the training of the knowledge graph diffusion. The loss function of the diffusion model can be expressed as:

[0098]

[0099] wherein, it contains two loss components: is used to learn the correct denoising distribution at each diffusion moment, ensures that the aggregated user-item embeddings can capture the correlation between the knowledge graph and user behavior, and the hyperparameter λ 1 is used to balance the two losses.

[0100] S2.4: Reconstruct the nodes and edges of the denoised data to generate the final denoised knowledge graph.

[0101] Specifically, in the denoised data each node (i.e., entity) has undergone diffusion denoising and has an updated representation. To reconstruct the edge structure of the knowledge graph, the present invention selects the k nearest neighbor nodes based on the Euclidean distance to ensure that the node pairs with high semantic relevance are retained. Specifically, the nearest neighbor set is defined as:

[0102]

[0103] wherein, represents the set of the k nearest neighbor nodes of item i, and d(·) represents the Euclidean distance between two nodes, and its calculation formula is as follows:

[0104]

[0105] Finally, the denoised edge set is calculated

[0106] To maintain the structural integrity of the knowledge graph, a symmetric edge construction strategy needs to be introduced. In many real applications, some relationships are symmetric, such as "friendship" in a social network or "bidirectional connection" between some entities. Therefore, it is necessary to ensure that the edge set of the denoised knowledge graph is bidirectional, that is, if the edge (h,t) exists, its reverse edge (t,h) should be added simultaneously. For this purpose, the present embodiment defines the symmetric edge set construction formula as follows:

[0107] ε sym = εdenoised ∪{(t, h) | (h, t) ∈ ε denoised},

[0108] where ε denoised refers to the edge set reconstructed based on the Euclidean distance of node representations after the denoising process. ε sym is a symmetric edge set, ensuring the symmetry of the edge relationships in the denoised knowledge graph. This process guarantees the connectivity and structural integrity of the knowledge graph, while enhancing its usability in the recommendation task. Finally, an undirected graph of the optimized knowledge graph is obtained

[0109] S3: Using the denoised knowledge graph, heterogeneous knowledge aggregation and contrastive reinforcement learning, train the recommendation model to obtain the trained recommendation model.

[0110] Step S3 of this embodiment specifically includes:

[0111] S3.1: Input the denoised knowledge graph, user-item interaction matrix, and the initial embedding vectors of users, items, and entities into the constructed recommendation model. Adopt the graph attention mechanism to update the embedding vectors of items, assign attention weights to different relationship types, capture the semantic information of the knowledge graph, and prevent overfitting.

[0112] To handle heterogeneous relationships in real-world knowledge graphs, the present invention designs a relation-aware knowledge embedding layer. This embedding layer uses the graph attention mechanism to capture the semantic information of different relationships in the knowledge graph, assigns different attention weights to each relationship type, and ensures that the recommendation model can distinguish the importance of different relationships. The message aggregation formula between an item and its connected entities, that is, the update formula of the item, is expressed as:

[0113]

[0114] where and respectively represent the embedding vectors of the item before and after update, x j represents the embedding vector of the item after update, represents the set of adjacent entities of item j belonging to different types of relationships r in the knowledge graph e,j , σ represents the random dropout operation to prevent overfitting, and W is a trainable parameter matrix. α(e, r e,j , j) represents the attention weight, and the calculation formula is:

[0115]

[0116] Among them, LeakyReLU is a non-linear activation function, and || represents the vector concatenation operation.

[0117] S3.2: Use a combination of a graph-based collaborative filtering framework and heterogeneous knowledge aggregation to update the embedding representations of users and items. The present invention employs a local graph embedding propagation layer, which can be described as follows:

[0118]

[0119] Among them, and respectively represent the embedding vectors of user u and item j input at the l-th layer of the graph propagation layer. and respectively represent the embedding vectors of user u and item j output after update at the l-th layer of the graph propagation layer. represents the number of adjacent users of user u, represents the number of adjacent items of item j. By stacking L graph propagation layers, the outputs and at the L-th layer can capture higher-order collaborative signals.

[0120] S3.3: Utilize the enhanced embedding vectors and of users and items to obtain the predicted scores of user-item preferences, and the calculation formula is as follows:

[0121]

[0122] represents the predicted score of the recommendation model for user-item preferences.

[0123] It should be noted that existing knowledge graph recommendation research usually relies on simple random augmentation methods or simple cross-view comparisons between the original knowledge graph view and the collaborative filtering view. However, these random augmentations may introduce unwanted noise, and the supplementary knowledge graph view may contain irrelevant information. The present invention recognizes that among the large number of semantic relationships existing in the knowledge graph, only a part may be truly relevant to the downstream recommendation task. To address these challenges, the present invention uses the knowledge graph subgraph reconstructed by the generative model as the contrast view.

[0124] For the two knowledge graph enhanced views and regard the node embedding corresponding to any node in one view as the positive sample in the other view, while the other node embeddings except this node in the two views are regarded as negative samples. A contrast loss function is defined, aiming to maximize the consistency between positive pairs while minimizing the consistency between negative pairs. The contrast loss can be expressed as:

[0125]

[0126] Among them, s(·) represents using the cosine similarity function to measure the similarity between two vectors, and x′ u represents the node embedding vector of user u in the knowledge graph enhanced view, and x′ v represents the node embedding vector of user v in the knowledge graph enhanced view. The hyperparameter τ is called the temperature and is used for the softmax operation. is the obtained contrastive loss on the user side. Similarly, use to calculate the contrastive loss on the item side. Combining these two losses, the loss function of the self-supervised task can be obtained and can be expressed as:

[0127] Furthermore, the Bayesian personalized loss is used to optimize the model to ensure that the preference score of the positive sample is higher than that of the negative sample. The specific calculation method is as follows:

[0128]

[0129] Among them, represents the Bayesian personalized loss, is the training data, is the observed historical interaction behavior of users and items, represents the unobserved interaction behavior obtained from the Cartesian product of the user set and the item set (excluding ).

[0130] Subsequently, the Bayesian personalized loss and the contrastive learning loss are jointly optimized to obtain the comprehensive optimization loss of the recommendation task:

[0131]

[0132] Among them, the learnable model parameters are represented by Θ, which includes the trainable variables in the model. In addition, λ 2 is a hyperparameter that determines the loss intensity based on contrastive learning.

[0133] It should be noted that the diffusion model and the recommendation model in this embodiment are trained simultaneously,

[0134] S4: Input the user to be predicted and the user-item interaction data into the trained recommendation model to obtain a relevant recommendation list for the user.

[0135] Step S4 of this embodiment includes:

[0136] S4.1: Using the trained recommendation model, input the user-item interaction feature data to predict the ratings of each user for the un-interacted items. The recommendation model will generate a rating matrix based on the user's historical preferences, item features, and semantic relationships extracted from the knowledge graph, representing the potential interest values of each user for different items.

[0137] S4.2: Sort the rating matrix and generate the top-N most relevant items for each user to form a personalized recommendation list.

[0138] Exemplarily, assume in a movie recommendation system:

[0139] The input includes:

[0140] - Target user ID: User A.

[0141] - Historical behavior of user A to be predicted: Watched "Avatar" and "Titanic".

[0142] Model inference:

[0143] - The recommendation model generates a user embedding vector x based on the historical behavior of user A u .

[0144] - The recommendation model generates an embedding vector x of candidate movies based on the semantic relationships in the knowledge graph j .

[0145] - The recommendation model calculates the interest ratings of user A for each movie

[0146] Output:

[0147] - Generate the top-N recommendation list for user A, for example:

[0148] 1. "The Terminator"

[0149] 2. "Interstellar"

[0150] 3. "Inception"

[0151] S4.2: Use Recall@N and NDCG@N as evaluation metrics to test the performance of the model in the top-N recommendation task.

[0152] Recall mainly focuses on whether the system can find the items that the user is truly interested in. It calculates the proportion of relevant items in the recommended results and emphasizes the coverage of the recommendations. NDCG (Normalized Discounted Cumulative Gain), on the other hand, pays more attention to the ranking order of the items. It not only considers the relevance of the recommended items but also their positions in the recommendation list. Relevant items ranked higher contribute more to the NDCG, reflecting the sorting quality of the recommendation system.

[0153] The following further illustrates the effect of the knowledge graph diffusion recommendation method based on information gain guidance of the present invention through simulation experiments:

[0154] (1) Simulation experiment conditions:

[0155] The hardware environment of the simulation experiment: 13th Gen Intel(R) Core(TM) i9-13900KF CPU, 32GB of memory, RTX 4090 GPU. The software environment: Ubuntu20.04 operating system, Python3.9, Pytorch1.13.0.

[0156] (2) Experimental content and result analysis:

[0157] The simulation experiment of the present invention performs the recommendation task on three publicly available knowledge graph recommendation datasets, Last-FM, MIND, and Alibaba-iFashion, and evaluates the performance of the method of the present invention through the evaluation metrics Recall@20 and NDCG@20. At the same time, the experimental results in multiple papers are cited and compared with the method proposed by the present invention to prove the advantages of the algorithm of the present invention. The main experimental parameters of the simulation experiment of the present invention are shown in Table 1.

[0158] Table 1 Simulation experiment parameters

[0159]

[0160]

[0161] Experiments were conducted on the recommended datasets Last-FM, MIND, and Alibaba-iFashion using the knowledge graph diffusion recommendation method proposed in the present invention. The experimental results compared with those in multiple papers are shown in Table 2. To further verify the advantages of the present invention, the experimental results of the knowledge graph recommendation algorithms for the above datasets in the cited papers were compared. According to the data in Table 2, compared with the current mainstream knowledge graph recommendation algorithms, the present invention achieved better performance in the evaluation metrics Recall@20 and NDCG@20. This indicates that the present invention is superior to traditional algorithms in terms of recommendation effect, ranking quality, and multi-scenario application ability.

[0162] Table 2 Main experimental results

[0163]

[0164] In contrast, traditional algorithms have significant deficiencies in dealing with knowledge graph recommendation tasks. Collaborative filtering methods (BPR and LightGCN) only rely on the similarity of user behavior and lack effective solutions to the sparse data problem; there are obvious performance gaps between other knowledge-aware models (CKE, KGAT, KGIN, MCCLK, and KGCL) and the method of the present invention, which indicates that knowledge graphs usually contain irrelevant relationships and will have a negative impact on recommendation quality. Compared with these methods, the present invention effectively suppresses the noise in the knowledge graph through a carefully designed diffusion model, uses a heterogeneous knowledge aggregation mechanism to better extract multi-level semantic information, and combines contrastive learning to further enhance the generalization ability of the model. Compared with traditional algorithms, the present invention can more efficiently handle the challenges of recommendation tasks in multi-relational heterogeneous graphs and significantly improve the accuracy and robustness of recommendations.

[0165] The present invention proposes a knowledge graph diffusion recommendation method guided by information gain, which realizes a significant improvement in the performance of the recommendation system by combining a knowledge graph diffusion model, a heterogeneous knowledge aggregation mechanism, and a contrastive learning strategy. The present invention proposes a denoising mechanism based on the knowledge graph, which effectively removes the noise in the knowledge graph through the forward diffusion and reverse diffusion processes, while retaining the high-quality information related to the recommendation task. By dynamically adjusting the node noise to protect important nodes and combining information gain to quantify the weights of features, the key relationships between user behavior and item attributes are accurately retained during the denoising process. Compared with the method of directly using the original knowledge graph, the present invention can generate a more pure and relevant subgraph, greatly improving the accuracy of the recommendation result and the generalization ability of the model.

[0166] The present invention designs a heterogeneous knowledge aggregation mechanism, which differentiates different relationships in the knowledge graph through the graph attention mechanism, assigns corresponding attention weights to each relationship, and captures diverse semantic information. Meanwhile, the local graph embedding propagation layer is used to mine high-order collaborative signals to generate more comprehensive and accurate embedding representations for users and items. By introducing a contrastive learning strategy, the present invention enhances on the basis of the denoised knowledge graph subgraph, effectively reducing the interference of irrelevant information. At the same time, through the optimization of positive and negative sample pairs, the model's ability to capture the semantics of the knowledge graph and the robustness of the recommendation results are further improved.

[0167] Another embodiment of the present invention provides a storage medium in which a computer program is stored, and the computer program is used to execute the steps of the knowledge graph diffusion recommendation method guided by information gain in the above embodiment. Another aspect of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the knowledge graph diffusion recommendation method guided by information gain in the above embodiment are implemented. Specifically, the above integrated modules implemented in the form of software function modules can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0168] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A knowledge graph diffusion recommendation method based on information gain guidance, characterized in that: include: S1: Preprocess the original knowledge graph and user-item interaction data, calculate information gain, and obtain the user feature weight matrix; S2: constructing a knowledge graph denoising process based on a diffusion model based on the user feature weight matrix, and training the diffusion model using the original knowledge graph and the user-item interaction data; S3: Use the denoised knowledge graph to train the recommendation model to obtain a trained recommendation model; S4: Input the user ID to be predicted and the user-item interaction data into the trained recommendation model to obtain a relevant recommendation list for the user.

2. The knowledge graph diffusion recommendation method based on information gain guidance according to claim 1 is characterized in that: The S1 includes: S1.1: Construct a user-item interaction matrix based on the interaction data between users and items in, Indicates the number of users, Represents the number of items, and the element Y in the user-item interaction matrix Y uj is 0 or 1, Y uj =1 indicates that user u and item j have interactive behavior, otherwise Y uj =0; S1.2: Represent the original knowledge graph in the form of triples as Among them, ε represents the entity set, Represents a set of relations, Represents a triple set, where items Mapping to entity e∈ε; based on triples Constructing the entity adjacency matrix Among them, e h Represents a triple (e h ,r,e t ) in the head entity, e t Represents a triple (e h ,r,e t ) in the tail entity, r represents e h With e t The relationship between A h,t =1 indicates the head entity e h With the tail entity e t There is a direct correlation. h,t =0 indicates the head entity e h With the tail entity e t No related relationship; S1.3: Based on the original knowledge graph and the interaction data between users and items, obtain the user feature weight matrix.

3. The knowledge graph diffusion recommendation method based on information gain guidance according to claim 2 is characterized in that: The S1.3 includes: S1.3.1: Calculate the global information entropy H(e) of each entity e in the original knowledge graph, and construct a semantic feature set for each item j: in, represents the semantic feature set of item j, H(e t ) indicates the tail entity e t The global information entropy of Indicates that item j is mapped to the head entity e h ; S1.3.2: Extract the interactive item associated entities of each user based on the user-item interaction data, and perform weighted aggregation on the interactive item associated entities to generate a user attribute context vector: Among them, C u represents the user attribute context vector of user u, H(e) represents the global information entropy of entity e, and H(e ′ ) represents entity e ′ The global information entropy, ε u represents the entity set corresponding to all items with which user u has historical interaction records, R e represents the embedding vector of entity e, η represents the entropy sensitivity coefficient; S1.3.3: Combining user behavior with knowledge graph semantics, construct positive sample datasets and negative sample datasets for each user, where the positive sample dataset of user u is shown as The negative sample data set of user u is randomly sampled from non-interactive items with user u, satisfying S1.3.4: Using global feature entropy and conditional entropy, define the information gain IG(f|u) of feature f: IG(f|u)=H(f)-H(f|C u ), Among them, H(f) represents the global feature entropy, H(f|C u ) represents conditional entropy; S1.3.5: Based on the information gain IG(f|u), assign a weight W to each user and feature u , get the user feature weight matrix Among them, the feature weight of user u is expressed as: in, Used to suppress popular feature weights, is the indicator function, when feature f belongs to user u ′ It is 1 when the feature set is , otherwise it is 0.

4. The knowledge graph diffusion recommendation method based on information gain guidance according to claim 1 is characterized in that: The S2 includes: S2.1: Input the entity adjacency matrix of the original knowledge graph, the user-item interaction data and the user feature weight matrix W into the diffusion model IG , based on the user feature weight matrix, adding Gaussian noise to the knowledge graph in T time steps to make the state of the knowledge graph approach the standard Gaussian distribution; S2.2: Use the denoising network to gradually denoise the results obtained in the forward process through conditional probability modeling, dynamically adjust the node weights based on user interaction information, and obtain denoised data; S2.3: The loss function of the diffusion model is constructed by maximizing the lower bound of evidence, so as to train and optimize the diffusion model in the denoising process and obtain the trained diffusion model; S2.4: Reconstruct the nodes and edges of the denoised data to generate the final denoised knowledge graph.

5. The knowledge graph diffusion recommendation method based on information gain guidance according to claim 4 is characterized in that: The S2.1 includes: In the forward process, Gaussian noise is gradually added to the original knowledge graph, where the noise addition process is expressed as: Among them, z t represents the state of the knowledge graph at time step t, z t-1 represents the knowledge graph state at time step t-1, ∈ t is standard Gaussian noise, α t is the attenuation coefficient used to control the noise level, is a standard Gaussian distribution with a mean of 0 and a variance of 1, W IG represents the user feature weight matrix; Get the state z of the diffusion model at any time step t The relationship between the initial state z0: Among them, the entity adjacency matrix of the original knowledge graph is used as the initial state z0, Use a linear noise scheduler to control the amount of noise added at each time step: Among them, the hyperparameters s∈[0,1], n low <n jigh ∈(0,1) are the lower and upper limits of the noise respectively.

6. The knowledge graph diffusion recommendation method based on information gain guidance according to claim 4 is characterized in that: The S2.2 includes: The denoising network is used to model the denoising process through conditional probability, from the knowledge graph state z that is completely covered by noise T Start to restore the semantic information of the knowledge graph in T time steps, where the denoising process is expressed as: Among them, ∈ θ (z t ,t,W IG ) is the noise of the current time step predicted by the denoising network, and the denoising network predicts the noise of the current time step using a multilayer perceptron: ∈ θ (With t ,t,W IG )=MLP(from t ,φ(t),ψ(W IG ), Where MLP represents multi-layer perceptron, φ(t) represents time step embedding, ψ(W IG )=WW IG +b indicates conditional embedding, and W and b are learnable parameters.

7. The knowledge graph diffusion recommendation method based on information gain guidance according to claim 6 is characterized in that: The loss function of the diffusion model is expressed as: in, is the denoising loss of the diffusion model, is the mutual information fusion loss, and the hyperparameter λ1 is used to balance the two losses. Among them, ∈ represents the noise actually injected in the forward process; Among them, Y represents the user-item interaction matrix, E u and E j represent user embedding and item embedding respectively.

8. The knowledge graph diffusion recommendation method based on information gain guidance according to claim 6 is characterized in that: The S3 includes: S3.1: The denoised knowledge graph, user-item interaction matrix, and the initial embedding vectors of the user, item, and entity are input into the constructed recommendation model, and the embedding vectors of the items are updated using the graph attention mechanism. The update formula is expressed as: in, and They represent the embedding vector of the item and the embedding vector of the entity before updating, respectively. j represents the embedding vector of the updated item, Represents item j in the knowledge graph In different types of relations e,j The set of adjacent entities, σ represents the random drop operation to prevent overfitting, and W is a trainable parameter matrix; S3.2: Through the L-layer graph propagation layers connected sequentially, the high-order collaborative signals of users and items are gradually aggregated to generate enhanced embedding vectors of users and items: in, and They represent the embedding vectors of user u and item j input in the l-th layer of graph propagation layer, and They represent the embedding vectors output by user u and item j after updating in the l-th layer of graph propagation layer, represents the number of adjacent users of user u, Represents the number of adjacent items of item j. S3.3: Leveraging enhanced embedding vectors for users and items and Get the predicted scores for user-item preferences: in, Represents the prediction score of the recommendation model for user-item preference; S3.4: Use the recommendation model loss function to update the parameters of the recommendation model.

9. The method for knowledge graph diffusion recommendation based on information gain guidance according to claim 8, characterized in that: The recommendation model loss function is: Among them, λ2 is a hyperparameter, represents the loss function of the self-supervised task, is the contrast loss on the user side, is the contrast loss on the object side; represents the Bayesian personalized loss, expressed as: in, represents the collaborative filtering loss, is the training data, is the observed historical interaction behavior of users and items, Indicates that from the user set and item sets The unobserved interaction behavior is obtained from the Cartesian product of .

Citation Information

Cited By

  • Emergency field multi-dimensional inference knowledge graph construction and real-time response method

    CN120851167A

  • Knowledge tracking data enhancement method based on adaptive diffusion model

    CN121009372A

  • Diffusion model training method and device based on spatial knowledge graph guidance, spatio-temporal data generation method and device, equipment and medium

    CN121436072A