Knowledge Graph Preference Prediction and Recommendation Method Based on Attention Mechanism

Through the knowledge graph preference prediction method based on attention mechanism, the problem of insufficient recommendation accuracy in the existing technology is solved, and more accurate personalized recommendations are achieved through preference propagation and path information integration.

CN115618009BActive Publication Date: 2025-07-11CHINA JILIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211169208.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-26
Publication Date
2025-07-11
Estimated Expiration
2042-09-26

AI Technical Summary

Technical Problem

The existing recommendation methods based on knowledge graphs have obvious shortcomings in accuracy, especially the connection and embedding methods are not effective in recommendations of different knowledge graphs, and improvements are needed to improve the recommended accuracy.

Method used

The knowledge graph preference prediction recommendation method based on attention mechanism is adopted, through symbol definition, preference propagation method design, representation of the above graph and loss function definition, the path information is integrated using the scaling dot product attention mechanism, to calculate the degree of interest of the candidates, and to perform personalized recommendations.

Benefits of technology

It improves the accuracy of recommendations, makes recommendation information more accurate, conforms to user interests, and improves recommendation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003862825550000033
    Figure BDA0003862825550000033
  • Figure BDA0003862825550000057
    Figure BDA0003862825550000057
  • Figure BDA0003862825550000059
    Figure BDA0003862825550000059
Patent Text Reader

Abstract

A knowledge graph preference prediction and recommendation method based on an attention mechanism, comprising the following steps: the first step, symbol definition; the second step, design of a preference propagation method; the third step, resource recommendation based on preference propagation; the fourth step, graph context representation based on scaled dot-product attention; the fifth step, loss function definition, and finally, the calculation result is used to determine whether to add item v # to the recommendation list of user u # . The present invention introduces the idea of preference propagation, obtains the interest set of users on the data set, adapts to complex paths composed of multiple head entities, then prunes the paths and introduces entity-relationship context pairs to enhance the preference interest of users and improve the accuracy of recommendations. Then, the scaled dot-product attention mechanism is used to integrate all path information and calculate the attention degree of users to candidate entities to provide personalized recommendations for users. The recommended information is made more precise to the parts that users are interested in, and the recommendation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a knowledge graph preference prediction and recommendation method based on an attention mechanism. Background Art

[0002] A knowledge graph describes concepts, entities and the relationships between them in the objective world in a structured form, expresses the information on the Internet in a form closer to the human cognitive world, and provides an ability to better organize, manage and understand the vast amount of information on the Internet. The knowledge graph has brought vitality to Internet semantic search and also shown great power in intelligent question answering, and has become the infrastructure of Internet knowledge-driven intelligent applications. The knowledge graph is usually represented in a basic structure of triples, and each triple (e h , r, e t ) includes a head entity e h , a tail entity e t and the mutual relationship r between the entities.

[0003] With the development of Internet technology, information technology has brought convenient life to humans. However, some problems have also arisen, among which the problem of information overload caused by data explosion is the most prominent. As an information filtering system, the recommendation system based on the knowledge graph can not only solve the problem of information overload, but also has important practical significance for promoting technological progress and improving the quality of life.

[0004] At first, the recommendation based on the knowledge graph mainly used the relationships in the knowledge graph for recommendation, that is, recommendation based on connections. For example, Shi et al. proposed the SemRec method, which weighted and calculated heterogeneous information networks and paths, described the semantic information of paths by distinguishing different attributes of paths, and finally used the user's preference for paths for personalized recommendation. Finally, SemRec weighted and summed the scoring functions of all users in the path as the preference degree of the target user for the item. However, the recommendation system based on connections highly depends on the connection pattern of the knowledge graph, and the recommendation effect for different knowledge graphs is not ideal.

[0005] Many researchers have begun to try to bring knowledge graph embedding methods into recommendation systems. For example, Wang et al. proposed the MKR method, which uses a multi-layer perceptron and cross-compression units to extract the features of users and items, and then judges the similarity between the entities in the knowledge graph and the item features. Finally, the entities with high similarity are recommended to users. Ma et al. used the knowledge graph completion method for recommendation and proposed the RecKGC method. RecKGC uses a soft attention mechanism to integrate the degree of user attention to items, then calculates the semantic association between the integrated vectors of users and items, and finally makes recommendations for users based on the association results. However, currently, in the methods of recommendation based on knowledge graphs, the use of embedding-based or connection-based methods has been proven to have obvious defects, and further research on knowledge graph recommendation methods is needed to improve the accuracy of knowledge graph recommendation. Summary of the Invention

[0006] To overcome the deficiencies of the prior art, the present invention proposes a knowledge graph preference prediction and recommendation method based on an attention mechanism, which uses the knowledge graph as the main information source for recommendation. The recommendation task is to predict the degree of interest of user u # in item v. The method takes the historical interest set of user u # and candidate item v # as inputs, and the probability P(u # likes item v # ),v # ,v # ) as the output. The present invention determines whether to add item v # to the recommendation list of user u # through the final calculation results of steps such as symbolic definition, preference propagation method design, resource recommendation based on preference propagation, graph context representation based on scaled dot product attention, and loss function definition.

[0007] To solve the above technical problems, the present invention provides the following technical solutions:

[0008] A knowledge graph preference prediction and recommendation method based on an attention mechanism, comprising the following steps:

[0009] The first step, symbolic definition, the process is as follows:

[0010] Given a user set U # ={u # 1,u # 2,...u # μ} and an item set V # ={v # 1,v # 2,...v #υ}, u # μ represents the μ-th user in the user set, v # υ represents the υ-th item in the item set, and the interaction behaviors between users and items are put into an interaction matrix where has only two values, 0 or 1. If user u # has interaction behaviors such as browsing and clicking with item v # , then otherwise it is equal to 0. Each user has a historical interest set represents the θ-th element in the historical interest set of user u # , and the elements in the historical interest set represent the items that have interaction behaviors with the user; U # and V # are both parts of the knowledge graph G. The knowledge graph G = {E, R, T}, where E represents the set of entities, R represents the set of relations, and T represents the set of positive example triples. For the triple (e h , r, e t ), use v # e h , v # r, v # e t to represent the embeddings of e h , r, e t in the item vector space respectively. Items are regarded as entities in the knowledge graph, V # ∈E, and the items in the item set V # are regarded as candidate items for user recommendation. If a candidate item has already interacted with the user in the interaction matrix, then this item is removed from the candidate items. The candidate item set is defined as: The goal of the present invention is to predict the probability that the user will like or click on the candidate item v # g ;

[0011] Step 2: Design of the preference propagation method;

[0012] Step 3: Resource recommendation based on preference propagation;

[0013] Step 4: Graph context representation based on scaled dot-product attention;

[0014] Step 5: Definition of the loss function, and finally calculate the result to determine whether to add the item v # to the recommendation list of user u # .

[0015] Furthermore, the process of the second step is as follows:

[0016] Step (2.1): Regard the historical data of the user as seeds, and in the knowledge graph, with the seeds as the center, spread outwards along the relationship path p = (r1, r2,... r l ), where r l represents the l-th relationship in the relationship path. Each iteration propagates to the next entity along the direction of the relationship. In order to describe the multi-level preferences of the user with the knowledge graph, a multi-hop relevant entity set is divided for the user u # . A k-hop relevant entity set is defined as:

[0017]

[0018] where represents the historical interest set of the user, that is, the set of items that have interacted with the user. H represents the maximum value of k. represents the range of values that e t can take in the relevant entity set;

[0019] Step (2.2): The relevant entities are regarded as the natural extension of the user's interests in the knowledge graph. These natural extensions are manifested as circles that spread out layer by layer from the historical interest items in the knowledge graph, simply called ripple circles. The set of triples composed of relevant entities in one layer of ripple circles is called a ripple set. The k-hop ripple set of the user u # is defined as follows:

[0020]

[0021] Step (2.3): During the process of the ripple set spreading outwards, the present invention will pay attention to the relevant entity set of each hop. If the existence of v is found in # g , the ripple set will continue to spread outwards until the H-th hop stops. If the existence of v is still not found in the # g at the H-th hop, then H = H + 1, and the ripple set will continue to spread outwards until v # g is included in , and at this time k * = H;

[0022] Step (2.4): After the ripple set stops spreading, sort out all the triples from to and, according to the order in The order of appearance in it reconstructs the entities and relationships into a user interest graph g composed of entities and relationships. g is a subset of the knowledge graph G, representing the user's interest set.

[0023] Furthermore, the process of the third step is as follows;

[0024] Step (3.1): In the knowledge graph, a user has multiple historical items. In order to adapt to a complex knowledge graph, among the historical items and the candidate item v # g The graph structure between them can be expressed as: S0→S1→S2...→S l , → means spreading from the previous node to the next node, where S i represents the i-th node between the historical item and the candidate item, and the entity The resources obtained from the previous ripple set are defined as:

[0025]

[0026] represents the ripple set All the tail entities in are of the triples, represents the ripple set All the head entities in are of the triples. Define the user's initial resource as R p (all)=1. In the recommendation system, the user's initial resource represents the user's interest in the item. Assuming that the user's interest in each historical item is the same, then the interest of each historical item

[0027] Step (3.2): The user's interest in the candidate item v # g can be defined as the user starting from all historical items and following and v # g The amount of resources finally reaching v # g In the process of resource flow, interactions are likely to occur between the paths starting from multiple historical items. There are entities at the intersections of multiple paths, just like the intersections of line segments. At this time, the resource amount of the intersection entity is the sum of the resource amounts of multiple direct predecessors. The entity at the intersection point represents that the user is very likely to prefer this related entity. Finally, record the interest resource amount of the candidate item v # g as Resv #g 。

[0028] Furthermore, the process of the fourth step is as follows:

[0029] Step (4.1): Further simplify the user interest graph g, only retaining the paths from historical items to candidate items. During the process of going from multiple historical items to candidate items, each historical item will generate one or more paths. After recording these paths, a set of paths is generated where p # i represents a historical item and the set of paths between candidate item v # g , p # i = {p i,1 , p i,2 ,... p i,j}, p i,j represents the j-th entity relationship set on the i-th path;

[0030] Step (4.2): The path ( represents the j * -th relationship in the i-th path, represents the j * -th entity in this path) can reflect the user's interest inference process to a certain extent. In path p i,j , use (h represents the entity in path p i,j , r represents the relationship in path p i,j ) to construct the entity-relationship context pair representation of path p i,j :

[0031]

[0032] Step (4.3): To measure the influence of the path on candidate item v # g , it is necessary to first construct a path pattern. Use an LSTM network (LSTM is a type of recurrent neural network for processing and predicting important events with relatively long intervals and delays in time series) to construct the path pattern. Input the entity-relationship context pairs in path p i,j into the LSTM network in sequence, and retain the current hidden state of the last LSTM as the context path vector representation of the candidate item. The calculation method of the current hidden state of the LSTM cell is as follows:

[0033]

[0034] Among them, f hidden is a non-linear activation function that includes ReLU operations. (v # e i,l , v # r i,l ) represents the l-th element in the i-th entity-relation context pair;

[0035] In step (4.4), in the context of the entity graph, each entity-relation context pair reflects partial features of the tail entity in the triple. At the same time, each entity-relation context pair contains different information and has different importance levels for the tail entity. For each entity-relation context pair, a scaled dot-product attention mechanism is used to calculate the attention coefficient of the tail entity to the entity-relation context pair, and the attention coefficients of multiple entity-relation context pairs in the path are multiplied to obtain the historical item 's degree of interest in the candidate item v # g ;

[0036] In step (4.5), the degree of attention of the historical item to the candidate item v # g along the path p i,j can be expressed as the product of the attention coefficients of multiple consecutive entity-relation context pairs. The degree of attention also represents the importance of the path p i,j to the recommendation, and the calculation method is as follows:

[0037]

[0038] Π represents the product symbol.

[0039] In step (4.6), the attention coefficient, as the weight of the path, is used to integrate path information. In the present invention, the historical item and the candidate item v # g are weighted and summed along the path, and the full-path pattern is calculated. The calculation formula is as follows:

[0040]

[0041] Among them, represents the above-path vector representation of the path p i,j , As the full-path pattern, it represents the aggregation of the user's preference interests spreading outward with the historical item as the center, and can be regarded as an approximate embedding of the user in terms of preference interests.

[0042] In step (4.7), finally, the aggregation of the user's preference interests and the candidate item v # gThe amount of interest resources and the embedded information are input into a linear function and the sigmoid function to calculate the probability P(u # ,v # g ) that the user will like the candidate item. The calculation formula is as follows:

[0043]

[0044] Preferably, the process of step (4.4) is as follows:

[0045] Step (4.4.1) The scaled dot-product attention mechanism can be described as a mapping from a query vector (Query) and a set of key-value pair vectors (Key-Value) to an output vector, denoted as Essentially, the output vector is generated by weighted combination of the input vector (Query), and the weight value of the input vector is the result of the scaled dot-product operation of the key vector (Key) and the query vector (Value);

[0046] Step (4.4.2), the embedding of the tail entity is regarded as the query vector, and the entity-relation context pair is used as both the key vector and the value vector at the same time. The calculation formula is as follows:

[0047] Q = W Q v # e t

[0048]

[0049]

[0050] Among them, W Q , is the linear transformation weight matrix, Q, respectively represent the parameterized representations of the query vector, the key vector and the value vector, W Q , d χ is the dimension of the column vector of the matrix W Q , d δ is the embedding dimension of the entity and the relationship;

[0051] Step (4.4.3), the degree of attention of the entity-relation context pair (e h ,r) to the tail entity, that is, the calculation formula of the attention coefficient is as follows:

[0052]

[0053] represents the transpose of the vector, Denote the scaling factor, which is used to keep the gradients stable during the backpropagation process. The softmax function, also known as the normalized exponential function. It is a generalization of the binary classification function sigmoid to multi-classification, aiming to present the results of multi-classification in the form of probabilities.

[0054] The process of the fifth step is as follows;

[0055] The loss function of the minimization method. To optimize the method, the triplet loss function based on self-adversarial negative sampling is adopted as the training objective of the method:

[0056]

[0057] The loss function consists of two parts. One is the loss caused during the prediction process, and the other is the loss caused during the embedding process. Among them, γ is a fixed margin, (e h ', r′, e t ) represents a negative example triplet constructed based on (e h , r, e t ). T represents the set of positive example triplets, and F represents the set of negative example triplets.

[0058] The beneficial effects of the present invention are as follows. A knowledge graph preference prediction and recommendation method based on the attention mechanism is proposed. The present invention introduces the idea of preference propagation, obtains the interest set of users on the data set, adapts to complex paths composed of multiple head entities, then prunes the paths and introduces entity-relationship context pairs to enhance the user's preference interest and improve the accuracy of recommendation. Then, the scaled dot-product attention mechanism is used to integrate all path information and calculate the degree of attention of the user to the candidate entity, so as to provide personalized recommendations for the user. Make the recommended information more accurate to the parts that the user is interested in, and improve the recommendation efficiency. Detailed implementation manners

[0059] The present invention will be further described below.

[0060] A knowledge graph preference prediction and recommendation method based on the attention mechanism includes the following steps:

[0061] The first step, symbol definition, the process is as follows:

[0062] Step (1.1), Given a user set U # ={u # 1, u # 2,... u # μ} and an item set V # ={v # 1, v # 2,... v # υ}, u# μ represents the μ-th user in the user set, v # υ represents the υ-th item in the item set, and the interaction behaviors between users and items are put into an interaction matrix where has only two values, 0 or 1. If user u # has interaction behaviors such as browsing and clicking with item v # , then otherwise it is equal to 0. Each user has a historical interest set represents the θ-th element in the historical interest set of user u # . The elements in the historical interest set represent the items that have interaction behaviors with the user. U # and V # are both parts of the knowledge graph G. The knowledge graph G = {E, R, T}, where E represents the set of entities, R represents the set of relations, and T represents the set of positive example triples. For the triple (e h , r, e t ), use v # e h , v # r, v # e t to represent the embeddings of e h , r, e t in the item vector space respectively. Items are regarded as entities in the knowledge graph, and V # ∈E. The items in the item set V # are regarded as candidate items for user recommendation. If a candidate item has already interacted with the user in the interaction matrix, then this item is removed from the candidate items. The candidate item set is defined as: The goal of the present invention is to predict the probability that the user will like or click on the candidate item v # g ;

[0063] Step 2: Design of the preference propagation method, the process is as follows:

[0064] Step (2.1): Regard the historical data of the user as seeds, and in the knowledge graph, with the seeds as the center, spread out along the relationship path p = (r1, r2,... r l ). r l represents the l-th relationship in the relationship path. Each iteration propagates along the direction of the relationship to the next entity. In order to describe the multi-level preferences of users using the knowledge graph, for user u #The relevant entity sets for multiple hops are partitioned. A relevant entity set for k hops is defined as:

[0065]

[0066] where represents the historical interest set of the user, that is, the set of items that have interacted with the user, and H represents the maximum value taken by k. represents the range of values that e t can take in the relevant entity set;

[0067] Step (2.2): The relevant entities are regarded as the natural extension of the user's interests in the knowledge graph. These natural extensions are manifested as circles spreading layer by layer from historical interest items in the knowledge graph. This invention is simply referred to as the ripple circle. The set of triples composed of relevant entities in one layer of the ripple circle is called the ripple set. For user u # the k-hop ripple set is defined as follows:

[0068]

[0069] Step (2.3): During the process of the ripple set spreading outwards, this invention will pay attention to the relevant entity sets of each hop. If the existence of v is found in # g , the ripple set will continue to spread outwards until the H-th hop stops. If the existence of v is still not found in the # g of the H-th hop, then H = H + 1, and the ripple set will continue to spread outwards until v # g is included in , and at this time k * = H;

[0070] Step (2.4): After the ripple set stops spreading, organize all the triples from to , and reconstruct the entities and relationships into a user interest graph g composed of entities and relationships according to the appearance order in . g is a subset of the knowledge graph G and represents the user's interest set;

[0071] The third step: Resource recommendation based on preference propagation, the process is as follows;

[0072] Step (3.1): In the knowledge graph, a user has multiple historical items. In order to adapt to a complex knowledge graph, the graph structure between the historical items and the candidate item v # g can be represented as: S0→S1→S2...→S l, → represents propagation from the previous node to the next node, where S i represents the i-th node between the historical item and the candidate item, the entity The resources obtained from the previous ripple set are defined as:

[0073]

[0074] represents the ripple set All the tail entities in of the triples, represents the ripple set All the head entities in of the triples, and define the user's initial resource as R p (all)=1. In the recommendation system, the user's initial resource represents the user's interest in items. Assuming that the user's interest in each historical item is the same, then the interest in each historical item

[0075] Step (3.2), the user's interest in the candidate item v # g can be defined as the user starting from all historical items and moving along and v # g to finally reach v # g of the resource volume. During the process of resource flow, interactions are likely to occur between paths starting from multiple historical items. There are entities at the intersections of multiple paths, just like the intersections of line segments. At this time, the resource volume of the intersection entity is the sum of the resource volumes of multiple direct predecessors. The entity at the intersection represents that the user is very likely to prefer this related entity. Finally, denote the interest resource volume of the candidate item v # g as Resv # g ;

[0076] Fourth step, the graph context representation based on scaled dot-product attention is as follows:

[0077] Step (4.1), further simplify the user interest graph g, only retain the paths from historical items to candidate items. During the process of going from multiple historical items to candidate items, each historical item will generate one or more paths. After recording these paths, a set of paths is generated where p # i represents the historical item and the candidate item v# g The set of paths between p i,j represents the j-th entity-relationship set on the i-th path;

[0078] Step (4.2), the path can reflect the user's interest reasoning process to a certain extent, represents the j-th * relationship in the i-th path represents the j-th * entity in this path. In path p i,j we use h to represent the entity in path p i,j and r to represent the relationship in path p i,j to construct the entity-relationship context representation of path p i,j as follows:

[0079]

[0080] Step (4.3), to measure the influence of the path on the candidate item v # g we need to first construct a path pattern. The present invention uses an LSTM network (LSTM is a type of recurrent neural network for processing and predicting important events with relatively long intervals and delays in time series) to construct the path pattern. The entity-relationship context pairs in path p i,j are sequentially input into the LSTM network, and the current hidden state of the last LSTM is retained as the context path vector representation of the candidate item. The calculation method of the current hidden state of the LSTM cell is as follows:

[0081]

[0082] where f hidden is a non-linear activation function containing ReLU operation. represents the l-th element in the i-th entity-relationship context pair,

[0083] Step (4.4) In the context of the entity graph, each entity-relationship context pair reflects some features of the tail entity in the triple. At the same time, each entity-relationship context pair contains different information and has different degrees of importance for the tail entity. For each entity-relationship context pair, a scaled dot-product attention mechanism is used to calculate the attention coefficient of the tail entity to the entity-relationship context pair, and the attention coefficients of multiple entity-relationship context pairs in the path are multiplied to obtain the degree of interest of the historical item in the candidate item v # g ;

[0084] Step (4.4.1) The scaled dot - product attention mechanism can be described as a mapping from a query vector (Query) and a set of key - value vector pairs (Key - Value) to an output vector, expressed as The output vector is essentially generated by the weighted combination of the input vector (Query), and the weight values of the input vector are the results of the scaled dot - product operation between the key vector (Key) and the query vector (Value);

[0085] Step (4.4.2), The embedding of the tail entity is regarded as the query vector, and the entity - relation context above is used as both the key vector and the value vector at the same time. The calculation formula is as follows:

[0086] Q = W Q v # e t

[0087]

[0088]

[0089] Among them, W Q , is the linear transformation weight matrix, Q, respectively represent the parameterized representations of the query vector, key vector, and value vector, W Q , d χ is the dimension of the column vector of the matrix W Q , d δ is the embedding dimension of the entity and the relation;

[0090] Step (4.4.3), The attention degree of the entity - relation context pair (e h , r) to the tail entity, that is, the calculation formula of the attention coefficient is as follows:

[0091]

[0092] represents the transpose of the vector, represents the scaling factor, which is used to keep the gradient stable during the back - propagation process. The softmax function, also known as the normalized exponential function. It is the generalization of the binary classification function sigmoid in multi - classification, and its purpose is to present the results of multi - classification in the form of probabilities;

[0093] Step (4.5), The historical item for the candidate item v # g along the path p i,jThe degree of attention can be expressed as the product of the attention coefficients of multiple consecutive entity-relationship context pairs, and the degree of attention also represents the importance of path p i,j to the recommendation, and the calculation method is as follows:

[0094]

[0095] ∏ represents the product symbol;

[0096] In step (4.6), the attention coefficient, as the weight of the path, is used to integrate path information. In the present invention, historical items and candidate item v # g The paths between are weighted and summed to calculate the full path pattern The calculation formula is as follows:

[0097]

[0098] where, represents the context path vector representation of path p i,j , As the full path pattern, it represents the aggregation of the user's preference interests spreading outwards with historical items as the center, and can be regarded as an approximate embedding of the user in terms of preference interests;

[0099] In step (4.7), finally, the aggregation of the user's preference interests and the interest resource amount of candidate item v # g The embedding information is input into a linear function and the sigmoid function to calculate the probability P(u # , v # g ) that the user will like the candidate item. The calculation formula is as follows:

[0100]

[0101] Step 5, Define the loss function, the process is as follows;

[0102] The loss function of the minimization method, in order to optimize the method, adopts the triplet loss function based on self-adversarial negative sampling as the training objective of the method:

[0103]

[0104] The loss function consists of two parts. One is the loss caused during the prediction process, and the other is the loss caused during the embedding process. Among them, γ is a fixed margin, (e h ', r′, e t ') represents based on (e h , r, e t)Constructed negative example triples, where \(T\) represents the set of positive example triples and \(F\) represents the set of negative example triples.

[0105] To evaluate the method, the present invention conducts experiments on three widely used recommendation benchmark datasets, namely MovieLens-1M, Book-Crossing, and Bing-News. MovieLens-1M is a widely used benchmark dataset in the movie recommendation field. Its main content is the ratings of related movies on the MovieLens website. The ratings range from 1 to 5, and there are more than 1 million rating data. Book-Crossing stores 1,149,780 ratings of different books in the Book-Crossing community. The book ratings range from 1 to 10, and the larger the rating, the more recommended the book is. Bing-News contains 1,025,192 news data collected from the Bing News server. Each news has a title and a part of the news snippet.

[0106] MovieLens-1M and Book-Crossing contain explicit interaction information between users and items. For example, if a user has rated a certain movie, it is considered that there is an explicit interaction behavior between the user and the movie. These interaction information can easily construct a user-item interaction matrix by building these interaction behaviors into an interaction matrix. The present invention uses Microsoft's knowledge graph construction tool Microsoft Satori to construct a knowledge graph for the dataset, and matches the head and tail entities of the triples in the same way to obtain the item id, and then excludes the items that have no match or multiple matches. An item without a corresponding id indicates that this is an incorrect data in the dataset, and there is no corresponding book or movie, which should be deleted to avoid interfering with the experimental results. When there are multiple items for an id match, if these items are retained, it is very likely to cause entity divergence and make the method give incorrect recommendation results. Therefore, the items in the match results are deleted from the dataset. After matching the id with the head and tail entities of the triples, select the triples with high confidence to form a new knowledge graph. The new graph is a subset of the original knowledge graph, and at the same time, the entity set is iteratively extended to a 4-hop range.

[0107] In the experiment where the user clicks on a candidate item, the present invention uses two metrics, Accuracy (abbreviated as ACC) and AUC, to measure the prediction ability of the method, and predicts the click-through rate (CTR) of the user clicking on the candidate item. ACC represents the accuracy of the method, and AUC represents the probability that when randomly selecting a positive sample and a negative sample, the current method ranks the positive sample before the negative sample. During the prediction process, according to the differences between the instance and the prediction result of the instance, the instance can be divided into 4 categories: True Positive (TP), True Negative (TN), False Positive (FP), and False Negative (FN). Among them, True Positive means that the instance is a positive example and is predicted as a positive class; True Negative means that the instance is a negative example and is predicted as a negative class; False Positive means that the instance is a negative example but is predicted as a positive class; False Negative means that the instance is a positive example but is predicted as a negative class; ACC calculates the proportion of all correctly predicted samples in the total samples, and the calculation formula is as follows:

[0108]

[0109] Calculating the value of AUC is to calculate the area under the ROC curve. The larger the area, the better the method. The ROC curve refers to the receiver operating characteristic (ROC). Each point on the ROC curve reflects the sensitivity to the same signal stimulus.

[0110] To verify the effectiveness of the proposed method, the present invention selects several representative and widely used knowledge graph recommendation methods for comparison, including CKE, DKN, PER, LibFM, RippleNet, and MKR. Table 1 reports the prediction results of the method of the present invention (hereinafter abbreviated as APPRM) and the comparison methods on three benchmark datasets. The numbers shown in bold in the table represent the optimal results in the control experiment, and the underlined numbers represent the sub-optimal results. - indicates that there is no data for the proposed method.

[0111] Table 1 CTR experiments on MovieLens-1M, Book-Crossing and Bing-News

[0112] Table 1 CTR experiments on MovieLens-1M, Book-Crossing and Bing-News

[0113]

[0114]

[0115] Table 1

[0116] As can be seen from Table 1, on the two datasets of MovieLens-1M and Book-Crossing, the APPRM method has better results, and multiple indicators are higher than those of methods such as CKE, DKN, and PER. Compared with the best RippleNet and MKR methods, the differences in various indicators are only 0.2% to 2.1%. On the Bing-News dataset, the APPRM method achieved the best results. The ACC indicator is generally 0.3% to 8.5% higher than that of each benchmark method, and the AUC indicator is 0.2% to 17.4% higher than that of each method. The APPRM method has obtained better results on all three datasets, which fully demonstrates the effectiveness of the method of the present invention in the CTR task. The indicators on the MovieLens-1M dataset show that the APPRM method achieved sub-optimal results. Compared with the benchmark method RippleNet, on the Book-Crossing dataset, the ACC indicator is 1.1% higher than that of RippleNet. On the Bing-News dataset, both indicators of the APPRM method are the highest. The ACC indicator is 0.3% higher than the sub-optimal MKR method and 1.6% higher than the RippleNet method. The AUC indicator is 0.2% higher than the sub-optimal MKR and 1.3% higher than the RippleNet. The method proposed in the present invention is not much different from RippleNet and MKR on the datasets of MovieLens-1M and Book-Crossing, but is higher than both on Bing-News. This is because the APPRM method, on the one hand, focuses on the user's interest in candidate items within the entire interest range, and on the other hand, focuses on the impact of the user's preferred interest on candidate items. This enables the method of the present invention to have better results on knowledge graph datasets with more domains and more complex structures, while having similar effects to conventional methods on datasets with more concentrated domains and more repetitive relationships between entities. Generally speaking, the APPRM method has achieved satisfactory results in the CTR experiment.

Claims

1. A knowledge graph preference prediction and recommendation method based on an attention mechanism, characterized in that The method includes the following steps: First step, symbol definition, the process is as follows: Given a user set U # ={u # 1,u # 2,...u # μ} and an item set V # ={v # 1,v # 2,...v # υ}, where u # μ represents the μ-th user in the user set, and v # υ represents the υ-th item in the item set. The interaction behaviors between users and items are put into an interaction matrix where has only two values, 0 or 1. If user u # has browsing or clicking interaction behavior with item v # , then otherwise it is equal to 0. Each user has a historical interest set representing the θ-th element in the historical interest set of user u # . The elements in the historical interest set represent the items that have interaction behaviors with the user; U # and V # are both parts of the knowledge graph G. The knowledge graph G = {E, R, T}, where E represents the set of entities, R represents the set of relationships, and T represents the set of positive example triples. For the triple (e h , r, e t ), let v # e h , v # r, v # e t represent the embeddings of e h , r, e t in the item vector space respectively. Items are regarded as entities in the knowledge graph, V # ∈ E. The items in the item set V # are regarded as candidate items for user recommendation. If a candidate item has already had an interaction with the user in the interaction matrix, then this item is removed from the candidate items. The candidate item set is defined as: The goal is to predict the probability that the user will like or click on the candidate item v # g ; Second step, design of the preference propagation method; Third step, resource recommendation based on preference propagation, the process is as follows; Step (3.1), in the knowledge graph, a user has multiple historical projects. To adapt to a complex knowledge graph, between the historical project and the candidate project v # g the graph structure is represented as: S0→S1→S2...→S l , → means propagating from the previous node to the next node, where S i represents the i-th node between the historical project and the candidate project, and the entity the resources obtained from the previous ripple set are defined as: Represents a ripple set All tail entities in Of the triples Represents a ripple set All head entities in Of the triples, defining the user's initial resource as R p (all)=1. In the recommendation system, the user's initial resource represents the user's interest in items. Assuming that the user's interest in each historical item is the same, then the interest in each historical item Step (3.2), the user's interest in candidate project v # g is defined as the amount of resources that the user starts from all historical projects and follows the path between and v # g to finally reach v # g . During the process of resource flow, interactions are likely to occur between paths starting from multiple historical projects. There are entities at the intersections of multiple paths, just like the intersections of line segments. At this time, the resource amount of the intersection entity is the sum of the resource amounts of multiple direct predecessors. The entity at the intersection represents that the user is very likely to prefer this related entity. Finally, denote the interest resource amount of candidate project v # g as Resv # g ; Fourth step, graph context representation based on scaled dot-product attention; Step 5: Define the loss function, and use the final calculation result to determine whether to add item v # to the recommendation list of user u # .

2. The method for predicting and recommending knowledge graph preferences based on the attention mechanism according to claim 1, wherein The process of the second step is as follows: Step (2.1): Use the user's historical data as seeds. In the knowledge graph, with the seeds as the center, spread outward along the relationship path p = (r1, r2,... r l ), where r l represents the l-th relationship in the relationship path. Each iteration spreads to the next entity along the direction of the relationship. To describe the multi-level preferences of users using the knowledge graph, for user u # a multi-hop relevant entity set is divided. A k-hop relevant entity set is defined as: Among them represents the user's historical interest set, that is, the set of items that have interacted with the user. H represents the maximum value that k can take. represents the range of values that e t can take in the relevant entity set; Step (2.2), the relevant entities are regarded as the natural extension of the user's interests in the knowledge graph. These natural extensions are manifested as circles that spread layer by layer outward from historical interest items in the knowledge graph, simply referred to as ripple circles. The set of triples composed of relevant entities in one layer of ripple circles is called a ripple set. For user u # 's k-hop ripple set is defined as follows: Step (2.3), during the process of the ripple set spreading outwards, pay attention to the relevant entity sets of each hop. If the existence of v is found in # g , the ripple set will continue to spread outwards until it stops at the H-th hop. If the existence of v is still not found in the # g of the H-th hop, then H = H + 1, and the ripple set will continue to spread outwards until v # g is included in . At this time, k * = H; Step (2.4), after the ripple set stops spreading, organize all the triples in and reconstruct the entities and relationships into a user interest graph g composed of entities and relationships according to the order of appearance in g is a subset of the knowledge graph G and represents the user's interest set.

3. The method for predicting and recommending knowledge graph preferences based on the attention mechanism according to claim 1 or 2, characterized in that, The process of the fourth step is as follows: Step (4.1): Further simplify the user interest graph g, only retaining the paths from historical items to candidate items. During the process of going from multiple historical items to candidate items, each historical item will generate one or more paths. After recording these paths, a set of paths is generated where p # i represents the set of paths between the historical item and the candidate item v # g , p # i = {p i,1 , p i,2 ,... p i,j}, p i,j represents the j-th entity relationship set on the i-th path; Step (4.2), path reflects the user's interest inference process, represents the j * th relationship in the i-th path, represents the j * th entity in this path. In path p i,j , use to construct the entity-relationship context representation of path p i,j . h represents the entity in path p i,j , and r represents the relationship in path p i,j : Step (4.3): To measure the impact of a path on candidate item v # g First, a path pattern needs to be constructed. An LSTM network is used to construct the path pattern. The entity-relationship context in path p i,j is sequentially input into the LSTM network, and the current hidden state of the last LSTM is retained as the context path vector representation of the candidate item. The calculation method of the current hidden state of the LSTM cell is as follows: where f hidden is a non-linear activation function including ReLU operation, (v # e i,l , v # r i,l ) represents the l-th element in the i-th entity-relation context pair; In step (4.4), in the context of the entity graph, each entity-relationship context pair reflects partial features of the tail entity in the triple. At the same time, each entity-relationship context pair contains different information and has different degrees of importance for the tail entity. For each entity-relationship context pair, the scaled dot-product attention mechanism is used to calculate the attention coefficient of the tail entity to the entity-relationship context pair, and the attention coefficients of multiple entity-relationship context pairs in the path are multiplied to obtain the historical item for the candidate item v # g the degree of interest; Step (4.5), historical project For candidate project v # g Along path p i,j The attention degree is expressed as the product of the attention coefficients of multiple consecutive entity-relationship context pairs, and the attention degree also represents the importance degree of path p i,j for recommendation, and the calculation method is as follows: ∏ represents the product symbol, which respectively represent the parameterized representations of the query vector, key vector, and value vector; Step (4.6), the attention coefficient, as the weight of the path, is used to integrate the path information and perform a weighted sum of the paths between the historical item and the candidate item v # g to calculate the full-path pattern The calculation formula is as follows: Among them, represents the upstream path vector representation of path p i,j ; As a full-path pattern, it represents the aggregation of the user's preference interests that spread outwards with the historical project as the center, and is regarded as an approximate embedding of the user in terms of preference interests. Step (4.7), finally, aggregate the user's preference interests and the interest resource amount of candidate item v # g and the embedding information are input into a linear function and a sigmoid function to calculate the probability P(u # ,v # g ) that the user will like the candidate item. The calculation formula is as follows:

4. The method for predicting and recommending knowledge graph preferences based on the attention mechanism according to claim 3, wherein The process of step (4.4) is as follows: Step (4.4.1) The scaled dot-product attention mechanism is described as a mapping from a query vector and a set of key-value pair vectors to an output vector, denoted as Essentially, the output vector is generated by a weighted combination of the input vectors, and the weight values of the input vectors are the result of the scaled dot-product operation between the key vectors and the query vector; Step (4.4.2), the embedding of the tail entity is regarded as the query vector, and the entity relationship context pair is used as the key vector and value vector at the same time. The calculation formula is as follows: Q = W Q v # e t Among them, is the linear transformation weight matrix, respectively represent the parameterized representations of the query vector, key vector, and value vector, d χ is the matrix dimension of the column vector, d δ is the embedding dimension of entities and relationships; Step (4.4.3), the degree of attention of the entity-relationship context pair (e h , r) to the tail entity, that is, the calculation formula of the attention coefficient is as follows: Denotes the transpose of a vector, Denotes the scaling factor used to maintain the stability of the gradient during backpropagation. The softmax function, also known as the normalized exponential function, is a generalization of the binary classification function sigmoid to multi-class classification, aiming to present the results of multi-class classification in the form of probabilities.

5. The method for predicting and recommending knowledge graph preferences based on an attention mechanism according to claim 3, wherein The process of the fifth step is as follows; Minimize the loss function of the method. To optimize the method, a triplet loss function based on self-adversarial negative sampling is used as the training objective of the method: The loss function consists of two parts. One is the loss incurred during the prediction process, and the other is the loss incurred during the embedding process. Among them, γ is a fixed margin, and (e h ', r′, e t ) represents a negative example triple constructed based on (e h , r, e t ). T represents the set of positive example triples, and F represents the set of negative example triples.

Citation Information

Patent Citations

  • Collaborative recommendation model construction method based on knowledge graph preference propagation

    CN113158033A