A knowledge graph construction method for knowledge-aware recommendation
By linking open domain knowledge graphs and recommendation data sets, a knowledge graph with enhanced connection is solved, and the problem of lack of rich information in the knowledge-aware recommendation system is improved, and the recommendation accuracy and interpretability of the recommendation system are improved.
Patent Information
- Application Number
- CN202210150860.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-02-18
AI Technical Summary
In the prior art, the knowledge graph used by the knowledge-aware recommendation system lacks rich project-related information, and the entity's triplet density is low, the relationship type is small, and the quality is poor, resulting in the recommendation accuracy and interpretability of the recommendation system being limited.
By linking the open domain knowledge graph and the recommended data set, the mapping between projects and entities is realized, and triples are iteratively extracted from the open domain knowledge graph according to the mapping relationship. After entity screening, relationship screening and other processing, a connection-enhanced knowledge graph for knowledge-aware recommendations is constructed.
It solves the problem of insufficient knowledge and poor relationship quality in the knowledge graph, improves the recommendation accuracy and interpretability of the recommendation system, and enables the recommendation system to make more efficient use of the knowledge in large open source knowledge graphs.
Smart Images

Figure CN114925207B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method for constructing a knowledge graph for knowledge-aware recommendation. Background Art
[0002] In recent decades, with the rapid development of the Internet, the amount of data on the Internet has increased exponentially. The excessive amount of information has caused information overload, and it is difficult for users to easily select the content they are interested in from the massive amount of data. In order to solve the problem of information overload and improve user experience, recommendation systems have been applied to various online application scenarios, such as music recommendations, movie recommendations, news recommendations, online shopping, etc. It can be said that almost all Internet services that provide content involve the application of recommendation systems.
[0003] The recommendation algorithm is the core part of the recommendation system. According to the different principles of the recommendation algorithm, the recommendation system is mainly divided into three types: collaborative filtering-based recommendation system, content-based recommendation system and hybrid recommendation system.
[0004] 1) Recommendation system based on collaborative filtering: mainly uses the similarity of users or items in interaction data to model user preferences.
[0005] 2) Content-based recommendation system: The features of the items are extracted through content analysis to calculate similarity.
[0006] Among them, the recommendation system based on collaborative filtering is widely used because it does not require manual feature extraction.
[0007] Although collaborative filtering algorithms have achieved many results in practical applications and academic research, they still face some challenges, such as data sparsity and cold start problems. In order to solve these problems, a hybrid recommendation system was proposed, which introduces "side information" on the basis of collaborative filtering, thereby simultaneously utilizing the similarity of the interaction layer and the content layer, combining the advantages of the above two recommendation systems. In recent years, hybrid recommendation systems have explored various types of side information, such as project attributes, project reviews, and users' social networks.
[0008] With the widespread application of knowledge graphs in information retrieval, knowledge question answering, artificial intelligence and other fields, the method of introducing knowledge graphs as side information into recommendation systems has gradually become a research hotspot. This method can not only alleviate many of the above problems and improve the accuracy of the recommendation system, but also provide recommendation explanations for the recommended items, realize explainable recommendation services, and enhance user recognition of recommendation results. This type of hybrid recommendation system that uses information in the knowledge graph to model user preferences is called a knowledge-aware recommendation system.
[0009] Knowledge-aware recommendation systems need to deeply mine and use project-related information in knowledge graphs to model user preferences. The quality of the knowledge graph itself largely determines the recommendation performance of the recommendation system. Therefore, in the implementation of knowledge-aware recommendation systems, a key issue is the data problem, that is, how to obtain rich and structured project-related knowledge information to build a high-quality knowledge graph. Depending on the source of data, the knowledge graphs used by existing knowledge-aware recommendation systems are mainly constructed in three ways: one is to use the side information in the original recommendation data set (usually only contains a small amount of useful information); the second is to use non-open source private knowledge bases, such as Microsoft's Satori; the third is to use the non-public project data of the recommendation service platform. After extensive research, a general knowledge graph construction method for knowledge-aware recommendation has not yet been found.
[0010] Technical solution of prior art 1
[0011] The prior art related to the present invention is the knowledge graph. The knowledge graph is a heterogeneous information network in which nodes can represent entities and edges can represent the relationship between entities. In the knowledge graph used for recommendation systems, items and their related attributes are usually regarded as entities, and the mutual relationship and high-level semantic relationship between different items can be revealed through the connectivity of attribute relationships. In addition, some researchers have also integrated users and related information of users into the knowledge graph. This knowledge graph can directly explore the relationship between users and items to capture users' preferences, which is called collaborative knowledge graph.
[0012] The commonly used knowledge graphs in recommendation systems are defined as follows:
[0013] G={(h,r,t)|h,t∈E,r∈R}
[0014] Among them, E and R represent the entity set and relationship set in the knowledge graph respectively, and the triple (h, r, t) indicates that there is a relationship r between the head entity h and the tail entity t. For example, in the field of movie recommendation, the triple (Leonardo DiCaprio, ActorOf, The Great Gatsby) represents a fact: Leonardo is an actor in the movie "The Great Gatsby". Obviously, this information can play a decisive role in recommending this movie to Leonardo's fans.
[0015] The definition of collaborative knowledge graph is similar to the above definition, except that user entities and user-project interactions are added to the original knowledge graph, thereby integrating user behavior and project knowledge into the same knowledge graph. Its definition is as follows:
[0016] G={(h,r,t)|h,t∈E',r∈R'},E'=E∪U,R'=R∪{Interact}
[0017] Among them, U represents the user entity set, E represents other types of entity sets (items and item attributes); R represents the item relationship, and Interact represents the interaction relationship between user-item entities. Figure 1 The figure shows an example of a collaborative knowledge graph in the field of movie recommendation. The watched relationship on the far left is the user-item interaction relationship. Through the user social information and item attribute information in the knowledge graph, the recommendation system can recommend the two movies on the right side of the figure to the target user Bob.
[0018] Disadvantages of the prior art 1
[0019] At present, the knowledge graphs used by some recommendation systems do not contain enough knowledge. The average triple density of entities is low, the relationship types are few, and the relationship quality is poor, such as the existence of bidirectional synonymous relationships, irrelevant or even semantically meaningless relationships. Obviously, such relationships and their related triples are invalid for the recommendation decision-making process of the recommendation system. A large amount of invalid information introduces noise into the knowledge graph, which will seriously interfere with the recommendation system's modeling of user preferences and reduce the accuracy of the recommendation system.
[0020] Prior art related to the present invention
[0021] Technical solution of prior art 2
[0022] Existing knowledge graph construction methods for recommendation systems can be mainly divided into two categories:
[0023] 1) Use the project-related data of the original recommendation service platform or the recommendation data set to collect side information, build a small structured knowledge base for the recommendation scenario, and finally build a knowledge graph based on the experimental data set.
[0024] 2) Establish a link between the recommendation dataset and the private knowledge base, and iteratively extract triples to build a knowledge graph.
[0025] The first type of method belongs to the general method of building knowledge graphs, and is not only for recommendation systems. The main workflow of this method includes: information extraction, knowledge extraction, knowledge fusion, knowledge processing, etc. Among them, information extraction is a technology that extracts structured information such as entities, relationships, and entity attributes from semi-structured or unstructured data, and it is also the first step in building a knowledge graph. Knowledge extraction is the conversion of extracted information into the form of structured knowledge, such as RDF data. Knowledge fusion includes entity linking and knowledge merging. It integrates the entities, relationships, and entity attributes obtained in the above steps into the same knowledge base, eliminates ambiguity, and obtains a series of basic facts, thereby achieving a complete description of the entity. Knowledge processing refers to further processing of facts on the basis of knowledge fusion to make them structured and networked knowledge. This part of the work includes: ontology construction, knowledge reasoning, and quality assessment.
[0026] The second method builds a knowledge graph based on an existing knowledge base, which is essentially a process of knowledge extraction. Therefore, it is simpler, faster, and less labor-intensive than the first method. Taking the classic knowledge-aware recommendation model RippleNet as an example, it uses Microsoft's Satori knowledge base to build knowledge graphs in three different recommendation fields. The specific process is as follows:
[0027] 1) For different recommendation fields, extract triples whose relation names contain keywords in specific fields from the entire knowledge base. For example, for the movie dataset MovieLens-1M, extract relation names containing "movie". Filter triplets with a confidence greater than 0.9 to form a triple subset.
[0028] 2) Collect all valid item IDs in the current triple subset by matching the item name with the tail entity of a specific triple. For example, the triples of the movie dataset are in the form of (head, film.film.name, tail), while the triples of the book dataset are (head, book.book.title, tail). For simplicity, items that are not matched or have multiple matching entities are directly discarded in this process.
[0029] 3) Use the valid project ID obtained in the above steps to match the head entity and tail entity of all triples in the knowledge base, extract well-matched triples, form a new subgraph, and repeatedly expand the entity set until four hops.
[0030] Disadvantages of the second prior art
[0031] The first type of method requires information extraction from a large amount of unstructured data and constructs a knowledge graph through knowledge fusion and processing. This method is labor-intensive and only applies to the current recommended data set, and is not universal. The second type of method obtains structured knowledge from an existing knowledge base, which is convenient and fast, but is heavily dependent on the original knowledge base. In addition, the private knowledge base used by this type of method is usually only used within the company and is not accessible to ordinary users. Summary of the invention
[0032] Based on the above research shortcomings, the present invention aims to use an open source knowledge base link dataset KB4Rec to link three widely used recommendation datasets in different recommendation fields with the open domain knowledge graph, realize the mapping between items and entities, and iteratively extract triples from the open domain knowledge graph according to the mapping relationship. After entity screening, relationship screening and other processes, a connection-enhanced knowledge graph for knowledge-aware recommendation is constructed.
[0033] The data sources of the present invention mainly include: recommendation data sets and open domain knowledge graphs. Among them, the recommendation data set is a data set obtained from a real Internet recommendation service platform for offline test experiments of the recommendation system, which usually mainly includes: user interaction data, basic attribute information of the project, and a small amount of user personal information. Among them, user interaction data is the main body of the recommendation data set, and is generally stored in the form of a triple of (user ID, project ID, interaction data value). Depending on the source of the data, the user interaction data value is also different. For example, in a movie recommendation data set, the interaction data is usually an explicit rating given by the user; in an e-commerce recommendation data set, in addition to the user's explicit rating, the interaction data may also include the user's click and browse operations on a certain product (also called implicit rating); and in a music recommendation data set, the user interaction data can be the total number of times the user listens to a song. The open domain knowledge graph is a large-scale knowledge graph that covers knowledge in multiple fields at the same time, also known as a knowledge base. Common open domain knowledge graphs are generally constructed from data from encyclopedia websites, including: Freebase, DBpedia, YAGO, etc. Since the open domain knowledge graph integrates structured knowledge from multiple fields, its practical application areas are very broad, including information retrieval, intelligent question and answer, intelligent recommendation, etc.
[0034] To achieve the purpose of the invention, the technical solution provided by the present invention is: a method for constructing a knowledge graph for knowledge-aware recommendation, which quickly constructs a knowledge graph for knowledge-aware recommendation by linking an open domain knowledge graph with a recommendation dataset, including the following steps:
[0035] Step 1) Obtain user interaction data through recommendation datasets; for general recommendations in multiple fields, the recommendation datasets are commonly used datasets in three different recommendation fields: movies, music, and e-commerce. Obtain real user interaction data through open source recommendation datasets in three popular recommendation fields: the dataset in the movie field is the Movielens-20M dataset, the dataset in the music field is the Last.fm-1b dataset, and the dataset in the e-commerce field is the Amazon-Book dataset.
[0036] Step 2) Sample the original recommendation dataset in step 1) to obtain a sub-dataset, and then go through the following three steps to obtain the binary scoring dataset for constructing the knowledge graph:
[0037] Step 2.1), K-core extraction: only retain users and projects with interaction records greater than K;
[0038] Step 2.2), interaction density control: make the interaction density within the range of human control;
[0039] Step 2.3), score binarization: The original data is binarized by artificially setting a score threshold. If the original score is greater than or equal to the threshold, the binary score is 1, otherwise it is 0;
[0040] Step 3), extracting the item set from the binary rating data set;
[0041] Step 4), using the open source knowledge base linking dataset KB4Rec, link the items in the recommendation dataset with the entities in the open domain knowledge graph to achieve mapping between items and entities;
[0042] Step 5), according to the mapping relationship obtained in step 4), triples are iteratively extracted from the open domain knowledge graph Freebase with the project entity as the initial seed set. This is an iterative triple extraction algorithm;
[0043] In step 6), after manually screening effective relationships and filtering low-frequency entities, a connection-enhanced knowledge graph for knowledge-aware recommendation is constructed.
[0044] The further preferred technical solution proposed by the present invention is:
[0045] In the step 5), the steps of extracting triples by the iterative triple extraction algorithm are:
[0046] Step 5.1), extracting item sets from the binary rating dataset;
[0047] Step 5.2), link the items in the item set with the entities in the general knowledge graph;
[0048] Step 5.3), according to the item-entity link relationship, use the item set as the seed set to extract triples from the open domain knowledge graph to form a subgraph;
[0049] Step 5.4), extract the entity set in the subgraph as a new seed set, and continue to extract triples from the open domain knowledge graph to expand the subgraph;
[0050] Step 5.5), filter the triplets related to low-frequency entities according to the K-core principle to form a knowledge graph with enhanced connections.
[0051] The beneficial effects of the present invention are:
[0052] The present invention uses the open source knowledge base link data set KB4Rec to link the projects in the recommendation data set with the entities in the open domain knowledge graph, realize the mapping between projects and entities, and iteratively extract triples from the open domain knowledge graph according to the mapping relationship. After entity screening, relationship screening and other processes, a connection-enhanced knowledge graph for knowledge-aware recommendation is constructed. This part of the work faces the background of lack of rich project-related information in the recommendation data set, and the unstructured data on the recommendation service platform or other Internet applications is not easy to extract knowledge. It solves the problem that the general knowledge graph construction method has a large workload and cannot be obtained from private knowledge bases, so that the recommendation system can use the knowledge in large open source knowledge graphs to improve the recommendation effect.
[0053] At the same time, the present invention solves the problems of poor quality of knowledge graph relationships, insufficient triples, low interaction density of data, etc. by manually screening relationship types, filtering low-frequency entities, and controlling interaction density, thereby ensuring the quality of project attribute relationships and user interaction information in the knowledge graph.
[0054] In addition, the present invention also realizes the mapping between entity and name, and provides a data basis for the generation of text explanations in explainable recommendation services.
[0055] Definitions of abbreviations and key terms used in this document:
[0056] Knowledge Graph: Essentially, a knowledge graph is a semantic network diagram used to describe various entities or concepts and their relationships in the real world. The nodes in the graph represent entities or concepts, and the edges represent attributes or relationships. A triple consisting of "two points and one edge" can be used as a representation of the relationship between entities, such as (head entity, relationship, tail entity), or (entity, attribute, attribute value).
[0057] Entity: refers to something that is distinguishable and exists independently. It is the most basic element in the knowledge graph.
[0058] Recommendation system: It is an information filtering system that uses recommendation algorithms to model user preferences and provide users with products and information that they may be interested in to help them make decisions.
[0059] Item: A product that is provided to users as a recommendation during the recommendation service.
[0060] Feature: A way of representing data. Item feature refers to representing items in the recommendation system as one or more feature vectors. Feature vectors are features. Features can represent some independent and distinguishable information of items, such as the director and actors of a movie project.
[0061] Interaction data: data generated by the interaction between users and items on the recommendation service platform, such as ratings and click operations of e-commerce users, and the number of times users on the music playback platform listen to a song.
[0062] Explainable recommendation: refers to the recommendation system that not only provides recommendation results to users or researchers, but also provides corresponding recommendation explanations to explain why the item is recommended. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a schematic diagram of the prior art 1 in the background technology;
[0064] Figure 2 This is a flowchart of a knowledge graph construction method for knowledge-aware recommendation according to the present invention. DETAILED DESCRIPTION
[0065] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings.
[0066] Figure 2 Flow chart of a knowledge graph construction method for knowledge-aware recommendation of the present invention. Figure 2 As shown, the present invention is a knowledge graph construction method for knowledge-aware recommendation, which constructs knowledge graphs for recommendation data sets in three different fields and introduces them into the knowledge-aware recommendation system as side information to improve the accuracy and interpretability of the recommendation system. The specific steps are as follows:
[0067] Step 1) Obtain user interaction data through recommendation datasets. The present invention selects three commonly used datasets in different recommendation fields, including movies, music, and e-commerce (this embodiment selects the book field): Movielens-20M, Last.fm-1b, and Amazon-Book datasets. The statistical results of the original interaction data of these three datasets are shown in the following table:
[0068]
[0069] Step 2), since the amount of data in the original data set is too large, the present invention obtains a sub-data set by sampling the original data set in step 1), and then obtains a binary scoring data set for constructing a knowledge graph through the following three steps:
[0070] Step 2.1), K-core extraction: only retain users and projects whose interaction records are greater than K.
[0071] Step 2.2), interaction density control: interaction density refers to the average number of interaction items per user in a data set. The present invention makes the interaction density within a suitable range of human control by sampling interaction data within a certain time span and discarding low-frequency users.
[0072] Step 2.3), binary rating: Since the original interaction data is an explicit rating, but most recommendation models based on deep learning technology need to use binary data samples for training, it is necessary to convert the original rating data into binary data, that is, 0 or 1 (0 means that the user has not observed this item; 1 means that the user has a positive interaction with the item). Specifically, the present invention binaryizes the original data by artificially setting a rating threshold. If the original rating is greater than or equal to the threshold, the binary rating is 1, otherwise it is 0.
[0073] Step 3), extracting the item set from the binary rating data set;
[0074] Step 4), using the open source knowledge base linking dataset KB4Rec, link the items in the recommendation dataset with the entities in the open domain knowledge graph to achieve mapping between items and entities;
[0075] Step 5), according to the mapping relationship obtained in step 4), triples are iteratively extracted from the open domain knowledge graph Freebase using the project entity as the initial seed set;
[0076] The steps to extract triples are:
[0077] Step 5.1), extracting item sets from the binary rating dataset;
[0078] Step 5.2), link the items in the item set with the entities in the general knowledge graph;
[0079] Step 5.3), according to the item-entity link relationship, use the item set as the seed set to extract triples from the open domain knowledge graph to form a subgraph;
[0080] Step 5.4), extract the entity set in the subgraph as a new seed set, and continue to extract triples from the open domain knowledge graph to expand the subgraph;
[0081] Step 5.5), filter the triplets related to low-frequency entities according to the K-core principle to form a knowledge graph with enhanced connections.
[0082] Step 6), after manual screening of effective relationships and low-frequency entity filtering steps, the knowledge graph finally constructed has clear relationship types and high triple density, contains rich and accurate project-related knowledge, and provides strong data support for the knowledge-aware recommendation system.
[0083] The following table shows the statistical information of the knowledge graph constructed by the present invention for three recommendation datasets:
[0084]
[0085]
[0086] Through the above description, the present invention generally has the following characteristics:
[0087] 1) Project-entity linking: Based on the KB4Rec dataset, the present invention obtains the project-entity linking relationship between the recommended dataset and the open domain knowledge graph Freebase, thereby constructing a connection-enhanced knowledge graph, providing a way to integrate user interaction information with project attribute information.
[0088] 2) Triple extraction algorithm: The present invention maps projects into entities in an open domain knowledge graph based on the project-entity link relationship, thereby iteratively extracting relevant triples in the open domain knowledge graph. This paper proposes a triple extraction algorithm for quickly constructing a knowledge graph for a recommendation system from existing structured knowledge.
[0089] 3) Connection-enhanced knowledge graph: Different from other methods of automatically constructing knowledge graphs, the knowledge graph construction method for knowledge-aware recommendation proposed in the present invention, in order to ensure the effectiveness of the knowledge graph, maximizes the richness and accuracy of the knowledge in the knowledge graph through means such as entity and project filtering and manual screening, so as to facilitate the recommendation system to use the effective information in the knowledge graph to model the connection between users and projects, thereby improving the performance of the recommendation system.
[0090] The described embodiments are only a part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
Claims
1. A knowledge graph construction method for knowledge-aware recommendation, characterized in that The following steps are involved: Step 1), obtain user interaction data through the recommended dataset; In the step 1), the recommendation data set is a commonly used data set including three different recommendation fields: movies, music, and e-commerce; Step 2) Sample the original recommendation dataset in step 1) to obtain a sub-dataset, and then go through the following three steps to obtain the binary scoring dataset for building the knowledge graph: Step 2.1), K-core extraction: only retain users and projects with interaction records greater than K; Step 2.2), interaction density control: make the interaction density within the range of human control; Step 2.3), score binaryization: the original data is binaryized by artificially setting the score threshold. If the original score is greater than or equal to the threshold, the binary score is 1, otherwise it is 0; Step 3), extracting the item set from the binary rating data set; Step 4) Use the open source knowledge base linking dataset KB4Rec to link the items in the recommendation dataset with the entities in the open domain knowledge graph to achieve mapping between items and entities. Step 5), according to the mapping relationship obtained in step 4), triples are iteratively extracted from the open domain knowledge graph Freebase using the project entity as the initial seed set; In step 6, after manually screening effective relationships and filtering low-frequency entities, a connection-enhanced knowledge graph for knowledge-aware recommendation is constructed.
2. A knowledge graph construction method for knowledge-aware recommendation according to claim 1, characterized in that: In the step 1), the dataset in the film field is the Movielens-20M dataset, the dataset in the music field is the Last.fm-1b dataset, and the dataset in the e-commerce field is the Amazon-Book dataset.
3. The method for constructing a knowledge graph for knowledge-aware recommendation according to claim 1, characterized in that: In the step 5), the step of extracting triples is: Step 5.1), extracting item sets from the binary rating dataset; Step 5.2), link the items in the item set with the entities in the general knowledge graph; Step 5.3), according to the item-entity link relationship, use the item set as the seed set to extract triples from the open domain knowledge graph to form a subgraph; Step 5.4), extract the entity set in the subgraph as a new seed set, and continue to extract triples from the open domain knowledge graph to expand the subgraph; Step 5.5), according to the K-core principle, filter the triplets related to low-frequency entities to form a knowledge graph with enhanced connections.
Citation Information
Patent Citations
Collaborative recommendation method based on knowledge graph representation learning and neural network
CN111582509A
Music recommendation method and system based on knowledge graph
CN113032618A