Knowledge Graph-Based Semantic Retrieval System for Literature and Books

By building a semantic search system for literature and book based on knowledge graphs, the problem that the literature and book search system in the prior art cannot efficiently search content is solved, and efficient semantic information management and accurate query effects are achieved, meeting readers' search needs for high information volume and high semanticity.

CN115563313BActive Publication Date: 2025-07-08SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211307718.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-07-08
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

The existing literature and book search system lacks efficient content search methods and cannot accurately search based on the content expected by readers. In addition, there is a problem of semantic mismatch in searches based on keywords.

Method used

The semantic search system for literature and books based on knowledge graphs is adopted to construct knowledge graphs through naming entity recognition, entity relationship extraction and reference digestion, and combined with natural language processing models, efficient semantic information modeling and querying of literature and books, and use the correlation between books to recommend and assist in search.

Benefits of technology

It realizes finer-grained document and book information management and precise semantic query, providing a high-information and high-semantic search experience, meeting readers' needs for rich semantic query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115563313B_ABST
    Figure CN115563313B_ABST
Patent Text Reader

Abstract

A semantic retrieval system for literature books based on a knowledge graph, comprising: a knowledge graph construction unit and a semantic query unit. The knowledge graph construction unit performs named entity recognition and relationship extraction based on data with semantic information such as the introductions and reviews of literature books, obtains a series of entities and entity-relationship triples, and completes the construction of the knowledge graph; the semantic query unit converts the natural language query statement input by the user into a set of structured query statements, sorts the query results of the knowledge graph of book literature, and returns them to the user. The present invention meets the requirements for an efficient, high-density, and high-information storage method for book knowledge, can efficiently store books and related classification, attribute information, content, etc. of books; can utilize the association information between books to meet the needs of readers for rich semantic queries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of data engineering, specifically a semantic retrieval system for literature books based on a knowledge graph. Background Art

[0002] Although there is an urgent need for literature book retrieval functions both at home and abroad, most current literature book retrieval systems still rely on exact matching of keyword fields, and usually only use the title, author, or index number of the literature book as the keyword fields, lacking an efficient storage and retrieval method for the content of the literature book. For the few literature book retrieval systems that provide content retrieval-related functions, they often rely on manually added keyword tags for the literature books. Limited by the length of the literature book and the number of keywords, it is almost impossible to achieve comprehensive coverage of the content of the literature book; or it is based on whether the keyword appears in the literature book. Since the same keyword may have different meanings according to different context, and different authors, versions, or translators may also result in the same concept appearing in the form of different keywords. Therefore, it is difficult to accurately retrieve relevant literature books according to the content expected by the reader only through keyword retrieval of books. Summary of the Invention

[0003] In view of the above deficiencies in the prior art, the present invention proposes a semantic retrieval system for literature books based on a knowledge graph, which meets the requirements of an efficient, high-density, and high-information storage method for book knowledge, and can efficiently store books, as well as related classification, attribute information, content, etc. of the books; emphasizes the associated information between books, and can use the associated information between books to provide services such as recommendation and assisted search for readers; meets the needs of readers for rich semantic queries, and readers hope to be able to use query statements with high information content and high semantics to retrieve literature books.

[0004] The present invention is implemented through the following technical solutions:

[0005] The present invention relates to a semantic retrieval system for literature books based on a knowledge graph, including: a knowledge graph construction unit and a semantic query unit. The knowledge graph construction unit performs named entity recognition and relationship extraction based on semantic information data such as the introduction and review of the literature book to obtain a series of entities and entity-relationship triples, and completes the construction of the knowledge graph; the semantic query unit converts the natural language query statement input by the user into a set of structured query statements, sorts the query results of the literature book knowledge graph, and returns them to the user.

[0006] The semantic retrieval of literature books mentioned above refers to: collecting knowledge information related to literature books, including titles, authors, tables of contents, introductions, reviews, etc., designing a knowledge graph framework for literature books according to their characteristics, and realizing its automatic construction. At the same time, reasoning based on existing knowledge to explore the relevance between literature books; constructing and training a natural language processing model to identify and extract semantic information such as entities, relationships, and attributes in natural language query statements, performing multi-directional expansions such as synonyms, near-synonyms, hypernyms, and hyponyms, and converting them into structured query statements, and further expanding the query results according to the relevance between books; constructing a sorting algorithm to sort the query results from multiple perspectives such as relevance and the number of queries. At the same time, according to the relevance between literature books, recommend literature books with a relatively high relevance to the existing retrieval results to users.

[0007] Technical effects

[0008] Through the modeling, extraction, management, and query of literature book information at the semantic granularity level, the present invention realizes a finer-grained modeling and management of literature book information compared with the prior art, and provides an efficient and accurate way and means for semantic and unstructured query of literature book data. Brief description of the drawings

[0009] Figure 1 It is a flowchart for the construction of the knowledge graph of literature books;

[0010] Figure 2 It is a flowchart for semantic query;

[0011] Figure 3 It is an example of the knowledge graph of literature books;

[0012] Figure 4 It is an illustration of the implementation scenario. Detailed implementation manners

[0013] This embodiment relates to a semantic retrieval system for literature books based on a knowledge graph, including: a knowledge graph construction unit and a semantic query unit. Among them, the knowledge graph construction unit performs named entity recognition and relationship extraction based on data with semantic information such as the introductions and reviews of literature books, obtains a series of entities and entity-relationship triples, and completes the construction of the knowledge graph; the semantic query unit converts the natural language query statement input by the user into a set of structured query statements, sorts the query results of the knowledge graph of literature books, and returns them to the user. As Figure 1 shown, it is the semantic retrieval process of literature books in this system, including:

[0014] Step 1) Extract semantic information from literature books: Perform knowledge extraction tasks on data with semantic information such as the introductions and reviews of literature books, and convert the semantic information therein into a series of entities and entity-relationship triples to facilitate the efficient storage and query of literature book knowledge. Specifically:

[0015] 1.1) Use named entity recognition technology to identify the named entities in the introductions and reviews of literature books. Specifically: First, manually mark the entities in the introductions and reviews of a small number of literature books. The marked content includes the entity positions and entity types. Then, adopt a training mode of fine-tuning a pre-trained language model combined with the manually marked data to obtain a named entity recognition model. Finally, input a large number of unmarked introductions and reviews of literature books into this model to predict the named entities and their entity types.

[0016] 1.2) Use entity relationship extraction technology to extract the relationships between entities in the introductions and reviews of literature books. Specifically: First, manually mark the relationships between entities in the introductions and reviews of a small number of literature books. The marked content includes the entity pairs with existing relationships, relationship directions, and relationship types. Then, adopt a training mode of fine-tuning a pre-trained language model combined with the manually marked data to obtain an entity relationship extraction model. Finally, input a large number of unmarked introductions and reviews of literature books and the entity positions and entity types therein into this model to predict the relationships, relationship directions, and relationship types between entities.

[0017] 1.3) Use coreference resolution technology to resolve the pronouns identified in step 1.1 and the coreference relationships extracted in step 1.2. Specifically: Determine the pronoun entity and the entity being referred to according to the coreference relationship direction, and replace the pronoun entity in the entity-relationship triple with the entity being referred to. If there are multiple coreferences, all pronoun entities are replaced with the entity being referred to at the beginning of the coreference chain.

[0018] Step 2) Construct a knowledge graph: Import the attribute information and knowledge information of the literature data into the database to complete the literature book knowledge graph as shown in Figure 3. Specifically:

[0019] 2.1) Import the attribute information such as the titles, authors, and types of literature books into the database in the form of tables.

[0020] 2.2) Import the semantic information of the introductions and reviews of literature books obtained in step 1 into the database in the form of a graph. Each named entity and each entity relationship carry an "ownership" attribute, and the attribute value is a list composed of literature book numbers, which is used to mark the subordinate relationship between the literature books and the named entities and entity relationships.

[0021] Step 3) Extract semantic information of natural language query statements: Perform a semantic information extraction task on the natural language query statements input by the user, and convert them into a series of entities and entity-relationship triples, which is convenient for the generation of structured query statements. Specifically:

[0022] 3.1) Use named entity recognition technology. Input the natural language query statement into the named entity recognition model trained in step 1.1 of the literature book knowledge graph construction process to predict the named entities and their entity types in the query statement.

[0023] 3.2) Use entity relationship extraction technology. Input the natural language query statement and the entity positions and entity types therein into the entity relationship extraction model trained in step 1.2 of the literature book knowledge graph construction process to predict the relationships, relationship directions, and relationship types between entities in the query statement.

[0024] 3.3) Use semantic extension technology to further extend the semantics of the natural language query. Through an external entity library, query the synonymous entities, near-synonymous entities, and hyponym entities of the entities obtained in step 1.1, add them to the entity list, and migrate the relationships between the original entities to the relationships between the corresponding synonymous entities, near-synonymous entities, and hyponym entities, and add them to the entity relationship triple list.

[0025] Step 4) Query literature books: According to the semantic information in the natural language query statements input by the user and the type of the database, convert the entities and entity-relationship triples obtained in step 1 into corresponding structured query statements, and further expand the query results returned by the database according to the relevance between literature books. Specifically:

[0026] 4.1) Since the attribute information and semantic information of literature books are stored in the database in the form of tables and graphs respectively, and table data can be stored in various relational and non-relational databases, and graph data can be stored in various graph databases, it is necessary to generate corresponding structured query statements according to the language information in the natural language query statement and the database type. Specifically: First, check whether the entity list obtained in step 1.1 contains literature book attribute keyword entities such as "title" and "author"; if the entity list contains attribute keyword entities, if so, further check whether the attribute keyword entity modifies the literature book to be queried in the entity-relationship triples obtained in step 1.2. If so, generate a corresponding table data query statement according to the database used; for non-attribute keyword entities and entity-relationship triples without attribute keyword entities, generate a corresponding graph data query statement according to the database used.

[0027] 4.2) Use keyword retrieval technology and graph connectivity algorithms to calculate the relevance of attribute information and knowledge information between literature books, further expand the query results of the literature book knowledge graph, and add some literature books with relatively high relevance to the current query results to the query result list.

[0028] Step 5) Sorting of query results: Sort the literature book query results returned in Step 2 according to indicators such as relevance, number of queries, and time of last query to improve the user's semantic query experience for literature books. Specifically: Use the Jaccard similarity algorithm to calculate the relevance between the natural language query statement input by the user and the literature books in the query results. Consider the named entities and entity relationships obtained in Step 1 as Graph A, and calculate the similarity between this graph and Graph B formed by the semantic information of the literature books in the query results respectively. Use the weighted summation method to calculate the importance score P of the query results. i = w j J i + w c C i + w t T i , and perform sorting, where: J is the relevance score, C is the number of queries, T is the time difference between the last query and the current query, and w i , w c , w t are the weights of the three respectively.

[0029] After specific actual experiments, the present invention uses the Bert model as the pre-trained model. Based on 1000 pieces of manually labeled literature book introduction information, the accuracy of the named entity recognition model trained is 0.9143, and the accuracy of the entity relationship extraction model is 0.9583, which can better predict the semantic features in the literature book introduction information. At the same time, in the experiment, the present invention randomly selects 5000 pieces of literature book introductions to construct a knowledge graph, integrates CN-Dbpedia as an external entity library, performs semantic expansion on the generated structured query statements, and selects w i = 0.8, w c = 0.1, w t = 0.1 as the importance weights of relevance, number of queries, and time of last query. Finally, 80 pieces are randomly selected from the literature book introductions used to construct the knowledge graph, and 20 pieces are randomly selected from the literature book introductions not used to construct the knowledge graph as experimental data. Randomly replace the synonyms or hyponyms in the above 100 introductions. The accuracy of the literature book semantic query results for the books already existing in the knowledge graph is 0.9625, and for some non-existent books, it can give recommended results of similar literature books.

[0030] Compared with the prior art, the present invention provides semantic granularity modeling and management for literature book information, realizing a storage structure with high semantic density for literature books. At the same time, the present invention realizes efficient semantic precise query for literature book data, meeting the needs of users for a retrieval mode with high information volume and high semanticity.

[0031] The above specific embodiments can be locally adjusted by those skilled in the art in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific embodiments. All implementation solutions within its scope are subject to the present invention.

Claims

1. A semantic retrieval system for literature and books based on a knowledge graph, characterized in that, Including: A knowledge graph construction unit and a semantic query unit. The knowledge graph construction unit performs named entity recognition and relationship extraction on data with semantic information based on the introductions and reviews of literature books to obtain a series of entities and entity-relationship triples, and completes the construction of the knowledge graph. The semantic query unit converts the natural language query statement input by the user into a set of structured query statements, sorts the query results of the knowledge graph of books and literature, and returns them to the user. The semantic retrieval of literature books mentioned above refers to: collecting relevant knowledge information of literature books, designing a knowledge graph framework for literature books according to their characteristics, and realizing its automatic construction, and reasoning based on the existing knowledge to discover the relevance between literature books. Construct and train a natural language processing model to identify and extract entity, relationship, and attribute semantic information in natural language query statements, perform multi-directional expansion of synonyms, near-synonyms, hypernyms, and hyponyms, and convert them into structured query statements, and further expand the query results according to the relevance between books; construct a sorting algorithm to sort the query results from multiple perspectives of relevance and the number of queries, and recommend literature books with a higher relevance to the existing retrieval results to the user according to the relevance between literature books. The specific query of literature books mentioned above includes: 4.1) Since the attribute information and semantic information of literature books are saved in the database in the form of tables and graphs respectively, and the table data is saved in multiple relational and non-relational databases, and the graph data is saved in multiple graph databases, it is necessary to generate corresponding structured query statements according to the language information in the natural language query statement and the database type. Specifically: first check whether the obtained entity list contains entity keywords of literature book attributes such as "title" and "author"; if the entity list contains entity keyword entities, if so, further check whether in the entity-relationship triples, the entity keyword entity modifies the literature book to be queried, if so, generate corresponding table data query statements according to the used database; for non-entity keyword entities and entity-relationship triples without entity keyword entities, generate corresponding graph data query statements according to the used database. 4.2) Use keyword retrieval technology and graph connectivity algorithms to calculate the relevance of attribute information and knowledge information between literature books, further expand the query results of the knowledge graph of literature books, and add some literature books with a higher relevance to the current query results to the query result list.

2. The semantic retrieval system for literature books based on a knowledge graph according to claim 1, characterized in that The extraction of semantic information from literature books mentioned above refers to: performing a knowledge extraction task on data with semantic information in the introductions and reviews of literature books, and converting the semantic information into a series of entities and entity-relationship triples for efficient storage and query of literature book knowledge.

3. The semantic retrieval system for literature books based on a knowledge graph according to claim 1 or 2, characterized in that, The specific extraction of semantic information from literature books mentioned above includes: 1.1) Use named entity recognition technology to identify named entities in the introductions and reviews of literature books. Specifically: First, manually mark the entities in a small number of literature book introductions and reviews, with the marked content including the entity positions and entity types; then adopt a training mode of fine-tuning a pre-trained language model combined with the manually marked data to obtain a named entity recognition model; finally, input a large number of unmarked literature book introductions and reviews into this model to predict the named entities and their entity types therein. 1.2) Use entity relationship extraction technology to extract the relationships between entities in the introductions and reviews of literature books. Specifically: First, manually mark the relationships between entities in a small number of literature book introductions and reviews, with the marked content including the entity pairs with existing relationships, the relationship directions, and the relationship types; then adopt a training mode of fine-tuning a pre-trained language model combined with the manually marked data to obtain an entity relationship extraction model; finally, input a large number of unmarked literature book introductions and reviews and the entity positions and entity types therein into this model to predict the relationships, relationship directions, and relationship types between entities. 1.3) Use coreference resolution technology to resolve the identified pronouns and the extracted coreference relationships. Specifically: Determine the pronoun entity and the entity being referred to according to the coreference relationship direction, replace the pronoun entity in the entity relationship triple with the entity being referred to. If there are multiple coreferences, all pronoun entities are replaced with the entity being referred to at the beginning of the coreference chain.

4. The semantic retrieval system for literature books based on a knowledge graph according to claim 1, characterized in that, The construction of the knowledge graph mentioned above refers to: Import the attribute information and knowledge information of the literature data into the database to complete the knowledge graph of literature books.

5. The semantic retrieval system for literature books based on a knowledge graph according to claim 1 or 4, characterized in that, The construction of the knowledge graph mentioned above includes: 2.1) Import the title, author, and type attribute information of the literature books into the database in the form of a table. 2.2) Import the semantic information of the literature book introductions and reviews into the database in the form of a graph; each named entity and each entity relationship carry an "ownership" attribute, and the attribute value is a list composed of literature book numbers, used to mark the subordination relationship between the literature books and the named entities and entity relationships.

6. The semantic retrieval system for literature books based on a knowledge graph according to claim 1, characterized in that, The extraction of the semantic information of natural language query statements mentioned above refers to: Perform a semantic information extraction task on the natural language query statements input by the user, and convert them into a series of entities and entity relationship triples to facilitate the generation of structured query statements.

7. The semantic retrieval system for literature books based on a knowledge graph according to claim 1 or 6, characterized in that, The extraction of the semantic information of natural language query statements specifically includes: 3.1) Use named entity recognition technology to input the natural language query statements into the named entity recognition model trained in the process of constructing the literature book knowledge graph to predict the named entities and their entity types in the query statements. 3.2) Use entity relationship extraction technology to input the natural language query statements and the entity positions and entity types therein into the entity relationship extraction model trained in the process of constructing the literature book knowledge graph to predict the relationships, relationship directions, and relationship types between entities in the query statements. 3.3) Use semantic extension technology to further extend the semantics of natural language queries; query for synonymous entities, near-synonymous entities, and hyponymous entities of the obtained entities through an external entity library, add them to the entity list, and migrate the relationships between the original entities to the relationships between the corresponding synonymous entities, near-synonymous entities, and hyponymous entities, and add them to the entity relationship triple list.

8. The semantic retrieval system for literature books based on a knowledge graph according to claim 1, characterized in that The query of literature books mentioned above means: according to the semantic information in the natural language query statement input by the user and the type of the database, convert the entities and entity relationship triples into corresponding structured query statements, and further extend the query results returned by the database according to the relevance between literature books.

9. The semantic retrieval system for literature books based on a knowledge graph according to claim 1, wherein Sorting the query results mentioned above means: sorting the query results of literature books according to the relevance, the number of queries, and the most recent query time indicators to improve the user's semantic query experience for literature books. Specifically, it includes: using the Jaccard similarity algorithm to calculate the relevance between the natural language query statement input by the user and the literature books in the query results; regarding the named entities and entity relationships obtained in step 1 as graph A, and calculating the similarity between this graph and graph B composed of the semantic information of the literature books in the query results , using the weighted summation method to calculate the importance score of the query results , and performing sorting, where: is the relevance score, is the number of queries, is the time difference between the last query and the current query, are the weights of the three respectively.