An Academic Knowledge Q&A Method, System, Device and Storage Medium Based on Semantic Components

By constructing knowledge graphs in the academic field and using semantic components to filter the knowledge graph sub-graphs of query sentences, the problem of inability to understand the intent of query in the existing technology is solved, and the rapid and accurate effect of academic information retrieval is achieved.

CN115344714BActive Publication Date: 2025-05-27HANGZHOU DIANZI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211018126.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2025-05-27
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

The existing academic information retrieval system cannot effectively understand the query intention, resulting in low correlation between the query results and the topic and cannot meet the needs of fast and accurate information acquisition in academic research.

Method used

Using a semantic component-based method, by constructing a knowledge graph, synonym mapping library and attribute knowledge graph in the academic field, combining intent and constraint semantic components, the knowledge graph sub-graph corresponding to the query sentence is filtered out, thereby achieving the correct judgment of the academic query intention.

Benefits of technology

It realizes an accurate understanding of academic query intentions, improves the relevance of query results, and allows academic researchers to quickly and accurately obtain the required information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344714B_ABST
    Figure CN115344714B_ABST
Patent Text Reader

Abstract

The present invention discloses an academic knowledge Q&A method based on semantic components, which comprises the following steps: S1, constructing an academic domain knowledge base; S2, constructing a knowledge graph subgraph for academic queries based on intent semantic components; S3, correcting the knowledge graph subgraph based on constraint semantic components; S4, generating answers. This method uses the academic domain information publicly available on the Internet as the initial data source, establishes a knowledge graph, a synonym mapping library, and an attribute knowledge graph within the domain; combines semantic components, filters out the knowledge graph subgraph corresponding to the academic query question sentence, realizes the correct determination of the intent in the academic query, and can be used in the basic academic information query scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data processing, and specifically refers to an academic knowledge Q&A method, system, device and storage medium based on semantic components. Background Art

[0002] In the era of artificial intelligence and big data, by using new-generation technologies such as big data and knowledge graphs to build a highly automated and human-computer interactive intelligent Q&A system, analyzing the huge and complex-structured data in the database and converting it into accurate and concise answers to answer natural language questions can meet people's needs for quickly and accurately obtaining information. The academic field knowledge Q&A system aims to store and process academic field data based on the knowledge graph, so as to answer the questions of academic researchers and help them obtain the information they need.

[0003] At present, the academic information retrieval method mainly relies on keyword matching, from searching in traditional databases to existing Q&A retrieval systems. Although the screening of data and the recognition query in natural language mode are realized, due to the inability to truly understand the query intention, the query results often have a low relevance to the searched topic and cannot meet the needs of quickly and accurately obtaining information in academic research. Therefore, how to correctly determine the intention in academic queries is an urgent problem to be solved in the current work of academic field Q&A systems. Summary of the Invention

[0004] To solve the above problems, the present invention provides an academic knowledge Q&A method and system based on semantic components. This method uses the academic field information publicly available on the Internet as the initial data source to establish a knowledge graph, a synonym mapping library and an attribute knowledge graph within the field; combined with semantic components, it filters out the knowledge graph subgraph corresponding to the academic query question sentence, realizes the correct determination of the intention in academic queries, and can be used in basic academic information query scenarios.

[0005] The present invention relates to an academic knowledge Q&A method and system based on semantic components. Among them, the knowledge graph adopts the data model of the Resource Description Framework (RDF), and uses the N-Triples data format to describe the relationship between academic field information and information, that is, by defining the entities in the knowledge graph for the academic field, the domains to which the entities belong, the topic types under the domains, and the relationships included in the topic types, a hierarchical data division is formed to realize the management of data.

[0006] The domains involved in the present invention are defined as "author" domain, "institution" domain, "field" domain, "journal" domain and "literature" domain. Among them, the "institution" domain is further divided into various types such as universities, research institutions and companies. Under the university type, there are various relationships such as name and research direction. Under the "literature" domain, there are various types such as academic paper types, master's and doctoral theses. The knowledge base involved in the present invention includes a domain knowledge graph, a synonym mapping library and an attribute knowledge graph. Among them, the domain knowledge graph is used to store triples of the relationships between entities in the five academic information domains. The synonym mapping library is a mapping table between standard words and alternative words in the domain. The attribute knowledge graph is used to store the attributes and attribute values of entities in the five academic information domains and represent them as constraint information.

[0007] The present invention relates to a subgraph construction technology based on semantic components. First, the query question proposed by the user is decomposed into semantic components; the pre-constructed knowledge graph is filtered by using the semantic components to obtain the knowledge graph subgraph corresponding to the query question, so as to realize the correct determination of the intention in the academic query question.

[0008] The technical solution adopted by the present invention is characterized in that: first, the academic information publicly available on the Internet by academic researchers is used to construct a domain knowledge graph, a synonym mapping library and an attribute knowledge graph in the domain. Then, through the synonym mapping library, the words in the query question are replaced with standard words, and then it is preprocessed and entity recognition is performed. Among them, the preprocessing is parsed into a constraint semantic component and an intention semantic component. The entity recognition obtains the entities belonging to the academic field in the query question, so as to obtain the relationship set connected to the entities; further, the similarity between the relationship and the query question is calculated by using the intention semantic component. If the similarity ranking is less than a predefined threshold, the relationship is regarded as relevant to the query intention and the relationship is retained as a path, otherwise the relationship is discarded. After iteratively expanding the path, a subgraph of the academic query question is obtained; the constraint semantic component is used to add constraint information to the subgraph from the attribute knowledge graph. Finally, the obtained subgraph is used as the academic information result of the query.

[0009] The present invention also provides an academic knowledge question-answering system based on semantic components, which is independently stored on various hardware such as terminals, database servers, application servers, etc., and the computing software program of the academic field knowledge question-answering system based on semantic components is called to execute. Those of ordinary skill in the art can combine the units and algorithm steps of each example described in the embodiments disclosed herein and can be implemented with electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0010] The technical solution of the present invention is as follows:

[0011] An academic knowledge Q&A method based on semantic components, comprising the following steps:

[0012] S1. Construct an academic domain knowledge base

[0013] S11. Obtain academic domain knowledge data, using the academic information publicly available on the Internet by academic researchers as the data source, and the publicly available data is crawled through a crawler tool;

[0014] S12. Define an academic information domain for the obtained academic domain knowledge data. The academic information domain includes "author" domain, "institution" domain, "field" domain, "journal" domain and "literature" domain, and apply hierarchical data management to form a relational semantic network of entities within the academic domain according to the academic information domain;

[0015] S13. Construct an academic domain knowledge graph based on the Resource Description Framework data model, perform annotation processing on the crawled data, and store it in the form of <entity, relationship, entity> triples;

[0016] S14. Construct an academic domain synonym mapping library;

[0017] S15. Construct an academic domain attribute knowledge graph;

[0018] S2. Construct a knowledge graph subgraph for academic queries based on intention semantic components

[0019] S3. Modify the knowledge graph subgraph based on constraint semantic components;

[0020] S4. Generate answers.

[0021] Preferably, the step S2 includes the following sub-steps:

[0022] S21. Preprocess the proposed academic query question, including using the synonym mapping library to replace the standard words in the question q, obtaining its syntactic parse graph G gram and semantic component g i , where the semantic component is the path where each leaf node is located;

[0023] S22. Perform academic named entity recognition on the question q

[0024] Perform entity sample annotation on the defined five academic information domains, and then realize named entity recognition within the domain. The entity recognition result is an entity set E=(e 1 ,e 2 ,e 3 ,…,e n ), n is the number of entities, ei is an entity;

[0025] S23. Encode the preprocessed question sentence and semantic components into feature vectors respectively;

[0026] S24. According to the pre-constructed academic domain knowledge graph, find the relationship set R associated with each entity e in the entity set E i The search result is (r 1 , r 2 , …, r n ), where each relationship r i and the entity e i both have a triple Then encode the relationship r in the relationship set R i into a feature vector, denoted as where d represents the embedding dimension;

[0027] S25. Filter the relationships and then construct a knowledge graph subgraph of the academic question sentence;

[0028] S26. Set a threshold k for the path, select the k relationships with the largest similarity, that is, (s 1 , s 2 , …, s k ), n 1 > k, retain the corresponding relationships (r 1 , r 2 , …, r k ), then perform path extension for the next hop, use the tail entity connected by the relationships (r 1 , r 2 , …, r k ) as the starting entity set for iterative extension, and repeat steps S25 and S26 in the process, where the relationships are replaced by the relationships r in the next hop relationship set j , and iterate and calculate n 1 times, which is the knowledge graph subgraph KG 1 of n sub .

[0029] Preferably, the step S3 includes the following sub-steps:

[0030] S31. Take out the entities in the knowledge graph subgraph as the entity set E 1 , and according to the pre-constructed attribute knowledge graph, find the attribute set A associated with each entity e in the entity set E 1 The search result is (a i , a 1 , …, a 2 ), where each relationship a f and the entity e i and the entity e iAll have triples where v is the attribute value, and then for the attribute a in the attribute set A i is encoded as a feature vector, denoted as where d represents the embedding dimension

[0031] S32. Filter attributes to correct the knowledge graph subgraph of the constructed academic query question

[0032] The method for correcting the knowledge graph subgraph is as follows

[0033] The filtering of attributes depends on the academic query question, the constraint semantic component, and the attribute a i , and calculate the semantic similarity between the attribute a i and the query question. The calculation formula is as follows

[0034]

[0035] Then sort the similarities of the attributes to obtain the sequence (S 1 , S 2 , …, S n ), where S j > S j+1 , where n 2 is the number of constraint semantic components

[0036] S33. Set the threshold S k . If the similarities of the attributes associated with the entity e i are all lower than the threshold, that is then consider this entity as irrelevant to the query and discard it. Otherwise, retain the entity node

[0037] Preferably, in the step S21, the semantic components include constraint semantic components and intention semantic components. The constraint semantic components represent the time constraints and numerical constraints in the query question, and the intention semantic components are the other semantic components in the query question, representing semantic information. The constraint semantic components are extracted by using the method of regular expression matching

[0038] Preferably, in the step S23, the encoding method of the feature vector is as follows

[0039] First, encode the query question q to obtain the complete semantic representation of each word where d represents the word embedding dimension and n represents the number of words in the query question

[0040] Then encode the semantic component g i as a feature vector, denoted as where d represents the embedding dimension, the query encoding represents the overall semantics of the query question, and the semantic component Represents part of the semantics of an interrogative sentence.

[0041] Preferably, in the step S25, the method for constructing the knowledge graph of the academic interrogative sentence is as follows:

[0042] The filtering of the relationship depends on the academic query interrogative sentence, the intention semantic component, and the relationship r j , perform the relationship r j Calculate the semantic similarity with the interrogative sentence. The calculation formula is as follows:

[0043]

[0044] Among them, Sim(·) is the similarity calculation function, and n 1 Is the number of intention semantic components, so as to filter the relationship r that completely contains the semantic information of the whole and the part j , then sort the similarity of the relationships to obtain a sequence Where S j > S j+1 .

[0045] The present invention discloses an academic knowledge question answering system based on semantic components, including an academic domain knowledge base, a synonym mapping library, an attribute knowledge graph, a subgraph construction module, an attribute constraint module, and an answer module.

[0046] The academic domain knowledge base is used to store the academic domain knowledge graph with the resource description framework as the data model;

[0047] The academic domain synonym mapping library is used to store the mapping dictionary from standard words to aliases in academic information;

[0048] The academic domain attribute knowledge graph has the same storage format as the domain knowledge graph, both in triple format, and the content is the attributes and attribute values of time and numerical values;

[0049] The subgraph construction module processes the query interrogative sentence and obtains the subgraph corresponding to the interrogative sentence from the pre-constructed domain knowledge graph;

[0050] The attribute constraint module uses the pre-constructed attribute knowledge graph to add constraint information to the subgraph and filters to obtain the final subgraph;

[0051] The answer generation module is used to display the filtered knowledge graph subgraph in a friendly form.

[0052] The present invention discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned academic knowledge question answering method based on semantic components are implemented.

[0053] The present invention discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above-mentioned academic knowledge Q&A method based on semantic components are implemented.

[0054] The present invention has the following characteristics and beneficial effects:

[0055] The present invention constructs an academic domain knowledge graph, a synonym mapping library, and an attribute knowledge graph; in addition, the academic knowledge Q&A method based on semantic components involved in the present invention utilizes semantic components to map the knowledge graph subgraph and the constrained attribute information of an academic query question sentence, so as to realize the correct determination of the intention in the academic query question sentence. In specific applications, it enables beginners in academic research to quickly understand and find academic information with high usability for themselves. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0057] Figure 1 The overall process of the academic knowledge Q&A method based on semantic components in the embodiments of the present invention.

[0058] Figure 2 The overall process of the academic knowledge Q&A system based on semantic components in the embodiments of the present invention.

[0059] Figure 3 The flowchart of the subgraph construction module in the embodiments of the present invention.

[0060] Figure 4 The flowchart of attribute constraints in the embodiments of the present invention.

[0061] Figure 5 The example diagram of syntax parsing in the embodiments of the present invention.

[0062] Figure 6 The example diagram of semantic components in the embodiments of the present invention.

[0063] Figure 7 The example diagram of relationship filtering in the embodiments of the present invention. Detailed Embodiments

[0064] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0065] The following will further clearly and completely elaborate on the specific implementation of the technical solution of the present invention in conjunction with the accompanying drawings and embodiments. The specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Equivalent alternative solutions made within the essence and scope of the present invention based on the embodiments in this specification all fall within the scope protected by this specification.

[0066] The present invention provides an academic knowledge Q&A method based on semantic components, as Figure 1 shown, the specific steps are as follows:

[0067] S1. Build an academic domain knowledge base,

[0068] S11. The academic domain knowledge base uses the academic information publicly available on the Internet by academic researchers as the data source, and the publicly available data is crawled by a crawler tool.

[0069] S12. Define information domains: "author" domain, "institution" domain, "field" domain, "journal" domain, and "literature" domain.

[0070] It can be understood that each domain is subdivided into many types, and each type is further subdivided into many attributes. For example, the institution domain is subdivided into many types such as university type, scientific research institution type, and company type. The university type has attributes such as name and research direction, and the literature domain has academic paper type and master's / doctoral thesis type. Therefore, the hierarchical data management constitutes the relationship semantic network of entities in the academic domain.

[0071] S13. Build an academic domain knowledge graph based on the Resource Description Framework data model, perform template annotation on the data, realize the annotation processing of the crawled data, and store it in the form of <entity, relationship, entity> triples.

[0072] It can be understood that existing knowledge graph construction tools are used to complete the annotation of the data and the construction of the academic domain knowledge graph.

[0073] S14. Build an academic domain synonym mapping library.

[0074] Specifically, the synonym mapping library is constructed by processing the academic information publicly available on the Internet and is constructed as a mutual mapping dictionary from standard words to alternative words. For example, the mapping between the standard word "Hangzhou Dianzi University" and the alternative words "HDU" and "Hangdian" in institution entities.

[0075] S15. Build an academic domain attribute knowledge graph.

[0076] It can be understood that since the questions in the academic field often involve time and numerical constraints, in this embodiment, the attribute knowledge graph is used to include time attributes and numerical attributes and is separated from the domain knowledge graph. The time attributes include attributes such as the publication time of papers, and the numerical attributes include attributes such as the number of citations of papers. The storage form of its data is equivalent to the triple form.

[0077] S2. Construct a knowledge graph subgraph for academic queries based on the intent semantic component

[0078] S21. Preprocess the proposed academic query question, including using a synonym mapping library to replace the standard words in the question q and obtain its syntactic parsing graph G gram and the semantic component g i , where the semantic component is the path where each leaf node is located.

[0079] It can be understood that the above technical solution can obtain the syntactic parsing graph by using natural language processing tools such as Stanza to perform dependency syntactic parsing on the academic query question. The syntactic parsing graph is only used to extract the semantic component.

[0080] Furthermore, the semantic components are divided into constraint semantic components and intent semantic components. Among them, the constraint semantic components represent the time constraints and numerical constraints in the question, and the intent semantic components are the other semantic components in the question, representing semantic information. As a possible implementation method, the constraint semantic components can be extracted by using the method of regular expression matching.

[0081] S22. Perform academic named entity recognition on the question q

[0082] Among them, academic named entity recognition includes the recognition of entities in the academic field knowledge graph such as person name entities, organization name entities, and literature name entities in the question. For example, the organization name entity "Hangzhou Dianzi University", the journal conference name entity "EMNLP", etc.

[0083] Specifically, as a possible implementation method, entity sample annotation can be performed on the defined five academic information domains to obtain a named entity recognition model for the academic field, and then named entity recognition within the domain can be realized. The entity recognition result is the entity set E=(e 1 ,e 2 ,e 3 ,…,e n ), n is the number of entities, and e i is an entity.

[0084] S23. Encode the preprocessed question and semantic component into feature vectors respectively.

[0085] Specifically, first, encode the question q to obtain the complete semantic representation of each word where d represents the word embedding dimension, and n represents the number of words in the question,

[0086] Then, for the semantic component g i encode it into a feature vector, denoted as where d represents the embedding dimension, and the question encoding represents the overall semantics of the question, and the semantic component represents the partial semantics of the question.

[0087] S24. According to the pre-constructed academic domain knowledge graph, find the set of relationships R associated with each entity e in the entity set E. The search result is (r i , r 1 , …, r 2 ), where each relationship r n and the entity e i both have a triple i Then, for the relationship r in the relationship set R i encode it into a feature vector, denoted as where d represents the embedding dimension.

[0088] S25. Filter relationships. Not all relationships associated with the entity e i are relevant to the query question. The purpose of filtering relationships is to eliminate unnecessary relationship paths for constructing a sub-graph of the knowledge graph for the academic query question. Specifically, the filtering of relationships depends on the academic query question, the intent semantic component, and the relationship r j , and calculate the semantic similarity between the relationship r j and the question. The calculation formula is as follows:

[0089]

[0090] where Sim(·) is the similarity calculation function, and n 1 is the number of intent semantic components, so as to select the relationship r j that completely contains the overall and partial semantic information, and then sort the similarity of the relationships to obtain a sequence where S j > S j+1 .

[0091] S26. Set a threshold k for the paths, and select the k relationships with the largest similarity, that is, (s 1 , s 2 , …, s k ), n 1 > k, and retain the corresponding relationships (r 1 , r 2 , …, r k), and then perform the path extension of the next hop, taking the tail entity connected by the relationship (r 1 , r 2 , …, r k ) as the starting entity set for iterative extension. The process repeats steps S25 and S26, where the relationship is replaced by the relationship r in the next hop relationship set j . Iteratively calculate 1 n 1 times, which is the knowledge graph subgraph KG sub of the n

[0092] S3. Correct the knowledge graph subgraph based on the constraint semantic component, specifically including:

[0093] S31. Take out the entities in the knowledge graph subgraph as the entity set E 1 . According to the pre-constructed attribute knowledge graph, find the attribute set A associated with each entity e 1 in the entity set E i . The search result is (a 1 , a 2 , …, a f ), where each relationship a i and the entity e i both have a triple where v is the attribute value. Then encode the attribute a i in the attribute set A into a feature vector, denoted as where d represents the embedding dimension

[0094] S32. Filter the attributes to correct the knowledge graph subgraph of the constructed academic question

[0095] The method for correcting the knowledge graph subgraph is as follows:

[0096] The filtering of attributes depends on the academic query question, the constraint semantic component, and the attribute a i . Calculate the semantic similarity between the attribute a i and the question. The calculation formula is as follows:

[0097]

[0098] Then sort the similarity of the attributes to obtain a sequence (S 1 , S 2 , …, S n ), where S j > S j+1 , where n 2 is the number of constraint semantic components

[0099] S33. Set a threshold S k . If it is related to the entity ei The similarity of associated attributes is lower than the threshold, that is then this entity is regarded as irrelevant to the query and discarded, otherwise the entity node is retained.

[0100] S4. Answer output. Since the retrieval in the academic field mainly provides multiple retrieval results and uses the retrieved information as auxiliary information, therefore, the present invention uses the knowledge graph subgraph obtained after filtering as the answer and displays it in a visual form.

[0101] This embodiment also discloses an academic field knowledge question-answering system based on semantic components, as Figure 2 shown, including an academic field knowledge base, a synonym mapping library, an attribute knowledge graph, a subgraph construction module, an attribute constraint module, and an answer module.

[0102] Academic field knowledge base: used to store the academic field knowledge graph with the resource description framework as the data model.

[0103] Academic field synonym mapping library: used to store the mapping dictionary from standard words to aliases in academic information.

[0104] Academic field attribute knowledge graph: The storage format is the same as that of the domain knowledge graph, both in triple format, and the content is the attributes and attribute values of time and numerical values.

[0105] Subgraph construction module: processes the query question sentence and obtains the subgraph corresponding to the question sentence from the pre-constructed domain knowledge graph.

[0106] Attribute constraint module: uses the pre-constructed attribute knowledge graph to add constraint information to the subgraph and filters to obtain the final subgraph.

[0107] Answer generation module: used to display the knowledge graph subgraph obtained after filtering in a friendly form.

[0108] Specifically, as Figure 3 shown, the subgraph construction module includes: a standard word replacement sub-module, a semantic component sub-module, a named entity recognition sub-module, a relationship search sub-module, a similarity calculation sub-module, and a relationship filtering sub-module. In the figure, hop refers to the number of hops of the path starting from the entity, and n refers to the number of semantic components.

[0109] Standard word replacement sub-module: performs synonym replacement on the user's academic query question sentence. The synonym replacement queries the synonym mapping library according to the keywords and replaces the aliases in the query question sentence with standard academic field entity vocabulary.

[0110] Semantic component sub-module: performs dependency syntactic analysis on the academic query question sentence to obtain a syntactic parsing graph, and then splits the syntactic parsing graph into multiple semantic components g i, where the semantic components are the paths where each leaf node is located, and the semantic components are divided into constraint semantic components and intention semantic components. In the present invention, the dependency relationship between words in the syntax parsing graph has no practical meaning and is only used for splitting semantic components. Taking "What are the papers on KGQA at the EMNLP2021 conference" as an example, the syntax parsing tree is as shown in Figure 5 shown, and the semantic components are as shown in Figure 6 shown.

[0111] Named entity recognition sub-module: Identify entities in academic query questions, including the identification of entity types in the knowledge graph of the academic field such as person name entities, organization name entities, and paper name entities in the question.

[0112] Relationship search sub-module: Search for relationships associated with entities in the current entity set, including the relationships in all triples with the entity as the head entity and the tail entity.

[0113] Similarity calculation sub-module: Calculate the similarity between the relationships in the retrieved relationship set and the question, specifically including the similarity between the relationship and the overall semantic information of the question and the partial semantic information of the question (i.e., semantic components) of the similarity. As a possible implementation method, the calculation of similarity can be implemented by cosine similarity or a function defined in the RoBERTa model.

[0114] Relationship filtering sub-module: Filter out relationships with a similarity ranking higher than a predefined threshold. If the similarity ranking is less than the defined threshold, the relationship is retained as a candidate path; otherwise, the relationship is discarded. The relationship filtering diagram is as shown in Figure 7 shown, and the relationships r 11 , r 12 , r 13 are associated with the entity node "EMNLP". In the first iteration, the first-hop relationships are filtered. r 12 has the greatest similarity with the sentence and its semantic components, and r 13 is the second. The relationship path r 11 is discarded. In the second iteration, the second-hop relationships are filtered, and the relationship path r 21 is "selected", and the path r 22 is discarded.

[0115] Furthermore, as shown in Figure 4 shown, the attribute constraint module includes an attribute search sub-module, a similarity calculation sub-module, and a constraint filtering sub-module.

[0116] Attribute search sub-module: Search for relevant attribute triples for the entity set in the sub-graph from the attribute knowledge graph.

[0117] Similarity calculation sub-module: Calculate the similarity between the attributes in the retrieved attribute set and the question sentence and the constraint semantic component.

[0118] Constraint filtering sub-module: Filter out some entity nodes in the sub-graph, that is, the similarity of the attributes in the attribute set associated with the discarded entity nodes is lower than the threshold.

[0119] Answer generation module: Used to visually display the knowledge graph sub-graph obtained after filtering.

[0120] This embodiment also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned semantic component-based academic knowledge question-answering method are implemented.

[0121] It can be understood that these computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including the instruction device, and the instruction device is to implement the process Figure 1 one process or multiple processes and / or boxes Figure 1 specified functions in one box or multiple boxes.

[0122] This embodiment also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above-mentioned semantic component-based academic knowledge question-answering method are implemented.

[0123] It can be understood that these computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the loaded computer or other programmable device to generate computer-implemented processing, so that the instructions executed on the loaded computer or other programmable device provide for implementing the process Figure 1 one process or multiple processes and / or boxes Figure 1 steps of the specified functions in one box or multiple boxes.

[0124] The above has described the embodiments of the present invention in detail with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principles and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments including components still fall within the protection scope of the present invention.

Claims

1. A method for academic knowledge question answering based on semantic components, characterized in that, it includes the following steps: S1. Construct an academic domain knowledge base; S2. Construct a knowledge graph subgraph for academic queries based on intent semantic components; S21. Preprocess the proposed academic query question sentence, including performing standard word replacement on the question sentence q using a synonym mapping library, and obtaining its syntactic parsing graph G gram and semantic component g i , where the semantic component is the path where each leaf node is located; S22. Perform academic named entity recognition on the question q Annotate entity samples for the five defined academic information domains, and then achieve named entity recognition within the domain. The entity recognition result is the entity set E = (e 1 , e 2 , e 3 , …, e n ), where n is the number of entities, and e i is an entity; S23. Encode the preprocessed question and semantic components into feature vectors respectively; S24. According to the pre-constructed knowledge graph of academic fields, find each entity e in the entity set E i associated relationship set R. The search result is (r 1 , r 2 , …, r n ), where each relationship r i and the entity e i both have a triple Then, encode the relationship r in the relationship set R i into a feature vector, denoted as where d represents the embedding dimension; S25. Filter the relationships, and then construct a knowledge graph subgraph for the academic question; S26. Set a threshold k for the path, and select k relationships with the largest similarity, i.e., (s 1 , s 2 , …, s k ), where n 1 > k. Retain the corresponding relationships (r 1 , r 2 , …, r k ), and then perform the path extension of the next hop. Use the tail entities connected by the relationships (r 1 , r 2 , …, r k ) as the starting entity set for iterative extension. Repeat steps S25 and S26 in the process, and replace the relationships with the relationships r j in the next hop relationship set. Iteratively calculate n 1 times, which is the knowledge graph subgraph KG 1 of n sub hops; S3. Correct the knowledge graph subgraph based on constraint semantic components; S31. Extract the entities in the knowledge graph subgraph as the entity set E 1 , and according to the pre-constructed attribute knowledge graph, find each entity e 1 in the entity set E i and its associated attribute set A. The search result is (a 1 , a 2 , …, a f ), where each relationship a i and the entity e i both have a triple where v is the attribute value. Then, encode the attribute a i in the attribute set A into a feature vector, denoted as where d represents the embedding dimension. S32. Filter the attributes to correct the constructed knowledge graph subgraph of the academic question, The correction method of the knowledge graph subgraph is: The filtering of attributes depends on academic query questions, constraint semantic components, and attribute a i , and perform attribute a i Calculate the semantic similarity with the query sentence, and then sort the similarity of attributes; S33. Set threshold S k If the similarity of the attributes associated with entity e i is lower than the threshold, the entity is considered irrelevant to the query and discarded; otherwise, the entity node is retained. S4. Answer generation.

2. The method for academic knowledge question answering based on semantic components according to claim 1, characterized in that, in the step S32, Attribute a i Calculation of the semantic similarity with the question sentence, and its calculation formula is as follows: Then, the similarities of the attributes are sorted to obtain a sequence (S 1 , S 2 , …, S n ), where S j > S j+1 , and where n 2 is the number of constraint semantic components, is the question encoding, and is the semantic component.

3. The method for academic knowledge question answering based on semantic components according to claim 1, characterized in that, in the step S21, the semantic components include constraint semantic components and intent semantic components. The constraint semantic components represent the time constraints and numerical constraints in the question, and the intent semantic components are other semantic components in the question, representing semantic information. The constraint semantic components are extracted by using the method of regular expression matching.

4. The method for academic knowledge question answering based on semantic components according to claim 3, characterized in that, in the step S23, the encoding method of the feature vector is: First, encode the question q to obtain the complete semantic representation of each word where d represents the dimension of word embedding, and n represents the number of words in the question Then, the semantic component g i is encoded as a feature vector, denoted as where d represents the embedding dimension, and the question encoding represents the overall semantics of the question, and the semantic component represents the partial semantics of the question.

5. The method for academic knowledge question answering based on semantic components according to claim 1, characterized in that, in the step S25, the method for constructing the knowledge graph of the academic question is: The filtering of the relationship depends on the academic query question, the intent semantic component, and the relationship r j , perform the relationship r j , calculate the semantic similarity between the relationship r and the question sentence, and its calculation formula is as follows: Among them, Sim(·) is a function for calculating similarity, and n 1 is the number of intent semantic components, so as to filter out the relationship r that completely contains the semantic information of the whole and the part j , and then sort the similarity of the relationships to obtain a sequence where S j > S j+1 .

6. The method for academic knowledge question answering based on semantic components according to claim 1, characterized in that, the step S1 includes the following sub-steps: S11. Obtain academic domain knowledge data; S12. Define an academic information domain for the obtained academic domain knowledge data; S13. Construct an academic domain knowledge graph based on the Resource Description Framework data model, perform annotation processing on the crawled data, and store it in the form of <entity, relationship, entity> triple data; S14. Construct an academic domain synonym mapping library; S15. Construct an academic domain attribute knowledge graph.

7. A system for implementing the method for academic knowledge question answering based on semantic components according to any one of claims 1-6, characterized in that, it is composed of an academic domain knowledge base, a synonym mapping library, an attribute knowledge graph, a subgraph construction module, an attribute constraint module and an answer module, The academic domain knowledge base is used to store the academic domain knowledge graph with the Resource Description Framework as the data model; The academic domain synonym mapping library is used to store the mapping dictionary from standard words to aliases in academic information; The academic domain attribute knowledge graph has the same storage format as the domain knowledge graph, both in triple format, and the content is the attributes and attribute values of time and numerical values; The subgraph construction module processes the query question and obtains the corresponding subgraph of the question from the pre-constructed domain knowledge graph; The attribute constraint module uses the pre-constructed attribute knowledge graph to add constraint information to the subgraph and filter to obtain the final subgraph; An answer generation module, configured to display the sub-graph of the knowledge graph obtained after filtering in a friendly form.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the semantic component-based academic knowledge Q&A method according to any one of claims 1-6 are implemented.

9. A computer-readable storage medium, having stored thereon a computer program, wherein when the program is executed by a processor, the steps of the semantic component-based academic knowledge Q&A method according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Method and system for translating user keywords into semantic queries based on a domain vocabulary

    US20140379755A1