A Structure-Priority Method and System for Implementing Knowledge Base Question Answering

Through the structure-first knowledge base question-and-answer method, through question structure analysis and SPARQL query structure diagram generation, the problems of natural language understanding and representation, structured knowledge inconsistency and question complexity are solved, and the high-accuracy SPARQL query generation and knowledge base question-and-answer system are achieved.

CN114090782BActive Publication Date: 2025-06-13NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010857287.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-24
Publication Date
2025-06-13
Estimated Expiration
2040-08-24

AI Technical Summary

Technical Problem

The existing knowledge base question and answer system has difficulties and challenges in natural language understanding and representation, inconsistency between natural language and structured knowledge, and the complexity of question questions, and it is difficult to effectively deal with complex questions.

Method used

A structure-first knowledge base question-and-answer method is proposed. Through question structure analysis and SPARQL query structure diagram generation, the entity description diagram and relaxation query diagram are constructed, and the mapping relationship between question and SPARQL query is learned to generate a highly accurate SPARQL query structure diagram.

Benefits of technology

It realizes effective processing of complex questions, the generated SPARQL query structure diagram accuracy reaches more than 80%, and the knowledge base question and answer system reaches an F1 value of 0.76 under ideal link conditions, which is better than the existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114090782B_ABST
    Figure CN114090782B_ABST
Patent Text Reader

Abstract

A method and system for implementing knowledge base question answering with structure priority, including two parts: question structure analysis and generation of SPARQL query structure diagrams. The technical method of question structure analysis performs syntactic parsing on natural language questions, and designs and constructs two graph models: entity description graph and relaxed query graph. The generation technology of SPARQL query structure diagrams starts from the relaxed query graph, uses the method of constructing templates to construct a query structure mapping library for the relaxed query graph and the SPARQL query graph, and then for the question to be answered, extracts templates from the mapping library and splices them to obtain candidates for the SPARQL query structure diagram corresponding to the question to be answered. The present invention can generate SPARQL query structure diagrams with high accuracy, and can construct a complete knowledge base-based question answering system by using entity linking and relationship linking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and relates to knowledge graph technology and natural language processing technology, and is a method and system for realizing knowledge base question answering with structural priority. Technical Background

[0002] In the field of computers, question answering refers to a machine answering natural language questions, involving fields such as natural language processing, information extraction, and knowledge representation, aiming to build an automated question answering system: the input is a natural language question, and by utilizing structured knowledge representation or unstructured information collection, the answer to the question is obtained and output. Among them, knowledge base question answering is a question answering system built on a knowledge base, aiming to answer natural language questions based on the knowledge base. Nowadays, more and more structured data is available on the network, including knowledge bases such as DBpedia, Freebase, and YAGO. How end users can conveniently and quickly access the knowledge base has become an important topic.

[0003] The Resource Description Framework (RDF), as the standard representation of a knowledge base, consists of triples <s, p, o>, where s = subject, p = predicate, o = object, and is usually expressed as a graph structure. The SPARQL language (SPARQL Protocol and RDF Query Language) is a query language and data acquisition protocol developed for RDF. SPARQL queries are the standard query method for accessing RDF data. However, SPARQL syntax and RDF schema have a high degree of complexity, and mastering the SPARQL language requires professional knowledge, making it difficult for non-professional users to use. In this regard, building a good knowledge base question answering system can enable non-professional users who do not understand SPARQL syntax and knowledge base construction to effectively use and access the knowledge base, which is a popular research topic that has received much attention from many researchers nowadays. The idea behind the knowledge base question answering system is to find the information requested by the user in natural language in the knowledge base, which is usually solved by converting natural questions into SPARQL queries, and SPARQL queries can be used to retrieve the required information.

[0004] In knowledge base question answering, the generation of formal queries is particularly important when answering complex questions. Given entity and relationship linking results, the purpose of formal query generation is to generate correct executable queries, such as SPARQL queries, from natural language questions. Existing methods include: methods based on pre-collected templates, methods based on semantic parsing, and methods based on information retrieval and neural networks, etc. As an important technique in query generation, templates are often used to simplify the semantic parsing of natural language descriptions and generate structured queries. Existing methods usually pair natural language description templates with query templates and map the components of the description to the components of the query. Many existing works rely on manually generated templates, which leads to limitations of the methods. Knowledge base question answering systems based on semantic parsing usually use combinatory categorial grammar to convert natural language into query graphs or logical expression forms. Question answering systems based on information retrieval do not perform semantic parsing on natural language, but select a series of answer candidates through methods such as relation extraction, and then evaluate and score them using different methods.

[0005] In existing research work on knowledge base question answering, there are still many difficulties and challenges, mainly including: the understanding and representation of natural language; the inconsistency between natural language and structured knowledge; the complexity of questions, etc.

[0006] Understanding and representation of natural language: The expressions of natural language are ever-changing, and the expression forms of sentences with the same meaning may be completely different. It is very difficult to correctly understand and express natural language.

[0007] Inconsistency between natural language and structured knowledge in the knowledge base: This is the most core difficulty in knowledge base question answering, which is reflected in both the syntactic organizational structure and semantic expression. In terms of semantic expression, the natural language words may be different from the entity / relationship names in the knowledge base; in terms of syntactic structure, it is difficult to analyze its structure in the knowledge base starting from natural language.

[0008] Complexity of questions: The performance of existing knowledge base question answering systems on complex questions is not ideal. This is because when the question structure is complex and the number of facts increases, it is far more difficult to correctly find and combine these facts to form a correct structure than to process simple questions; at the same time, some special question patterns, such as comparative sentences and superlative sentences, often require more targeted analysis and processing.

[0009] Based on the above challenges in knowledge base question answering, the present invention proposes a structure-priority knowledge base question answering technology, which provides solutions to the above challenges. Summary of the Invention

[0010] The problems to be solved by the present invention are as follows: There are still many difficulties and challenges in the existing research work on knowledge base question answering, including the understanding and representation of natural language, the inconsistency between natural language and structured knowledge, the complexity of questions, etc. It is necessary to conduct research on this and propose solutions.

[0011] The technical solution of the present invention is as follows: A method for implementing structure - priority knowledge base question answering includes two parts: question structure analysis and generation of SPARQL query structure diagrams.

[0012] Question structure analysis obtains a syntactic tree by parsing a natural language question, and then constructs an entity description graph and a relaxed query graph. The entity description graph is a graph - structure representation of entities and corresponding descriptions in the question. The question is divided into sub - questions by entity blocks composed of entity + description, expressing the structural hierarchy of the question. The relaxed query graph is the embodiment of natural language in the query graph structure.

[0013] The generation of SPARQL query structure diagrams includes the construction of a query structure mapping library and the construction of SPARQL query structure diagram candidates. Starting from the relaxed query structure diagram, the SPARQL query structure diagram learns the mapping relationship between the relaxed query structure diagram of the question and the SPARQL query structure diagram. In the way of constructing templates, a query structure mapping library for the relaxed query structure diagram and the SPARQL query structure diagram is constructed. Then, for the question to be solved, a template that can cover the relaxed query structure diagram of the question to be solved is extracted from the mapping library, and the SPARQL query structure diagram candidate corresponding to the question to be solved is obtained by rule splicing. Based on the entity and relationship linking of the knowledge graph, the slots of the knowledge base entities and relationships of the SPARQL query structure diagram candidate are filled to obtain a SPARQL query, and the answer is queried and returned from the knowledge base.

[0014] The present invention also provides a structure - priority knowledge base question - answering system. The system has a data processor and a memory and is configured with a computer program. The computer program is configured as a question structure analysis module and a SPARQL query structure diagram generation module. The SPARQL query structure diagram generation module further includes a query structure mapping library construction module and a SPARQL query structure diagram candidate construction module. When the computer program is executed, the above - mentioned knowledge base question - answering method is implemented.

[0015] The present invention can construct a complete knowledge - base - based question - answering system, which is specifically manifested as follows: The SPARQL query structure diagram candidate corresponding to the question to be solved is obtained by the SPARQL query structure diagram generation technology. By using the entity and relationship linking technology, the slots of the knowledge base entities and relationships of the SPARQL query structure diagram are filled to obtain a SPARQL query, and the answer is queried and returned from the knowledge base.

[0016] The beneficial effects of the present invention are as follows: The entity description graph EDG given by the present invention emphasizes the "answer purpose" of natural language questions, innovatively proposes a graph expression form of entities and corresponding descriptions, divides the questions into nested sub-questions, and clearly expresses the structural hierarchy of the questions; the relaxed query graph RQG is a query graph model for natural language questions. Different from existing query graphs, RQG is independent of the knowledge base and has greater universality. The two graph models can simultaneously play a good bridging role between natural language and formal queries, and have broad application prospects.

[0017] The SPARQL query structure graph generation technology given by the present invention can automatically learn the correspondence between RQG and SPARQL queries by constructing templates, and supports dynamic template splicing, which can significantly improve the practicality and effectiveness of the method. This technology can generate SPARQL query structure graphs with high accuracy. As Figure 11 shown, when taking the top 10 structure graphs, the accuracy can reach 80%, and reach 90% at 20.

[0018] Based on the above two technologies, the present invention implements a knowledge base question answering system SFQA. The system achieves an F1 value result of 0.76 under ideal link conditions, which is better than existing methods Sina, NLIWOD, SQG, as Figure 12 shown. Description of the Drawings

[0019] Figure 1 It is an example of an entity description pattern in an embodiment of the present invention.

[0020] Figure 2 It is an example of a relaxed query pattern in an embodiment of the present invention.

[0021] Figure 3 It is a flowchart for generating an entity description graph and a relaxed query graph of the present invention.

[0022] Figure 4A It is a relaxed query structure graph in an embodiment of the present invention.

[0023] Figure 4B It is a SPARQL query graph in an embodiment of the present invention.

[0024] Figure 4C It is a SPARQL query structure graph in an embodiment of the present invention.

[0025] Figure 5 It is a template schematic diagram in an embodiment of the present invention.

[0026] Figure 6 It is a schematic diagram of the mapping relationship between a relaxed query graph and a SPARQL query graph in an embodiment of the present invention.

[0027] Figure 7This is the set coverage example diagram of the relaxation query structure diagram for the embodiments of the present invention.

[0028] Figure 8 This is the template splicing example diagram for the present invention.

[0029] Figure 9 This is the flowchart of the question - answering system for the present invention.

[0030] Figure 10 This is the schematic diagram of the method for the present invention.

[0031] Figure 11 This is the accuracy schematic diagram of the SPARQL query structure diagram generated by the method of the present invention.

[0032] Figure 12 This is the effect comparison between the present invention and the existing question - answering system. Detailed implementation manners

[0033] In order to further elaborate on the technical features and effects of the present invention, the following further describes the present invention in combination with the accompanying drawings and specific implementation manners.

[0034] As Figure 10 shown, the present invention proposes a method for implementing knowledge - base question - answering with structure priority, which adopts two methods: semantic parsing and based on pre - collected templates, and includes two parts: question - sentence structure analysis and SPARQL query structure diagram generation.

[0035] The question - sentence structure analysis obtains a syntactic tree by parsing a natural - language question - sentence, and then constructs an entity description graph and a relaxation query graph; the entity description graph is a graph - structure representation of the entities and corresponding descriptions in the question - sentence, and divides the question - sentence into sub - question nestings through entity blocks composed of entity + description, expressing the structural hierarchy of the question - sentence; the relaxation query graph is the embodiment of the natural - language structure in the query.

[0036] The generation of the SPARQL query structure diagram includes the construction of a query structure mapping library and the construction of SPARQL query structure diagram candidates. Starting from the relaxation query structure diagram, the SPARQL query structure diagram learns the mapping relationship between the relaxation query structure diagram and the SPARQL query structure diagram of the question - sentence, and uses the method of constructing templates to construct a query structure mapping library for the relaxation query structure diagram and the SPARQL query structure diagram. Then, for the question - sentence to be solved, extract the template that can cover the relaxation query structure diagram of the question - sentence to be solved from the mapping library, and obtain the SPARQL query structure diagram candidate corresponding to the question - sentence to be solved through rule splicing. Based on the entity and relationship linking of the knowledge graph, fill the slots of the entities and relationships in the knowledge base for the SPARQL query structure diagram candidate to obtain the SPARQL query, and query and return the answer from the knowledge base.

[0037] The following illustrates the implementation of the present invention through embodiments.

[0038] Taking the question "What is the deathplace of the rugby player who is the relatives of Anton Oliver?" as an example, Figure 1 and Figure 2 are the entity description graph and the relaxation query graph corresponding to the question respectively. Among them, Figure 2 in the entity vertices, class and name are the class and name attributes of the entity respectively.

[0039] Figure 3 describes the process of generating the entity description graph and the relaxation query graph from the question.

[0040] Step 1: Replace the parts with quotes such as famous quotes and long entities (token count greater than 3) in the question with markers <quote> , <entity>; Replace "What jobs did the artist who recorded"The Incredible"hold?" with "What jobs did the artist who recorded <quote1>hold?

[0041] Step 2: Starting from the syntactic tree, process all non-terminal tags in the syntactic tree, such as SQ, NP, PP, etc. Different tags correspond to different processing rules, and recursively generate the vertices and edges of the entity description graph. The specific processing rules are as follows: For each tag, select the generation method according to the current tag, the tag names of its child nodes, and the vertices of the entity description graph generated in the upper layer. For example, the SQ tag will generate a verb phrase description vertex, pointing to the entity vertex in the upper layer. If the child node contains NP / VP / PP tags, enter the corresponding program of NP / VP / PP for further processing. Another example, when the current tag is NP, when the upper layer is an entity vertex, generate a new non-verbal phrase description vertex pointing to the entity vertex; when the upper layer is a non-verbal phrase description vertex, do not generate a new non-verbal phrase description vertex, but judge the sub-tags. If a clause or verb phrase modification situation is encountered, generate a new entity block, an entity vertex, and further generate possible description vertices.

[0042] Step 3: After constructing the entity description graph, substitute back the references and long entity markers, that is, replace the markers with the original references and long entity content again to obtain the complete entity description graph.

[0043] Step 4: Starting from the entity description graph, use different methods to identify and process the verb phrase descriptions and non-verbal phrase descriptions in each entity block respectively: From the non-verbal phrase descriptions, use the named entity recognition technology to identify the name and class of the entity, namely name / class, and put them inside the entity in the relaxation query graph. name / class is used as the internal attribute of the entity vertex. From the verb phrase descriptions, extract the verbs / relationships according to the short sentence syntactic tree structure, add out-edges to the entity vertices, use the verbs / relationships as the attributes of the relaxation query graph edges, and at the same time generate new possible entity vertices. The new entity vertices are generated according to the syntactic tree structure. For example, a new entity vertex should be generated at the object position, and identify whether the description of this part is the class or name of the entity, and put it inside the newly generated entity vertex as an attribute.

[0044] Step 5: Merge the entity vertices of the graph structures obtained from each entity block of the entity description graph to obtain the final relaxation query graph.

[0045] Figure 4A -C are respectively the relaxation query structure graph, SPARQL query graph, and SPARQL query structure graph corresponding to the question "What is the deathplace of the rugby player who is the relatives of Anton Oliver?" This method names the relaxation query graph and the above three graphs as RQG, Rstruct, QG, and Qstruct respectively. Figure 4A It is the relaxed query structure graph Rstruct. The internal attributes hasClass and hasName of the entity vertex respectively represent the attributes that the entity vertex has a class and has a name. Figure 4B It is a SPARQL query graph. dbr, dbp, and dbo are the namespace prefixes of the DBpedia knowledge base, which are respectively "http: / / dbpedia.org / resource / ", "http: / / dbpedia.org / property / ", and "http: / / dbpedia.org / ontology / ". Figure 4C It is a SPARQL query structure graph.

[0046] The generation of the SPARQL query structure graph is divided into two parts, namely the construction of the query structure mapping library and the construction of the SPARQL query structure graph candidates.

[0047] The query structure mapping library consists of templates of multiple question types. The question types are divided into: general questions, questions asking about quantity, and questions asking about the entity itself. The templates are as follows Figure 5 shown. Each template consists of a relaxed query structure graph, n SPARQL query structure graphs, the mapping relation function map from this relaxed query structure graph to the n SPARQL query structure graphs, and the scores corresponding to the n SPARQL query structure graphs. The construction method of the mapping library is as follows

[0048] Step 1: Classify the questions in the question dataset into three categories: general questions, questions asking about quantity, and questions asking about the entity itself.

[0049] Step 2: Establish a training set from the question dataset and learn the mapping relation between the relaxed query structure graph and the SPARQL query structure graph in the training set. Figure 6 shows the mapping relation result obtained from a question. On the left is the relaxed query graph RQG corresponding to the question, on the right is the SPARQL query structure graph corresponding to the question, and the dashed line in the middle represents the mapping relation.

[0050] Step 3: Extract the structural information of the question relaxed query graph and the SPARQL query structure graph to obtain the relaxed query structure graph Rstruct and the SPARQL query structure graph Qstruct.

[0051] Step 4: Extract one or more original templates (Rstruct’, Qstruct’, map’) from (Rstruct, Qstruct, map), where Rstruct’ and Qstruct’ are subgraphs of Rstruct and Qstruct respectively, and map’ is the relationship function of the corresponding subgraph. The original template refers to the template that has not undergone isomorphic merging and statistical induction. There may be one or more extracted original templates.

[0052] Step 5: Merge the isomorphic Rstruct’ and Qstruct’ in the original templates of the same question type to obtain the final template, including the merged relaxed query structure graph Rstruct m and the SPARQL query structure graph Qstruct m , and each template includes a relaxed query structure graph, n SPARQL query structure graphs, the mapping relationship between them, and the corresponding scores of the SPARQL query structure graphs, to construct a query structure mapping library.

[0053] The construction of the SPARQL query structure graph candidates includes two parts: set covering and template splicing.

[0054] Set covering is to perform graph covering on the relaxed query structure graph Rstruct quest of the question to be solved. As Figure 7 shown, assume that the relaxed query structure graphs of two templates, Template 1 and Template 2, are Rstruct1 and Rstruct2 respectively. Rstruct1 covers the subgraph generated by Entity1 and Entity2 of Rstruct quest , and Entity1 and Entity2 of Rstruct1 correspond to Entity1 and Entity2 of Rstruct quest respectively; Rstruct2 covers the subgraph generated by Entity2, Entity3, and Entity4 of Rstruct quest , and Entity1, Entity2, and Entity3 of Rstruct2 correspond to Entity2, Entity3, and Entity4 of Rstruct quest respectively. The specific method is to uniformly regard the point set, edge set, and logical constraint set of Rstruct quest as elements in the universal set, and use an improved set covering greedy algorithm to select the template set that can cover the relaxed query structure graph of the question to be solved. The traditional set covering greedy algorithm is to select the set that minimizes the total loss function in the remaining sets each time and put it into the result until the universal set can be covered. Here, the main differences between the improved set covering greedy algorithm of the present invention and the traditional algorithm are: 1) Regarding Rstruct quest The point set, edge set, and logical constraint set are uniformly regarded as elements in the universal set, and the Rstruct in the template m The elements in it are used as a set to perform set covering on Rstruct quest The elements in it; 2) The goal is changed to selecting the top k set coverings instead of one set covering. For this, we rank the sets. Suppose there are n sets, and the traditional set covering algorithm is repeated n times. In the i-th time, only the sets starting from the i-th set are selected, i = 1,..., n; 3) The loss function is related to the approximation degree of the relaxed query structure graph, that is, the set corresponding to Rstruct m and Rstruct quest The similarity with Rstruct is inversely proportional to the loss function.

[0055] Figure 7 The templates where Rstruct1 and Rstruct2 are located in it form a template set that meets the requirements.

[0056] Template splicing is to perform template splicing on all template sets that meet the requirements respectively. The specific steps are as follows:

[0057] Step 1: Perform the operations in Step 2 on each template set respectively.

[0058] Step 2: For each template in the template set, sequentially try to select the relaxed query structure graphs in each template for splicing, and remove the contradictory situations. The pruning operation is to record all pairs of ((template 1, Qstruct1), (template 2, Qstruct2)) that generate contradictions, where Qstruct1 and Qstruct2 are the SPARQL query structure graphs in template 1 and template 2 respectively.

[0059] Step 3: Score the SPARQL query structure graphs obtained from all template sets, and select the top k as the candidates for the SPARQL query structure graph of the question to be solved.

[0060] Figure 8 Shows a situation of template splicing. Rstruct1 and Rstruct2 are the relaxed query structure graphs Rstruct of template 1 and template 2 respectively, Qstruct i1 and Qstruct i2 are the SPARQL query structure graphs selected from template 1 and template 2 respectively. Rstruct quest is the relaxed query structure graph of the question to be solved, and Qstruct is the Qstruct selected from template 1 and 2 i1 and Qstruct i2 The SPARQL query structure diagram of the to-be-solved query sentence obtained by post splicing. According to the mapping relationship, as shown by the dotted box in the figure, entity vertex 2 in Rstruct1 and entity vertex 1 in Rstruct2 are both mapped to entity vertex 2 in Rstruct quest and respectively correspond to variable 1 in Qstruct i1 and variable 1 in Qstruct i2 The overlapping parts are not contradictory, so Qstruct i1 and Qstruct i2 corresponding variable vertices are combined to form a new SPARQL query structure diagram, that is, the Qstruct of the to-be-solved query sentence.

[0061] Figure 9 is the flow schematic diagram of the method and the corresponding system of the present invention. The query structure mapping library construction module and the SPARQL query structure diagram construction module are as described above. The SPARQL query construction module performs knowledge base entity and relationship linking on the SPARQL query structure diagram candidates of the to-be-solved query sentence, and further searches the knowledge base for trial and error and scores to obtain the final SPARQL query, and returns the answer by querying the knowledge base. < / entity> < / quote>

Claims

1. A method for implementing knowledge base question answering with structure priority, characterized in that it includes two parts: question structure analysis and generation of SPARQL query structure diagrams. Question structure analysis obtains a syntactic tree by parsing a natural language question, and then constructs an entity description graph and a relaxed query graph. The entity description graph is a graph structure representation of entities and corresponding descriptions in the question. The question is divided into sub-questions by entity blocks composed of entity + description, expressing the structural hierarchy of the question. The relaxed query graph is the embodiment of natural language in the query graph structure. Generation of SPARQL query structure diagrams includes construction of a query structure mapping library and construction of SPARQL query structure diagram candidates. Starting from the relaxed query structure diagram, the SPARQL query structure diagram learns the mapping relationship between the relaxed query structure diagram of the question and the SPARQL query structure diagram. Using the method of constructing templates, a query structure mapping library for the relaxed query structure diagram and the SPARQL query structure diagram is constructed. Then, for the question to be answered, a template that can cover the relaxed query structure diagram of the question to be answered is extracted from the mapping library, and the SPARQL query structure diagram candidate corresponding to the question to be answered is obtained by rule splicing. Based on entity and relationship linking in the knowledge graph, the slots of entities and relationships in the SPARQL query structure diagram candidate are filled to obtain a SPARQL query, and the answer is queried and returned from the knowledge base. Among them, the query structure mapping library consists of templates of multiple question types. The question types are divided into: general questions, questions asking about quantity, and questions asking about entities themselves. Each template consists of a relaxed query structure diagram, n SPARQL query structure diagrams, n mapping relationship functions map between a relaxed query structure diagram and n SPARQL query structure diagrams, and scores corresponding to n SPARQL query structure diagrams. The construction method of the mapping library is as follows: Step 1: Classify the questions into three categories: general questions, questions asking about quantity, and questions asking about entities themselves. Step 2: Establish a training set from the question data set and learn the mapping relationship between the relaxed query structure diagram and the SPARQL query structure diagram in the training set. Step 3: Extract the structural information of the relaxed query structure diagram and the SPARQL query structure diagram of the question to obtain the relaxed query structure diagram Rstruct and the SPARQL query structure diagram Qstruct. Step 4: Extract the original template (Rstruct’, Qstruct’, map’) from (Rstruct, Qstruct, map), where Rstruct’ and Qstruct’ are subgraphs of Rstruct and Qstruct respectively, and map’ is the relationship function of the corresponding subgraph. The original template refers to the template that has not undergone isomorphic merging and statistical induction. Step 5: Merge the isomorphic Rstruct’ and Qstruct’ in the original templates of the same question type to obtain a template and construct a query structure mapping library.

2. The method for implementing knowledge base question answering with structure priority according to claim 1, characterized in that the steps of constructing an entity description graph and a relaxed query graph from a syntactic tree are as follows: Step 1: Replace the reference and long entity parts in the question sentence, where the long entity refers to the part with more than 3 tokens, and replace the reference part with a tag <quote>, Replace the long entity part with a marker <entity> ;< / entity> < / quote> Step 2: Starting from the syntactic tree, process all non-terminal tags in the syntactic tree, recursively generate the vertices and edges of the entity description graph, and for different tags, determine the generation methods of vertices and edges according to the current tag, the tag names of its child nodes, and the vertices of the entity description graph generated in the upper layer; Step 3: After constructing the entity description graph, substitute back the references and long entity tags <quote> , <entity>, and obtain the complete entity description graph;< / entity> < / quote> Step 4: Starting from the entity description graph, use different methods to identify and process the verb phrase descriptions and non-verb phrase descriptions in each entity block respectively: from the non-verb phrase descriptions, use the named entity recognition method to identify the name and class of the entity, and put them inside the entity in the relaxation query graph, with name / class as the internal attribute of the entity vertex; From the verb phrase descriptions, extract the verbs / relationships according to the syntactic tree structure of the sentence, add outgoing edges to the entity vertices, use the verbs / relationships as the attributes of the edges in the relaxation query graph, and at the same time the edges point to the newly generated entity vertices; Step 5: Merge the entity vertices of the graph structures obtained from each entity block of the entity description graph to obtain the final relaxation query graph.

3. A method for implementing knowledge base question answering with structure priority according to claim 1, characterized in that SP The construction of candidates for the ARQL query structure diagram includes two parts: set covering and template splicing. Set covering is to perform graph covering on the relaxed query structure diagram Rstruct of the question to be answered, and use an improved greedy algorithm for set covering to select the template set that can cover the relaxed query structure diagram of the question to be answered Rstruct in the query result mapping library. In the improved greedy algorithm for set covering: 1) The point set, edge set, and logical constraint set of Rstruct are uniformly regarded as elements in the universal set, and the elements in the relaxed query structure diagram in the template are used as sets to perform set covering on the elements in Rstruct; 2) The goal is to select the top k set coverings. Rank the sets, and repeat the traditional greedy algorithm for set covering n times. Only select the sets starting from the i-th set in the i-th time; 3) The loss function is related to the approximation degree of the relaxed query structure diagram, that is, the similarity between the template relaxed query structure diagram corresponding to the set and Rstruct is inversely proportional to the loss function. quest Perform graph covering on it, and use an improved greedy algorithm for set covering to select the template set that can cover the relaxed query structure diagram Rstruct of the question to be answered in the query result mapping library. quest In the improved greedy algorithm for set covering: 1) Regard the point set, edge set, and logical constraint set of Rstruct quest as elements in the universal set, and use the elements in the relaxed query structure diagram in the template as sets to perform set covering on the elements in Rstruct quest ; 2) The goal is to select the top k set coverings. Rank the sets, and repeat the traditional greedy algorithm for set covering n times. Only select the sets starting from the i-th set in the i-th time; 3) The loss function is related to the approximation degree of the relaxed query structure diagram, that is, the similarity between the template relaxed query structure diagram corresponding to the set and Rstruct quest is inversely proportional to the loss function. Template splicing is to perform template splicing on all template sets that meet the requirements respectively, and the specific steps are as follows: Step 1: Perform the operations in Step 2 on each template set respectively; Step 2: For each template in the template set, sequentially try to select the relaxation query structure diagrams in each template for splicing, and remove contradictory situations through pruning operations; Step 3: Score the SPARQL query structure diagrams obtained from all template sets, and select the top k with the highest scores as candidates for the SPARQL query structure diagrams of the question to be asked.

4. A knowledge base question answering system with structure priority, characterized in that the system has a data processor and a memory and is configured with a computer program, the computer program is configured as a question structure analysis module and a SPARQL query structure diagram generation module, and the SPARQL query structure diagram generation module further includes a query structure mapping library construction module and a SPARQL query structure diagram candidate construction module, and when the computer program is executed, the method described in any one of claims 1-3 is implemented.

Citation Information

Patent Citations

  • Mapping knowledge domain questioning and answering system and method based on template matching technique

    CN105868313A

  • Question sentence and knowledge graph structure analysis-based natural-language question answering method and system

    CN108052547A