An Entity and Relationship Linking Method Based on Abstract Semantic Representation in Knowledge Base Question Answering
Through the dual encoder and cross encoder model based on abstract semantic representation, the accuracy of entity and relationship links in knowledge base questions and answers is solved, efficient links are achieved in multi-knowledge base scenarios, and link accuracy and generalization capabilities are improved.
Patent Information
- Application Number
- CN202310729848.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-06-19
AI Technical Summary
There is no method based on abstract semantic representation in the prior art to extract entities and relationship links in knowledge base questions and answers, resulting in insufficient accuracy when dealing with ambiguity.
Using an abstract semantic representation method, a link model composed of dual encoder and cross encoder is used to extract entities and relationship candidates through semantic analytical graphs, and efficient links of entities and relationships are combined with Wikipedia information.
It realizes efficient entity and relationship links in multi-knowledge base scenarios, improves link accuracy and generalization capabilities, and is suitable for different knowledge base Q&A scenarios.
Smart Images

Figure CN116821292B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology for a computer to link entity mentions and relationship representations extracted from natural language questions to corresponding entities and relationships in a knowledge base, and belongs to the field of information processing technology. Background Art
[0002] Entity Linking aims to link entities (such as people, locations, organizations, etc.) in text to corresponding entities in a knowledge base. Entity Linking needs to match entity mentions that appear in the text with entities in the knowledge base and assign a unique identifier to them. This identifier can be a URL, an ID, or a URI. Entity Linking mainly includes named entity recognition and entity disambiguation. The named entity recognition task needs to identify potential entities in the text and determine which words are entity mentions. In the entity disambiguation step, each identified entity mention needs to be corresponded to an entity in the knowledge base. The difficulty of the entity linking task lies in dealing with ambiguity. Existing entity linking work mainly uses supervised and unsupervised learning and other technologies to identify and link entities, and improves the accuracy of recognition and linking through various context features and embeddings.
[0003] Relation Linking aims to identify the relationships between two or more entities in text and link them to predefined or knowledge base relationship types. Relation Linking usually builds on entity linking. By determining the semantic relationships between entities, generating candidate relationships, excluding irrelevant relationships, etc., the relationship expressions in the text are linked to the knowledge base relationships, so as to use the knowledge base entities and relationships to assist in generating query statements for answer retrieval. Current relation linking methods include models based on semantic similarity, candidate generation and selection based on a knowledge base, methods based on logical rules, methods based on semantic parsing, etc.
[0004] Currently, no method has been found for entity and relationship linking in knowledge base question answering by obtaining extraction rules for entity mentions and relationship representations based on an abstract semantic representation and using the semantic parsing graph of this representation. However, there are entity linking methods using dual encoders and relation linking methods based on an abstract semantic representation, and the present method is completely different from these methods. Summary of the Invention
[0005] To overcome the deficiencies in the prior art, the present invention provides a method for entity and relationship linking based on abstract semantic representation in knowledge base question answering. This method uses abstract semantic representation to parse questions and obtain their logical expressions. Based on this expression, a linking model composed of a dual encoder and a cross encoder consisting of pre-trained language models is flexibly used to complete entity linking and relationship linking in knowledge base question answering, facilitating the development of a series of subsequent applications (such as question answering systems).
[0006] To achieve the above object, the technical solution of the present invention is as follows: A method for entity and relationship linking based on abstract semantic representation in knowledge base question answering, comprising the following steps:
[0007] Step 1, use a semantic parsing component to obtain the semantic logical representation of the sentence, and extract potential entity nodes in the semantic parsing graph of the abstract semantic representation of the question according to the structural information as entity candidates to be linked. To extract potential entity mention information in the question, use a string matching method to search for each "name" node in the semantic parsing graph, and splice all concept nodes connected to each "name" node from left to right, and use the convergence result as the entity candidate mention to be matched.
[0008] Step 2, use a dual encoder to preliminarily screen a large number of entities in the knowledge base to reduce the number of entity candidates. Then use a cross encoder to rank and score the entities mentioned in the question to determine the knowledge base entities corresponding to the entity mentions in the question. To further improve the accuracy of linking, entities and their descriptions in Wikipedia are introduced as external knowledge.
[0009] Step 3, use a semantic parsing component to obtain the semantic logical representation of the sentence, and extract the relationship substructure in the semantic parsing graph of the abstract semantic representation according to the structural information as the relationship candidate to be linked. To extract the potential relationship representation substructure in the question, first find the shortest path between the linked entity node and the "amr-unknown" target node, and combine relationship-related nodes, delete irrelevant nodes and compress the path through multiple rules to obtain a relatively complete relationship representation substructure based on the semantic parsing graph.
[0010] Step 4, use a cross encoder to comprehensively score and rank the relationship substructure in the semantic parsing graph and the candidate relationship set. Subsequently, combine the local dictionary with the highly relevant candidate relationships after scoring to supplement the knowledge base relationships that may be filtered out. Encode the question and relationship candidates respectively through a dual encoder, and score and select the encoded vectors to determine the knowledge base relationship corresponding to the relationship representation substructure.
[0011] Preferred: Extraction of entity mentions in the semantic parsing graph of the abstract semantic representation in step 1. Aggregate the entity mention candidate word nodes using the "name" node in the graph to obtain the entity mention list M in the question Q ={m1,…,m n}.
[0012] Preferred: Entity linking for each mentioned entity in the entity mention list in step 2. First, use the string matching method to determine the position of the mention in the question in the list, and re-divide the question into three parts according to the position using special tags, which is expressed as:
[0013] [CLS] Left text of mention [M s Mention [M e Right text of mention [SEP]
[0014] [CLS] and [SEP] represent the start and end positions of the sequence respectively, and divide the input sequence into multiple segments. The boundary of the mention is marked by the tags [M s and [M e . This part is used as the representation τ of the mention m , which is the input of one end of the dual encoder.
[0015] Based on the Wikidata knowledge base and Wikipedia documents, each entity knowledge is composed of the entity itself and the relevant description. By concatenating the description of the Wikidata entity in Wikipedia after the entity, its composition form is:
[0016] [CLS] Entity [ENT] Entity description [SEP]
[0017] [ENT] is used as a special tag to mark the entity in the sentence. This part is regarded as the representation τ of the Wikidata knowledge base entity with attached Wikipedia knowledge e as the input of the other end of the dual encoder.
[0018] These two parts of the input sequence are encoded by two encoders based on BERT to obtain their respective vector representations:
[0019] y m =red(BERT1(τ m ))
[0020] y e =red(BERT2(τ e ))
[0021] Among them, BERT1 and BERT2 are two BERT-base encoders, and red(·) is a function that reduces the vector sequence generated by BERT to a single vector. The candidate entities injected with Wikipedia information are encoded as The candidate mention in the question is y m , and the score s(m, e i ) of the entity-mention pair is given by the scoring function:
[0022]
[0023] Then, a cross-encoder is used for refined ranking. The input of the encoder is composed of the concatenation of the question and the candidate entity explanations after the initial screening, that is, composed of τ m and τ e combined, where the [CLS] label before τ e is removed. The concatenated representation of the entity and the mention is y m,e , and the score s cross (m, e) of each concatenated result is obtained by linear transformation with the linear layer W:
[0024] s cross (m, e) = y m,e W
[0025] Thus, the correct entity linking results in the Wikidata knowledge base are evaluated. If other knowledge base entities have triple link addresses linking to Wikidata knowledge base entities, this link address helps the retrieval results of the dual encoder in the dense embedding space composed of Wikidata entities and Wikipedia knowledge to correspond to other knowledge base entities, realizing the alignment of multi-knowledge base entities.
[0026] Preferably: In step 3, based on the semantic parsing graph and the linked entities of the abstract semantic representation, the shortest path from the linked entity node to the target node in the graph is retrieved, and the relationship substructure corresponding to the knowledge base relationship is extracted using the screening rules. For a single predicate in the shortest path in the graph, the "amr-unknown" node is regarded as a special placeholder, which also has an implicit relationship meaning and should be included in the relationship substructure; for a path that does not contain any predicates, the node with two non-core role edges is selected as the center of the relationship substructure. If there are multiple nodes in this path that meet this condition, the node closest to the entity is selected.
[0027] Preferably: In step 4, the candidate relationships come from all relevant relationships of the linked entities in the knowledge base. These relationships are combined with the extracted relationship substructure and expressed as:
[0028] [CLS][AMR] Relationship substructure [REL] Candidate relationship [SEP]
[0029] Where [AMR] and [REL] are special tags used to separate the relational substructure and the relational representation. Use v s,r to represent the combined relational embedding:
[0030] v s,r = red(BERT cross (τ s,r ))
[0031] Where τ s,r is the input representation of the connection between the relational representation substructure and the knowledge base relation, and BERT cross is a cross-encoder model based on BERT-base. Similar to entity linking, red(·) is selected to obtain the first output of the last layer of the model. And a linear layer W is used to score the embedding v s,r :
[0032] S cross (s,r) = v s,r W
[0033] After the initial screening, the parameters of the two encoders are initialized using the pre-trained BERT-base model, and then fine-tuned. The inputs τ q and τ r of the two encoders are respectively:
[0034] [CLS][AMR]Relational substructure[TEXT]Natural language question[SEP]
[0035] [CLS]Candidate relation[SEP]
[0036] On different encoders, the relational substructure is represented with the question embedding and the knowledge base relation embedding as:
[0037] v q = red(BERT question (τ q ))
[0038] v r = red(BERT relation (τ r ))
[0039] BERT question and BERT relation are two encoder models based on BERT-base. After encoding the inputs τ q and τ r by the encoders, red(·) is still used to obtain the output embeddings v q and v r , and then a multi-layer perceptron MLP is used to integrate and score the embeddings:
[0040] S(q, r) = MLP(v q , v r )
[0041] Finally, select the relationship in the question - relationship pair S(q, r) with the highest score as the knowledge - base relationship link
[0042] Beneficial effects: Compared with the prior art, the advantages of the present invention are as follows: The present invention does not require designing specific extraction rules based on a specific knowledge base, but proposes an entity and relationship linking method based on abstract semantic representation applicable to multiple knowledge bases; In view of different situations in the knowledge - base question - answering scenario, the present invention flexibly uses an encoder - composed model to complete entity linking and relationship linking tasks, and the method is simple and efficient; The model proposed by the present invention has strong generalization ability in scenarios with sufficient training data. For the situation of lack of training data, by introducing external knowledge and constructing a local dictionary, good linking results can still be obtained. In short, the present invention uses a highly abstract semantic parsing method and an entity - relationship representation extraction method with clear rules, can better extract entity and relationship representations in natural - language questions, and then uses various methods (such as introducing external knowledge and local dictionary, etc.) to improve the linking ability of the pre - trained language model for entities and relationships in multiple knowledge bases, thus facilitating the further development of subsequent knowledge - base question - answering tasks. Description of the Drawings
[0043] Figure 1 It is a schematic diagram of the control flow of the present invention. Detailed Embodiments
[0044] To deepen the understanding of the present invention, the solution will be introduced in detail below in conjunction with embodiments.
[0045] Embodiment 1: Refer to Figure 1 , an entity and relationship linking method based on abstract semantic representation in knowledge - base question - answering, including the following steps:
[0046] Step 1, use a semantic parsing component to obtain the semantic - logical representation of a sentence, and extract potential entity nodes in the semantic - parsing graph of the abstract semantic representation of the question as entity candidates to be linked. To extract potential entity - mention information in the question, use a string - matching method to search for each "name" node in the semantic - parsing graph, and splice all concept nodes connected to each "name" node from left to right, and use the convergence result as the entity - candidate mention waiting to be matched.
[0047] Step 2, use a dual encoder to preliminarily screen a large number of entities in the knowledge base to narrow down the number of entity candidates. Then use a cross encoder to rank and score the entities mentioned in the question to determine the knowledge base entities corresponding to the entity mentions in the question. To further improve the accuracy of linking, entities and their descriptions in Wikipedia are introduced as external knowledge.
[0048] Step 3, use a semantic parsing component to obtain the semantic logical representation of the sentence, and extract the relationship substructure in the semantic parsing graph of the abstract semantic representation according to the structural information as the relationship candidates to be linked. To extract the potential relationship representation substructure in the question, first find the shortest path between the linked entity node and the "amr-unknown" target node, and combine the relationship-related nodes, delete the irrelevant nodes and compress the path through multiple rules, so as to obtain a relatively perfect relationship representation substructure based on the semantic parsing graph.
[0049] Step 4, use a cross encoder to comprehensively score and rank the relationship substructure in the semantic parsing graph and the candidate relationship set. Subsequently, combine the local dictionary with the highly relevant candidate relationships after scoring to supplement the knowledge base relationships that may be filtered out. Encode the question and relationship candidates separately through a dual encoder, and score and select the encoded vectors to determine the knowledge base relationship corresponding to the relationship representation substructure.
[0050] Among them, the extraction of entity mentions in the semantic parsing graph of the abstract semantic representation in Step 1. Aggregate the candidate word nodes of each entity mention using the "name" node in the graph to obtain the entity mention list M Q ={m1,…,m n}。
[0051] Among them, the entity link of each mention in the entity mention list in Step 2. First, use the string matching method to determine the position of the mention in the question in the list, and re-divide the question into three parts according to the position using special tags, which are expressed as:
[0052] [CLS] Left text of mention [M s Mention [M e Right text of mention [SEP]
[0053] [CLS] and [SEP] represent the start and end positions of the sequence respectively, and divide the input sequence into multiple segments. The boundaries of the mention are marked by the tags [M s and [M e . This part is used as the representation τ m of the mention and is the input at one end of the dual encoder.
[0054] Based on the Wikidata knowledge base and Wikipedia documents, for each entity's knowledge, the description of the Wikidata entity in Wikipedia is added to the entity itself and related descriptions through concatenation. Its composition form is:
[0055] [CLS] Entity [ENT] Entity description [SEP]
[0056] [ENT] is used to label the entity in the sentence as a special tag. This part is regarded as the Wikidata knowledge base entity representation τ with appended Wikipedia knowledge e As the input at the other end of the dual encoder.
[0057] These two parts of the input sequence are encoded by two BERT-based encoders to obtain their respective vector representations:
[0058] y m = red(BERT1(τ m ))
[0059] y e = red(BERT2(τ e ))
[0060] Where BERT1 and BERT2 are two BERT-base based encoders, and red(·) is a function that reduces the vector sequence generated by BERT to a single vector. The candidate entity injected with Wikipedia information is encoded as The mention candidate in the question is y m , and the scoring function of the model for the entity and the mention:
[0061]
[0062] Then, a cross encoder is used for refined ranking. The input of the encoder is composed of the concatenation of the question sentence and the candidate entity explanations after the initial screening, that is, composed of τ m and τ e combined, where the [CLS] tag before τ e is removed. The concatenated representation of the entity and the mention is y m,e , and a linear transformation is performed on each concatenated result:
[0063] s cross (m, e) = y m,e W
[0064] Thus, the correct entity linking results in the Wikidata knowledge base are evaluated. If other knowledge base entities have triple link addresses that link to Wikidata knowledge base entities, these link addresses help the retrieval results of the dual encoder in the dense embedding space composed of Wikidata entities and Wikipedia knowledge to correspond to other knowledge base entities, achieving the alignment of multi-knowledge base entities.
[0065] Preferably: In step 3, based on the semantic parsing graph and the linked entities of the abstract semantic representation, the shortest path from the linked entity node to the target node in the graph is retrieved, and the relationship substructure corresponding to the knowledge base relationship is extracted using the screening rules. For a single predicate in the shortest path in the graph, the "amr-unknown" node is regarded as a special placeholder, which also has an implicit relationship meaning and should be included in the relationship substructure; for a path that does not contain any predicates, the node with two non-core role edges is selected as the center of the relationship substructure. If there are multiple nodes in this path that meet this condition, the node closest to the entity is selected.
[0066] Preferably: In step 4, the candidate relationships come from all relevant relationships of the linked entities in the knowledge base. These relationships are combined with the extracted relationship substructure and expressed as:
[0067] [CLS][AMR]Relationship substructure[REL]Candidate relationship[SEP]
[0068] Where [AMR] and [REL] are special tags used to separate the relationship substructure and the relationship representation. Use v s,r To represent the combined relationship embedding:
[0069] v s,r = red(BERT cross (τ s,r ))
[0070] Where τ s,r Is the input representation of the connection between the relationship representation substructure and the knowledge base relationship, and BERT cross Is a cross-encoder model based on BERT. Similar to entity linking, red(·) is selected to obtain the first output of the last layer of the model. And the embedding is scored using a linear transformation:
[0071] S cross (s,r) = v s,r W
[0072] After the initial screening, the parameters of the two encoders are initialized using the pre-trained BERT-base model, and then fine-tuned. The inputs τ q And τ r Of the two encoders are respectively:
[0073] [CLS][AMR]Relationship sub-structure[TEXT]Natural language question[SEP]
[0074] [CLS]Candidate relationship[SEP]
[0075] On different encoders, the relationship sub-structure, question embedding, and knowledge base relationship embedding are represented as:
[0076] v q = red(BERT question (τ q ))
[0077] v r = red(BERT relation (τ r ))
[0078] BERT question and BERT relation are two encoder models based on BERT-base. After encoding the input τ q and τ r , still use red(·) to obtain the output embedding v of the last layer of the model q and v r . Subsequently, use a multi-layer perceptron MLP to integrate and score the embeddings:
[0079] S(q, r) = MLP(v q , v r )
[0080] Finally, select the relationship in the question-relationship pair S(q, r) with the highest score as the knowledge base relationship link result.
[0081] Example 2: An entity and relationship linking method based on abstract semantic representation in knowledge base question answering, including the following steps:
[0082] Step 1, the AMR semantic parsing graph of the question "How many famous people are born in Long Island?" contains two entity mention nodes "Long" and "Island", and these two nodes are connected by the "name" node, indicating that these two mention nodes can be combined as an entity mention, so as to perform entity linking on this mention subsequently.
[0083] Step 2, entity mentions are initially screened by a dual encoder. By calculating and caching the entity representations of all candidate entities, 5.9 million entities in Wikidata are embedded into the same vector space. Among them, the question can be represented as "[CLS]How many famous people are born in[Ms]Long Island[Me]?[SEP]", and the entity description can be represented as "[CLS]Long Island[ENT]is a densely populated island in the southeastern region of the U.S. state of New York, part of the New York metropolitan area...". Through encoder encoding, the maximum dot product between the mention representation (i.e., the representation vector of the question) and the entity candidate representation (the vectors of all entities) is found. The question after the initial screening is then concatenated with the entity description and represented as "[CLS]How many famous people are born in[Ms]Long Island[Me]?Long Island[ENT]is a densely populated island in the southeastern region of the U.S. state of New York, part of the New York metropolitan area...[SEP]" for input to the cross encoder. After encoding, this input is transformed by a linear layer to obtain the final score, and the highest score is considered to be the knowledge base entity corresponding to the mention.
[0084] Step 3, it is necessary to find the shortest path between the linked entity node and the "amr-unknown" target node, then delete the irrelevant nodes, compress the path and extract the relation slots. For the shortest path from the entity node to the target node in the AMR graph of the question "How many famous people are born in Long Island?", which is "(amr-unknown|quant|person|ARG1|bear-02|location|state|Long_Island)", in order to obtain a general substructure representation, it is necessary to remove the perceptual labels from the slots ("bear-02" becomes "bear") and convert it into a linearized representation form through a top-down graph traversal method. Based on this rule, the slots corresponding to the entity can be obtained (bear|ARG1|person|location|state).
[0085] Step 4: First, link all the relationships obtained from the relationship slots and the linked entities, and use the model to sort each combined score. The relationship substructure of the question "How many famous people are born in Long Island?" and the candidate relationships obtained from the knowledge base entity "Long_Island" are combined into "[CLS][AMR]bear:ARG1
[0086] person:location state[REL]birth place[SEP]", and then input it into the cross-encoder model. Use the linear layer scoring to sort and obtain the filtered candidate relationship set. Expand the filtered candidate relationship set through the local dictionary, and then connect the question sentence with the relationship substructure "[CLS][AMR]bear:ARG1person:location state[TEXT]How many famous people are born in Long Island?[SEP]" as one end input of the dual encoder, and the candidate relationship set as the other end input. Calculate the final score through encoding and dot product to determine the knowledge base relationship corresponding to the relationship substructure.
[0087] It should be noted that the above embodiments are not used to limit the protection scope of the present invention. Any equivalent transformation or substitution made on the basis of the above technical solutions falls within the protection scope of the claims of the present invention.
Claims
1. An entity and relationship linking method based on abstract semantic representation in knowledge base question answering, characterized in that, The method includes the following steps: Step 1: Use a semantic parsing component to obtain the semantic logical representation of a sentence, and extract potential entity nodes in the semantic parsing graph of the problem abstract semantic representation according to the structural information as entity candidates to be linked. To extract potential entity mention information in the problem, use a string matching method to search for each "name" node in the semantic parsing graph, and splice all the concept nodes connected to each "name" node from left to right. The convergence result is used as the entity candidate mention waiting to be matched; Step 2: Use a dual encoder to preliminarily screen a large number of entities in the knowledge base to reduce the number of entity candidates, and then use a cross encoder to rank and score the entities mentioned in the problem to determine the knowledge base entities corresponding to the entity mentions in the problem. To further improve the accuracy of linking, entities and their descriptions in Wikipedia are introduced as external knowledge; Step 3: Use a semantic parsing component to obtain the semantic logical representation of a sentence, and extract the relationship substructure in the semantic parsing graph of the abstract semantic representation according to the structural information as the relationship candidate to be linked. To extract the potential relationship representation substructure in the problem, first find the shortest path between the linked entity node and the "amr-unknown" target node, and combine relationship-related nodes, delete irrelevant nodes and compress the path through multiple rules to obtain a relatively complete relationship representation substructure based on the semantic parsing graph; Step 4: Use a cross encoder to comprehensively score and rank the relationship substructure in the semantic parsing graph and the candidate relationship set, and then combine the local dictionary with the highly relevant candidate relationships after scoring to supplement the knowledge base relationships that may be filtered out. Encode the problem and the relationship candidates separately through a dual encoder, and score and select the encoded vectors to determine the knowledge base relationship corresponding to the relationship representation substructure.
2. The entity and relationship linking method based on abstract semantic representation in knowledge base question answering according to claim 1, wherein: In the extraction of entity mentions in the semantic parsing graph of the abstract semantic representation in step 1, the "name" nodes in the graph are used to aggregate the entity mention candidate word nodes, so as to obtain the entity mention list M in the question Q ={m1,…,m n}.
3. The entity and relationship linking method based on abstract semantic representation in knowledge base question answering according to claim 1, wherein: For each entity mention link in the entity mention list in Step 2, first use a string matching method to determine the position of the mention in the question sentence, and re-divide the question into three parts according to the position using special tags, which is expressed as: [CLS]Mention the left text[M s Mention[M e Mention the right text[SEP] [CLS] and [SEP] respectively represent the start and end positions of the sequence, and divide the input sequence into multiple segments. The mentioned boundaries are labeled by the tags [M s and [M e , and this part is used as the mentioned representation τ m , which is the input at one end of the dual encoder. Based on the Wikidata knowledge base and Wikipedia documents, each entity's knowledge is composed of the entity itself and relevant descriptions. By concatenating the descriptions of Wikidata entities in Wikipedia after the entities, its composition form is: [CLS] Entity [ENT] Entity description [SEP] [ENT] is used as a special tag to mark the entity in the sentence, which is regarded as the Wikidata knowledge base entity representation τ with Wikipedia knowledge e As the input of the other end of the dual encoder, these two parts of the input sequence are encoded by two BERT-based encoders to obtain their respective vector representations: y m = red(BERT1(τ m )) y e = red(BERT2(τ e )) Among them, BERT1 and BERT2 are two BERT-base-based encoders, and red(·) is a function that reduces the vector sequence generated by BERT to a vector; The candidate entities injected with Wikipedia information are encoded as The candidate mention in question is y m , the scoring functions of the model for entities and mentions: Then, a cross-encoder is used for fine ranking. The input of the encoder is composed of the concatenation of the question and the candidate entity explanations after initial screening, that is, composed of τ m and τ e Combined, where the [CLS] label before τ e is removed, and the connection representation of the entity and the mention is denoted as y m,e , and a linear transformation is performed on each concatenation result: s cross (m, e) = y m,e W Thereby, the correct entity link result in the Wikidata knowledge base is evaluated. If other knowledge base entities have triple link addresses linking to Wikidata knowledge base entities, this link address helps the retrieval result of the dual encoder in the dense embedding space composed of Wikidata entities and Wikipedia knowledge to correspond to other knowledge base entities, realizing the alignment of multiple knowledge base entities.
4. The entity and relationship linking method based on abstract semantic representation in knowledge base question answering according to claim 1, characterized in that: In step 3, based on the semantic parsing graph and the linked entities in the abstract semantic representation, retrieve the shortest path from the linked entity nodes to the target node in the graph, and use the screening rules to extract the relational sub-structures corresponding to the knowledge base relationships. For a single predicate in the shortest path in the graph, the "amr-unknown" node is regarded as a special placeholder, which also has implicit relational meaning and should be included in the relational sub-structure. For a path that does not contain any predicates, select the node with two non-core role edges as the center of the relational sub-structure. If there are multiple nodes in this path that meet this condition, select the node closest to the entity.
5. The entity and relationship linking method based on abstract semantic representation in knowledge base question answering according to claim 1, characterized in that: In step 4, the candidate relationships come from all relevant relationships of the linked entities in the knowledge base. These relationships are combined with the extracted relational sub-structures and represented as: [CLS][AMR]Relational sub-structure[REL]Candidate relationship[SEP] where [AMR] and [REL] are special tags used to separate the relational substructure and the relational representation, and use v s,r to represent the combined relational embedding: v s,r = red(BERT cross (τ s,r )) where τ s,r is the input representation of the relationship between the relationship representation substructure and the knowledge base connection, and BERT cross is a cross-encoder model based on BERT-base. Select red(·) to obtain the first output of the last layer of the model, and use linear transformation to score the embeddings: S cross (s, r) = v s,r W After the initial screening, the parameters of two encoders are initialized using the pre-trained BERT-base model, and then fine-tuned. The inputs τ q and τ r are respectively: [CLS][AMR]Relational sub-structure[TEXT]Natural language question[SEP] [CLS]Candidate relationship[SEP] On different encoders, the relational sub-structure, the question embedding, and the knowledge base relationship embedding are represented as: v q = red(BERT question (τ q )) v r = red(BERT relation (τ r )) BERT question and BERT relation are two encoder models based on BERT-base. After encoding the input τ q and τ r through the encoder, still use red(·) to obtain the output embedding v of the last layer of the model q and v r , and then use a multi-layer perceptron MLP to integrate and score the embedding: S(q, r) = MLP(v q , v r ) Finally, select the relationship in the question-relationship pair S(q, r) with the highest score as the knowledge base relationship linking result.
Citation Information
Patent Citations
Question and answer method based on knowledge map
CN107748757A
Knowledge graph questioning and answering method and device based on semantic blocks
CN111930906A