Document-level relation extraction method based on machine reading comprehension
By reconstructing the document-level relationship extraction task into machine reading comprehension tasks and using the answer extraction model of mixed pointer-sequence labeling, the problem of difficult to capture information between sentences with a large span in the prior art is solved, and multiple or zero answer extraction for documents is achieved, improving the accuracy of relationship extraction.
Patent Information
- Application Number
- CN202310986563.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-07
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-08-07
AI Technical Summary
In the document-level relationship extraction, there is a difficult time to capture inter-sentence information with a large span in the document-level relationship extraction, making it difficult to construct complete document structure information.
By reconstructing the relationship extraction task into machine reading comprehension tasks, using the question template to characterize entities and relationship types, and extracting models through mixed pointer-sequence annotated answers, extracting multiple or zero answers from the context, improving the model's inference ability.
It realizes the extraction of zero or multiple answers of documents, improves the accuracy of document-level relationship extraction, and enhances the model's reading, memory and reasoning ability of documents.
Smart Images

Figure CN117151118B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a document-level relationship extraction method based on machine reading comprehension. Background Art
[0002] Knowledge graph construction is an important task in the field of natural language processing. Relation extraction, as a key step in the automatic construction of knowledge graphs, aims to extract the relationship between entity pairs from unstructured texts through natural language processing technology and describe them in the form of triples (Subject, Relation, Object), thereby forming a large-scale knowledge Internet network.
[0003] Earlier work mainly focused on sentences, that is, extracting semantic relations between entity pairs from a single sentence to obtain relational facts. However, in real scenarios, a large number of relational facts exist in complex documents, and it is very necessary to extend the research on relation extraction from the sentence level to the document level.
[0004] Previous methods for document-level relationship extraction mainly focused on hierarchical networks and graph neural networks. The prior art discloses a document-level relationship extraction model based on transformer multi-instance and multi-task learning, which can simultaneously predict the relationship between multiple mention pairs. It also includes the use of CNN with additional character-level embedding to extract the relationship between chemicals and diseases in biomedical texts. A large-scale human-annotated document-level relationship extraction dataset DocRED. At the same time, a neural prediction model is proposed to obtain the representation of the context through BiLSTM, and then each entity pair is input into a bilinear function to predict the relationship. The prior art also includes the use of BERT as an encoder to complete the relationship extraction in stages: the first step is to predict whether two entities have a relationship, and the second step is to predict a specific relationship. A Star-BiLSTM-LAN method based on a neural network combines star transformation and bidirectional long short-term memory network to capture document-level semantic and grammatical information from different aspects. A document-level entity relationship extraction model that integrates bidirectional simple recurrent network and capsule network. The bidirectional simple recurrent network realizes the fusion representation of the relationship between multiple sentences, and the shortest dependency path is modeled. The capsule network optimizes the relationship representation of learning entity relationships in multiple dimensions such as space and direction. The hierarchical network-based relation extraction model uses neural networks at different levels to achieve hierarchical feature extraction from words, sentences to documents, thereby obtaining local and global feature information. Since these sequence models have inherent defects in the ability to capture long-range information, especially the information between sentences with a large span, it is difficult to construct complete document structure information.
[0005] In contrast, graph-based methods build graph structures through entities and their references, and explicitly learn the associations between entities through graph propagation, which can better establish connections between entities at the document level, thereby realizing document-level relational reasoning. The prior art also discloses the first introduction of a document graph for document-level relation extraction, exploring cross-sentence extraction and cross-sentence extraction. In order to capture text structure information, a unified method is proposed to integrate graph structure LSTMs of different sequence models in combination with various intra-sentence and inter-sentence dependencies. Based on Peng, a graph state LSTM model is proposed to model a graph as a whole. State transitions can be completed periodically on the graph, and information exchange of word states can be achieved through dependency edges and discourse edges. Christopoulou deviates from the existing graph model and proposes an edge-oriented graph neural model for document-level relation extraction, learning edge representations to predict relationships between entities in text, rather than node representations. By further refining the graph iteration strategy, a document-level graph is constructed in an end-to-end manner for reasoning without relying on common references or rules. By constructing a heterogeneous graph and using R-GCN for feature propagation, local information and overall information are spliced together as the representation of entity pairs. The prior art also proposed a dual-graph network for graph aggregation and reasoning. After the document obtains word representations with contextual information through the encoder, a heterogeneous graph is constructed to capture the complex interactions between different mentions. This solves the problem that the same entity may be in different sentences and entities of the same relationship may be in different sentences. The prior art pays more attention to related entity pairs, reconstructs the path, and uses a graph attention network to encode each node of the heterogeneous graph. Summary of the invention
[0006] To solve the above technical problems, the present invention proposes a document-level relationship extraction method based on machine reading comprehension, which is combined with the machine reading comprehension task to incorporate more information into the question query, making the model's document reading, memory and reasoning process more mature and natural. A new extraction model is designed to improve the model's reasoning ability while realizing the extraction of zero answers or multiple answers to the document.
[0007] To achieve the above object, the present invention provides a document-level relationship extraction method based on machine reading comprehension, comprising:
[0008] Constructing a question-answer pair reading comprehension dataset, inputting the dataset into an answer extraction model based on hybrid pointer-sequence labeling, vectorizing the input question-answer pair at the input layer, and obtaining first vector data;
[0009] The first vector data is decoded through a pointer network based on a multi-span extraction model, the pointer network is optimized, and the answer in the data set is identified based on the optimized pointer network.
[0010] Preferably, constructing the question-answer pair reading comprehension dataset includes:
[0011] Constructing the question-answer pair reading comprehension dataset based on questions and documents;
[0012] The process of constructing the problem includes:
[0013] According to the dataset annotation guide as the basis for describing the label categories, annotators ask questions to the documents, and a crowdsourcing method is used to collect and verify questions for each relationship, wherein the number of questions generated by each relationship is greater than 1, and each question contains the subject and relationship in the entity pair.
[0014] Preferably, after constructing the question-answer pair reading comprehension dataset, the quality of the generated questions is verified, including:
[0015] The verifier is provided with a document and the corresponding question of the document. The validity of the question is verified by judging whether the verifier can answer the question correctly. There are no less than 5 verifiers for each document, and the probability that the verifier correctly answers the question is no less than 3 / 5. The verified question is then judged to be valid. Otherwise, the generated questions and answers will be considered to be reasonable. If not, they will be adjusted. Thus, the construction of the question-answer pair reading comprehension dataset is completed.
[0016] Preferably, the answer extraction model based on hybrid pointer-sequence labeling is used to vectorize the input question-answer pairs, and adopts the BERT pre-trained model as the encoder to learn deep representation, and realizes the extraction of multiple answers and zero answers through the pointer network and sequence labeling method.
[0017] Preferably, vectorizing the input question-answer pair at the input layer includes:
[0018] The input document and question are vectorized respectively, and the paragraphs in the question and document are tokenized using WordPiece segmentation. y ={Tok1, Tok2, ..., Tok n}, document passage = {Tok1, Tok2, ..., Tok m};
[0019] The question q y and paragraphs Connection, input embedding Each Token contains wordpiece embedding, position embedding, segment embedding, D is the hidden size, and T is the sequence length.
[0020] Preferably, the method of inputting the H0 as an input vector into the bidirectional encoder is:
[0021]
[0022] Among them, H i represents the context representation of the i-th layer input paragraph, H i =(h1, ..., h n ), L represents the number of Transformer blocks.
[0023] Preferably, the pointer network includes a feedforward neural network unit, which is used to calculate the score of each Token and determine whether the Token is the starting position or the ending position of the extraction sequence according to the score.
[0024] Preferably, the calculation method of the Token is:
[0025]
[0026]
[0027]
[0028] Among them, f start (h i ), f end (h i ) is the function of the feedforward neural network unit, p start Formula (2) and formula (3) calculate the probability that Token is the start and end position of the sequence respectively, and formula (4) calculates the match of the start and end positions of the pointer.
[0029] Preferably, optimizing the pointer network includes:
[0030] The same binary classifier is used to predict the start and end positions of the answer respectively, and each token is assigned a binary label 0 or 1, where 1 indicates the start or end position of the answer and the rest of the positions are marked as 0.
[0031] Preferably, the answer span for a given paragraph is identified and labeled based on a maximum likelihood function:
[0032]
[0033] Where L represents the length of a given paragraph, I{z}=1 means z is true, otherwise it is 0, A binary marker indicating the start or end position of the answer to the jth token, θ = {W start , β start , W end , β end}, a represents the answer.
[0034] Compared with the prior art, the present invention has the following advantages and technical effects:
[0035] The present invention reconstructs the relationship extraction task into a machine reading comprehension task, uses question templates to characterize the types of entities and relationships, extracts answers from the context by given questions, and realizes the joint modeling of document structure and context reasoning; proposes a new hybrid pointer-sequence annotation extraction model, which improves the model's reasoning ability while realizing the extraction of zero or multiple answers to documents, ultimately improving the accuracy of document-level relationship extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0037] Figure 1 This is an overall architecture diagram of a document-level relationship extraction method based on machine reading comprehension in an embodiment of the present invention. DETAILED DESCRIPTION
[0038] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0039] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] The present invention proposes a document-level relation extraction method based on machine reading comprehension, such as Figure 1 ,include:
[0041] Constructing a question-answer pair reading comprehension dataset, inputting the dataset into an answer extraction model based on hybrid pointer-sequence labeling, vectorizing the input question-answer pair, and obtaining first vector data;
[0042] The first vector data is decoded through a pointer network based on a multi-span extraction model, the pointer network is optimized, and the answer in the data set is identified based on the optimized pointer network.
[0043] The answer extraction model based on hybrid pointer-sequence labeling includes:
[0044] Input layer: input the data set into the answer extraction model based on hybrid pointer-sequence labeling, vectorize the input question-answer pair, and obtain first vector data;
[0045] Encoding layer: Use the BERT pre-trained model as the encoder to learn deep representation.
[0046] Multi-span extraction layer: After receiving the output sequence from the encoder BERT, the model uses a pointer network to calculate the score of each Token and uses two identical binary classifiers to predict the start and end positions of the answer respectively to extract multiple answers. It also uses a 0 / 1 sequence labeling method to additionally mark each Token to determine the start and end positions of each answer.
[0047] from Figure 1 As can be seen on the left side, questions and documents serve as the input part of the model. Since the questions encode knowledge related to entity and relationship types, the quality of question generation is very important, so question construction is performed first. Based on the dataset annotation guide as the basis for describing label categories, the present invention adopts a crowdsourcing method to collect and verify questions for each relationship. Annotators must fully understand the semantics of finding relationships in documents and ask creative questions. However, due to the differences in each annotator's understanding of the text and different mentions of the same entity, this embodiment requires that an average of 2-3 questions be generated for each relationship during the questioning process, and each question must contain the subject and relationship in the entity pair. At the same time, for this problem, the document Multiple valid answers in compose the answer set A.
[0048] Examples of question construction are shown in Table 1. Each instance contains a relation, a question, and the corresponding document. and an answer set A. The question explicitly mentions an entity e, which also appears in For brevity, the answers are underlined rather than listed separately.
[0049] Table 1
[0050]
[0051] In the verification stage of the quality of generated questions, a document and a corresponding question are provided to the verifier (not involved in annotation). On average, 5 people are assigned to verify each document, and the validity of the question is verified by judging whether they can answer correctly. If the answers of more than (including) 3 people out of 5 people match any answer in the valid answer set A, the question is valid. Otherwise, whether the generation of the question and the setting of the answer are reasonable will be considered. If not, corresponding adjustments will be made. This embodiment ultimately discards about 2 / 5 of the invalid examples, forming a reading comprehension data set containing nearly 30,000 question-answer pairs.
[0052] As the input part of the model, the present invention needs to use the WordPiece word segmentation algorithm to perform subword tagging on the input questions and paragraphs, and use the [CLS] and [SEP] special tags to connect the questions and paragraphs into a string and perform vectorized representation.
[0053] Tokenize the questions and paragraphs using WordPiece segmentation, where question q y ={Tok1, Tok2, ..., Tok n}, document passage = {Tok1, Tok2, ..., Tok m}.
[0054] It not only achieves a good balance between the efficiency of the “character” level and the “word” level models, but also naturally handles the translation of rare words, improving the overall accuracy. y and paragraphs Concatenate to form a string {[CLS]Tok1, ..., Tok n [SEP]Tok1,…,Tok m [SEP]}, where [CLS] and [SEP] are special tags.
[0055] Input embedding Each Token contains wordpiece embedding, position embedding, segment embedding, D is the hidden size, and T is the sequence length.
[0056] Then H0 is fed into Bert as the input vector:
[0057]
[0058] Among them, H i represents the context representation of the i-th layer input paragraph, H i =(h1, ..., h n), L represents the number of Transformer blocks.
[0059] As can be seen from the right side of the model diagram, the multi-span extraction module directly decodes the vector H generated by the encoding layer through the pointer network i , to identify all possible answers in the input paragraph. The pointer network consists of two feedforward neural networks, both of which are used to calculate the score of each Token, and decide whether the Token is the beginning or end of the extraction sequence based on the score.
[0060] The calculation formula for each Token is as follows:
[0061]
[0062]
[0063]
[0064] where f start (h i ), f end (h i ) is the function of the feedforward neural network, p start ,
[0065] Formulas (2) and (3) calculate the probability that Token is the start and end position of the sequence respectively. Formula (4) calculates the match of the start and end positions of the pointer.
[0066] The pointer network performs two softmax calculations on the sequence to extract the start and end positions of the answer. However, since the softmax function is used for multi-classification tasks, the predicted categories are mutually exclusive and only one answer can be extracted. However, for the text corpus itself, there are cases where an entity has the same relationship with multiple entities or no relationship. Therefore, the present invention converts the relationship extraction task into extracting answers from the context, which corresponds to the problem of multiple answer and zero answer extraction in machine reading comprehension. In order to achieve the relationship extraction of multiple answers and zero answers, improvements are made to the pointer network, and two identical binary classifiers are used to predict the start and end positions of the answer respectively. By assigning a binary label 0 or 1 to each token, where 1 represents the start or end position of the answer, and the remaining positions are marked as 0.
[0067] The detailed operation of this method on each token is as follows:
[0068]
[0069]
[0070]
[0071]
[0072] Among them, α(.) and W(.) represent learning matrices, β(.) is the bias, σ represents the sigmoid function for binary classification, and h j Represents the jth token in the encoded sequence. and Respectively represent the probability of identifying the jth token in the input sequence as the starting position and the ending position of the answer. If the probability exceeds the threshold, the corresponding mark is 1, otherwise the mark 0 is used.
[0073] The model is optimized using the maximum likelihood function to identify and tag a given paragraph. The answer span:
[0074]
[0075] Among them, L represents The length of I{z}=1 means z is true, otherwise it is 0; A binary mark indicating the start or end position of the answer to the jth token; θ = {W start , β start , W end , β end}; a represents the answer.
[0076] In paragraph In the text, a question may have multiple answers, and an answer may consist of multiple words. Inspired by extracting a variable number of spans from the input text in sequence tagging, in order to achieve the extraction of multiple answers, each token needs to be additionally marked to determine the start and end positions of each answer. Traditional sequence tagging methods tend to bring too much redundant information to the tags. By designing the simplest 0 / 1 tagging method, a binary tag 0 / 1 is assigned to each token. This means that multiple scattered tokens marked as 1 and tokens marked as 0 may be generated. Since there are almost no overlapping entities in the answers, the proximity principle is selected, that is, the two nearest tokens marked as 1 constitute an answer. Of course, if the start and end positions of any answer span in a given paragraph are correctly detected, then this matching strategy can maintain the integrity of any answer span.
[0077] In order to demonstrate the advantages of using this method in document-level relation extraction, the present invention will conduct experiments on the CDR (Chemical-Disease Relations) and GDA (Gene-Disease Relations) datasets. These two datasets are the most commonly used relation extraction datasets currently available. The specific statistical results of the data used are shown in Table 2.
[0078] Table 2
[0079]
[0080] In order to more clearly demonstrate the advantages of this method, the present invention is compared with the following mainstream baseline methods: 1) BRANs, 2) CNN+CNNchar, 3) GCNN, 4) EoG, 5) LSR, 6) ATLOP, 7) path+U-net.
[0081] To ensure consistency with related work, this example uses the F1 value to evaluate model performance, with a maximum value of 1 and a minimum value of 0. The higher the value, the better. In order to measure the model's ability to reason between sentences, intra-F1 and inter-F1 are introduced to evaluate the relationship within a sentence and the relationship between sentences.
[0082] Experimental setup: This example builds its own model on two public BERT versions: Multi-Span BERT-Base and Multi-Span SciBERT The specific settings of the experimental parameters are shown in Table 3.
[0083] Table 3
[0084]
[0085] Table 4 lists the comparison results of different models on the CDR dataset.
[0086] Table 4
[0087]
[0088] Table 5 lists the comparison results of different models on the GDA dataset.
[0089] Table 5
[0090]
[0091] From Tables 4 and 5, we can see that the proposed method achieves the best results on different datasets.
[0092] In addition, the present invention also designs an ablation experiment to observe the zero-shot learning ability of the model. Zero-shot learning hopes that the model can extract relations that have not appeared in the training set. For a reading comprehension model trained with N labeled relation types, when encountering a relation type that has never been encountered during the training process in the test set, there is no need to label data for the new relation. It is only necessary to define a question template for the relation to complete the extraction of the new relation. To verify the superiority of the model, the present invention trains the model on the CDR dataset and tests it using the test set of the GDA dataset. Table 6 lists the results of different models under zero-shot learning.
[0093] Table 6
[0094]
[0095] From Table 6, we can see that our method outperforms the most widely used methods, EoG and LSR, by 25.9% and 32.5% in zero-shot learning ability, respectively.
[0096] The present invention reconstructs the relationship extraction task into a machine reading comprehension task. Each pair of entities and relationships is characterized by a question template, and the extraction of entities and relationships is converted into identifying answers from the context. The structural information contained in the question provides great help for contextual reasoning, which not only realizes the joint modeling of document structure and contextual reasoning, but also makes the model's reading, memory and reasoning process of documents more mature and natural. For multi-label and multi-entity questions in documents, the present invention proposes an answer extraction model based on hybrid pointer-sequence annotation, which improves the model's reasoning ability while realizing the extraction of zero answers or multiple answers to documents.
[0097] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A document-level relationship extraction method based on machine reading comprehension, characterized in that: include: Constructing a question-answer pair reading comprehension dataset, inputting the dataset into an answer extraction model based on hybrid pointer-sequence labeling, vectorizing the input question-answer pair at the input layer, and obtaining first vector data; Decoding the first vector data through a pointer network based on a multi-span extraction model, optimizing the pointer network, and identifying an answer in the data set based on the optimized pointer network; The step of optimizing the pointer network includes: Use the same binary classifier to predict the start and end positions of the answer respectively, assigning a binary label 0 or 1 to each token, where 1 indicates the start or end position of the answer and the rest of the positions are marked as 0; Based on the maximum likelihood function, identify and tag the given paragraph The answer spans: Where L represents the length of a given paragraph, I{z}=1 means z is true, otherwise it is 0, A binary marker indicating the start or end position of the answer to the jth token, θ = {W start ,β start ,W end ,β end }, a represents the answer.
2. The document-level relationship extraction method based on machine reading comprehension according to claim 1 is characterized in that: Construct the question-answer pair reading comprehension dataset, including: Constructing the question-answer pair reading comprehension dataset based on questions and documents; The process of constructing the problem includes: According to the dataset annotation guide as the basis for describing the label categories, annotators ask questions to the documents, and a crowdsourcing method is used to collect and verify questions for each relationship, wherein the number of questions generated by each relationship is greater than 1, and each question contains the subject and relationship in the entity pair.
3. The document-level relationship extraction method based on machine reading comprehension according to claim 2 is characterized in that: After constructing the question-answer pair reading comprehension dataset, the quality of the generated questions is verified, including: The verifier is provided with a document and the corresponding question of the document. The validity of the question is verified by judging whether the verifier can answer the question correctly. There are no less than 5 verifiers for each document, and the probability that the verifier correctly answers the question is no less than 3 / 5. The verified question is then judged to be valid. Otherwise, the generated questions and answers will be considered to be reasonable. If not, they will be adjusted. Thus, the construction of the question-answer pair reading comprehension dataset is completed.
4. The document-level relationship extraction method based on machine reading comprehension according to claim 1 is characterized in that: The hybrid pointer-sequence labeling-based answer extraction model is used to vectorize the input question-answer pairs, and uses the BERT pre-trained model as an encoder to learn deep representations, and realizes the extraction of multiple answers and zero answers through pointer networks and sequence labeling methods.
5. The document-level relationship extraction method based on machine reading comprehension according to claim 1, characterized in that: The input question-answer pair is vectorized at the input layer, including: The input document and question are vectorized respectively, and the paragraphs in the question and document are tokenized using WordPiece segmentation. y ={Tok1,Tok2,...,Tok n }, document passage = {Tok1, Tok2, ..., Tok m }; The question q y and paragraphs Connection, input embedding Each Token of contains wordpiece embedding, position embedding, segment embedding, D is the hidden size, and T is the sequence length.
6. The document-level relationship extraction method based on machine reading comprehension according to claim 5 is characterized in that: The method of inputting H0 as the input vector into the bidirectional encoder is: Among them, H i represents the context representation of the i-th layer input paragraph, H i =(h1,...,h n ), L represents the number of Transformer blocks.
7. The document-level relationship extraction method based on machine reading comprehension according to claim 1, characterized in that: The pointer network includes a feedforward neural network unit, which is used to calculate the score of each Token and determine whether the Token is the starting position or the ending position of the extraction sequence according to the score.
8. The document-level relationship extraction method based on machine reading comprehension according to claim 7 is characterized in that: The calculation method of the Token is: Among them, f start (h i ), f end (h i ) is the function of the feedforward neural network unit, Formula (2) and formula (3) calculate the probability that Token is the start and end position of the sequence respectively, and formula (4) calculates the match of the start and end positions of the pointer.