A Method for Extracting Relational Triples Based on a Cascade Binary Annotation Framework

By adopting a cascaded binary annotation framework and BERT pre-trained model in entity relationship joint extraction, the problems of sample imbalance and context information learning in the existing technology are solved, and efficient relationship triple extraction is achieved, achieving high recall and accuracy.

CN114297408BActive Publication Date: 2025-05-30ZHONGKE GUOLI (ZHENJIANG) INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111658767.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-05-30
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The prior art has sample imbalance in the joint extraction of entity relationships, the inability to effectively learn sentence context information, and the difficulty in identifying overlapping relationships, resulting in low efficiency and accuracy.

Method used

The cascading binary annotation framework is adopted to model the relationship as a Subject entity mapped to the Object entity in the sentence, and the BERT pre-trained model is used as the Encoder end to solve the sample imbalance problem through multi-label binary annotation, and rich context information is integrated to improve the extraction efficiency and accuracy.

Benefits of technology

It effectively solves the problems of sample imbalance and context information learning, and improves the recall and accuracy of relational triple extraction, reaching 89.9% recall and 91.3% accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114297408B_ABST
    Figure CN114297408B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for extracting relational triples based on a cascaded binary annotation framework, comprising the following steps: obtaining the semantic feature representation H in the sentence after processing the extracted sentence through a BERT pre-trained model N encoding vector; decoding the output H N encoding vector, identifying the start and end position labels of the Subject entity, so as to obtain the feature vector matrix V of all possible Subject entities and their corresponding Tokens in the sentence sub ; taking the average of the vectors corresponding to the Tokens of the feature vector matrix V sub to obtain the Subject entity feature vector V K sub , fusing the output H N decoding vector to obtain the fused vector V. According to the fused vector V and in combination with a specific set of relationship sets, identify the start and end position labels of the Object entity corresponding to the relationship, so as to identify all the relationships and Object entities related to the Subject entity, and finally extract the relational triples
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to natural language processing technology in the field of computers, and particularly to a method for extracting relational triples based on a cascaded binary annotation framework. Background Art

[0002] With the rapid development of information processing technology and the Internet, the amount of data processed by people has increased sharply. How to quickly and efficiently extract entity and relationship information between entities from these texts in the open domain has become an important problem that urgently needs to be solved. Entity relationship extraction is a core task for information extraction from unstructured data. Its main goal is to extract entities from text and identify the semantic relationships between entity pairs, and it is widely used in knowledge graph construction, information retrieval, dialogue generation, question answering systems, and other aspects.

[0003] Entity relationship extraction is an important basic task in natural language processing. The traditional method is to use a pipeline model, that is, entity relationship extraction is divided into two tasks. First, entity recognition is performed, and then relationship extraction is performed. These two tasks are independent, ignoring the internal connection and dependency relationship between the two tasks. Errors in entity recognition will affect the performance of the next relationship extraction and cause error propagation and accumulation. Joint entity relationship extraction is a key issue in entity relationship extraction. Existing joint entity relationship extraction methods adopt a feature-structured system and an end-to-end model (Encoder-Decoder). The feature-structured system method is relatively complex to process, requiring a large number of complex feature engineering and NLP toolkits. Complex feature engineering will increase the labor cost, and over-reliance on NLP toolkits will cause error propagation and accumulation. The end-to-end model is based on single-label annotation. The Encoder and Decoder ends use LSTM or a variant of the LSTM neural network model for encoding and decoding, thus transforming the joint extraction problem into an annotation problem (a machine learning problem), and realizing the assignment of relationships to discrete labels of entity pairs, that is, f(s,o)=r. Although the extraction problem is transformed into a machine learning problem, in most of the extracted entity pairs, effective relationships cannot be formed, resulting in a large number of negative examples and causing sample imbalance; when the same entity pair participates in multiple effective relationships, the classifier will be confused, so overlapping relationships cannot be recognized; using the LSTM neural network also cannot learn richer context information in the sentence, resulting in low efficiency and accuracy of joint entity relationship extraction. Therefore, in this case, a method is studied to extract relational triples through an end-to-end algorithm according to the cascaded binary annotation framework.

[0004] The following problems need to be solved in this method:

[0005] (1) The single-label annotation model assigns discrete labels of relationships to entity pairs, generating a large number of negative examples and resulting in sample imbalance.

[0006] (2) Using an LSTM neural network cannot learn richer context information in the sentence, leading to low efficiency and accuracy in relation triple extraction.

[0007] (3) When the same entity pair participates in multiple valid relationships, the classifier will be confused, resulting in the inability to recognize overlapping relationships. Summary of the Invention

[0008] Aiming at the problems existing in the prior art, the present invention provides a method for modeling relationships by mapping the Subject entity to the Object entity in the sentence, i.e., fr(s)=o, which solves the problem of relationship overlap. Moreover, it abandons the discrete labels of relationships assigned to entity pairs by the single-label annotation model and uses multi-label binary annotation to label the Start and End positions of entities, solving the problem of sample imbalance. In particular, the BERT pre-trained model used in the Encoder end can learn richer context information, effectively improving the efficiency and accuracy of relation triple extraction, which is a relation triple extraction method based on a cascaded binary annotation framework.

[0009] The object of the present invention is achieved by the following technical solutions.

[0010] A relation triple extraction method based on a cascaded binary annotation framework, comprising the following steps:

[0011] Step 1): The Encoder end of the cascaded binary annotation framework uses the BERT pre-trained model to obtain the semantic feature representation H N encoding vector of the extracted sentence after being processed by the BERT pre-trained model;

[0012] Step 2): Decode the output H N encoding vector, identify the Start and End position labels of the Subject entity, so as to obtain the feature vector matrix V of all possible Subject entities in the sentence and their corresponding Tokens; sub ;

[0013] Step 3): Take the average of the vectors corresponding to the tokens in the feature vector matrix V sub to obtain the Subject entity feature vector V K sub , fuse the output H N decoding vector to obtain the fused vector V.

[0014] Step 4): Based on the fused vector V, combined with a specific set of relationship sets, identify the Start and End position tags of the Object entities corresponding to the relationships, so as to identify all the relationships and Object entities related to the Subject entity, and finally extract the relationship triples.

[0015] The specific steps of the described step 1) include:

[0016] Step 11): The input is a text sentence, and the word embedding representation and position embedding representation of the input are obtained through embedding lookup.

[0017] Step 12): Input all the obtained embedding layer representations into the BERT pre-trained model, that is, through 12 layers of encoders. In each layer of encoder, the self-attention mechanism is used to learn information, and then the previously learned information is processed through a fully connected layer and passed to the next layer of encoder; BERT will add a [CLS] flag at the beginning of the sentence, and the [CLS] of the last layer is used as the semantic information of the entire sequence or the entire text, so as to obtain the semantic encoding vector H. N 。

[0018] The specific steps of the described step 2) include:

[0019] Step 21): Decode the output semantic encoding vector H N , and extract the representation of each Token from it;

[0020] Step 22): Use two identical binary label systems to assign a binary mark (0 / 1) for the Start and End positions to each Token, and obtain the binary marks (0 / 1) of the Start and End positions of all Tokens in the sentence;

[0021] Step 23): Adopt the principle of proximity of Start-End positions to identify all possible Subject entities and the decoding vector matrix V corresponding to all Tokens included in them sub 。

[0022] The specific steps of the described step 3) include:

[0023] Step 31): For the decoding vector matrix V of the Tokens corresponding to the Subject entity sub , take the average of all the vectors in the matrix to get V K sub ;

[0024] Step 32): The average vector V obtained in C1 K sub , fuse the semantic encoding vector H N , and obtain the fused vector V.

[0025] The specific steps of step 4) include:

[0026] Step 41): According to the fusion vector V, in combination with a set of specific relationship sets, two identical binary tag systems are used to assign a binary mark (0 / 1) of the Start and End positions to each Token.

[0027] Step 42): Adopting the principle of proximity of Start-End positions, identify all Object entities of specific relationships that may be related to the Subject entity, so as to extract relationship triples.

[0028] Compared with the prior art, the advantages of the present invention are as follows: Through operation, the present invention can effectively extract relationship triples from existing sentences. We verified the relationship triples extracted by the method of the present invention through experiments. The experiments show that the recall rate of the relationship triples can reach 89.9%, and the accuracy rate is 91.3%, thus verifying the effectiveness and rationality of the present invention. Description of the Drawings

[0029] Figure 1 It is a schematic diagram of the module of the present invention.

[0030] Figure 2 It is a flow chart of the present invention. Detailed Embodiment

[0031] The present invention will be described in detail below in conjunction with the drawings in the specification and specific embodiments.

[0032] As Figure 1 shown, the present invention is a method for extracting relationship triples based on a cascaded binary framework, including the following modules:

[0033] Module A: The Encoder end of the cascaded binary annotation framework uses the BERT pre-trained model to replace the traditional LSTM to obtain the semantic feature representation H N Coding vector.

[0034] Module B: Decode the H N coding vector output by Module A, identify the Start and End position tags of the Subject entity, so as to obtain the feature vector matrix V of all possible Subject entities and their corresponding Tokens in the sentence sub .

[0035] Module C: The feature vector matrix V of the Tokens corresponding to the Subject entity output by Module B sub , take the average of the Token feature vectors of the matrix V sub to obtain the Subject entity feature vector V Ksub , fuse the H output by Module A N decoding vector and input it into Module D.

[0036] Module D: Based on the fused vector V of Module C and combined with a specific set of relationship sets, identify the Start and End position tags of the Object entity corresponding to the relationship, thereby identifying all relationships and Object entities related to the Subject entity, that is, the relationship triple (s, r, o).

[0037] As Figure 2 shown, the workflow of the present invention includes the following steps:

[0038] Step 1) Input the sample, that is, the extracted sentence, into Module A. After being processed by BERT, obtain the output, that is,

[0039] the semantic feature representation H in the sentence N encoding vector.

[0040] Step 2) Decode the semantic feature representation H obtained in Step 1 N encoding vector, input it into Module B, and identify the feature vector matrix V of the Subject entity in the sentence and its corresponding token sub .

[0041] Step 3) Input the Token feature vector matrix V in Step 2 sub and the semantic feature representation H in the sentence output in Step 1 N encoding vector into Module C to obtain the fused vector V.

[0042] Step 4) Input the fused vector V and a specific set of relationship sets into Module D, and finally obtain all relationships and Object entities related to the Subject entity output in Step 2, that is, the relationship triple (s, r, o).

[0043] Next, for the above steps, combined with the corresponding legends, the following will give a detailed elaboration.

[0044] Module A: Use BERT at the Encoder end of the cascaded binary annotation framework to obtain the semantic feature representation H N encoding vector in the sentence.

[0045] For the convenience of explanation, take the sample [A is a French painter born in 1824] as an example.

[0046] The constructed input format is: [CLS] sentence [SEP]

[0047] Find the index id corresponding to the word (token) in the sentence in the vocabulary. The maximum length is 128, and all those less than 128 are padded to 128.

[0048] Step A1) Perform an embedding lookup on the index id through the embedding information of the entire vocabulary to obtain all the word embeddings of the sample;

[0049] Step A2) Process the output of the previous step. Add the word embeddings and the position embeddings together and input them into a 12-layer encoder. Each layer of the encoder uses the self-attention mechanism to learn information, and then passes the information learned previously through a fully connected layer to the next layer of the encoder, layer by layer to the last layer. The [CLS] of the last layer

[0050] serves as the semantic information of the entire sequence or the entire sentence. Since [CLS] is just a label without explicit semantic information, compared with other input words, it more fairly integrates the semantic information of each input word, so choosing to use [CLS] also better represents the semantics of the entire sentence, and finally obtains the semantic feature representation H N Encoded vector.

[0051] Module B decodes the H output by Module A N Encoded vector, identifies the start and end position labels of the Subject entity, and finally obtains the vector matrix V of all Subject entities and their corresponding Tokens according to the start-end proximity principle sub 。

[0052] Step B1) Decode the semantic feature representation H N Encoded vector to obtain the Token vector representation of each word;

[0053] Step B2) According to the Token vector representation of each word, use two identical binary annotation frameworks to identify the start and end position labels (0 / 1) of the Subject entity;

[0054] Step B3) According to the start and end position labels of the Subject obtained in Step B2, use the start-end proximity principle to obtain all Subject entities and output the vector matrix V of their corresponding Tokens sub 。

[0055] Module C fuses the H output by Module A N Encoded vector and the Subject entity feature vector V K sub to obtain the fused vector V.

[0056] Step C1) Take the average of the Token vector matrix V corresponding to the Subject entity output by Module B to obtain the Subject entity feature vector V sub to obtain the Subject entity feature vector V K sub ;

[0057] Step C2) Fuse the Subject entity feature vector V obtained in Step C1 K sub and the encoding vector H output by Module A N to obtain the fused vector V

[0058] Module D, based on the fused vector V and in combination with a specific set of relationships, identifies the Start and End position labels (0 / 1) of the Object entity corresponding to the relationship, and based on the Start-End proximity principle, identifies the Object entity corresponding to the relationship, obtaining the relationship and Object entity associated with the Subject entity, i.e., the relationship triple (s, r, o).

[0059] Step D1) Using the fused vector V output by Module C and a specific set of relationships, two identical binary frameworks are used to identify the Start and End position labels (0 / 1) of the Object entity corresponding to each relationship, and finally, the Start and End position labels (0 / 1) of the Object entity corresponding to all relationships are obtained;

[0060] Step D2) Based on the Start and End position labels (0 / 1) of the Object entity corresponding to all relationships output in Step D1, the Start-End proximity principle is used to obtain all the Object entities corresponding to the relationships, and the relationship and Object entity associated with the Subject entity are output, i.e., the relationship triple (s, r, o).

[0061] Experimental results

[0062] By running the algorithm of the present invention, the relationship triple extraction of existing sentences can be effectively performed

[0063] The following table is a sample of relationship triple extraction for some sentences

[0064] Sentence Relational triple A is a French painter born in 1824 (A, Date of birth, 1824) (A, Nationality, France) "E" is a book published by F Publishing House in 2007, and the author is B (E, Publishing house, F Publishing House) (E, Author, B) "G" is included in the music album "H" of singer C, written and composed by D, and was first released on October 15, 2010 (G, The album, H) (G, Singer, C) (G, Lyricist, D) (G, Composer, D) On September 9, 2009, I was officially opened for registration (I, Date of establishment, September 9, 2009)

[0065] We verified the relationship triples extracted by the method of the present invention through experiments. The experiments show that the recall rate of the relationship triples can reach 89.9% and the accuracy rate is 91.3%, thus verifying the effectiveness and rationality of the present invention

Claims

1. A method for extracting relationship triples based on a cascaded binary annotation framework, characterized in that: It includes the following steps: Step 1): The Encoder end of the cascaded binary annotation framework uses the BERT pre-trained model. After processing the extracted sentences through the BERT pre-trained model, the semantic feature representation H of the sentences is obtained N Encoded vector; Step 2): Decode the output H N encoding vector, identify the Start and End position tags of the Subject entity, so as to obtain the feature vector matrix V of all possible Subject entities and their corresponding Tokens in the sentence sub ; Step 3): Take the average of the vectors corresponding to the Tokens of the feature vector matrix V sub to obtain the Subject entity feature vector V K sub , fuse the output H N decoding vector to obtain the fused vector V; Step 4): According to the fused vector V, combined with a specific set of relationship sets, identify the Start and End position tags of the Object entities corresponding to the relationships, so as to identify all the relationships and Object entities related to the Subject entity, and finally extract the relationship triples.

2. The method for extracting relationship triples based on a cascaded binary annotation framework according to claim 1, characterized in that: The specific steps of the said step 1) include: Step 11): The input is a text sentence, and the word embedding representation and position embedding representation of the input are obtained through embedding lookup. Step 12) Input all the obtained embedding layer representations into the BERT pre-trained model together, that is, through 12 layers of encoders. The self-attention mechanism is adopted at each layer of the encoder to learn information, and then the information learned previously is processed by a fully connected layer and passed to the next layer of the encoder; BERT adds a [CLS] flag at the beginning of the sentence, and the [CLS] of the last layer serves as the semantic information of the entire sequence or the entire text, thereby obtaining the semantic encoding vector H N .

3. The method for extracting relationship triples based on a cascaded binary annotation framework according to claim 1, characterized in that: The specific steps of the said step 2) include: Step 21) Decode the output semantic encoding vector H N , and extract the representation of each Token from it; Step 22): Use two identical binary label systems to assign a binary mark for the Start and End positions to each Token, and obtain the binary marks for the Start and End positions of all Tokens in the sentence. Step 23) Using the principle of proximity of Start-End positions, identify all possible Subject entities and the decoding vector matrix V corresponding to all Tokens contained therein sub .

4. The method for extracting relationship triples based on a cascaded binary annotation framework according to claim 1, characterized in that: The specific steps of the said step 3) include: Step 31) Decoding vector matrix V of the Token corresponding to the Subject entity sub , take the average of all vectors in the matrix to obtain V K sub ; Step 32) Average vector V obtained by C1 K sub , fuse semantic encoding vector H N , to obtain the fused vector V.

5. The method for extracting relationship triples based on a cascaded binary annotation framework according to claim 1, characterized in that: The specific steps of the said step 4) include: Step 41): According to the fused vector V, combined with a specific set of relationship sets, use two identical binary label systems to assign a binary mark for the Start and End positions to each Token; Step 42): Adopt the principle of proximity of Start-End positions to identify all possible Object entities of specific relationships related to the Subject entity, so as to extract the relationship triples.

Citation Information

Patent Citations

  • An entity relationship joint extraction method and system based on an attention mechanism

    CN109902145A

  • Model fusion triad representation learning system and method based on deep learning

    CN111581395A