A method for multi-level feature fusion extraction using multi-head attention

Through the multi-head attention mechanism and multi-level feature fusion method, the redundancy and insufficient coverage of the relationship triple extraction results in OpenIE6 system are solved, and a more compact and practical relationship triple extraction is achieved, which improves the downstream task application effect of open domain information extraction.

CN117312487BActive Publication Date: 2025-07-25NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311277971.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-07-25
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

The existing open domain information extraction system OpenIE6 independently extracts multiple relational triples, ignoring the inherent dependencies between multiple relational triples, resulting in redundant extraction results, unable to cover all relational triples in the entire sentence, and ignoring the practicality and compactness of the extraction results, limiting its application in downstream tasks.

Method used

The multi-head attention is used to perform multi-level feature fusion extraction, and the sentence context representation is obtained through pre-training language model. The word score vector is calculated using the dual affine attention module, the Softmax function calculates the label probability, and the multi-level feature fusion device is used to connect the features, iteratively extract the relationship triplets in the sentence, model the dependencies between each extraction, and use the Transformer model to obtain context embeddings, reduce redundancy and improve coverage and practicality.

Benefits of technology

Through the multi-head attention mechanism, the sentence and predicate characteristics are integrated, the dependencies between different triplets are focused, the redundancy of the extracted results is reduced, the coverage and practicality are improved, and the application effect of the extracted results in downstream tasks is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117312487B_ABST
    Figure CN117312487B_ABST
Patent Text Reader

Abstract

A method for multi-level feature fusion extraction using multi-head attention, which relates to the field of information extraction, includes: obtaining the context representation of each word in a sentence and performing recognition; calculating the scoring vector of each pair of words using the recognition result; calculating the label probability; filling and training a two-dimensional table; decoding the predicate and filtering out other components; using a multi-level feature fusion device to concatenate features and use them as the input for iterative extraction; using iterative extraction to model the inherent dependencies between each extraction, that is, using the multi-head attention module for parameter extraction, label classification, and label embedding, obtaining the context embedding of each word in this extraction, concatenating it with other features through the multi-level feature fusion device and using it as the input for the next iterative extraction, and repeating the extraction until all predicates are extracted. The present invention reduces the redundancy of the extraction result, improves the coverage and practicality of the extraction result, and helps the application of the extraction result in downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information extraction, and particularly relates to a method for multi-level feature fusion extraction using multi-head attention. Background Art

[0002] With the rapid development of modern Internet technology, a vast amount of unstructured text data has been generated. Information extraction technology in natural language processing is used to extract structured information from the vast amount of unstructured text data, and these structured information is commonly represented in the form of relational triples (entity 1; relationship; entity 2). Traditional information extraction methods focus on answering narrow and well-defined requests through a predefined set of target relationships on a small homogeneous corpus. For this reason, traditional information extraction methods generally use the target relationships and manually extracted patterns or patterns learned from manually labeled training examples as input. When applying traditional information extraction methods to new fields, not only do users need to name the target relationships themselves, but they also need to manually define new extraction rules or manually annotate new training data. Therefore, traditional information extraction methods rely on extensive human participation.

[0003] Currently, in order to reduce the manual operations required by traditional information extraction methods, a new extraction paradigm: OpenIE (Open Domain Information Extraction) has been introduced. Different from traditional information extraction methods, open domain information extraction is not limited to a small set of predefined target relationships, but extracts all types of target relationships found in the text. Open domain information extraction can use information such as domain-independent syntactic features to extract relationships from the text and extend to large heterogeneous corpora such as the Web. Most early open domain information extraction methods used templates automatically learned from annotated text or manually constructed, and relied on the dependency features of sentences to extract relational triples. Due to the use of information such as domain-independent syntactic features, open domain information extraction methods can be applied to different fields and relationship types.

[0004] Some researchers also believe that the lack of complete context information in relational triples is not conducive to the understanding of downstream tasks and may extract non-factual and hypothetical relational triples. Therefore, some researchers have also explored how to extract relational triples with complete context information. However, since structurally complex sentences pose a huge challenge to open domain information extraction methods, it is difficult to extract relational triples from complex sentences using methods such as rules. Therefore, in order to improve the accuracy of relational triple extraction, some researchers have proposed converting complex sentences into simple clauses and using simple templates to extract triples in these simple clauses to improve the accuracy of relational triple extraction.

[0005] With the continuous development of deep learning methods in recent years, the open-domain information extraction method based on deep learning has become the mainstream. The closest prior art to the present invention is the OpenIE6 system, which is an extraction-based system that uses the IGL network to perform open-domain information extraction through two-dimensional grid annotation. Its main idea is: converting it into a two-dimensional grid marking task (IGL) + iterative marking to improve the metrics and accelerate inference. However, the OpenIE6 system has the following problems: Although it can independently extract multiple relational triples, it ignores the inherent dependency relationships between multiple relational triples, resulting in redundant extraction results and an inability to highly cover all the relational triples in the entire sentence, ignoring the practicality and compactness of the extracted relational triples, and restricting the application of the open-domain information extraction results in downstream tasks. Summary of the Invention

[0006] The technical problem that the present invention aims to solve is that the existing open-domain information extraction system, namely the OpenIE6 system, independently extracts multiple relational triples, ignoring the inherent dependency relationships between multiple relational triples, resulting in redundancy in the final extraction results and an inability to highly cover all the triples in the entire sentence. In addition, while focusing on high coverage and low redundancy of the extraction results, the existing open-domain information extraction system, namely the OpenIE6 system, often ignores the practicality and compactness of the extracted relational triples, which restricts the application of the open-domain information extraction results in downstream tasks. Therefore, the present invention provides a method for multi-level feature fusion extraction using multi-head attention.

[0007] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0008] A method for multi-level feature fusion extraction using multi-head attention according to the present invention includes the following steps:

[0009] S1. Input the given sentence into a pre-trained language model to obtain the context representation of each word and identify it;

[0010] S2. Input the above identification result into a bi-affine attention module to calculate the scoring vector for each pair of words;

[0011] S3. Input the above scoring vector into the Softmax function to calculate the label probability;

[0012] S4. Use the above label probability to fill a two-dimensional table and train it;

[0013] S5. Use a predicate component extractor to calculate the distances between adjacent rows and columns of the two-dimensional table to find the boundaries of the constituent components, assign a label to each span, decode the predicates, and filter out other components;

[0014] S6. Concatenate features using a multi-level feature fuser Feature H p and feature E pos , where the feature represents the final hidden state of the sentence after passing through the predicate extraction module. Feature H p represents the arithmetic mean vector in the predicate position hidden state, and feature E pos represents the predicate position embedding of binary values;

[0015] S7. Use the concatenated feature X as the Query of the multi-head attention module, i.e., X q , and use the subset derived from the concatenated feature X at the predicate position as the Key-Value of the multi-head attention module, i.e., X k and X v ;

[0016] S8. For sentences containing multiple predicates, use an iterative extraction method to model the inherent dependencies between each extraction without re-encoding, and obtain the attention for each head;

[0017] S9. Connect the attention outputs of each head above and perform a linear transformation; input the final output result into a label classifier for label classification;

[0018] S10. After step S9, label the relationship between each word and the predicate used in this iterative extraction, and obtain the context embedding of each word in this extraction. For each word h n , each predicate in each extraction has a different context embedding, which is denoted as the feature where m is the number of predicates. Concatenate this feature with other features using the multi-level feature fuser and use it as the input for the next iterative extraction. Repeat the above process until all predicates are extracted.

[0019] Furthermore, in step S1, the language model is a Bert model.

[0020] Furthermore, in step S1, two MLPs are used to identify the head and tail of the context representation of each word.

[0021] Furthermore, in step S2, a bi-affine attention mechanism is used to learn the interaction between word pairs to identify nested relationship words; then a bi-affine scoring function is used to calculate the scoring vector for each pair of words.

[0022] Further, in step S3, the scoring vector obtained in step S2 is input into the Softmax function to calculate the probabilities of each tag belonging to P, A, S, O, and N; where P represents a relational word, A represents a parameter, S represents the subject in a relational triple, O represents the object in a relational triple, and N represents other components that do not belong to the relational triple.

[0023] Further, in step S4, when training the two-dimensional table, the ultimate goal is to minimize the training function, and the formula of the training function is as follows:

[0024]

[0025] where n represents the number of words in the sentence, and Y i,j represents the gold label, i.e., the correct label, of the cell (i, j), and P(y i,j = Y i,j |s) represents the probability that each label is correctly labeled.

[0026] Further, in step S4, the following constraints are imposed during the training process of the two-dimensional table:

[0027] 1) The two-dimensional table is square and symmetric about the diagonal; the constraint loss L2 is:

[0028]

[0029] where s represents the sentence, t represents the word component in the set {P, A, S, O}, and ρ i,j,t and ρ j,i,t both represent the stack of each word pair in each sentence;

[0030] 2) For each word, the probability of it being A and P is not lower than the probability of it being S and O; the constraint loss L3 is:

[0031]

[0032] where s represents the sentence, ρ i,:,t represents the stack of word pairs in each row, ρ :,i,t represents the stack of word pairs in each column, ρ i,i,t represents the stack of word pairs on the diagonal, P represents the relational word, A represents the parameter, S represents the subject in the relational triple, O represents the object in the relational triple, and N represents other components that do not belong to the relational triple;

[0033] 3) In a relational triple, the probability of S appearing is not lower than the probability of O appearing; the constraint loss L4 is:

[0034]

[0035] Among them, ρ i,:,O represents the stack of word pairs belonging to component O in each row, and ρ i,:,s represents the stack of word pairs belonging to component S in each row, and ρ :,i,O represents the stack of word pairs belonging to component O in each column, and ρ :,i,S represents the stack of word pairs belonging to component S in each column, represents the union of the spans of the relational word P in the sentence s; finally, jointly optimize L1 + L2 + L3 + L4.

[0036] Furthermore, in step S8, first convert X q , X k and X v into Q = X q W q , K = X k W k and V = X v W v , where W q , W k and W v all represent weight matrices; calculate the attention for each head using Equation (9):

[0037]

[0038] Among them, Z h represents the calculation result of the attention for each head, and the subscript h is the index of the head, and d h = d mh / n h , where d mh represents the dimension of the multi-head attention, and n h is the number of heads.

[0039] Furthermore, in step S10, obtain the context embedding of each word in this extraction through a two-layer Transformer model.

[0040] The beneficial effects of the present invention are:

[0041] 1) The present invention uses multi-head attention to replace bidirectional LSTM for multi-level feature fusion, can fuse sentence and predicate features, and can pay attention to the inherent dependencies and shared components between different triples.

[0042] 2) The present invention reduces the redundancy of the extraction results, improves the coverage of the extraction results, improves the practicality of the extraction results, and helps the application of the extraction results in downstream tasks.

[0043] 3) The present invention uses a bi-affine attention mechanism to generate shared components and triples, which is beneficial to extracting compact triples. Description of the Drawings

[0044] Figure 1 This is a flowchart of a method for multi - level feature fusion extraction using multi - head attention in the present invention. Detailed Implementation Manner

[0045] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] Refer to Figure 1 for description. A method for multi - level feature fusion extraction using multi - head attention in the present invention specifically includes the following steps:

[0047] Step 1: First, given a sentence s, obtain the context representation of each word in the sentence s through the pre - trained language model Bert; subsequently, use two MLPs (Multi - Layer Perceptrons) to identify the head and tail of the context representation of each word. The calculation formulas involved in this process are as follows:

[0048] h i head = MLP head (h i ) (1)

[0049] h i tail = MLP tail (h i ) (2)

[0050] Among them, h i head represents the head of the context representation h i of word i, h i tail represents the tail of the context representation h i of word i, and MLP represents the identification operation of the multi - layer perceptron.

[0051] Step 2: Next, input the head and tail of the context representation of each word obtained in Step 1 into the bi - affine attention module, and learn the interaction between word pairs through the bi - affine attention mechanism to identify nested relationship words; then use the bi - affine scoring function to calculate the scoring vector of each pair of words. The calculation formulas involved in this process are as follows:

[0052] t i,j=(h i head ) T W (1) (h j tail )+(h i head ⊕h j tail ) T W (2) +b(3)

[0053] where W (1) and W (2) both represent weight parameters, ⊕ represents concatenation, b represents the bias term, the superscript T represents transpose, t i,j represents the scoring vector for each pair of words i and j, h i head represents the head of the context representation h i of word i, h j tail represents the tail of the context representation h j of word j.

[0054] Step 3. Subsequently, input the scoring vector obtained in Step 2 into the Softmax (normalized exponential) function to calculate the probabilities of each label belonging to P, A, S, O, N; where P represents the relational word, A represents the parameter, S represents the subject in the relational triple, O represents the object in the relational triple, and N represents other components not belonging to the relational triple. The calculation formula involved in this process is as follows:

[0055] P(y i,j |s) = Softmax(t i,j ) (4)

[0056] where P(y i,j |s) represents the probabilities of each label Y i,j belonging to P, A, N, S, O obtained by inputting the scoring vector into the Softmax function, y i,j represents the gold label of cell (i, j), s represents the sentence, Softmax represents the normalization operation, and t i,j represents the scoring vector for each pair of words i and j.

[0057] Step 4. Use the label probabilities obtained in Step 3 to fill the two-dimensional table and train this two-dimensional table to minimize the following training function. The formula of this training function L1 is as follows:

[0058]

[0059] Among them, n represents the number of words in the sentence, and Y i,j represents the gold label, i.e., the correct label, of the cell (i, j). P(y i,j = Y i,j |s) represents the probability that each label is correctly marked.

[0060] During the process of training the filled two-dimensional table, in order to enhance the final extraction effect, the following constraints are required:

[0061] 1) The two-dimensional table is square and symmetric about the diagonal. The constraint loss L2 is:

[0062]

[0063] Among them, s represents the sentence, t represents the word component in the set of word components {P, A, S, O}, and ρ i,j,t and ρ j,i,t both represent the stack of each word pair in each sentence.

[0064] 2) Relationships do not appear unless there are components of the relationship in the two-dimensional table. That is, for each word, the probability of it being an A (parameter) and a P (relationship word) is not lower than the probability of it being an S (subject in the relationship triple) and an O (object in the relationship triple). The constraint loss L3 is:

[0065]

[0066] Among them, s represents the sentence, ρ i,:,t represents the stack of word pairs in each row, ρ :,i,t represents the stack of word pairs in each column, ρ i,i,t represents the stack of word pairs on the diagonal, P represents the relationship word, A represents the parameter, S represents the subject in the relationship triple, O represents the object in the relationship triple, and N represents other components that do not belong to the relationship triple.

[0067] 3) An S (subject in the relationship triple) must exist in a relationship triple, but an O (object in the relationship triple) can be absent. That is, the possibility of the appearance of S (subject in the relationship triple) is not lower than the possibility of the appearance of O (object in the relationship triple). The constraint loss L4 is:

[0068]

[0069] Among them, ρ represents the stack of P(y i,j |s) of all word pairs in the sentence s. Specifically, ρ i,:,O represents the stack of word pairs belonging to component O in each row, ρ i,:,s represents the stack of word pairs belonging to component S in each row, ρ:,i,O A stack representing word pairs of component O for each column, ρ :,i,S A stack representing word pairs of component S for each column Represents the union of the spans of P(relational words) in sentence s.

[0070] Finally, during the process of training the filled two-dimensional table, jointly optimize L1+L2+L3+L4, where L1 is the training function and L2, L3, and L4 are all constraint losses.

[0071] Step Five: Finally, use the predicate component extractor to calculate the distances between adjacent rows and columns of the two-dimensional table to find the boundaries of the components, then assign a label to each span, and finally decode the predicates and filter out other components.

[0072] Step Six: Use a multi-level feature fusion device to concatenate the following three features and use them as the input for iterative extraction: H p and E pos . Among them, the feature Represents the final hidden state of the sentence passing through the predicate component extractor, the feature H p Represents the arithmetic mean vector in the predicate position hidden state, and the last feature E pos Represents the predicate position embedding of binary values, which indicates whether each token is included in the predicate range.

[0073] Step Seven: The present invention represents the concatenated features as X, and uses X itself as the Query of the multi-head attention module, denoted as X q , and uses the subset of X derived at the predicate position as the Key-Vaule of the multi-head attention module, which are denoted as X k and X v respectively.

[0074] Step Eight: Parameter extraction;

[0075] Generally, the extracted sentence contains more than one predicate. For a sentence containing multiple predicates, the present invention uses an iterative extraction method to model the inherent dependencies between each extraction without re-encoding. First, convert X q , X k and X v into Q = X q W q , K = X k W k and V = X v W v , where W q , W k and W vBoth represent the weight matrix, and then attention calculation is performed on each head.

[0076]

[0077] Among them, Z h represents the calculation result of the attention of each head. The subscript h is the index of the head, and d h = d mh / n h d mh represents the dimension of the multi-head attention, and n h is the number of heads.

[0078] Step Nine: Connect the attention outputs of each head obtained above and perform a linear transformation; finally, input the final output result into the label classifier for label classification.

[0079] Step Ten: Label embedding;

[0080] At this point, a relationship label (S, O, or N) between each word and the predicate extracted in this iteration will be marked, and a 2-layer Transformer model is used to obtain the context embedding of each word in this extraction. For each word h n , each extracted predicate has a different context embedding, so it is denoted as a feature where m is the number of predicates (that is, the number of extractions, the number of the iteration extraction). The feature is concatenated with other features and used as the input for the next iteration extraction. Repeat the above process until all predicates are extracted. The present invention can obtain relevant information before this extraction each time. In principle, this will maintain the information of the extraction output so far, so the dependency relationship between the labels of different extractions can be captured.

[0081] The present invention uses iterative extraction to replace traditional repeated encoding, and at the same time realizes iterative extraction through the method of filling a two-dimensional table, so as to improve the extraction accuracy, capture the inherent dependency relationship and shared components between extractions while ensuring speed, and can also identify and extract compact relationship triples.

[0082] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for multi-level feature fusion extraction using multi-head attention, characterized in that It includes the following steps: S1. Input the given sentence into a pre-trained language model to obtain the context representation of each word and identify it; S2. Input the above recognition result into a bi-affine attention module to calculate the scoring vector for each pair of words; S3. Input the above scoring vector into the Softmax function to calculate the label probability; S4. Use the above label probability to fill a two-dimensional table and train it; S5. Use a predicate component extractor to calculate the distances between adjacent rows and columns of the two-dimensional table to find the boundaries of the constituent components, assign a label to each span, decode the predicates and filter out other components; S6. Concatenate features using a multi-level feature fuser Feature H p and Feature E pos , where the feature represents the final hidden state of the sentence after passing through the predicate component extractor, Feature H p represents the arithmetic mean vector in the predicate position hidden state, and Feature E pos represents the predicate position embedding of the binary value; S7. Use the concatenated feature X as the Query of the multi-head attention module, i.e., X q , and use the subset derived from the concatenated feature X at the predicate position as the Key-Value of the multi-head attention module, i.e., X k and X v ; S8. For a sentence containing multiple predicates, use an iterative extraction method to model the inherent dependencies between each extraction without re-encoding to obtain the attention for each head; S9. Connect the attention outputs of each above head and perform a linear transformation; input the final output result into a label classifier for label classification; S10. After step S9, relationship labels between each word and the predicate extracted in this iteration are marked, and the context embedding of each word in this extraction is obtained. For each word h n , each extracted predicate has a different context embedding, which is denoted as a feature m is the number of predicates. This feature is concatenated with other features through a multi-level feature fuser and used as the input for the next iteration of extraction. The above process is repeated until all predicates are extracted.

2. A method for multi-level feature fusion extraction using multi-head attention, as described in claim 1, wherein In step S1, the language model is a Bert model.

3. A method for multi-level feature fusion extraction using multi-head attention, as described in claim 1, wherein In step S1, two MLPs are used to identify the head and tail of the context representation of each word.

4. A method for multi-level feature fusion extraction using multi-head attention, as described in claim 1, wherein In step S2, a bi-affine attention mechanism is used to learn the interaction between word pairs to identify nested relational words; then a bi-affine scoring function is used to calculate the scoring vector for each pair of words.

5. A method for multi-level feature fusion extraction using multi-head attention, as described in claim 1, wherein In step S3, the scoring vector obtained in step S2 is input into the Softmax function to calculate the probability that each label belongs to P, A, S, O, N; where P represents a relational word, A represents an argument, S represents the subject in a relational triple, O represents the object in a relational triple, and N represents other components that do not belong to the relational triple.

6. A method for multi-level feature fusion extraction using multi-head attention, as claimed in claim 1, wherein In step S4, when training the two-dimensional table, minimizing the training function is the ultimate goal, and the formula of the training function is as follows: Among them, n represents the number of words contained in the sentence, and Y i,j represents the gold label, i.e., the correct label, of the cell (i, j). P(y i,j = Y i,j |s) represents the probability that each label is correctly annotated.

7. A method for multi-level feature fusion extraction using multi-head attention according to claim 1, characterized in that In step S4, the following constraints are imposed during the training process of the two-dimensional table: 1) The two-dimensional table is square and symmetric about the diagonal; the constraint loss L2 is: where s represents a sentence, t represents a word component in the set of word components {P, A, S, O}, ρ i,j,t and ρ j,i,t both represent the stack of each word pair in each sentence; 2) For each word, the probability of it being A and P is not less than the probability of it being S and O; the constraint loss L3 is: Among them, s represents a sentence, ρ i,:,t represents the stack of word pairs in each line, ρ :,i,t represents the stack of word pairs in each column, ρ i,i,t represents the stack of diagonal word pairs, P represents a relational word, A represents an argument, S represents the subject in a relational triple, O represents the object in a relational triple, and N represents other components that do not belong to the relational triple; 3) The probability of S appearing in a relational triple is not less than the probability of O appearing; the constraint loss L4 is: where ρ i,:,O represents the stack of word pairs belonging to component O for each row, ρ i,:,s represents the stack of word pairs belonging to component S for each row, ρ :,i,O represents the stack of word pairs belonging to component O for each column, ρ :,i,S represents the stack of word pairs belonging to component S for each column, and ζ represents the union of the spans of the relational word P in sentence s; finally, jointly optimize L1+L2+L3+L4.

8. A method for multi-level feature fusion extraction using multi-head attention according to claim 1, characterized in that, In step S8, first, X q 、X k and X v are respectively converted into Q = X q W q , K = X k W k and V = X v W v , where W q , W k and W v all represent weight matrices; Use equation (9) to calculate the attention for each head: Among them, Z h represents the calculation result of each head attention, where the subscript h is the index of the head, and d h = d mh / n h , and d mh represents the dimension of the multi-head attention, and n h is the number of heads.

9. A method for multi-level feature fusion extraction using multi-head attention according to claim 1, characterized in that, In step S10, a 2-layer Transformer model is used to obtain the context embedding of each word in this extraction.