A Joint Entity Relation Extraction Method Combining Attention Mechanism and Segment Arrangement
By integrating attention mechanism and fragment arrangement, the problem of task dependence and excessive negative samples in entity and relationship extraction in the prior art is solved, and more efficient and accurate joint extraction of entity relationships is achieved.
Patent Information
- Application Number
- CN202210341776.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-02
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-04-02
AI Technical Summary
The prior art ignores the intrinsic connection between tasks in entity and relationship extraction, resulting in entity extraction errors affecting the relationship extraction performance, and too many negative samples lead to performance degradation.
The entity relationship joint extraction method that combines attention mechanism and fragment arrangement is adopted, word vectors are obtained through pre-training language models, candidate fragments are enumerated, and the attention mechanism is used to prune to reduce the number of negative samples of entities, thereby improving the performance of named entity recognition and relationship extraction.
It effectively solves the problem of task dependence and too many negative samples in entity and relationship extraction, and improves the accuracy and efficiency of joint entity relationship extraction.
Smart Images

Figure CN114757192B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer natural language processing, and particularly relates to a method for jointly extracting entity relationships by integrating an attention mechanism and segment permutation. Background Art
[0002] With the development of science and technology, more and more important information appears around us in the form of text, such as paper materials, newspaper news, social chats, and blogs. These text information have problems such as large amount of information, complex content, and inconsistent structures, making it difficult for people to quickly obtain useful information from these text information. In such a modern society with information explosion, how to quickly and effectively extract useful information from these information redundant and structurally chaotic documents, and store these useful information in a fixed form so that subsequent users can accurately and quickly utilize these information has become an urgent challenge to be solved. Facing this challenge, people have proposed information extraction. And entity and relationship extraction are one of the key tasks of information extraction, which have received extensive attention from the academic and industrial circles in recent years. It can provide support for downstream tasks such as automatic question answering, information retrieval, knowledge base filling, and knowledge reasoning.
[0003] The research methods of named entity recognition and relationship extraction are mainly divided into two categories: pipeline methods and joint extraction methods. Pipeline methods usually require training two models, one for named entity recognition and the other for relationship extraction. The joint extraction method jointly models these two tasks of named entity recognition and relationship extraction, either projects them into a structured prediction framework or performs multi-task learning by sharing representations. Although pipeline methods are easy to implement, the flexibility of these two extraction models is high, and the entity extraction model and the relationship extraction model can use independent data sets without the need for a data set that simultaneously annotates entities and relationships. However, the pipeline method ignores the internal connection and dependency relationship between these two tasks, and the error of entity extraction will affect the performance of the next relationship extraction. Summary of the Invention
[0004] The present invention is proposed to solve the above problems, and aims to provide an accurate and highly adaptable method for jointly extracting entity relationships by integrating an attention mechanism and segment permutation. The present invention adopts the following technical solutions:
[0005] The present invention provides a method for jointly extracting entity relations by integrating an attention mechanism and segment permutation, which is characterized by the following steps: Step S1, input a text sentence and perform token parsing on the text sentence; Step S2, based on the token parsing, use a pre-trained language model for encoding to obtain word vectors of the input text; Step S3, based on the word vectors, enumerate all candidate segments in a segment permutation manner; Step S4, input the candidate segments into a neural network model of the attention mechanism and obtain attention scores of each candidate segment; Step S5, based on the attention scores, arrange the candidate segments into an ordered queue; Step S6, retain the candidate segments in the front of the ordered queue and delete the remaining candidate segments; Step S7, input the retained candidate segments into an entity classifier to predict entity types and obtain entity segments predicted to be true; Step S8, match the entity segments predicted to be true in pairs to obtain relationship representations of each pair of entity segments, and input the relationship representations into a relationship classifier for prediction to obtain the relationship types between each pair of entity segments.
[0006] The method for jointly extracting entity relations by integrating an attention mechanism and segment permutation provided by the present invention may also have the following technical feature: The token parsing of the input text sentence in Step S1 refers to parsing the text sentence into the most basic units of natural language processing.
[0007] The method for jointly extracting entity relations by integrating an attention mechanism and segment permutation provided by the present invention may also have the following technical feature: In Step S5, arranging the candidate segments into an ordered queue is performed in descending order of the attention scores.
[0008] The method for jointly extracting entity relations by integrating an attention mechanism and segment permutation provided by the present invention may also have the following technical feature: The calculation formula for the number of segments retained in the front in Step S6 is as follows: n = λN, where N is the total number of candidate segments and λ is the threshold of the retention factor.
[0009] The method for jointly extracting entity relations by integrating an attention mechanism and segment permutation provided by the present invention may also have the following technical feature: The entity classifier in Step S7 and the relationship classifier in Step S8 are both deep neural networks.
[0010] Functions and effects of the invention
[0011] The entity relation joint extraction method integrating the attention mechanism and segment permutation according to the present invention converts the input text into word vectors, enumerates all possible candidate segments based on the segment permutation method, inputs all the candidate segments into the neural network model of the attention mechanism, and performs pruning according to the attention scores to reduce the number of entity negative samples, thereby performing named entity recognition and relation extraction. Based on the segment permutation method, the present invention can enumerate all possible segments, and each selected segment is independent, and the segment-level features can be directly extracted to solve the overlapping entity problem. Aiming at the problem of too many entity negative samples, the present invention also adds an attention mechanism, and according to the attention scores, some negative samples can be effectively deleted to improve the performance of entity relation joint extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a schematic flowchart of the entity relation joint extraction method integrating the attention mechanism and segment permutation in an embodiment of the present invention;
[0013] Figure 2 is a schematic diagram of the principle of the entity relation joint extraction method integrating the attention mechanism and segment permutation in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the entity relation joint extraction method integrating the attention mechanism and segment permutation of the present invention will be specifically described below in conjunction with embodiments and drawings.
[0015] <Embodiment>
[0016] Figure 1 is a schematic flowchart of the entity relation joint extraction method integrating the attention mechanism and segment permutation in an embodiment of the present invention;
[0017] Figure 2 is a schematic diagram of an embodiment of the present invention.
[0018] As Figure 1 shown, the entity relation joint extraction method integrating the attention mechanism and segment permutation mainly includes the following steps:
[0019] Step S1, input a text sentence and perform token parsing on the text sentence;
[0020] In this embodiment, taking the text sentence: Joe Biden is the President of the United States. as an example, the text sentence is parsed into the most basic units of natural language processing, such as Figure 2As shown; specifically, the text sentence "Joe Biden is the President of the United States." is segmented into the most basic units of natural language processing such as 'Joe', 'Biden', 'is', 'the', 'President', 'of', 'the', 'United', 'States', and '.'.
[0021] Step S2: Based on the token parsing, use a pre-trained language model for encoding to obtain the word vectors of the input text.
[0022] In this embodiment, the tokens are input into the pre-trained language model to obtain the word vectors of each token, as Figure 2 shown; specifically, the tokens after token parsing: 'Joe', 'Biden', 'is', 'the', 'President', 'of', 'the', 'United', 'States', and '.' are input into the pre-trained language model to obtain the word vectors of each token.
[0023] Step S3: Based on the word vectors, enumerate all candidate segments in a fragment permutation manner.
[0024] In this embodiment, enumerating all candidate segments includes the following steps: Set the maximum fragment length L max = 8; if L = 1, enumerate all segments with a length equal to L; if L < L max , then L = L + 1, and enumerate all segments with a length equal to L. Specifically, as Figure 2As shown, fragments with a fragment length L = 1: 'Joe', 'Biden', 'is', 'the', 'President', 'of', 'the', 'United', 'States', and '.'; In addition, since L < 8, then L = 1 + 1 = 2, fragments with a fragment length L = 2: 'Joe Biden', 'Biden is', 'is the', 'the President', 'President of', 'ofthe', 'the United', 'United States', and 'States.'; Since L < 8, then L = 2 + 1 = 3, fragments with a fragment length L = 3: 'Joe Biden is', 'Biden is the', 'is the President', 'the President of', 'President of the', 'of the United', 'the United States', and 'UnitedStates.';...; Since L < 8, then L = 7 + 1 = 8, fragments with a fragment length L = 8: 'Joe Biden is thePresident of the United', 'Biden is the President of the United States', and 'isthe President of the United States.'
[0025] Step S4: Input the candidate fragments into the neural network model of the attention mechanism, and obtain the attention scores of each candidate fragment.
[0026] Step S5: Based on the attention scores, arrange the candidate fragments into an ordered queue.
[0027] Step S6: Retain the candidate fragments at the front of the ordered queue, and delete the remaining candidate fragments.
[0028] In this embodiment, the specific steps are as follows: First, set the threshold value λ of the retention factor to 0.05. Second, input all the candidate fragments into the neural network model of the attention mechanism to obtain the attention score α i of each candidate fragment s i , where the neural network model of the attention mechanism uses a one-layer linear network with a single output, and the attention score α i of each candidate fragment s iis a single value of the linear network; then, according to the attention scores α from large to small, the candidate segments are arranged in an ordered queue; finally, the first n = λN segments in the ordered queue of candidate segments are retained, and the candidate segments behind the ordered queue are deleted. The number of retained candidate segments is n = 0.05×N, where N is the total number of all candidate segments before all pruning.
[0029] Step S7, input the retained candidate segments into an entity classifier to predict the entity type, and obtain the entity segments predicted to be true.
[0030] In this embodiment, the entity classifier is a single layer of deep neural network to predict the entity type of the candidate segments through the neural network, as Figure 2 shown. Specifically, the entity representation includes three parts in total: the start and end word vectors of the segment and the length embedding of the segment. Input the entity representation into the entity classifier to obtain the logits score vector of the entity type, and after the softmax operation, the predicted type of the entity can be obtained.
[0031] Step S8, pair the entity segments predicted to be true pairwise, obtain the relationship representation of each pair of the entity segments, and input the relationship representation into a relationship classifier for prediction, so as to obtain the relationship type between each pair of the entity segments.
[0032] In this embodiment, the entity segments predicted to be true by the entity classifier are paired pairwise, and then the features of each pair of entity segments are used as the relationship representation between the two entities, and the relationship representation is input into the relationship classifier to predict their relationship type; among them, the relationship classifier is a single layer of deep neural network to predict the relationship type between each pair of entities through the neural network, as Figure 2 shown. Specifically, first pair all the entities predicted to be true pairwise. The relationship representation between two entities includes four parts: the start and end word vectors of the head entity and the start and end word vectors of the tail entity, including a total of four word vectors; then input the relationship representation into the relationship classifier to obtain the logits score vector of the relationship type, and finally after the softmax operation, the predicted type of the relationship can be obtained.
[0033] Functions and effects of the embodiment
[0034] According to the entity relation joint extraction method that combines the fusion attention mechanism and segment arrangement of the present invention, the input text is converted into word vectors, and all possible candidate segments are enumerated based on the segment arrangement method. By inputting all the candidate segments into the neural network model of the attention mechanism and pruning according to the attention scores, the number of entity negative samples is reduced, so as to perform named entity recognition and relation extraction. Based on the segment arrangement method, the present invention can enumerate all possible segments, and each selected segment is independent, and the segment-level features can be directly extracted to solve the overlapping entity problem. Aiming at the problem of too many entity negative samples, the present invention also adds an attention mechanism. According to the attention scores, some negative samples can be effectively deleted to improve the performance of entity relation joint extraction.
[0035] The above embodiments are only used to illustrate the specific implementation manners of the present invention, and the present invention is not limited to the description scope of the above embodiments.
Claims
1. A joint entity relation extraction method integrating attention mechanism and segment arrangement, characterized in that It includes the following steps: Step S1, input a text sentence and perform token parsing on the text sentence; Step S2, based on the token parsing, use a pre-trained language model for encoding to obtain the word vectors of the input text; Step S3, based on the word vectors, enumerate all candidate segments in a segment permutation manner; Step S4, input the candidate segments into a neural network model of the attention mechanism and obtain the attention scores of each candidate segment; Step S5, based on the attention scores, arrange the candidate segments into an ordered queue; Step S6, retain the candidate segments at the front of the ordered queue and delete the remaining candidate segments; Step S7, input the retained candidate segments into an entity classifier for entity type prediction and obtain the entity segments predicted to be true; Step S8, pair up the entity segments predicted to be true, obtain the relationship representation of each pair of entity segments, and input the relationship representation into a relationship classifier for prediction to obtain the relationship type between each pair of entity segments.
2. The entity relationship joint extraction method integrating the attention mechanism and segment permutation according to claim 1, wherein: Among them, The token parsing of the input text sentence in step S1 means parsing the text sentence into the most basic units of natural language processing.
3. The entity relationship joint extraction method integrating the attention mechanism and segment permutation according to claim 1, wherein: Among them, In step S5, arranging the candidate segments into an ordered queue is performed in descending order of the attention scores.
4. The entity relationship joint extraction method integrating the attention mechanism and segment permutation according to claim 1, wherein: Among them, The calculation formula for the number of candidate segments retained in step S6 is as follows: n = λN, wherein, N is the total number of candidate segments, and λ is the threshold of the retention factor.
5. The entity relationship joint extraction method integrating the attention mechanism and segment permutation according to claim 1, wherein: Among them, Both the entity classifier in step S7 and the relationship classifier in step S8 are deep neural networks.
Citation Information
Patent Citations
Unstructured text extraction multi-task joint training method based on pointer network
CN111488726A
Word processing method and device based on multi-task model
CN113887225A