A Named Entity Recognition Method Based on Syntactic Dependency Relations
By introducing a named entity recognition method based on syntactic dependency in deep learning models, using Bi-LSTM and self-attention mechanism to calculate features, combined with CRF refinement, the problem of low accuracy in named entity boundary recognition is solved, and the accuracy of named entity recognition is significantly improved.
Patent Information
- Application Number
- CN202010556881.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-16
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-06-16
AI Technical Summary
The existing deep learning models have low accuracy in the recognition of named entities in text, resulting in frequent false positive and false negative examples, affecting the performance of natural language processing systems.
The named entity recognition method based on syntactic dependency is adopted, and the local and global features of each word in the sentence are calculated through pre-training of the Word2vec model, bidirectional long and short-term memory network (Bi-LSTM), syntactic analysis technology and self-attention mechanism, and refined it with CRF to improve the accuracy of boundary recognition.
Through the self-attention mechanism, weaken the connection between entities and external words, enhance the relationship between words within entities, significantly improve the accuracy of naming entity recognition, especially in boundary judgment.
Smart Images

Figure CN111783461B_ABST
Abstract
Description
Technical Field:
[0001] The present invention relates to the field of deep learning and the named entity recognition technology in text. Background Art:
[0002] Traditional named entity recognition methods rely on a large number of manually defined features. However, such methods of manually defining features are not only time-consuming and laborious, but also require professionals with knowledge in the fields and languages. In recent years, deep learning, relying on its powerful data mining ability, has minimized the cost of manually constructing features and achieved remarkable achievements in fields such as image classification, speech recognition, and natural language processing. Therefore, using deep learning methods for named entity recognition has great research significance.
[0003] In text, accurately identifying the named entity type and its entity boundary has a great impact on the development of complex natural language systems, such as information extraction, question answering, text summarization, etc. In named entity recognition, only when the entity boundary and type recognized by the model match the boundary and type of the annotated entity, it is considered a true positive example (TP). In most test samples, false positive examples (FP) and false negative examples (FN) are often caused by incorrect judgment of the entity boundary, that is to say, boundary recognition is much more difficult than type recognition. And in most deep network models, there is no specific function for boundary recognition, making the model often have a higher accuracy rate in type judgment but a lower accuracy rate in boundary judgment. Summary of the Invention:
[0004] The purpose of the present invention is to provide a method that can more accurately identify the named entity boundary and type in text.
[0005] To solve the above technical problems, the present invention provides a named entity recognition method based on syntactic dependency relationships, including the following steps:
[0006] Step S1, in the model training stage, first use the pre-trained Word2vec to map the one-hot word vector to the defined low-dimensional space to obtain the word vector of each word;
[0007] Step S2, use a bidirectional long short-term memory network (Bi-LSTM) to perform forward and backward encoding on the word vectors at each time step in the sentence respectively, and splice them to obtain the global features with context information;
[0008] Step S3, use syntactic analysis technology to obtain the syntactic dependency tree of each sentence, and calculate the shortest dependency path between two words on the tree;
[0009] Step S4: Obtain the top - down and bottom - up feature sequences of each word according to the shortest dependency path and input them into the LSTM network to calculate the local features of the words;
[0010] Step S5: Calculate the relationship weights between pairwise words through local feature dot - product and perform normalization;
[0011] Step S6: Use the self - attention mechanism to integrate the local relationship features between words into the global features with the normalized relationship weights to obtain the fused features;
[0012] Step S7: Initially predict the sequence labels according to the fused features and use CRF to refine the predicted sequence to obtain the final label sequence;
[0013] Step S8: In the model testing stage, use the network trained in the above steps to perform named entity recognition.
[0014] Furthermore, in step S1 during the model training stage, first use the pre - trained Word2vec to map the one - hot word vectors into a defined low - dimensional space to obtain the word vectors of each word, including:
[0015] Denote the size of the dictionary as V. Use the pre - trained Word2vec to map the one - hot word vectors of dimension V into a defined low - dimensional space, and denote the dimension of the output word vectors as d. For the input sample sequence {w 1 , w 2 ,... w T} of length T, the output of the embedding layer is denoted as {x 1 , x 2 ,... x T}, where x t ∈R 1×d ;
[0016] Furthermore, in step S2, use a bidirectional long short - term memory network (Bi - LSTM) to encode the word vectors at each time step in the sentence forward and backward respectively, and concatenate them to obtain the global features with context information, including:
[0017] Use a bidirectional long short - term memory network (Bi - LSTM 1 ) with the number of hidden units h 1 to encode the input x t at a given time step t forward and backward, and denote the forward hidden state at this time step as and the backward hidden state as Then, concatenate the hidden states in the two directions and to obtain the hidden state is the global feature with the context information of the given time step t. For the input sequence {x 1 , x 2 ,... x T}, denote the output feature of Bi-LSTM 1 as
[0018] Furthermore, in step S3, the syntactic dependency tree of each sentence is obtained by using syntactic analysis technology. Calculating the shortest dependency path between two words on the tree includes:
[0019] For the input sample sequence {w 1 , w 2 ,... w T}, use dependency grammar analysis technology to perform syntactic analysis on it to obtain the dependency syntactic tree of the sample sequence. For any two words a and b in the input sequence, their shortest dependency path (SDP) is {a, a 1 ,..., a m , c, b n ,..., b 1 , b}, where c represents their lowest common ancestor in the dependency syntactic tree, a 1 ,..., a m represents the words between a and c on the SDP, and b 1 ,..., b n represents the words between b and c. If a and b represent the same word, the SDP is denoted as {a, b}.
[0020] Furthermore, in step S4, the top-down and bottom-up feature sequences of each word are obtained according to the shortest dependency path and input into the LSTM network. The calculated word local features include:
[0021] For any two words a and b in the input text sequence {w 1 , w 2 ,..., w T}, their shortest dependency path (SDP) can be divided into two parts: the bottom-up sequence {a, a 1 ,... a m , c} and {b, b 1 ,..., b n , c}; the top-down sequences {c, a m ,..., a 1 , a} and {c, b n ,..., b 1 , b}. If a and b represent the same word, the SDP is divided into: {a}; {b} two parts.
[0022] The number of hidden units used is h 2 Bi-LSTM 2 ) extracts local relationship features between words from these two sequences. Each LSTM 2 The input of the unit is the series connection of two parts, consisting of
[0023]
[0024] Indicates that It is the word w t In Bi-LSTM 1 The output of emb(d t ) represents word w t and the dependency relationship type d between the dominant words on its dependency syntactic tree t Distributed expression of .
[0025] Forward LSTM 2 According to the bottom-up sequence {a, a 1 , ..., a m ,c} and {b,b 1 , ..., b n , c} calculate the forward hidden state and Backward LSTM 2 According to the top-down sequence {c, a m , ..., a 1 , a} and {c, b n , ..., b 1 , b} calculate the backward hidden state and Connect the hidden states in both directions↑h t and ↓h t To get the word w t Local features
[0026] Further, in step S5, calculating the relationship weights between two words by using the local feature dot product and normalizing them includes:
[0027] For local features With local features Do the dot product and get the word w i With the word w j The closeness coefficient of
[0028]
[0029] The same method is used to calculate the relationship coefficients between two words in the text sequence, and all the relationship coefficients are organized into a matrix R∈RT×T where the i-th row of the matrix represents the word w i and the tightness coefficient of the relationship between {w 1 , w 2 ,... w T}, and then normalize R row by row to obtain the self-attention weight matrix
[0030] Q = Softmax(R)
[0031] Furthermore, in step S6, the self-attention mechanism is used to integrate the local relationship features between words into the global features with the normalized relationship weights, and the fused features obtained include:
[0032] First, perform a linear transformation on the global features 1 output by the Bi-LSTM and left-multiply the normalized self-attention weight matrix Q to obtain the word features with enhanced entity boundary information
[0033] S = QH 1 W v
[0034] S ∈ R T×s , where s is the length of the fused features is the linear transformation parameter matrix
[0035] Furthermore, in step S7, the sequence labels are initially predicted based on the fused features, and the CRF is used to refine the predicted sequence to obtain the final label sequence including:
[0036] Use the fused features S for sequence label prediction, and adjust the initially predicted label sequence through the CRF to obtain the final label sequence
[0037] The beneficial effect of the present invention is that through the self-attention mechanism, the connection between entities and words outside the entities is weakened, and the relationship between words within the entities is strengthened, making the network more accurate in identifying entity boundaries and improving the accuracy of named entity recognition Description of the Drawings:
[0038] The present invention will be further described below in conjunction with the drawings and embodiments
[0039] Figure 1 is the method flow chart of a named entity recognition method based on syntactic dependency relationships of the present invention
[0040] Figure 2 is the dependency syntactic tree of the sample sequence Specific Embodiments:
[0041] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0042] Embodiment 1
[0043] As Figure 1 shown, Embodiment 1 of the present invention provides a named entity recognition method based on syntactic dependency relationships, including the following steps:
[0044] Step S1, in the model training stage, first use the pre-trained Word2vec to map the one-hot word vectors into a defined low-dimensional space to obtain the word vectors of each word;
[0045] Step S2, use a bidirectional long short-term memory network (Bi-LSTM) to encode the word vectors at each time step in the sentence forward and backward respectively, and splice them to obtain the global features with context information;
[0046] Step S3, use syntactic analysis technology to obtain the syntactic dependency tree of each sentence, and calculate the shortest dependency path between two words on the tree;
[0047] Step S4, obtain the top-down and bottom-up feature sequences of each word according to the shortest dependency path and input them into the LSTM network to calculate the local features of the word;
[0048] Step S5, calculate the relationship weights between two words through local feature dot products and normalize them;
[0049] Step S6, use the self-attention mechanism to integrate the local relationship features between words into the global features with the normalized relationship weights to obtain the fused features;
[0050] Step S7, preliminarily predict the sequence tags according to the fused features, and use CRF to refine the predicted sequence to obtain the final tag sequence;
[0051] Step S8, in the model testing stage, use the network trained in the above steps to perform named entity recognition.
[0052] In Chinese and English named entity recognition tasks, to accurately recognize an entity, it is necessary to judge both the type of the entity and the boundary of the entity. According to a large amount of experimental data, the accuracy of the named entity recognition task often depends on the accuracy of entity boundary judgment, that is, the judgment of entity boundary is much more difficult than the judgment of entity category. In the past, most models assisted the model to judge entity boundaries by simply adding part-of-speech tags in the embedding layer, and there was no special module or mechanism in the deep neural network to strengthen the judgment of entity boundaries.
[0053] Based on the above problems, this patent provides a named entity recognition method based on syntactic dependency relationships to strengthen the judgment of entity boundaries, thereby improving the accuracy of named entity recognition. Dependency is a language method that can structure sentences hierarchically. Dependency grammar directly represents the relationships between words while retaining the phrase structure information of the sentence, which is very beneficial for further semantic analysis. Dependency grammar believes that the verb is the central word and other words are dominated by it, which facilitates clarifying the relationships between words in a sentence. In summary, it is considered that using a named entity recognition method based on syntactic dependency relationships is conducive to analyzing the relationships between words in an entity and between an entity and words outside the entity, and thus can better judge entity boundaries.
[0054] To address the above problems, in step S1 of Embodiment 1, the size of the dictionary is denoted as V. During the model training phase, first, the pre-trained Word2vec is used to map the one-hot word vectors of dimension V into a defined low-dimensional space (the dimension of the output word vectors is denoted as d).
[0055] For example, for an input sample sequence of length 6, "Tsinghua University is located in Beijing", the output of the embedding layer is denoted as {x 1 , x 2 ,... x 6}, where the vector x t ∈ R 1×d .
[0056] Furthermore, in step S2, a bidirectional long short-term memory network (Bi-LSTM) is used to encode the word vectors at each time step in the sentence both forward and backward, and the concatenated result is a global feature with context information, including:
[0057] Using a bidirectional long short-term memory network (Bi-LSTM 1 ) with the number of hidden units being h 1 to encode the input x t at a given time step t both forward and backward, and denoting the forward hidden state at this time step as and the backward hidden state as Then, connect the hidden states in the two directions and to obtain the hidden state which is the global feature with context information at the given time step t. For the input sequence {x 1 , x 2 ,... x 6}, denote the output feature of Bi-LSTM 1 as
[0058] Further, in step S3, the syntactic dependency tree of each sentence is obtained by using syntactic analysis technology. Calculating the shortest dependency path between every two words on the tree includes:
[0059] For the input sample sequence "Tsinghua University is located in Beijing", use dependency grammar analysis technology to perform syntactic analysis on it, and obtain the dependency syntactic tree of the sample sequence, as Figure 2 shown.
[0060] For any two words in the input sequence, such as "Tsinghua" and "Beijing", the shortest dependency path (SDP) between them is {"Tsinghua", "University", "located", "Beijing"}, where "located" is their lowest common ancestor in the dependency syntactic tree, and "University" is the word between "Tsinghua" and "located" on the SDP. Here, the SDP of a word (such as "University") to itself is denoted as {"University", "University"}.
[0061] Further, in step S4, according to the shortest dependency path, the top-down and bottom-up feature sequences of each word are obtained and input into the LSTM network, and the local features of the word are calculated, including:
[0062] For the words "Tsinghua" and "Beijing" in the input text sequence "Tsinghua University is located in Beijing", the shortest dependency path (SDP) between them, {"Tsinghua", "University", "located", "Beijing"}, can be divided into two parts: the bottom-up sequences {"Tsinghua", "University", "located"} and {"Beijing", "located"}; the top-down sequences {"located", "University", "Tsinghua"} and {"located", "Beijing"}. The SDP of the word "University" to itself, {"University", "University"}, can be divided into: {"University"}; {"University"} in two parts.
[0063] Use a bidirectional long short-term memory network (Bi-LSTM) with the number of hidden units being h 2 2)Extract the local relationship features between words from these two sequences. Each LSTM 2 The input of the cell is
[0064]
[0065] is composed of two parts spliced together, where is the output of the t-th word in the Bi-LSTM 1 emb(d t ) represents the dependency relationship type d t between the t-th word and its governor on the dependency syntax tree, for example: d 1 = compound, which is the dependency relationship type between the first word "Tsinghua" and its governor "University". The distributed representations of emb(d 1 ) and other relationship types will be randomly initialized and trained together with the network model parameters.
[0066] Calculate the relationship weight between the words "Tsinghua" and "Beijing": The forward LSTM 2 Calculates the forward hidden states and according to the bottom-up sequences {"Tsinghua", "University", "located"} and {"Beijing", "located"} 2 The backward LSTM and
[0067] Concatenate the two-direction hidden states ↑h 1 and ↓h 1 of the word "Tsinghua" to obtain the local relationship feature between the word "Tsinghua" and the word "Beijing" 6 Concatenate the two-direction hidden states ↑h 6 and ↓h
[0068] of the word "Beijing" to obtain the local relationship feature
[0069] Further, calculating the relationship weights between pairwise words through local feature dot product and normalizing in step S5 includes:
[0070] Taking the words "Tsinghua" and "Beijing" as examples again, for the local features and the local features perform a dot product to obtain the relationship tightness coefficient r between the words "Tsinghua" and "Beijing" 16 :
[0071]
[0072] Calculate the relationship tightness coefficients between every two words in the text sequence in the same way, and organize all the relationship tightness coefficients into a matrix R ∈ R 6×6 , where the i-th row of the matrix represents the relationship tightness coefficient between the i-th word and each word in the sentence, and then normalize R row by row to obtain the self-attention weight matrix
[0073] Q = Softmax(R)
[0074] Furthermore, in step S6, use the self-attention mechanism to integrate the local relationship features between words into the global features with the normalized relationship weights, and the obtained fused features include:
[0075] First, perform a linear transformation on the global features 1 output by the Bi-LSTM and left-multiply by the normalized self-attention weight matrix Q to obtain the word features with enhanced entity boundary information
[0076] S = QH 1 W v S ∈ R 6×s , where s is the length of the fused features, and is the linear transformation parameter matrix.
[0077] Furthermore, in step S7, initially predict the sequence labels according to the fused features, and use CRF to refine the predicted sequence to obtain the final label sequence including:
[0078] Use the fused feature S to predict the sequence labels, and adjust the initially predicted label sequence through CRF to obtain the final label sequence.
[0079] Inspired by the ideal embodiments of the present invention described above, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and must be determined according to the scope of the claims.
Claims
1. A named entity recognition method based on syntactic dependency relationships, characterized in that, it includes the following steps: Step S1, in the model training stage, first use the pre-trained Word2vec to map the one-hot word vectors into a defined low-dimensional space to obtain the word vectors of each word; Step S2, use a bidirectional long short-term memory network to encode the word vectors at each time step in the sentence forward and backward respectively, and splice them to obtain the global features with context information; Step S3, use syntactic analysis technology to obtain the syntactic dependency tree of each sentence, and calculate the shortest dependency path between two words on the tree; Step S4, obtain the top-down and bottom-up feature sequences of each word according to the shortest dependency path and input them into the LSTM network to calculate the local features of the words; Step S5, calculate the relationship weights between two words through local feature dot products and normalize them; Step S6, use the self-attention mechanism to integrate the local relationship features between words into the global features with the normalized relationship weights to obtain the fused features; The self-attention mechanism used in Step S6 to integrate the local relationship features between words into the global features with the normalized relationship weights to obtain the fused features includes: First, perform a linear transformation on the global features 1 output by Bi-LSTM and left-multiply the normalized self-attention weight matrix Q to obtain word features with enhanced entity boundary information S = QH 1 W v S ∈ R T×s , where s is the length of the fused feature, is the linear transformation parameter matrix; Step S7, preliminarily predict the sequence labels according to the fused features, and use CRF to refine the predicted sequence to obtain the final label sequence; Step S8, in the model testing stage, use the network trained in the above steps to perform named entity recognition; The use of a bidirectional long short-term memory network in Step S2 to encode the word vectors at each time step in the sentence forward and backward respectively, and splice them to obtain the global features with context information includes: The number of hidden units used is h 1 Bidirectional long short-term memory network Bi-LSTM 1 For the input x at a given time step t t Perform forward and backward encoding, and record the forward hidden state of this time step as The reverse hidden state is recorded as Then, concatenate the hidden states in both directions and To get the hidden state It is a global feature with context information at a given time step t. For the input sequence {x 1 ,x 2 ,…x T }, remember Bi-LSTM 1 The output characteristics are The use of syntactic analysis technology in Step S3 to obtain the syntactic dependency tree of each sentence, and calculate the shortest dependency path between two words on the tree includes: For the input sample sequence {w 1 , w 2 , … w T}, use the dependency grammar analysis technology to perform syntactic analysis on it to obtain the dependency syntax tree of the sample sequence; for any two words a and b in the input sequence, the shortest dependency path SDP between them is {a, a 1 ,..., a m , c, b n ,..., b 1 , b}, where c represents their lowest common ancestor in the dependency syntax tree, a 1 ,..., a m represents the words between a and c on the SDP, and b 1 ,..., b n represents the words between b and c; if a and b represent the same word, the SDP is denoted as {a, b}; The obtaining of the top-down and bottom-up feature sequences of each word according to the shortest dependency path in Step S4 and inputting them into the LSTM network to calculate the local features of the words includes: For any two words a and b in the input text sequence {w 1 , w 2 ,... w T}, the shortest dependency path SDP between them is divided into two parts: the bottom-up sequence {a, a 1 ,..., a m , c} and {b, b 1 ,..., b n , c}; the top-down sequences {c, a m ,..., a 1 , a} and {c, b n ,..., b 1 , b}; if a and b represent the same word, then SDP is divided into: {a}; {b} two parts; Use a bi-directional long short-term memory network Bi-LSTM with the number of hidden units being h 2 to extract local relationship features between words from these two sequences; the input of each LSTM 2 unit is the concatenation of two parts, which is composed of 2 indicates that is the word w t in the Bi-LSTM 1 output, emb(d t ) represents the dependency relation type d t between the word w t and its governor on the dependency parse tree; Forward LSTM 2 According to the bottom-up sequences {a, a 1 ,..., a m , c} and {b, b 1 ,..., b n , c}, calculate the forward hidden states and Backward LSTM 2 According to the top-down sequences {c, a m ,..., a 1 , a} and {c, b n ,..., b 1 , b}, calculate the backward hidden states and Concatenate the hidden states in two directions ↑h t and ↓h t to obtain the local feature of the word w t The calculation of the relationship weights between two words through local feature dot products and normalization in Step S5 includes: For local features and local features perform a dot product to obtain word w i and word w j to obtain the closeness coefficient Calculate the relationship tightness coefficients between every two words in the text sequence in the same way, and organize all the relationship tightness coefficients into a matrix \(R\in\mathbb{R}\) T×T , where the \(i\)-th row of the matrix represents the relationship tightness coefficient between the word \(w\) i and each word in \(\{w\) 1 , \(w\) 2 ,... \(w\) T \}, and then normalize \(R\) row by row to obtain the self-attention weight matrix Q = Softmax(R).
2. The named entity recognition method based on syntactic dependency relationships according to claim 1, characterized in that, The use of the pre-trained Word2vec to map the one-hot word vectors into a defined low-dimensional space to obtain the word vectors of each word in Step S1 in the model training stage includes: Let the size of the dictionary be \(V\). Use the pre-trained Word2vec to map the one-hot word vector of dimension \(V\) to the defined low-dimensional space, and denote the dimension of the output word vector as \(d\). Specifically, for the input sample sequence \(\{w 1 ,\cdots,w t ,\cdots,w T \}\) of length \(T\), where \(w t \in\mathbb{R} 1×V , the output of the embedding layer is denoted as \(\{x 1 ,\cdots,x t ,\cdots,x T \}\), where \(x t \in\mathbb{R} 1×d .
3. The named entity recognition method based on syntactic dependency relationships according to claim 1, characterized in that, The preliminary prediction of the sequence labels according to the fused features in Step S7 and the refinement of the predicted sequence using CRF to obtain the final label sequence includes: Use the fused feature S to predict the sequence labels, and adjust the preliminarily predicted label sequence through CRF to obtain the final label sequence.
4. The named entity recognition method based on syntactic dependency relationships according to claim 1, characterized in that, In the model testing stage of step S8, using the network trained by the above steps for named entity recognition includes: After training the network parameters with supervised text, in the testing stage, use this network to identify the categories and boundaries of named entities.
Citation Information
Patent Citations
Word expression learning method based on syntax dependence relationship
CN108628834A
Knowledge graph relational data classification method based on syntactic attention neural network
CN111177394A