A Named Entity Recognition Method Based on Semantic and Syntactic Dependency Information
Through the BiLSTM-AELGCN-CRF model, combined with semantic and syntactic dependency information, the problems of missing information and erroneous recognition in the existing nomenclature recognition model are solved, and the accuracy of nomenclature recognition is improved.
Patent Information
- Application Number
- CN202210645695.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-06-08
AI Technical Summary
While using semantic information, the existing nomenclature recognition model ignores syntactic dependency information, resulting in missing information and incorrect entity recognition.
The BiLSTM-AELGCN-CRF model is used to combine semantic and syntactic dependency information, and syntactic dependency information is extracted through attention mechanism and graph convolution network to improve the accuracy of nomenclature recognition.
Effective use of syntactic dependency information improves the accuracy of nomenclature recognition and avoids entity recognition errors caused by the lack of semantic information and the propagation of incorrect syntactic information.
Smart Images

Figure CN114997170B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information extraction in natural language processing, and relates to a named entity recognition method based on semantic and syntactic dependency information, which is a named entity recognition method for making up for insufficient semantic information. Background Art
[0002] With the development of computer science and technology, the field of natural language processing has also made progress with practical significance and application prospects in the aspect of deep learning. For natural language processing, to achieve fine-grained and in-depth semantic understanding, simply relying on manual methods for data annotation and computing power investment cannot solve the essential problems.
[0003] Named entity recognition is a task of recognizing entities in unstructured text, such as tasks, locations, and organizations, etc., which plays a crucial role in the entity recognition of unstructured text. According to the research of previous researchers, named entity recognition has been widely applied in multiple scenarios such as relation extraction, question answering, event extraction, information retrieval, and knowledge graph construction.
[0004] For existing named entity recognition models, most of them focus on how to make full use of semantic information to improve the accuracy of named entity recognition and have achieved advanced results in deep learning-based methods. However, some research shows that syntactic dependency information can further improve the accuracy of named entity recognition to a certain extent and can make up for insufficient semantic information. And the existing technologies generally believe that the weights of each syntactic dependency information are the same, which will lead to the propagation of incorrect syntactic information dependencies and cause incorrect entity recognition. At the same time, the existing methods ignore the use of information in syntactic dependency labels, such as the information expressed by nominal subjects and prepositions themselves. Summary of the Invention
[0005] The first object of the present invention is to overcome the deficiencies of the prior art and propose a named entity recognition method based on semantic and syntactic dependency information. On the basis of the BiLSTM-CRF model, it integrates the Attention and Edge-Label based Graph Convolutional Network (AELGCN) model, called BiLSTM-AELGCN-CRF, which can effectively extract syntactic dependency information, introduce syntactic dependency relation types and attention mechanisms, and can effectively solve the problems of missing syntactic information and insufficient utilization rate existing in existing named entity recognition models. At the same time, it can avoid the problems of missing semantic information and incorrect syntactic information propagation to a certain extent, resulting in incorrect entity recognition, so as to improve the accuracy of named entity recognition.
[0006] The technical solution adopted by the present invention is as follows:
[0007] Step (1): Perform text analysis on the text, where the text analysis includes part-of-speech analysis and syntactic dependency analysis;
[0008] Step (2): Preprocess the part-of-speech information and syntactic dependency information, convert all part-of-speech information and syntactic dependency relationship types into one-hot vectors, and construct an adjacency matrix according to the dependency relationship directions between different words;
[0009] Step (3): Construct a named entity recognition model BiLSTM-AELGCN-CRF and train it. After the model parameters converge, obtain the optimal parameter model;
[0010] Step (4): Use the trained named entity recognition model BiLSTM-AELGCN-CRF to achieve entity prediction.
[0011] The second object of the present invention is to provide a named entity recognition device based on semantic and syntactic dependency information, which is characterized by including a memory, a processor, and a program of a named entity recognition method based on semantic and syntactic dependency information stored on the memory and executable on the processor. When the program of the named entity recognition method based on semantic and syntactic dependency information is executed by the processor, the steps of the named entity recognition method based on semantic and syntactic dependency information are realized.
[0012] The third object of the present invention is to provide a storage medium, which is characterized by storing a program of a named entity recognition method based on semantic and syntactic dependency information. When the program of the named entity recognition method based on semantic and syntactic dependency information is executed by a processor, the steps of the named entity recognition method based on semantic and syntactic dependency information are realized;
[0013] Another object of the present invention is to provide a computing device, including a memory and a processor, where an executable code is stored in the memory. When the processor executes the executable code, the above method is realized.
[0014] The beneficial effects of the present invention are:
[0015] 1. The present invention effectively utilizes additional syntactic dependency information. By introducing syntactic dependency relationship types and using the attention mechanism, it can effectively obtain syntactic information beneficial to entity recognition at different levels, and at the same time can, to a certain extent, avoid the problem of partial entity recognition errors caused by semantic information loss and incorrect syntactic information propagation.
[0016] 2. The present invention proposes a new hybrid deep neural network model, which fuses an Attention and Edge-Label based Graph Convolutional Network (AELGCN) model on the basis of the BiLSTM-CRF model, called BiLSTM-AELGCN-CRF. The BiLSTM model is used to extract semantic information, and the AELGCN model is used to fully and effectively extract important syntactic dependency information, effectively solving the problems of missing semantic and syntactic information and insufficient utilization rate in existing named entity recognition models. Finally, the transfer characteristics of the CRF model are used to perform annotation probability sorting on the output state sequence to obtain the result of entity recognition, complete entity recognition, and improve the accuracy of named entity recognition to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart related to the present invention;
[0018] Figure 2 is the overall model structure diagram;
[0019] Figure 3 is an example diagram of syntactic dependency relationship. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following further describes in detail the specific implementation of the present invention in conjunction with the drawings. The specific process is described as Figure 1 shown, where:
[0021] Step (1): Perform text analysis on the text data to obtain part-of-speech information and syntactic dependency information; the text analysis includes part-of-speech analysis and syntactic dependency analysis;
[0022] The part-of-speech information includes proper nouns, verbs, articles, and prepositions.
[0023] The syntactic dependency information includes the type of syntactic dependency relationship and the dependency relationship direction between different words; the types of syntactic dependency relationships are, for example, nsubj (nominal subject), advmod (adverbial modifier), tmod (temporal modifier), mmod (modal verb), etc. Figure 3 is an example diagram of syntactic dependency relationship.
[0024] Step (2): Preprocess the part-of-speech information and syntactic dependency information, convert all part-of-speech information and syntactic dependency relationship types into one-hot (one-hot encoding) vectors, and construct an adjacency matrix according to the dependency relationship direction between different words.
[0025] Step (3): Construct a named entity recognition model BiLSTM-AELGCN-CRF and train it.
[0026] The overall structure of the named entity recognition model BiLSTM-AELGCN-CRF is as Figure 2 shown.
[0027] The named entity recognition model BiLSTM-AELGCN-CRF includes an input representation layer, a semantic extraction layer, a syntactic dependency extraction layer, and an output layer.
[0028] (1) Input Representation Layer:
[0029] The text data of the entity to be recognized is converted into a text data one-hot vector by one-hot encoding, and then word embedding processing is performed to obtain the word vector of each word; at the same time, word embedding processing is performed on the one-hot vectors of the part-of-speech information and the syntactic dependency relationship type of the text data of the current entity to be recognized, obtaining a part-of-speech information embedding vector and a syntactic dependency relationship type embedding vector; the word embedding processing is to map the one-hot vector to a defined low-dimensional space using a pre-trained word embedding vector algorithm;
[0030] For example, for an input sample sequence of length 6 "Abramov had an accident in Moscow", the output of the input representation layer is denoted as {x0, x1,..., x5}, where the t-th vector x t ∈R 1×d .
[0031] (2) Semantic Extraction Layer (Context Encoder), which extracts semantic information through BiLSTM. BiLSTM encodes the word vectors at each time step forward and backward respectively, and concatenates them to obtain the global feature of the context information; specifically including:
[0032] Use a BiLSTM with m hidden units to encode the word vector x t at a given time step t forward and backward, and denote the forward hidden state of this time step as h lt ∈R 1×m , the backward hidden state is denoted as h rt ∈R 1×m , and then concatenate the forward and backward hidden states to obtain the global feature h t = [h lt , h rt ∈R 2×m .
[0033] For example, for the input sequence {x0, x1, …, x5}, the output of the BiLSTM is {h0, h1, …, h5}.
[0034] (3) Syntactic dependency extraction layer. The graph convolutional network GCN is used to perform weighted aggregation according to the syntactic information and semantic information between two words with syntactic dependency relationships, so as to obtain word embedding vectors with syntactic and semantic information. The syntactic dependency extraction layer includes N layers of cascaded AELGCN. The AELGCN includes a node joint update module, an edge update module, and M layers of cascaded Attention Guided GCN (AGGCN). Its internal structure is as Figure 2 shown on the right half, which includes the following steps:
[0035] a) Node joint update module (Edge-Aware Node Joint Update Module, EANJU), which is used to perform weighted aggregation on the neighbor node information according to the global features of the context information output by the BiLSTM and the adjacency matrix information in step (2). Specifically:
[0036] The output of the semantic extraction layer (BiLSTM) within a continuous period of time is denoted as H l-1 , where l represents the current layer number of the AELGCN. When l - 1 is 0, it means that the current input vector is the output vector of the BiLSTM, that is, H 0 ={…, h t ,…}.
[0037] Update the information of its own node according to the adjacency matrix information. The formula is as follows:
[0038]
[0039] where Pool represents the aggregation operation, A is the adjacency matrix, W l-1 is the weight matrix of AH l-1 , E l-1 represents the set of syntactic dependency relationship type embedding vectors, H l-1 is the output vector of the previous layer of AELGCN, and σ represents the activation function, usually RELU. is the node information of the i-th dimension after adding the syntactic dependency relationship type information. p represents the dimension size of the syntactic dependency relationship type embedding vector, where The calculation method of is as follows:
[0040]
[0041] where represents the set of syntactic dependency relationship type embedding vectors of the i-th vector dimension, i = 1, …, p, and W is weight matrix
[0042] For example, for the output of BiLSTM is denoted as H overall l-1 , for where i represents the i-th word, l represents the current layer number of AELGCN. When l - 1 is 0, it means the current input vector is the output vector of BiLSTM, that is, H 0 = H0, where H0 is {h0, h1, …, h5}
[0043] b) Edge Update Module (Node-Aware Edge Update Module), which is used to input the output H of the node joint update module l into the edge update module to update the syntactic dependency type embedding vector from the i-th word to the j-th word:
[0044]
[0045] where is the output vector of the i-th node of the node joint update module in step a), W u is the weight matrix represents the embedding vector of all dimensions of the syntactic dependency type from the i-th word to the j-th word is the concatenation operation of matrices
[0046] Combined by to obtain the set E of the syntactic dependency type embedding vectors of the current layer l , and then input it into the node joint update module of the next layer of AELGCN
[0047] c) Attention Guided GCN (AGGCN) includes an attention guide layer, a dense connection layer, and a linear combination layer:
[0048] In the attention guide layer (Attention Guide Layer), the adjacency matrix is converted into an attention-guided adjacency matrix, and each represents the weight of the edge from node i to node j can be constructed through the self-attention mechanism and used as the input for subsequent calculations The calculation method is as follows:
[0049]
[0050] where and represent the weight matrices of Query and Key of H M , H M-1Represents the output of the AGGCN of the (M - 1)-th layer. When M = 1, it is H M-1 Is H l , H l Represents the output of the node joint update module, Represents the t-th attention-guided adjacency matrix, d head Represents The calculated vector dimension.
[0051] In the Densely Connected Layer, which contains multiple sub-layers, and the number of these sub-layers is equal to the number of attention-guided adjacency matrices;
[0052] Based on the attention-guided adjacency matrix Calculate the node vectors in different feature spaces The specific calculation formula is as follows:
[0053]
[0054] Where Is the output vector of the i-th node of the t-th sub-layer, i represents the i-th word, j represents the j-th word, Represents the weight vector representation of word i and word j in the t-th attention-guided adjacency matrix, Is the weight matrix, Is the bias vector, σ is the activation function, usually RELU, t = 1,..., N.
[0055] For Its calculation formula is as follows:
[0056]
[0057] Among them, Is obtained by combining the vectors of the j-th node in N After vectors, Represents the output vector H of the node joint update module l The vector of the j-th node in it.
[0058] In the Linear Combination Layer, this layer combines the final output of the densely connected layer in the way of a linear layer. Its calculation method is as follows:
[0059] h out = W out h out + b out
[0060] Where W out Is the weight matrix, hout is the concatenation of the outputs of a total of N densely connected layers, i.e., h out = [a 0 , …, a M , and b out is the bias vector.
[0061] Finally, after passing through the repeated N layers of AELGCN, the final output h final of the syntactic dependency extraction layer is obtained.
[0062] (4) Output layer: Use the conditional random field (CRF) for prediction. Among them, CRF is a sequential labeling algorithm. This layer uses the output of CRF for sequence label prediction, and adjusts the initially predicted label sequence through CRF to obtain the final label sequence. The formula is as follows:
[0063]
[0064]
[0065] where P(y|x) represents the probability that x is the y label, represents the transition score from label y i to label y i+1 , which is learned through the gradient descent algorithm. exp represents the exponential function with the natural constant e as the base, and F x is the emission score matrix, represents the score of label y i at the i-th label, and its value is the vector representation h final of the output of the syntactic dependency extraction layer.
[0066] The parameter learning process is based on maximizing the log-likelihood function to solve the model parameters, and the loss function is as follows:
[0067] log(P(x|y)) = P(x|y) - log(score(x, y))
[0068] The minimum value of the loss function is iteratively found through the gradient descent optimization algorithm to complete the parameter training process of the neural network. The trained model can be used for prediction, and the prediction process of the conditional random field is based on the Viterbi algorithm to solve the optimal prediction sequence y * , y * is the entity label result sequence corresponding to each input word, i.e.,
[0069] y * = argmax(score(x, y))
[0070] where argmax is the function to find the maximum of the independent variable.
[0071] Step 4: Use the trained named entity recognition model BiLSTM-AELGCN-CRF to achieve entity prediction.
[0072] The performance of the present invention was evaluated on four benchmark datasets: the Catalan and Spanish datasets of SemEval 2010 Task 1, and the English and Chinese datasets of OntoNotes 5.0. We selected these four datasets because they have explicit syntactic dependency annotations, which allow us to evaluate the effectiveness of our method when using dependency trees of different qualities. For the SemEval 2010 Task 1 dataset, there are 4 entity types: PER, LOC, ORG, and MISC. For the OntoNotes 5.0 dataset, there are a total of 18 entity types, such as ORG, LOC, PER, PRODUCT, TIME, etc. The following table shows the number of sentences and entities in the four datasets:
[0073]
[0074]
[0075] The following table shows the experimental results of entity prediction of the present invention on the above four datasets:
[0076]
[0077] SemEval2010 Spanish P R F1 BiLSTM-CRF 78.33 69.89 73.87 BiLSTM-GCN-CRF 84.10 79.88 81.93 GCN-BiLSTM-CRF 84.36 79.48 81.85 DGLSTM-CRF 84.05 82.90 83.47 Syn-LSTM-CRF 86.22 84.24 85.09 BiLSTM-AELGCN-CRF 88.75 87.52 88.13
[0078]
[0079]
[0080] Ontonotes5.0 English P R F1 Chiu and Nichols 86.04 86.53 86.28 Li et al. 88.00 86.50 87.21 BiLSTM-CRF 87.21 86.93 87.07 BiLSTM-GCN-CRF 88.30 88.06 88.18 GCN-BiLSTM-CRF 88.56 88.76 88.66 DGLSTM-CRF 88.53 88.50 88.50 Syn-LSTM-CRF 88.96 89.13 89.04 BiLSTM-AELGCN-CRF 88.72 89.79 89.25
[0081] In the above table of experimental results of named entity recognition prediction, BiLSTM-AELGCN-CRF is a named entity recognition method based on semantic and syntactic dependency information in the present invention. The F1 metric is used as the evaluation metric for named entity recognition performance. The formula for the F1 evaluation metric is as follows:
[0082]
[0083] Among them, P represents precision, and R represents recall. They represent the entity recognition precision rate and recall rate respectively, and evaluate whether the entity classification of the model is accurate and the proportion of positive examples judged by the classifier in all positive examples. It can be seen from the above formula that the F1-score is an evaluation metric that combines the precision rate and recall rate of the classifier.
Claims
1. A named entity recognition method based on semantic and syntactic dependency information, characterized in that It includes the following steps: Step (1): Perform text analysis on the text data of the entity to be recognized to obtain part-of-speech information and syntactic dependency information; the text analysis includes part-of-speech analysis and syntactic dependency analysis; wherein the syntactic dependency information includes the type of syntactic dependency relationship and the dependency relationship direction between different words; Step (2): Preprocess the part-of-speech information and syntactic dependency information, convert all part-of-speech information and syntactic dependency relationship types into one-hot vectors, and construct an adjacency matrix according to the dependency relationship direction between different words; Step (3): Construct a named entity recognition model BiLSTM-AELGCN-CRF and train it; The named entity recognition model BiLSTM-AELGCN-CRF includes an input representation layer, a semantic extraction layer, a syntactic dependency extraction layer, and an output layer; (1) Input representation layer: Convert the text data of the entity to be recognized into a text data one-hot vector by one-hot encoding, and then perform word embedding processing to obtain the word vector of each word; at the same time, perform word embedding processing on the one-hot vectors of the part-of-speech information and syntactic dependency relationship type of the current text data of the entity to be recognized to obtain a part-of-speech information embedding vector and a syntactic dependency relationship type embedding vector; (2) Semantic extraction layer, extract semantic information through BiLSTM, and the BiLSTM encodes the word vectors at each time step forward and backward respectively, and splices them to obtain the global feature of the context information; (3) Syntactic dependency extraction layer, use the graph convolutional network GCN to perform weighted aggregation according to the syntactic information and semantic information between two words with syntactic dependency relationships to obtain a word embedding vector with syntactic and semantic information; the syntactic dependency extraction layer includes N layers of cascaded AELGCN, where AELGCN includes a node joint update module, an edge update module, and M layers of cascaded Attention Guided GCN, specifically as follows: The node joint update module is used to perform weighted aggregation on the neighbor node information according to the global feature of the context information output by BiLSTM and the adjacency matrix information in step (2); Specifically: Denote the output of the semantic extraction layer over a continuous period of time as a whole as H l-1 , where l represents the current layer number of AELGCN. When l - 1 is 0, it means the current input vector is the output vector of BiLSTM, that is, H 0 = {…, h t ,…}. Update the information of its own node according to the adjacency matrix information, and the formula is as follows: Among them, EANJU represents the output of the node joint update module, Pool represents the aggregation operation, A represents the adjacency matrix, and W l-1 represents AH l-1 's weight matrix, E l-1 represents the set of syntactic dependency type embedding vectors, H l-1 represents the output vector of the previous layer of AELGCN, and σ represents the activation function; is the node information of the i-th dimension after adding syntactic dependency type information, p represents the dimension size of the syntactic dependency type embedding vector, where is calculated as follows: Among them represents the set of syntactic dependency relation type embedding vectors for the i-th vector dimension, where i = 1, …, p, and W represents weight matrix of Edge update module, which is used to update the output H of the node joint update module l Update the syntactic dependency type embedding vector of the i-th word to the j-th word; specifically: Among them is the output vector of the i-th node of the node joint update module in step a), W u is the weight matrix, is the concatenation operation of matrices, represents the embedding vectors of all dimensions of the syntactic dependency relationship type from the i-th word to the j-th word output by the l-th layer of AELGCN; Combined from to obtain the set E of syntactic dependency relation type embedding vectors of the current layer l , and then input it into the node joint update module of the next layer of AELGCN; Attention Guided GCN includes an attention guidance layer, a dense connection layer, and a linear combination layer: The attention guidance layer is used to convert the adjacency matrix into an attention-guided adjacency matrix; specifically: Among them and represent the weight matrices of Query and Key of H M ; H M-1 represents the output of the AGGCN of the (M - 1)-th layer. When M is 1, H M-1 is H l ; H l represents the output of the node joint update module, represents the t-th attention-guided adjacency matrix, and d head represents the calculated vector dimension; The dense connection layer contains multiple sub-layers, and the number of these sub-layers is equal to the number of attention-guided adjacency matrices; according to the attention-guided adjacency matrix output by the attention guidance layer, calculate the node vectors in different feature spaces; specifically: Among them is the output vector of the i-th node of the t-th sub-layer, where i represents the i-th word and j represents the j-th word. represents the weight vector representation of word i and word j in the t-th attention-guided adjacency matrix. is the weight matrix. is the bias vector, σ is the activation function, and t = 1, …, N; For Its calculation formula is as follows: Among them, is the vector of the j-th node in the merged N posterior vectors, denotes the vector of the j-th node in the output vector H l of the node joint update module The linear combination layer combines the final output of the dense connection layer in a linear layer manner; Finally, after repeating N layers of AELGCN, the final output of the syntactic dependency extraction layer is obtained; (4) Output layer: Use conditional random field for prediction to obtain the final label sequence; Step 4: Use the trained named entity recognition model BiLSTM-AELGCN-CRF to realize entity prediction.
2. The method according to claim 1, characterized in that The semantic extraction layer in the named entity recognition model BiLSTM-AELGCN-CRF described in step (3) is specifically as follows: Encode the word vector x at a given time step t using a BiLSTM with m hidden units t in both the forward and backward directions, and denote the forward hidden state at this time step as h lt ∈R 1×m , and the backward hidden state as h rt ∈R 1×m . Then concatenate the forward and backward hidden states to obtain the global feature h with the context information of the given time step t t = [h lt , h rt ∈ R 2×m .
3. The method according to claim 1, characterized in thatThe specific linear combination layer in the Attention Guided GCN of the syntactic dependency extraction layer in the named entity recognition model BiLSTM-AELGCN-CRF described in step (3) is: h out = W out h out + b out where W out is the weight matrix, h out is the concatenation of the outputs of a total of N densely connected layers, i.e., h out = [a 0 , …, a M , and b out is the bias vector.
4. The method according to claim 3, wherein The output layer in the named entity recognition model BiLSTM-AELGCN-CRF described in step (3) is specifically as follows: where P(y|x) represents the probability that x is the label of y, represents y i the transition score from label i+1 to label y, learned through the gradient descent algorithm; exp represents the exponential function with the natural constant e as the base, and F x is the emission score matrix, represents the score of label y i on the i-th label, whose value is the vector representation h of the output of the syntactic dependency extraction layer final .
5. The method according to claim 4, wherein The parameter learning process of the named entity recognition model BiLSTM-AELGCN-CRF described in step (3) is to solve the model parameters based on maximizing the log-likelihood function, and the loss function is as follows: log(P(y|x)) = P(x|y) - log(score(x,y)) The minimum value of the loss function is iteratively found through the gradient descent optimization algorithm to complete the parameter training process of the neural network; The trained model can be used for prediction. The prediction process of the conditional random field is based on the Viterbi algorithm to solve the optimal prediction sequence y * , y * is the entity label result sequence corresponding to each input word, that is y * = argmax(score(x, y)) where argmax is the function to find the maximum of the independent variable.
6. A named entity recognition device based on semantic and syntactic dependency information, wherein It includes a memory, a processor, and a program of a named entity recognition method based on semantic and syntactic dependency information stored in the memory and executable on the processor. When the program of the named entity recognition method based on semantic and syntactic dependency information is executed by the processor, it implements the steps of a named entity recognition method based on semantic and syntactic dependency information as described in any one of claims 1-5.
Citation Information
Patent Citations
Text named entity information identification method based on syntactic guidance
CN112989796A
Associated information fused program language recognition system and method
CN114330338A