Named Entity Recognition Method, System, Electronic Device, and Storage Medium
By introducing gated-conditional random field (GCRF) in named entity recognition, dynamically adjusting the proportion of tag transmission scores and transfer scores, the error propagation problem when entity proximity is solved and the accuracy of named entity recognition is improved.
Patent Information
- Application Number
- CN202110220352.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-02-26
AI Technical Summary
When existing named entity recognition methods appear adjacent to entities, the recognition accuracy is greatly reduced, and there is a problem of error propagation.
A named entity recognition method based on gating-conditional random field (GCRF) is adopted to extract the context characteristics of the text, calculate the emission score, transfer score and gating coefficient, and dynamically adjust the specific gravity of the tag emission score and transfer score to alleviate the problem of error propagation.
Improves the accuracy of named entity recognition, especially in the case of entity proximity.
Smart Images

Figure CN113076751B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and in particular, to a named entity recognition method and system, an electronic device, and a storage medium. Background Art
[0002] The named entity recognition (NER) task is to identify entities with specific meanings in a text, which belongs to the category of sequence labeling problems.
[0003] So far, conditional random fields (CRF) have been used as the last step of the model in most NER tasks. CRF decodes the prediction label sequence with the highest probability using the Viterbi algorithm based on emission scores and transition scores. The transition score constraint makes the final predicted label smoother and the label transition more natural and reasonable.
[0004] In most scenarios, CRF can well identify "isolated" entities in a text. However, when entities appear adjacent to each other, the recognition accuracy of the entities decreases significantly. Part of the reason is that there is an error propagation problem when entities are adjacent. That is, when the previous entity is misrecognized, it is very likely to affect the recognition of adjacent entities, resulting in a significant decrease in recognition accuracy.
[0005] Therefore, how to provide a named entity recognition method and system, an electronic device, and a storage medium to improve the named entity recognition accuracy when there are adjacent entities in a text has become an urgent problem to be solved. Summary of the Invention
[0006] In view of the defects in the prior art, the present invention provides a named entity recognition method and system, an electronic device, and a storage medium.
[0007] The present invention provides a named entity recognition method, including:
[0008] Inputting the text information to be recognized into a named entity recognition model to obtain the entity recognition result output by the named entity recognition model;
[0009] wherein, the named entity recognition model is trained by sample text information and the corresponding label sequence; the text information includes: a text word vector sequence;
[0010] The named entity recognition model is used to determine the text feature sequence to be recognized and the emission score corresponding to the text information to be recognized, determine the gating coefficient based on the gated-conditional random field; and determine the entity recognition result based on the emission score, the gating coefficient, and the transition score.
[0011] The gating coefficient is the relative prediction confidence of the previous time step and the current time step in the feature sequence of the text to be recognized.
[0012] According to the named entity recognition method provided by the present invention, the named entity recognition model includes: a feature extraction layer, a feature processing layer, a gating processing layer, and a probability prediction layer;
[0013] The feature extraction layer is used to determine the context features of each time step in the word vector sequence of the text to be recognized, and determine the feature sequence of the text to be recognized based on the context features of each time step;
[0014] The feature processing layer is used to determine the emission score corresponding to each time step according to the feature sequence of the text to be recognized;
[0015] The gating processing layer is used to determine the prediction confidence of each time step according to the feature sequence of the text to be recognized, and determine the gating coefficient based on the prediction confidence of each time step;
[0016] The probability prediction layer is used to determine the entity label sequence and the corresponding probability of the text to be recognized according to the emission score, the transition score, and the gating coefficient, as the entity recognition result.
[0017] According to the named entity recognition method provided by the present invention, inputting the text information to be recognized into the named entity recognition model to obtain the entity recognition result output by the named entity recognition model specifically includes:
[0018] Inputting the word vector sequence of the text to be recognized into the feature extraction layer to obtain the feature sequence of the text to be recognized output by the feature extraction layer;
[0019] Inputting the feature sequence of the text to be recognized into the feature processing layer to obtain the emission score corresponding to each time step output by the feature processing layer;
[0020] Inputting the feature sequence of the text to be recognized into the gating processing layer to obtain the gating coefficient of each time step output by the gating processing layer;
[0021] Inputting the emission score, the transition score, and the gating coefficient into the probability prediction layer to obtain the entity recognition result output by the probability prediction layer.
[0022] According to the named entity recognition method provided by the present invention, the gating processing layer includes: a linear processing layer and a coefficient calculation layer;
[0023] The linear processing layer is used to transform the text features to be recognized at the current time step and the previous time step in the text feature sequence to be recognized into dimension 1, and determine the prediction confidence at the current time step and the previous time step through the Sigmoid activation function;
[0024] The coefficient calculation layer is used to calculate the gating coefficient at the current time step according to the prediction confidence at the current time step and the previous time step.
[0025] According to the named entity recognition method provided by the present invention, the feature extraction layer includes: a hidden information extraction layer and a feature sequence determination layer;
[0026] The hidden information extraction layer is used to determine the forward information and backward information of the word vectors at each time step in the text word vector sequence to be recognized, and determine the context features according to the forward information and backward information;
[0027] The feature sequence determination layer is used to determine the text feature sequence to be recognized according to the context features at each time step.
[0028] According to the named entity recognition method provided by the present invention, the text information further includes: text;
[0029] Correspondingly, the named entity recognition model further includes: a text preprocessing layer;
[0030] The text preprocessing layer is used to process the text to be recognized and determine the text word vector sequence corresponding to the text to be recognized.
[0031] According to the named entity recognition method provided by the present invention, the named entity recognition model further includes: a recognition result output layer;
[0032] The recognition result output layer is used to determine the entity label sequence with the best output probability in the entity label sequence as the target entity recognition result.
[0033] The present invention also provides a named entity recognition system, including:
[0034] A text recognition unit, configured to input text information to be recognized into a named entity recognition model, and obtain an entity recognition result output by the named entity recognition model;
[0035] Wherein, the named entity recognition model is trained by sample text information and a corresponding label sequence; the text information includes: a text word vector sequence;
[0036] The named entity recognition model is used to determine the to-be-recognized text feature sequence and emission score corresponding to the to-be-recognized text information, and determine the gating coefficient based on the gated conditional random field; based on the emission score, the gating coefficient and the transition score, determine the entity recognition result;
[0037] The gating coefficient is the relative prediction confidence between the previous time step and the current time step in the to-be-recognized text feature sequence.
[0038] The present invention also provides an electronic device, including a memory and a processor, and the processor and the memory complete communication with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute each step of the above-mentioned named entity recognition method by calling the program instructions.
[0039] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each step of the above-mentioned named entity recognition method is implemented.
[0040] The named entity recognition method, system, electronic device and storage medium provided by the present invention strengthen the judgment of entity boundaries during the recognition process through the named entity recognition model based on the gated conditional random field, and let the gating coefficient determine the proportion of the label emission score and the transition score, alleviating the problem of error propagation caused by an overly large incorrect label transition score, thereby improving the accuracy of named entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of the named entity recognition method provided by the present invention;
[0043] Figure 2 It is a schematic structural diagram of the named entity recognition model provided by the present invention;
[0044] Figure 3 It is a schematic structural diagram of the named entity recognition system provided by the present invention;
[0045] Figure 4 It is a schematic entity structure diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] In the Chinese-English named entity recognition task, to accurately recognize an entity, it is necessary to judge both the type of the entity and the boundary of the entity. According to a large amount of experimental data, when using CRF for sequence prediction, the accuracy of the named entity recognition task often depends on the accuracy of entity boundary judgment. That is to say, the judgment of entity boundary is much more difficult than the judgment of entity category. One of the reasons is that CRF uses the Viterbi algorithm to decode the prediction label sequence with the highest probability based on emission scores and transition scores. The constraint of the transition scores makes the final prediction label smoother and the label transition more natural and reasonable. However, when multiple entities appear adjacent to each other and the emission score of a certain entity is relatively large, the constraint of the transition scores may cause the prediction results of adjacent entities to also go wrong, resulting in error propagation.
[0048] To solve the problem of error propagation existing in the named entity recognition model, we propose a named entity recognition method based on gated-conditional random field (GCRF for short). The gated-conditional random field (GCRF) can replace the existing model using conditional random field (CRF).
[0049] Based on the above problems, the present invention provides a named entity recognition method based on gated conditional random field to strengthen the judgment of entity boundaries, thereby improving the accuracy of named entity recognition. The detailed method steps of the named entity recognition method based on gated-conditional random field in the process of named entity recognition are described as follows.
[0050] Figure 1 For the flowchart of the named entity recognition method provided by the present invention, as Figure 1 shown, the present invention provides a named entity recognition method, including:
[0051] Step S1, input the text information to be recognized into the named entity recognition model to obtain the entity recognition result output by the named entity recognition model;
[0052] Among them, the named entity recognition model is trained by sample text information and the corresponding label sequence; the text information includes: text word vector sequence.
[0053] The named entity recognition model is used to determine the to-be-recognized text feature sequence and emission score corresponding to the to-be-recognized text information, determine the gating coefficient based on the gated conditional random field; determine the entity recognition result based on the emission score, the gating coefficient and the transition score;
[0054] The gating coefficient is the relative prediction confidence of the previous time step and the current time step in the to-be-recognized text feature sequence.
[0055] Specifically, in step S1, it is necessary to input the to-be-recognized text information into a pre-trained named entity recognition model. The named entity recognition model is used to determine the to-be-recognized text feature sequence and emission score corresponding to the to-be-recognized text information, determine the gating coefficient corresponding to the to-be-recognized text feature sequence based on the gated conditional random field, and then obtain the entity recognition result output by the named entity recognition model according to the emission score, the transition score and the gating coefficient. The to-be-recognized text information includes: the to-be-recognized text word vector sequence.
[0056] Among them, when the named entity recognition model processes text information, the complete text information is divided into multiple time steps according to preset rules, and the gating coefficient is the relative prediction confidence of the previous time step and the current time step in the to-be-recognized text feature sequence.
[0057] It should be noted that before recognizing the to-be-recognized text information, it is also necessary to pre-train the named recognition model in advance. The named entity recognition model is trained by the sample text information and the corresponding label sequence, and the parameters of the model and the transition score matrix are determined based on a large number of samples. The sample text information includes: the sample text word vector sequence.
[0058] Among them, the judgment condition for determining that the named recognition model has been trained can be to determine that the named recognition model is updated according to the parameters obtained by training the named recognition model and the named recognition model converges, input the test sample into the named recognition model, determine that the test error of the input of the named recognition model is less than the preset value, or determine that the number of training iterations of the named recognition model meets the preset threshold. The specific method can be adjusted according to actual needs, and the present invention does not limit this.
[0059] The named entity recognition method provided by the present invention strengthens the judgment of entity boundaries in the recognition process through the named entity recognition model based on the gated conditional random field, allows the gating coefficient to determine the proportion of the label emission score and the transition score, and alleviates the problem of error propagation caused by an overly large incorrect label transition score, thereby improving the accuracy of named entity recognition.
[0060] Figure 2 This is the structural schematic diagram of the named entity recognition model provided by the present invention, as Figure 2As shown, optionally, according to the named entity recognition method provided by the present invention, the named entity recognition model includes: a feature extraction layer, a feature processing layer, a gating processing layer, and a probability prediction layer;
[0061] The feature extraction layer is used to determine the context features of each time step in the word vector sequence of the text to be recognized, and determine the feature sequence of the text to be recognized based on the context features of each time step;
[0062] The feature processing layer is used to determine the emission score corresponding to each time step according to the feature sequence of the text to be recognized;
[0063] The gating processing layer is used to determine the prediction confidence of each time step according to the feature sequence of the text to be recognized, and determine the gating coefficient based on the prediction confidence of each time step;
[0064] The probability prediction layer is used to determine the entity label sequence and the corresponding probability of the text to be recognized according to the emission score, the transition score, and the gating coefficient, as the entity recognition result.
[0065] Specifically, the named entity recognition model includes: a feature extraction layer, a feature processing layer, a gating processing layer, and a probability prediction layer;
[0066] The feature extraction layer is used to determine the context features of each time step in the word vector sequence of the text to be recognized, and combine the context features of each time step in the text to be recognized to form a feature sequence of the text to be recognized. For the input word vector sequence {x1, x2,... x n}, the output context feature sequence is denoted as {h1, h2,..., h n}.
[0067] It should be noted that the context features (context information) reflect the dependency relationship between words within a sentence. The specific extraction method can construct a forward LSTM and a backward LSTM to extract forward and backward feature information respectively, and combine them to form a BiLSTM, which can effectively use past and future input information and extract context features. Other feature extraction methods can also be used, such as: Bi-RNN, Transformer, etc. The specific method in actual use can be selected according to the actual situation, and the present invention does not limit this.
[0068] The feature processing layer is used to determine the emission score corresponding to each time step according to the feature sequence of the text to be recognized.
[0069] Perform a linear transformation on the context features of the word and normalize through Softmax to obtain the label emission score sequence predicted by the model where, Et The label emission score representing the word at the t-th time step, and E t ∈R m×1 , where m is the number of label types, and W e is the named entity recognition model parameter during normalization.
[0070]
[0071]
[0072] Denote the label transition score matrix as T ∈ R m×m .
[0073] The gating processing layer is used to determine the prediction confidence of each time step according to the feature sequence of the text to be recognized, and determine the gating coefficient based on the prediction confidence of each time step.
[0074] The feature sequence of the text to be recognized not only contains information of adjacent time steps, but also includes prediction tendency and prediction confidence. Based on the feature sequence of the text to be recognized, determine the prediction confidence of each time step, and calculate the ratio of the prediction confidence of the previous time step to the sum of the prediction confidences of the previous time step and the current time step in adjacent time steps as the gating coefficient.
[0075] The probability prediction layer is used to determine the entity label sequence corresponding to the text to be recognized and the corresponding probability according to the emission score, transition score and gating coefficient, as the entity recognition result.
[0076] Before recognizing the information of the text to be recognized, it is also necessary to pre-train the named entity recognition model. During model training, calculate the prediction probability of the true label Y of the named entity recognition model for the given text sequence X and optimize it. The calculation method of the prediction probability P(Y|X) of the true label Y is as shown in the formula:
[0077]
[0078] Among them, the true label sequence y t ∈{l1, l2, ···, l m} represents the true label corresponding to the t-th word, and P n represents all label sequences from the 1st word to the nth word including the true label path. Take the negative log-likelihood of the prediction probability P(Y|X) to obtain the loss function of the model:
[0079]
[0080] Denote the label transition score matrix as T ∈ R m×m , and the calculation method of the score s(X, Y) is:
[0081]
[0082] Among them, is the transition score with y1 as the start-of-sentence tag, is y t-1 to y t 's transition score, is the transition score as the end-of-sentence tag.
[0083] Calculate the output probability of the true tag sequence through the forward-backward algorithm, and optimize the output probability through the optimization algorithm to achieve the purpose of training the network parameters of the named entity recognition model.
[0084] Compared with the CRF network, based on the emission score, transition score and gating coefficient of the gated-conditional random field (GCRF), the prediction probability of the true tag can be determined when identifying the text to be recognized. The gating coefficient is used to determine the proportion of the tag emission score and the tag transition score at the current time step. It can effectively alleviate the problem of error propagation caused by an overly large incorrect tag transition score.
[0085] The named entity recognition method provided by the present invention extracts the context feature sequence of the text information to be recognized based on the named entity recognition model of the gated-conditional random field, determines the emission score, transition score and gating coefficient, and allows the gating coefficient to determine the proportion of the tag emission score and the transition score, strengthens the judgment of the entity boundary during the recognition process, alleviates the problem of error propagation caused by an overly large incorrect tag transition score, and thus improves the accuracy of named entity recognition.
[0086] Optionally, according to the named entity recognition method provided by the present invention, the inputting the text information to be recognized into the named entity recognition model to obtain the entity recognition result output by the named entity recognition model specifically includes:
[0087] Input the word vector sequence of the text to be recognized into the feature extraction layer to obtain the feature sequence of the text to be recognized output by the feature extraction layer;
[0088] Input the feature sequence of the text to be recognized into the feature processing layer to obtain the emission score corresponding to each time step output by the feature processing layer;
[0089] Input the feature sequence of the text to be recognized into the gating processing layer to obtain the gating coefficient of each time step output by the gating processing layer;
[0090] Input the emission score, the transition score and the gating coefficient into the probability prediction layer to obtain the entity recognition result output by the probability prediction layer.
[0091] Specifically, when inputting the text information to be recognized into the named entity recognition model to obtain the entity recognition result output by the named entity recognition model, the specific processing steps for the text information to be recognized are as follows:
[0092] Input the text word vector sequence {x1, x2, …, x n} of the text to be recognized into the feature extraction layer to obtain the text feature sequence {h1, h2, …, h n} output by the feature extraction layer.
[0093] Input the text feature sequence {h1, h2, …, h n} of the text to be recognized into the feature processing layer to obtain the emission scores corresponding to each time step output by the feature processing layer (emission score sequence) and the transition score T ∈ R m×m (transition score matrix).
[0094] Input the text feature sequence {h1, h2, …, h n} of the text to be recognized into the gating processing layer to obtain the gating coefficients at each time step output by the gating processing layer; g t represents the gating coefficient at time step t.
[0095] Input the emission scores the trained transition score matrix T ∈ R m×m and the gating coefficient g t into the probability prediction layer to obtain the entity recognition result output by the probability prediction layer.
[0096] It should be noted that after performing named entity recognition on the text information to be recognized, there are several different possibilities for the entity recognition result, and the entity label sequence and the corresponding prediction probability are different in each possibility. One can choose to use all possibilities as the entity recognition result, or further screen and only output some results, which can be adjusted according to actual needs, and the present invention does not make any limitations in this regard.
[0097] The named entity recognition method provided by the present invention extracts the context feature sequence of the text information to be recognized based on the named entity recognition model of gated conditional random field, determines the emission score, transition score, and gating coefficient, allows the gating coefficient to determine the proportion of the label emission score and the transition score, strengthens the judgment of the entity boundary during the recognition process, alleviates the problem of error propagation caused by an overly large incorrect label transition score, and thus improves the accuracy of named entity recognition.
[0098] Optionally, according to the named entity recognition method provided by the present invention, the gating processing layer includes: a linear processing layer and a coefficient calculation layer;
[0099] The linear processing layer is used to transform the text features to be recognized at the current time step and the previous time step in the text feature sequence to be recognized into dimension 1, and determine the prediction confidence at the current time step and the previous time step through the Sigmoid activation function;
[0100] The coefficient calculation layer is used to calculate the gating coefficient at the current time step according to the prediction confidence at the current time step and the previous time step.
[0101] Specifically, the gating processing layer in the named entity recognition model can be subdivided into a linear processing layer and a coefficient calculation layer.
[0102] The linear processing layer is used to transform the text features to be recognized at the current time step and the previous time step in the text feature sequence to be recognized into dimension 1, and compress the value range of the real number obtained after the dimensionality reduction transformation to between 0 and 1 through the Sigmoid activation function to determine the prediction confidence at the current time step and the previous time step.
[0103] For the prediction confidence c at time step t t ,
[0104] where W g is the named recognition body model parameter during the dimensionality reduction transformation.
[0105] According to the prediction confidence c at time step t t , and the prediction confidence c at time step t-1 t-1 determine the gating coefficient g at time step t t ,
[0106] During recognition, the transition gate coefficient g at time step t t represents the relative prediction confidence of time step t-1 compared to time step t. The higher g t is, the more accurate the model's prediction at time step t-1 is compared to the prediction at time step t. At this time, a higher weight should be assigned to the transition score at time step t to transfer the higher prediction confidence to time step t. Otherwise, a higher weight should be assigned to the emission score to avoid passing on the wrong prediction at time step t-1.
[0107] When the label prediction at the previous time step is incorrect, resulting in an incorrect transition score, even if this incorrect transition score is much larger than the emission score at the current time step, the gating coefficient will reduce the proportion of the incorrect transition score among the emission score and the transition score, reducing the influence of the previous time step on the current time step and alleviating the problem of error propagation.
[0108] The named entity recognition method provided by the present invention extracts the context feature sequence of the text information to be recognized based on the named entity recognition model of gated-conditional random field, determines the emission score and the transition score, calculates the prediction confidence of each time step, determines the gating coefficient based on the prediction confidence of the previous time step and the prediction confidence of the current time step, and allows the gating coefficient to determine the proportion of the label emission score and the transition score, strengthens the judgment of entity boundaries during the recognition process, alleviates the problem of error propagation caused by an overly large incorrect label transition score, and thus improves the accuracy of named entity recognition.
[0109] Optionally, according to the named entity recognition method provided by the present invention, the feature extraction layer includes: a hidden information extraction layer and a feature sequence determination layer;
[0110] The hidden information extraction layer is used to determine the forward information and backward information of the word vectors of each time step in the word vector sequence of the text to be recognized, and determine the context features according to the forward information and the backward information;
[0111] The feature sequence determination layer is used to determine the feature sequence of the text to be recognized according to the context features of each time step.
[0112] Specifically, in most named entity recognition tasks, the most commonly used solution at present is to use a model of a deep bidirectional temporal network connected to a conditional random field (BiLSTM-CRF). Although this classical model can solve most problems, it also has some disadvantages, such as error propagation.
[0113] To solve the error propagation problem existing in the BiLSTM-CRF model, we propose a named entity recognition method based on gated-conditional random field (abbreviated as GCRF). Based on the typical deep bidirectional temporal network (BiLSTM), it combines gated-conditional random field.
[0114] In the named entity recognition model, the feature extraction layer includes: a hidden information extraction layer and a feature sequence determination layer;
[0115] The hidden information extraction layer is used to process the word vector sequence {x1, x2, … x n} of the text to be recognized with dimension d, where x i ∈R 1×d . Use a BiLSTM with the number of hidden units h to perform forward and backward encoding on the input x t at a given time step t, and denote the forward hidden state of this time step as (forward information), the backward hidden state as (backward information), and concatenate the hidden states in both directions and to obtain the hidden state h t which is the global feature (context feature) of the context information at the given time step t.
[0116] The feature sequence determination layer is used to arrange the context features in the order of time steps after determining the context features of each time step in the sequence of text word vectors {x1, x2, …, x n} to determine the text feature sequence {h1, h2, …, h n} of the text to be recognized.
[0117] For the named entity recognition method provided by the present invention, after the named entity recognition model based on the gated conditional random field embeds the text using word vectors, it extracts the context feature sequence of the text information to be recognized based on the BiLSTM model, calculates the label emission scores, and then dynamically adjusts the ratio of the emission scores to the transition scores at each time step through the gating mechanism of the GCRF. Since the hidden state features of the BiLSTM not only contain information of adjacent time steps but also include the prediction tendency and prediction confidence, using the BiLSTM can effectively reflect the internal connection between the contexts of the text to be recognized. Determining the gating coefficient in this way helps to reflect the internal connection before and after label propagation and alleviates the problem of error propagation caused by the CRF in named entity recognition.
[0118] Optionally, according to the named entity recognition method provided by the present invention, the text information further includes: text;
[0119] Correspondingly, the named entity recognition model further includes: a text preprocessing layer;
[0120] The text preprocessing layer is used to process the text to be recognized and determine the sequence of text word vectors corresponding to the text to be recognized.
[0121] Specifically, when performing named entity recognition on the text to be recognized, the sequence of text word vectors to be recognized can be processed first, and the processed sequence of word vectors can be directly used as the input and training samples of the model. The method for determining word vectors based on text can be selected according to the actual situation, such as word2vec, Glove, FastText, Elmo, etc., and the present invention does not limit this.
[0122] In addition, it is also possible not to preprocess the text to be recognized in advance and directly use the text to be recognized and the sample text as the input and training samples of the model.
[0123] Correspondingly, at this time, the named entity recognition model further includes: a text preprocessing layer;
[0124] The text preprocessing layer is used to process the text to be recognized and determine the sequence of word vectors of the text to be recognized corresponding to the text to be recognized.
[0125] Specifically, the text preprocessing layer can use the pre-trained Word2vec to map the one-hot word vectors to a defined low-dimensional space to obtain the word vectors of each word.
[0126] Denote the size of the dictionary as V. Use the pre-trained Word2vec to map the one-hot word vectors of dimension V to a defined low-dimensional space, and denote the dimension of the output word vectors as d. For the input text sequence {w1, w2, …, w n} of length n to be recognized, the sequence of word vectors of the text to be recognized output by the text preprocessing layer is denoted as X = {x1, x2, …, x n}, where x i ∈R 1×d .
[0127] The named entity recognition method provided by the present invention realizes the preprocessing of the text to be recognized by adding a text preprocessing layer, which can transform the input of the named entity recognition model from the sequence of word vectors of the text to be recognized into the text to be recognized, and the model directly performs the operation of converting the text to be recognized into the corresponding sequence of word vectors, reducing the operation complexity of named entity recognition.
[0128] Optionally, according to the named entity recognition method provided by the present invention, the named entity recognition model further includes: a recognition result output layer;
[0129] The recognition result output layer is used to determine the entity label sequence with the best output probability in the entity label sequence as the target entity recognition result.
[0130] Specifically, the named entity recognition model further includes: a recognition result output layer;
[0131] The recognition result output layer is used to determine the entity label sequence with the best output probability in the entity label sequence as the target entity recognition result.
[0132] Preferably, after performing named entity recognition on the text information to be recognized, there are several different possibilities for the recognition results of the entities, and the entity label sequences and the corresponding prediction probabilities are different in each possibility. After determining each entity label sequence, preferably, the Viterbi algorithm is used to derive the label sequence with the best output probability as the prediction result, that is, the target entity recognition result.
[0133] After introducing the gating coefficient, when deriving the sequence label using the Viterbi algorithm, the sequence score is jointly determined by the emission score, the transition score, and the gating coefficient:
[0134]
[0135] In contrast, when deriving sequence tags using the Viterbi algorithm, without introducing a gating coefficient, the sequence score of the CRF is calculated from emission scores and transition scores. If the tag prediction at the previous time step is incorrect, the transition score from the previous time step to the current time step is obviously also incorrect. If this incorrect transition score is much larger than the emission score at the current time step, it will cause the tag at the current time step to be predicted incorrectly, that is, error propagation occurs.
[0136] The named entity recognition method provided by the present invention realizes the screening of several recognition result entity tag sequences through the recognition result output layer, and jointly determines to select the tag sequence with the best output probability as the prediction result, that is, the target entity recognition result, based on the emission score, the transition score, and the gating coefficient, ensuring that the final output of the named entity recognition model is the best result, without the need for manual screening, avoiding the occurrence of error propagation in the output result, and improving the recognition accuracy.
[0137] An example of processing a specific sentence in combination with the named entity recognition method provided by the present invention is described as follows:
[0138] For example, in a trained named entity recognition model, input the word vector sequence of the sentence "Eight dogs experienced ventricular tachycardia.", extract the context features of each time step through BiLSTM, transform the context features to obtain emission scores and gating coefficients, obtain the probabilities of all tag paths according to the trained transition scores, and finally calculate the optimal path, that is, the final tag sequence "O O O B-Disease I-Disease O" according to the Viterbi algorithm.
[0139] It should be noted that the above method is only a specific example to illustrate the present invention. In actual use, the method of extracting features and the algorithm for determining the optimal path can be adjusted according to the actual situation, and the present invention does not limit this.
[0140] Figure 3 For the structural schematic diagram of the named entity recognition system provided by the present invention, as Figure 3 shown, the present invention also provides a named entity recognition system, including:
[0141] A text recognition unit 310, configured to input the text information to be recognized into the named entity recognition model to obtain the entity recognition result output by the named entity recognition model;
[0142] Among them, the named entity recognition model is trained by sample text information and the corresponding tag sequence; the text information includes: text word vector sequence;
[0143] The named entity recognition model is used to determine the to-be-recognized text feature sequence and emission score corresponding to the to-be-recognized text information, and determine the gating coefficient based on a gated conditional random field; based on the emission score, the gating coefficient, and the transition score, determine the entity recognition result;
[0144] The gating coefficient is the relative prediction confidence of the previous time step and the current time step in the to-be-recognized text feature sequence.
[0145] Specifically, the text recognition unit 310 is configured to input the to-be-recognized text information into a pre-trained named entity recognition model. The named entity recognition model is used to determine the to-be-recognized text feature sequence corresponding to the to-be-recognized text information, determine the emission score, transition score, and gating coefficient corresponding to the to-be-recognized text feature sequence based on a gated conditional random field, and then obtain the entity recognition result output by the named entity recognition model according to the emission score, transition score, and gating coefficient. The to-be-recognized text information includes: a to-be-recognized text word vector sequence.
[0146] Among them, when the named entity recognition model processes text information, the complete text information is divided into multiple time steps according to a preset rule, and the gating coefficient is the relative prediction confidence of the previous time step and the current time step in the to-be-recognized text feature sequence.
[0147] It should be noted that before recognizing the to-be-recognized text information, it is also necessary to pre-train the named recognition model in advance. The named entity recognition model is trained by sample text information, and the sample text information includes: a sample text word vector sequence.
[0148] The named entity recognition system provided by the present invention strengthens the judgment of entity boundaries during the recognition process through a named entity recognition model based on a gated conditional random field, allows the gating coefficient to determine the proportion of the label emission score and the transition score, alleviates the problem of error propagation caused by an overly large incorrect label transition score, and thus improves the accuracy of named entity recognition.
[0149] It should be noted that the named entity recognition system provided by the embodiments of the present invention is used to execute the above-mentioned named entity recognition method, and its specific implementation manner is consistent with the method implementation manner, and will not be elaborated herein.
[0150] Figure 4 It is a schematic entity structure diagram of the electronic device provided by the present invention, as Figure 4As shown in the figure, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the above-mentioned named entity recognition method, including: inputting the text information to be recognized into the named entity recognition model to obtain the entity recognition result output by the named entity recognition model; wherein, the named entity recognition model is trained by the sample text information and the corresponding tag sequence; the text information includes: a text word vector sequence; the named entity recognition model is used to determine the text feature sequence to be recognized corresponding to the text information to be recognized and the emission score, and determine the gating coefficient based on the gated conditional random field; based on the emission score, the gating coefficient, and the transition score, determine the entity recognition result; the gating coefficient is the relative prediction confidence of the previous time step and the current time step in the text feature sequence to be recognized.
[0151] In addition, when the logical instructions in the above-mentioned memory 430 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0152] On the other hand, an embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the named entity recognition method provided by each of the above method embodiments. The method includes: inputting the text information to be recognized into a named entity recognition model to obtain an entity recognition result output by the named entity recognition model; wherein, the named entity recognition model is trained by sample text information and a corresponding tag sequence; the text information includes: a text word vector sequence; the named entity recognition model is used to determine a text feature sequence to be recognized and an emission score corresponding to the text information to be recognized, determine a gating coefficient based on a gated conditional random field; determine an entity recognition result based on the emission score, the gating coefficient, and a transition score; the gating coefficient is the relative prediction confidence between the previous time step and the current time step of the text feature sequence to be recognized.
[0153] On another aspect, an embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the named entity recognition method provided by each of the above embodiments. The method includes: inputting the text information to be recognized into a named entity recognition model to obtain an entity recognition result output by the named entity recognition model; wherein, the named entity recognition model is trained by sample text information and a corresponding tag sequence; the text information includes: a text word vector sequence; the named entity recognition model is used to determine a text feature sequence to be recognized and an emission score corresponding to the text information to be recognized, determine a gating coefficient based on a gated conditional random field; determine an entity recognition result based on the emission score, the gating coefficient, and a transition score; the gating coefficient is the relative prediction confidence between the previous time step and the current time step of the text feature sequence to be recognized.
[0154] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.
[0155] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A named entity recognition method, characterized in that, Including: Inputting the text information to be recognized into a named entity recognition model to obtain the entity recognition result output by the named entity recognition model; Among them, the named entity recognition model is trained by sample text information and the corresponding tag sequence; the text information includes: a text word vector sequence; The named entity recognition model is used to determine the text feature sequence to be recognized and the emission score corresponding to the text information to be recognized, determine the gating coefficient based on a gated conditional random field; and determine the entity recognition result based on the emission score, the gating coefficient, and the transition score; The gating coefficient is the relative prediction confidence of the previous time step and the current time step in the text feature sequence to be recognized; The named entity recognition model includes: a feature extraction layer, a feature processing layer, a gating processing layer, and a probability prediction layer; The feature extraction layer is used to determine the context feature of each time step in the text word vector sequence to be recognized, and determine the text feature sequence to be recognized based on the context feature of each time step; The feature processing layer is used to determine the emission score corresponding to each time step according to the text feature sequence to be recognized; The gating processing layer is used to determine the prediction confidence of each time step according to the text feature sequence to be recognized, and determine the gating coefficient based on the prediction confidence of each time step; the gating coefficient is used to determine the proportion of the label emission score and the label transition score at the current time step; The probability prediction layer is used to determine the entity label sequence and the corresponding probability of the text to be recognized according to the emission score, the transition score, and the gating coefficient as the entity recognition result.
2. The named entity recognition method according to claim 1, characterized in that, The step of inputting the text information to be recognized into the named entity recognition model to obtain the entity recognition result output by the named entity recognition model specifically includes: Inputting the text word vector sequence to be recognized into the feature extraction layer to obtain the text feature sequence to be recognized output by the feature extraction layer; Inputting the text feature sequence to be recognized into the feature processing layer to obtain the emission score corresponding to each time step output by the feature processing layer; Inputting the text feature sequence to be recognized into the gating processing layer to obtain the gating coefficient of each time step output by the gating processing layer; Inputting the emission score, the transition score, and the gating coefficient into the probability prediction layer to obtain the entity recognition result output by the probability prediction layer.
3. The named entity recognition method according to claim 1, wherein The gating processing layer includes: a linear processing layer and a coefficient calculation layer; The linear processing layer is used to transform the text feature to be recognized at the current time step and the previous time step in the text feature sequence to be recognized to dimension 1, and determine the prediction confidence of the current time step and the previous time step through a Sigmoid activation function; The coefficient calculation layer is used to calculate the gating coefficient of the current time step according to the prediction confidence of the current time step and the previous time step.
4. The named entity recognition method according to claim 2, wherein The feature extraction layer includes: a hidden information extraction layer and a feature sequence determination layer; The hidden information extraction layer is used to determine the forward information and backward information of the word vectors at each time step in the to-be-recognized text word vector sequence, and determine the context features according to the forward information and backward information; The feature sequence determination layer is used to determine the to-be-recognized text feature sequence according to the context features at each time step.
5. The named entity recognition method according to any one of claims 1-4, characterized in that The text information further includes: text; Correspondingly, the named entity recognition model further includes: a text preprocessing layer; The text preprocessing layer is used to process the to-be-recognized text and determine the to-be-recognized text word vector sequence corresponding to the to-be-recognized text.
6. The named entity recognition method according to any one of claims 2-4, characterized in that The named entity recognition model further includes: a recognition result output layer; The recognition result output layer is used to determine the entity label sequence with the best output probability in the entity label sequence as the target entity recognition result.
7. A named entity recognition system, characterized in that, Including: A text recognition unit, configured to input the to-be-recognized text information into the named entity recognition model to obtain the entity recognition result output by the named entity recognition model; Wherein, the named entity recognition model is trained by sample text information and the corresponding label sequence; the text information includes: text word vector sequence; The named entity recognition model is used to determine the to-be-recognized text feature sequence and emission score corresponding to the to-be-recognized text information, determine the gating coefficient based on the gated-conditional random field; based on the emission score, the gating coefficient and the transition score, determine the entity recognition result; The gating coefficient is the relative prediction confidence between the previous time step and the current time step in the to-be-recognized text feature sequence; The named entity recognition model includes: a feature extraction layer, a feature processing layer, a gating processing layer and a probability prediction layer; The feature extraction layer is used to determine the context features at each time step in the to-be-recognized text word vector sequence, and determine the to-be-recognized text feature sequence based on the context features at each time step; The feature processing layer is used to determine the emission score corresponding to each time step according to the to-be-recognized text feature sequence; The gating processing layer is used to determine the prediction confidence of each time step according to the to-be-recognized text feature sequence, and determine the gating coefficient based on the prediction confidence of each time step; the gating coefficient is used to determine the proportion of the label emission score and the label transition score at the current time step; The probability prediction layer is used to determine the entity label sequence corresponding to the to-be-recognized text and the corresponding probability according to the emission score, the transition score and the gating coefficient, as the entity recognition result.
8. An electronic device, characterized in that, Including a memory and a processor, the processor and the memory complete communication with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute the named entity recognition method according to any one of claims 1 to 6 by calling the program instructions.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the named entity recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Deep neural network-based legal language named entity identification method
CN109871535A
Chinese named entity recognition method fusing character and word features
CN111310470A