A named entity recognition method fusing entity boundary information
Patent Information
- Application Number
- CN202410063445.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-01-17
AI Technical Summary
[0005]本发明为克服上述的不足之处,目的在于提供一种融合实体边界信息的命名实体识别方法,本发明设计了一种融合实体边界信息的NER方法,通过增加中文语义边界信息来提高抽取实体的准确性,可以提高了中文实体抽取的准确率和鲁棒性,从而解决了以往NER模型在中文数据集上准确率低、可用性低的问题
[0041]本发明的有益效果在于:(1)本发明针对小领域数据集,取得了显著的成果;采用Focal Loss显著提升了数据标签较少类别的准确率,一定程度上解决了标签不平衡问题;成功将边界语义信息与字符语义信息相结合,提高了方法在整体预测结果上的准确率;(2)本发明提高了中文实体抽取的准确率和鲁棒性,从而解决了以往NER模型在中文数据集上准确率低、可用性低的问题。
Smart Images

Figure CN117744659B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of named entity recognition technology based on deep learning, and in particular to a named entity recognition method that integrates entity boundary information. Background Technology
[0002] Extracting entity information from existing text plays a crucial role in various applications of natural language processing, such as question answering, search engine optimization, and voice assistants. In these applications, dictionary-based and rule-based methods are commonly used, achieving good results for texts in specific domains or formats. However, these methods suffer from limited applicability and maintenance difficulties, thus limiting their widespread adoption.
[0003] Compared to traditional dictionary- and rule-based NER methods, deep learning-based methods have proven to achieve better results. A typical approach utilizes Bidirectional Long-Short-Term Memory (BiLSTM) networks to capture the relationships between sentences, then uses a softmax function to output the probability of each character's corresponding label. However, this method suffers from the inability to capture the dependencies between labels, resulting in poor performance in Chinese NER. Subsequent research has proposed models combining BiLSTM and Conditional Random Fields (CRF), but these models have still not yielded satisfactory results in Chinese NER tasks. Recent work has used BERT (Bidirectional Encoder Representations from Transformers) pre-trained models for word embedding, then inputs the resulting word embeddings into recurrent neural network models for further processing. These models have significantly improved accuracy compared to previous models, but still cannot effectively handle Chinese NER tasks.
[0004] In summary, existing models have not achieved satisfactory results for Chinese Named Entity Recognition (NER) tasks. This is because these models are designed for English, while Chinese has a larger vocabulary, lacks natural word separators, and its word segmentation and annotation standards are more complex. Therefore, traditional NER methods struggle to achieve good results on Chinese. Consequently, a named entity recognition method that incorporates entity boundary information is urgently needed to address these issues. Summary of the Invention
[0005] To overcome the above-mentioned shortcomings, this invention aims to provide a named entity recognition method that integrates entity boundary information. This invention designs a NER method that integrates entity boundary information, which improves the accuracy of entity extraction by adding Chinese semantic boundary information. This can improve the accuracy and robustness of Chinese entity extraction, thereby solving the problems of low accuracy and low usability of previous NER models on Chinese datasets.
[0006] This invention achieves the above objective through the following technical solution: a named entity recognition method that integrates entity boundary information, comprising the following steps:
[0007] (1) Input the Chinese text into the BERT model and obtain the corresponding word embedding information;
[0008] (2) Input the word embedding information into the boundary information extraction network model to obtain the boundary information and obtain the output feature feature_left;
[0009] (3) Input the word embedding information into the character feature extraction network model, learn the semantic information through the attention mechanism, and use the BiLSTM model to capture the dependency relationship, and output the obtained feature_right;
[0010] (4) The features left and right obtained above are fused and input into the CRF for final processing to learn the dependency relationship between the labels and output entity prediction information, i.e. the labels corresponding to the Chinese text.
[0011] Preferably, in step (1), the Chinese text is converted into the corresponding index in vocab using the vocab.txt file provided with the BERT model, and the obtained index is input into the BERT model to obtain the corresponding word embedding vector; specifically, the steps are as follows:
[0012] (1.1) First, the data needs to be preprocessed: use the Chinese vocab.txt provided in the BERT model as the vocabulary, and convert the input Chinese text into the index of the corresponding word in the vocabulary; at the same time, obtain the maximum length of all sentences in the batch, and input a batch of Chinese sentences into the BERT model;
[0013] (1.2) The BERT pre-trained model is used as the word embedding model, where the hidden layer dimension of the model is 768. In the word embedding process, the Chinese text obtained in step (1.1) is input into the BERT model. The BERT model is used to capture the complex dependencies and semantic associations between words and output word embedding vectors containing rich language information, with dimensions of [batch_size, max_sen_len, hidden_dim].
[0014] Preferably, step (2) specifically includes the following steps:
[0015] (2.1) Modify the original labels of the dataset by labeling the start and end positions of entities as 1, i.e., the start and end labels are 1; while other text is labeled as 0, i.e., the remaining labels are 0; (2.2) Input the word embedding vectors obtained in step (1.2) into the BiLSTM model to capture long-distance dependencies, as shown in the following formula:
[0016]
[0017]
[0018]
[0019]
[0020] Formulas (1) and (2) represent the outputs of the forward and backward layers of the biLSTM model, and formula (3) represents the fusion of the outputs in the two directions. Finally, the output feature of the boundary information extraction network model can be obtained, which is called feature_left. The dimension of this feature is [batch_size, max_sen_len, hidden_dim].
[0021] (2.3) Set up a fully connected layer with a hidden layer dimension of 2, and output the feature_left obtained in step (2.2) into the fully connected layer. The output vector represents the probability of each character being the start or end label. The output dimension after passing through the fully connected layer is [batch_size, max_sen_len, 2].
[0022] (2.4) Input the new label obtained in step (2.1) and the character prediction result obtained in step (2.3) into the loss function, and use the gradient descent optimization algorithm to update the model parameters to minimize the loss function; Focal Loss is used as the loss function, and the formula is as follows:
[0023]
[0024]
[0025] FL(P t )=-α t (1-P t ) γ log(P t (7)
[0026] Formula (5) is the binary classification cross-adsorption loss function, where p is the character label prediction result and y is the label value; combining formulas (5) and (6), the binary classification cross-adsorption loss function is rewritten as CE(p,y)=CE(P t ) = -log(P t ); by adding α t and (1-P) t ) γ To adjust the weights of positive and negative samples, because P t ∈[0,1], γ∈[0,5], when P t →1, Adjust (1-P) t ) γ →0, the calculated Focal Loss will be automatically used by PyTorch's backward() function to calculate all gradients, and all gradients will be automatically accumulated for optimization of model parameters.
[0027] Preferably, step (3) specifically includes the following steps:
[0028] (3.1) Input the word embedding vectors obtained through the BERT model into the attention layer of the character feature extraction network model, and output word vectors with more semantic information; the formula used in the attention layer is as follows:
[0029]
[0030] Where Q, K, and V represent Query, Key, and Value, respectively. These three parameters are obtained by multiplying the embeds obtained in step 1.2) by W. Q W k and W V The three trainable parameter matrices are obtained; the output size after passing through the attention layer is [batch_size, max_sen_len, hidden_dim].
[0031] (3.2) Input the embedding information output by the attention layer into the BiLSTM layer to learn long-range dependencies and output feature_right. The size of feature_right is [batch_size, max_sen_len, hidden_dim].
[0032] Preferably, step (4) specifically includes the following steps:
[0033] (4.1) The feature_left obtained in step (2) and the feature_right obtained in step (3) are fused and concatenated to obtain a new feature representation vector. The new feature representation vector integrates the boundary information learned by the boundary information extraction network model and the character semantic information learned by the character feature extraction network model, which enhances the semantics learned by the model from two aspects and increases the accuracy of the model. The feature fusion formula is as follows:
[0034] feature=β1*feature_left+β2*feature_right (9)
[0035] Where β1 and β2 are weight parameters, β1 and β2∈[0,1], and feature_left and feature_right represent the features learned by the models on both sides;
[0036] (4.2) Construct a fully connected layer with a hidden layer dimension of the label size (target_size) to output the probability of each character corresponding to the label; and input the feature output in step (4.1) into the fully connected layer to predict the probability of each character in the fully connected layer. The output size is [batch_size, max_sen_len, target_size], and the output represents the probability of each character corresponding to different labels.
[0037] (4.3) Input the output obtained in step (4.2) into the CRF layer to predict the label; in the CRF layer, the CRF layer maintains an emission matrix and a transition matrix. The emission matrix represents the probability of each character corresponding to each label, and the transition matrix represents the score of transitioning from label1 to label2; combine the above two matrices and train the CRF using maximum conditional likelihood estimation, as follows:
[0038]
[0039]
[0040] Where Y = (y1, y2, ..., y n Let X = (x1, x2, ..., xn) represent the label sequence of the model. n ) represents a sentence, x i and y i Represents a character and its corresponding label, y′ i The character x represents the model's prediction.i Corresponding tags.
[0041] The beneficial effects of this invention are as follows: (1) This invention has achieved remarkable results for small domain datasets; Focal Loss significantly improves the accuracy of data with fewer categories and solves the label imbalance problem to a certain extent; it successfully combines boundary semantic information with character semantic information, which improves the accuracy of the method in the overall prediction results; (2) This invention improves the accuracy and robustness of Chinese entity extraction, thereby solving the problems of low accuracy and low availability of previous NER models on Chinese datasets. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the method flow framework of the present invention;
[0043] Figure 2 This is a schematic diagram of the model prediction process in an embodiment of the present invention. Detailed Implementation
[0044] The present invention will be further described below with reference to specific embodiments, but the scope of protection of the present invention is not limited thereto:
[0045] Example: Figure 1 As shown, a named entity recognition method that integrates entity boundary information includes the following steps:
[0046] 1) Using the vocab.txt file that comes with the BERT model, convert Chinese sentences into the corresponding indices in the vocab file; input the obtained sentence indices into the BERT model to obtain the corresponding word embedding vectors;
[0047] 2) The boundary information of Chinese semantics is extracted from the word embedding vector input boundary information network model, and Focal Loss is used to solve the label imbalance problem;
[0048] 3) Embed words into the input character feature extraction network model, learn more semantic information through the attention mechanism, and obtain long-distance dependencies between words through the BiLSTM model;
[0049] 4) The obtained features are fused and input into the CRF layer for training to learn the dependencies between labels and finally output the labels corresponding to the text.
[0050] Step 1) Convert the Chinese text into word embedding vectors, specifically including:
[0051] 1.1) During the data processing stage, the batch data with an input size of [batch_size, max_sen_len] needs to be transformed, where batch_size represents the batch size and max_sen_len represents the length of the longest sentence in the batch. Each character in the input batch is converted into its corresponding index in vocab.txt, which is the Chinese vocabulary included with the BERT model, containing 21,129 characters, which can meet the requirements of Chinese NER tasks;
[0052] 1.2) The obtained sentence indexes are input into the BERT model to obtain the corresponding word embeddings. This invention uses the BERT-Base model, which has 12 layers, each with a hidden layer dimension of 768, and 12 attention heads, with a total parameter count of approximately 110M. We chose BERT-Base considering the requirements of scale and computational resources. In practice, we found that using BERT-Large did not yield better results, while using BERT-Base could save computational resources while reducing training time. After the output of the BERT layers, we obtain the word embeddings, whose dimensions are [batch_size, max_sen_len, hidden_dim]. The embeddings can obtain more semantic information, providing more information for subsequent model training.
[0053] Step 2) The model learns Chinese semantic boundary information and solves the label imbalance problem, specifically including:
[0054] 2.1) To enable the model to learn boundary information, the original labels need to be modified when calculating the loss function. Before conducting the experiment, this invention sets the head and tail labels of each entity in the dataset to 1, and the remaining labels to 0. This method effectively helps the model learn boundary information.
[0055] 2.2) Input the embeds obtained in step 2.1) into the BiLSTM model to capture long-range dependencies, as shown in the formula below:
[0056]
[0057]
[0058]
[0059]
[0060] Equations (1) and (2) represent the outputs of the forward and backward layers of the biLSTM model, respectively. Equation (3) represents the fusion of the outputs from the two directions, ultimately yielding the output feature of the boundary information extraction network model, called feature_left, with dimensions [batch_size, max_sen_len, hidden_dim]. Through the BiLSTM layer, word embeddings learn features from distant locations, which helps enhance the robustness and accuracy of the model.
[0061] 2.3) To obtain the probability of the label corresponding to each character in the dataset, this invention uses a fully connected layer with a dimension of 2 for calculation. We input the feature_left obtained in step 2.2) into the fully connected layer to predict the label probability of each character. These labels consist of 0 and 1, used to predict whether the character belongs to the first or last label. The output dimension after the fully connected layer is [batch_size, max_sen_len, 2];
[0062] 2.4) Input the new labels obtained in step 2.1) and the character prediction results obtained in step 2.3) into the loss function to calculate the loss function. We use the gradient descent optimization algorithm to update the model parameters to minimize the loss function. Considering that the number of new labels belonging to the head and tail labels is relatively small, Focal Loss is used as the loss function, as shown in the following formula:
[0063]
[0064]
[0065] FL(P t )=-α t (1-P t ) γ log(P t (7)
[0066] Formula (5) is the cross-wrap loss function commonly used in binary classification problems, where p is the predicted label result of the character and y is the label value. Combining formulas (5) and (6), we can rewrite the binary classification cross-wrap loss function as CE(p,y) = CE(P t ) = -log(P t Traditional cross-entropy loss has achieved good results in binary classification problems, but it still doesn't work well in binary classification tasks with imbalanced labels. Therefore, by adding α... t and (1-P) t ) γ To adjust the weights of positive and negative samples, because P t ∈[0,1], γ∈[0,5], when Pt →1, Adjust (1-P) t ) γ →0, so well-classified examples are weighted downwards, thereby adjusting the weights of positive and negative samples and solving the label imbalance problem in binary classification. This invention uses PyTorch's `backward()` function to automatically calculate all gradients from the calculated Focal Loss, automatically accumulating all gradients for optimization of model parameters.
[0067] Step 3) The model learns character-level semantic information, specifically including:
[0068] 3.1) Input the word embeddings obtained in step 1.2) into the attention layer. The attention layer helps the model focus on important information in Chinese semantics more quickly, thereby improving the model's performance and generalization ability. The formula used in the attention layer is shown below:
[0069]
[0070] Where Q, K, and V represent Query, Key, and Value, respectively. These three parameters are obtained by multiplying the embeds obtained in step 1.2) by W. Q W k and W V The model is obtained from three trainable parameter matrices. Each multiplication of these matrices is equivalent to a linear transformation, which enhances the model's fitting ability. The output size after the attention layer is [batch_size, max_sen_len, hidden_dim].
[0071] 3.2) The word embedding representation after the attention layer does not learn long-range dependencies well, so the output of the attention layer needs to be input into the BiLSTM layer. This is similar to step 2.2), where the model is input into the BiLSTM layer to learn long-range dependencies and outputs feature_right, the size of which is [batch_size, max_sen_len, hidden_dim].
[0072] Step 4) Learn the dependencies between tags through the CRF layer, specifically including:
[0073] 4.1) Fuse the feature_left obtained in step 2.3) and the feature_right obtained in step 3.2) to obtain a new feature representation vector. This new feature representation vector integrates the boundary information learned by the boundary information extraction network model and the character semantic information learned by the character feature extraction network model, enhancing the semantics learned by the model from two aspects and increasing the model's accuracy. The feature fusion formula is as follows:
[0074] feature=β1*feature_left+β2*feature_right (9)
[0075] Where β1 and β2 are weight parameters, β1 and β2∈[0,1], and feature_left and feature_right represent the features learned by the models on both sides. This invention enhances semantic information and improves model accuracy through fusion.
[0076] 4.2) To facilitate CRF label prediction, this invention constructs a fully connected layer with a hidden layer dimension equal to the label size (target_size) to output the probability of each character corresponding to a label. We input the feature output in step 4.1) into the fully connected layer, predict the probability of each character in the fully connected layer, and output the probability of each character corresponding to a different label in the batch_size, max_sen_len, target_size.
[0077] 4.3) Input the output obtained in step 4.2) into the CRF layer. The CRF layer adds constraints to the final predicted label to ensure that the predicted label is valid. During training, the CRF layer maintains an emission matrix and a transition matrix. The emission matrix represents the probability of each character corresponding to each label, and the transition matrix represents the score of transitioning from label1 to label2, such as W. label1→label2 =0.2 indicates that the score for the transition from label1 to label2 is 0.2. This method helps to learn the dependencies between labels and avoids unreasonable label predictions. Combining the two matrices above, the CRF is trained using maximum conditional likelihood estimation, as follows:
[0078]
[0079]
[0080] Where Y = (y1, y2, ..., y n Let X = (x1, x2, ..., xn) represent the label sequence of the model. n ) represents a sentence, xi and y i represent characters and their corresponding labels, y' i represent the character x predicted by the model i corresponding labels.
[0081] In this embodiment, the NER task for friction and wear testing machines is described as an example of the present embodiment: based on the steps given above, the sentence "A friction and wear testing machine is a device industrially used to test the wear resistance of objects." is input first, and the corresponding labels are [B-ENY I-ENY I-ENY I-ENY I-ENY I-ENY I-ENY O O O O O O O O O O O O OO O O O O O O], that is, "friction and wear testing machine" needs to be predicted as an entity label, specifically as Figure 2 shown below:
[0082] Through step 1.1), the sentence is converted into the corresponding indexes in vocab.txt, that is [102, 3040, 3092, 4836, 2938, 6407, 7741, 3322, 3221, 671, 4905, 1762, 2339, 689, 677, 4500, 754, 4289, 860, 5447, 4836, 2595, 4638, 6392, 1906, 511, 103], the labels are also converted correspondingly, at this time the data size is [1,27], where 1 represents batch_size and 27 represents max_sen_len;
[0083] Through step 1.2), that is, the data with a size of [1,27] is input into BERT. The hidden_dim of the BERT model is 748, so the output size is [1, 27, 748], which represents the word embedding of the sentence; the word embedding vector is input into the boundary information extraction network model;
[0084] Through step 2.1), the labels of "摩" and "机" are set to 1, and the remaining labels are set to 0;
[0085] Through step 2.2), long-distance dependencies are learned through BiLSTM. The hidden_dim of BiSLTM is set to 748, and the output vector feature_left with an output dimension of [1,27,748] is obtained;
[0086] Through step 2.3), feature_left is input into a fully connected layer with a dimension of 2, and the label probability corresponding to each character is output, that is, the output dimension is [1,27,2];
[0087] After step 2.4), the Focal Loss is calculated by combining the output of step 2.3) and the labels obtained in step 2.1).
[0088] After step 3.1), the word embeddings from step 1.2) are input into the character feature extraction network model. The character feature information is enhanced by an attention layer with 8 attention heads, and the output vector has dimensions [1, 27, 748].
[0089] After step 3.2), the vector output from step 3.1) is input into a BiLSTM model with a dimension of 748 to learn long-range dependencies, and the output feature_right has a dimension of [1,27,748].
[0090] After step 4.1), feature_left and feature_right are concatenated, and parameters w1 and w2 are both set to 1 to obtain the fused feature.
[0091] Following step 4.2), the feature is input into the CRF layer to predict the label corresponding to each character. Experimental results demonstrate that, compared to the baseline model BERT-BiLSTM-CRF, this invention improves the overall prediction accuracy by nearly 2%, and also significantly improves the prediction probability of small labels.
[0092] In summary, this invention mainly comprises four parts: a BERT layer for word embedding, a boundary information extraction network model for acquiring semantic boundary information, a character feature extraction network model for acquiring more character-level semantic information, and a CRF layer for acquiring and learning label relationships. This invention employs Focal Loss to address the label imbalance problem when acquiring boundary information and introduces an attention mechanism to better learn character information. By fusing boundary and character information, this invention successfully solves problems such as the complexity of Chinese semantics and the lack of natural separators. Significant results have been achieved on both our self-constructed functional unit extraction dataset and large public datasets. The research results show that this invention significantly improves label accuracy with a smaller number of data labels without affecting the accuracy of the benchmark model BERT-BiLSTM-CRF.
[0093] The above description describes specific embodiments of the present invention and the technical principles employed. Any changes made in accordance with the concept of the present invention that do not exceed the spirit of the specification and drawings should still fall within the protection scope of the present invention.
Claims
1. A named entity recognition method that integrates entity boundary information, characterized in that, Includes the following steps: (1) Input the Chinese text into the BERT model and obtain the corresponding word embedding information; (2) Input the word embedding information into the boundary information extraction network model to obtain the boundary information and obtain the output feature_left; specifically, the following steps are included: (2.1) Modify the original labels of the dataset by labeling the start and end positions of entities as 1, i.e., the start and end labels are 1; while other text is labeled as 0, i.e., the remaining labels are 0. (2.2) Input the word embedding vectors obtained in step (1.2) into the BiLSTM model to capture long-range dependencies, as shown in the following formula: , , , , Formulas (1) and (2) represent the outputs of the forward and backward layers of the biLSTM model, and formula (3) represents the fusion of the outputs in the two directions. Finally, the output feature of the boundary information extraction network model can be obtained, which is called feature_left. The dimension of this feature is [batch_size, max_sen_len, hidden_dim]. (2.3) Set up a fully connected layer with a hidden layer dimension of 2, and output the feature_left obtained in step (2.2) into the fully connected layer. The output vector represents the probability of each character being the start or end label. The output dimension after passing through the fully connected layer is [batch_size, max_sen_len, 2]. (2.4) Input the new label obtained in step (2.1) and the character prediction result obtained in step (2.3) into the loss function, and use the gradient descent optimization algorithm to update the model parameters to minimize the loss function; Focal Loss is used as the loss function, and the formula is as follows: , , , Formula (5) is the binary classification cross-adsorption loss function, where p is the character's predicted label result and y is the label value; combining formulas (5) and (6), the binary classification cross-adsorption loss function can be rewritten as CE(p,y) = CE( ) = -log( ); By joining and To adjust the weights of positive and negative samples, because , ,when ,adjust The calculated Focal Loss is automatically used to calculate all gradients using PyTorch's backward() function, and all gradients are automatically accumulated for optimization of model parameters. (3) Input the word embedding information into the character feature extraction network model, learn the semantic information through the attention mechanism, and use the BiLSTM model to capture the dependency relationship, and output the obtained feature_right; specifically including the following steps: (3.1) Input the word embedding vectors obtained by the BERT model into the attention layer of the character feature extraction network model, and output word vectors with more semantic information; The formula used in the attention layer is shown below: , Where Q, K, and V represent Query, Key, and Value, respectively. These three parameters are obtained by multiplying the embeds obtained in step 1.2) by... , and The three trainable parameter matrices are obtained; the output size after passing through the attention layer is [batch_size, max_sen_len, hidden_dim]. (3.2) Input the embedding information output by the attention layer into the BiLSTM layer to learn long-range dependencies and output feature_right, the size of feature_right is [batch_size, max_sen_len, hidden_dim]; (4) The features left and right obtained above are fused and input into the CRF for final processing to learn the dependency relationship between the labels and output entity prediction information, i.e. the labels corresponding to the Chinese text.
2. The named entity recognition method that integrates entity boundary information according to claim 1, characterized in that: In step (1), the Chinese text is converted into the corresponding index in vocab using the vocab.txt file that comes with the BERT model, and the obtained index is input into the BERT model to obtain the corresponding word embedding vector; specifically, the steps are as follows: (1.1) First, the data needs to be preprocessed: use the Chinese vocab.txt provided in the BERT model as the vocabulary, and convert the input Chinese text into the index of the corresponding word in the vocabulary; at the same time, obtain the maximum length of all sentences in the batch, and input a batch of Chinese sentences into the BERT model; (1.2) The BERT pre-trained model is used as the word embedding model, where the hidden layer dimension of the model is 768. In the word embedding process, the Chinese text obtained in step (1.1) is input into the BERT model. The BERT model is used to capture the complex dependencies and semantic associations between words and output word embedding vectors containing rich language information, with dimensions of [batch_size, max_sen_len, hidden_dim].
3. The named entity recognition method that integrates entity boundary information according to claim 1, characterized in that: Step (4) specifically includes the following steps: (4.1) The feature_left obtained in step (2) and the feature_right obtained in step (3) are fused and concatenated to obtain a new feature representation vector. The new feature representation vector integrates the boundary information learned by the boundary information extraction network model and the character semantic information learned by the character feature extraction network model, which enhances the semantics learned by the model from two aspects and increases the accuracy of the model. The feature fusion formula is as follows: , in and These are weight parameters. and feature_left and feature_right represent the features learned by the models on both sides; (4.2) Construct a fully connected layer with a hidden layer dimension of the label size target_size to output the probability of each character corresponding to the label; and input the feature output in step (4.1) into the fully connected layer to predict the probability of each character in the fully connected layer. The output size is [batch_size, max_sen_len, target_size], and the output represents the probability of each character corresponding to different labels. (4.3) Input the output obtained in step (4.2) into the CRF layer to predict the label; in the CRF layer, the CRF layer maintains the emission matrix and the transition matrix. The emission matrix represents the probability of each character corresponding to each label, and the transition matrix represents the score of transitioning from label1 to label2; combine the above two matrices and train the CRF using maximum conditional likelihood estimation, as follows: , Where, Y = ( ) represents the label sequence of the model, X = ( () represents a sentence. and Represents characters and their corresponding labels. Characters representing model predictions Corresponding tags.