Named entity identification method based on multiple features and gating mechanism

By introducing Gate-BiLSTM and IDCNN modules into the named entity recognition task, multi-level text features are extracted and fusion is solved, and the problem of existing methods only considering single features is significantly improved.

CN119990128APending Publication Date: 2025-05-13BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510067736.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing deep learning methods only consider single features in named entity recognition tasks, ignoring the importance of text features for label recognition, resulting in poor recognition results.

Method used

A named entity recognition method based on multi-features and gating mechanism is proposed. Text features are extracted from global and local perspectives through Gate-BiLSTM and IDCNN modules, and deep feature extraction is performed using gating mechanisms, integrating global and local features to improve the recognition effect.

Benefits of technology

Through multi-feature fusion, the model's prediction accuracy of most character labels in sentences and character labels within entities is significantly improved, and the effect of naming entity recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990128A_ABST
    Figure CN119990128A_ABST
Patent Text Reader

Abstract

The invention discloses a named entity recognition method based on multiple features and a gating mechanism, which is designed for the field of middle school mathematics education, and comprises the following steps: constructing a middle school mathematics test question named entity recognition data set for training a model; a traditional named entity recognition model is trained, and a BERT + BiLSTM + CRF traditional model serves as a baseline model and is improved; introducing an IDCNN (iterative expansion convolutional neural network) to extract local features of the text; introducing a gating weight mechanism to expand text global features extracted by the BiLSTM; and fusing BiLSTM and IDCNN feature training models and generating an entity tag sequence. According to the method, the global features and the local features are fused, the recognition accuracy of the internal tags of the single entity and most character tags of the whole sentence is improved, and finally the overall tag recognition effect of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to natural language processing technology in the field of deep learning, and in particular to a named entity recognition technology based on multiple features and a gating mechanism. Background Art

[0002] With the advancement and development of computer technology, more and more researchers are involved in natural language processing tasks. As a core task in NLP, named entity recognition is widely used in machine translation, information extraction, and question-answering systems. Many researchers have also begun to study Chinese named entity recognition technology. With the development of technology, the research methods used in named entity recognition tasks have also kept pace with the times. Initially, the method based on dictionaries and rules was widely used, but this method has certain limitations. With the development of machine learning, methods based on statistical machine learning have gradually been developed. In recent years, methods based on deep learning have become the mainstream method and have dominated the market.

[0003] (1) Dictionary- and rule-based NER methods;

[0004] Early research on Chinese NER mainly dealt with how to identify entities such as names of people, places, and institutions in Chinese texts. Based on this, a rule-based and dictionary-based NER method was developed. This method does not require labeled data. It mainly uses rule templates and special domain dictionaries manually constructed by experts and linguistic experts based on language characteristics and data set characteristics. In 1995, KRUPKA et al. proposed matching strings with templates to achieve named entity extraction from text. In 1995, Ralph et al. developed a rule-based dictionary-based named entity recognition system with the help of a name dictionary containing country names, major cities, and company names. In 2011, Yan Ping started from the internal structure of the entity and the context of the text, and constructed recognition rules for names, which greatly improved the accuracy of Chinese name recognition.

[0005] (2) NER methods based on statistical machine learning;

[0006] Machine learning is a technology that can learn autonomously without relying on humans, which is equivalent to the process of human beings summarizing historical experience of things. Named entity recognition models based on statistical machine learning have their own strengths, including the Hidden Markov Modele (HMM) proposed by Morwals et al. in 2012, the Maximum Entropy Model (ME) proposed by Lin YF et al. in 2013, the Support Vector Machine (SVM) proposed by Surya BB et al. in 2014, and the Conditional Random Fields (CRF) proposed by Li K et al. in 2015. Morwal et al. used the HMM method to build a system based on named entity recognition for the Hindi system in 2012, and trained and tested it in languages ​​such as Urdu and Punjabi. Meenachisundaram et al. used the SVM method for biomedical named entity recognition in the GENIA corpus in 2013, and found that SVM has a good performance in NER in the biomedical field. McCallum et al. first introduced CRF in the named entity recognition task and achieved good results in the Hindi named entity recognition task through a large number of tests. The disadvantage is that the training time is relatively long.

[0007] (3) NER methods based on deep learning;

[0008] According to the classification of the main network model architecture, the models of named entity recognition tasks can be divided into three categories, namely RNN-based models, CNN-based models, and Transformer-based models. Zhang et al. designed the Lattice LSTM structure in 2018. This model first introduced dictionary information to handle CNER tasks. By using the grid structure to integrate the potential word information in the text into the BiLSTM+CRF model structure, the performance on four data sets is very excellent, but the model has some problems: the number of grids is inconsistent, which makes the model training unable to be parallelized. In addition, this structure is only applicable to LSTM models and has poor transferability. Xu et al. proposed the BiGRU+CRF model structure in 2019. This model introduced radicals and word features, organically combined contextual semantic information of different granularities, and achieved good results in sequence labeling tasks. Kong et al. designed a Chinese clinical named entity recognition model based on CNN. It captured contextual feature information through multi-layer CNN, and used a residual structure to fuse multi-scale information, while avoiding the problem of gradient disappearance, significantly improving the model performance. In 2017, Strubell et al. proposed the Iterative Convolutional Neural Network (ID-CNN), which expanded the receptive field of the convolution kernel through dilated convolution, while retaining the advantage of CNN for batch parallel computing, and has a faster training speed than LSTM. In 2018, Google proposed the BERT model structure, which is a model that uses a bidirectional Transformer network structure for pre-training and has better performance than a single-layer Transformer model. Wang et al. combined BERT with the BILSTM+CRF model and achieved better experimental results on the People's Daily dataset without adding any features, proving the effectiveness of this method.

[0009] Compared with the first two methods, the method based on deep learning occupies the leading position in current research with its superior performance. It uses deep neural network processing to automatically learn hidden features from a large number of training samples, so that it can achieve better recognition results in NER tasks. Existing deep learning methods only consider a single feature in the feature extraction module, while the feature of the text is very important for the recognition of the label, so it is necessary to model it in the feature extraction module. Summary of the invention

[0010] To solve the above problems, the present invention proposes a named entity recognition method based on multiple features and a gating mechanism. By extracting text features from both global and local perspectives, rich feature representations are obtained, and a gating mechanism weight mechanism is used to perform deep feature extraction on the extracted global features, and the global features are improved. Finally, the global features and local features are fused to improve the entity recognition effect of the model.

[0011] The scheme of this method is as follows:

[0012] First, the screened data set is manually annotated; then the data set is divided into a training set and a test set for training and testing the model proposed in the present invention; finally, the model prediction results are obtained, and the prediction results are evaluated to verify the F1 value of the model.

[0013] The specific steps of the technical solution of the present invention are as follows:

[0014] Step 1. Dataset construction: A correctly annotated and rich data set is necessary for training deep neural networks, verifying algorithm effects, and testing systems. This project obtains middle school mathematics test questions from the September 1 website (http:www.cn901.com), including test questions from different textbooks such as the People's Education Press, Jiangsu Education Press, and Beijing Normal University Press. This project selects some test questions from them, cleans and manually annotates the collected test question resources; divides the annotated data set into a training set and a test set;

[0015] Step 2. Train the traditional model: Since the BERT+BILSTM+CRF based on the traditional model has achieved good training results on the People's Daily dataset and the Chinese resume dataset, and these two datasets also contain special entity types in their respective fields, this model is selected as the baseline model and trained, and then improvements are made based on this baseline model;

[0016] Step 3. Local feature extraction module: In the traditional model mentioned in step 2, multiple DCNN blocks are introduced to form IDCNN, that is, iterative dilated convolutional neural network to extract local features of sentences;

[0017] Step 4. Improve the extracted global feature module: Since the intermediate cell unit of the BiLSTM model only integrates part of the forward and backward information of the current character, but does not integrate the complete sentence representation, that is, the intermediate cell lacks global information, and since modifying the internal structure of BiLSTM is more complicated, this model adjusts the output information of BiLSTM externally to supplement the hidden state information of the internal cell unit;

[0018] Step 5. Integrate the modules mentioned in Step 3 and Step 4 into the traditional model mentioned in Step 2, and input the training set in Step 1 into the model for training. Finally, input the test set into the trained model for prediction and evaluate the F1 index of the model.

[0019] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0020] The present invention provides a named entity recognition method based on multiple features and a gating mechanism. The global and local features of a text are extracted by Gate-BiLSTM and IDCNN modules, and the two are fused to obtain rich feature information, so that the model can improve the label prediction effect of most characters in the entire sentence and the label prediction effect of characters inside the entity. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a structural diagram of a named entity recognition model based on multiple features and gating mechanisms;

[0022] Figure 2 is a sample annotation example;

[0023] Figure 3 This is the structure diagram of the IDCNN local feature extraction module;

[0024] Figure 4 This is the structure diagram of the Gate-BiLSTM global feature extraction module based on the gated weight mechanism;

[0025] Figure 5 It is a comparison chart of the experimental results of the model provided by the present invention and the traditional model;

[0026] Figure 6 It is a comparison chart of experimental results of the model provided by the present invention and other models;

[0027] Figure 7 It is the experimental result diagram of local feature and global feature extraction module and traditional model. DETAILED DESCRIPTION

[0028] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings so that those skilled in the art can better understand the present invention. The present invention proposes a named entity recognition method based on multiple features and a gating mechanism, which can fully combine local and global features to analyze entity labels and improve the prediction accuracy of text labels.

[0029] Figure 2 The overall flow chart of the method of the present invention can be decomposed into the following steps:

[0030] Step 1. Dataset construction and data set division;

[0031] Step 2. Train the traditional named entity recognition model;

[0032] Step 3. Introduce the IDCNN local feature extraction module;

[0033] Step 4. Introduce the Gate-BiLSTM global feature extraction module based on the gated weight mechanism;

[0034] Step 5. Use the training set to train the model, test the model effect, and evaluate the recognition accuracy.

[0035] The following is a detailed description of each step:

[0036] Step 1:

[0037] Correctly annotated and rich data sets are necessary for training deep neural networks, verifying algorithm effects, and testing systems. This project obtains middle school mathematics test resources from the September 1 website (http:www.cn901.com), including test questions from different textbooks such as the People's Education Press, Jiangsu Education Press, and Beijing Normal University Press. These test questions almost cover all the knowledge points of junior high school mathematics. This project selects test questions corresponding to the four major knowledge points of sets, equations, functions, and geometry, and cleans and manually annotates the collected test resources.

[0038] The initially selected test question resources are HTML strings, which contain many HTML tags, such as line breaks. , paragraph mark And other escape characters marked by "&", etc., the present invention escapes them into corresponding Chinese characters according to the requirements. In addition, it is also detected that these html strings also contain many common superscript and subscript tags in mathematics, such as 以及,本发明将其分别转化为_{}与^{}。经过初步的数据清洗之后,将重复出现的试题数据进行删除,为了保证数据的准确性以及可靠性,接着又进行了人工检查,去掉一些不符合要求的数据,比如乱码数据。在数据清洗之后,本发明将合格的未标注的试题资源送入LabelStudio中进行人工标注。根据试题数据所包含的知识点的特点,本文定义了9大实体类型:集合类、函数类、方程类、几何类、不等式类、数值类、操作类、关系类以及知识点类。其中函数类又分为n次函数、指数函数、幂函数、对数函数、三角函数以及组合函数类;方程类又分为m元n次方程类、包含其它函数的方程以及等式类;几何类又分为几何图形类、坐标类以及属性类。细化之后的实体类型一共有20类。本发明选用的是BIO标注体系,其中B表示实体的起始字符、I表示实体的中间以及结尾字符、O表示非实体部分,按照上述所划分的20类实体,该体系的标签共有41个。

[0039] 最终本发明共标注了6281条数据,其中训练集有5652条,测试集有629条。训练集和测试集中的实体标签分布、实体对应的BIO标签命名以及部分实体样例如下表1所示,标注的数据样例如图2所示。

[0040] 表1部分实体统计信息

[0042] 步骤2:

[0043] 由于命名实体识别任务还未在中学数学领域做过研究,近几年的模型几乎都是在公开数据集上做的实验,且公开数据集数据量比较大,模型结构也会比较复杂,本发明构造的数据集样本少,训练复杂结构的模型容易过拟合;此外,基于传统模型的BERT+BILSTM+CRF在人民日报数据集以及样本少的中文简历数据集上均取得较好的训练效果,且这些数据集不止包含了传统的人名地名等常见的实体类型,还包含了所属领域的特殊实体类型,表明该传统模型能够捕捉到不同领域的文本特征,具有领域自适应性,故而本发明选择基于传统模型的BERT+BILSTM+CRF模型作为本发明的基线模型,并在此基础上做出改进。

[0044] 下面对本文所选择的基线模型的各个模块进行具体介绍。

[0045] (1)编码层

[0046] 常用的词向量模型有静态词向量模型以及上下文词向量模型。其中静态词向量模型有One-hot、Word2Vec以及Glove,上下文词向量模型有ELMO、BERT。顾名思义,上下文词向量能够根据上下文信息对词进行编码,可以解决一词多义的问题。在中学数学试题语料库中,同一字符在不同的文本中可以表示不同的含义,即字符多义问题。例如,"一元二次方程”以及"一元二次函数”,其中"一元二次”分别表示方程类的实体以及函数类的实体。BERT可以根据不同的上下文提供不同的词向量,能够更好地识别和提取中学数学知识特征信息。

[0047] (2)特征提取层

[0048] 循环神经网络RNN可以将上一个时间片的输出作为下一个时间片的输入,从而可以很好的处理时间序列数据。中学数学试题语料库的文本也是一个序列文本。然而,它的语料库长度比较长,有些文本的句子长度超过160个字符,使用RNN更容易产生长距离依赖的问题。这些将导致模型无法有效地学习语料的全局上下文特征,可能会忽略掉很多重要的信息,从而导致较差的实体识别效果。

[0049] 长短期记忆网络LSTM是一种特殊的RNN,它能够学习长距离依赖关系,并且通过引入门机制(输入门、遗忘门以及输出门)来克服传统RNN在处理长序列文本时的梯度消失和梯度爆炸的问题。其中,遗忘门决定丢弃多少来自先前时刻的信息,将留下来的实际有用的信息传递给输出门。相较于单向的LSTM,BiLSTM能够同时学习到序列的前后文信息,对需要了解全局上下文的任务非常有利。

[0050] (3)解码层

[0051] 在序列标注任务中,标签之间是存在一定的制约关系的,比如实体的开头一定是"B-xx”开头,而不是"I-xx”开头的、"O”也一定不会出现在"I-xx”标注前面等。若仅使用BiLSTM则是无法学习到上述约束的,会存在很多无效预测标签,即只使用BiLSTM,序列中各元素的预测标签之间是相互独立的,没有依赖关系。鉴于此,模型需要引入CRF层对标签的输出进行限制。CRF将BiLSTM的输出作为字符与标签的概率矩阵,为了实现标签之间的约束效果,CRF又引入了一个标签转移概率矩阵,通过模型的训练,对两个矩阵进行学习,最终将概率最大的路径所对应的标签序列作为最终的标签预测结果。

[0052] 步骤3:

[0053] BiLSTM模型主要用于全局特征的提取,容易忽略文本的局部特征,而一个实体一般只是一段文本中的一小部分,若只考虑全局特征,则容易忽略实体内部字符的特征,不利于整个实体标签的识别。因此,我们需要在BiLSTM的基础上引入CNN来提取局部特征,也可以说是对全局特征的"细化”。可以理解为BiLSTM是为了让模型初步了解整段文本是在讲什么的,而CNN是为了让模型了解某个文本片段在集中讲什么,两者对于实体识别同等重要。

[0054] 然而对于序列标注任务来说,普通的CNN有一个缺点,就是在卷积操作之后,末层的神经元可能只是得到了部分原始输入数据中的信息,而对当前需要标注的字符来讲,与该字符相邻的每个字都有可能会对其造成影响,为了覆盖到更多的字符信息就需要更多的卷积层,那么卷积层数和参数也就越来越多,使得模型变得庞大且难以训练。于是就有了DCNN,即扩张卷积。通过给卷积核增加扩张宽度,指的是kernel的间隔数量,以此在不增加卷积核大小的情况下,增大了感受野,让每个卷积的输出都包含较大范围的信息。

[0055] 迭代扩张卷积神经网络(IDCNN)由多层不同扩张宽度的DCNN组成,然后利用前一层扩张卷积计算电流扩张卷积的特征向量。计算结果如式(1)所示。其中,D为卷积之后的特征向量,H为卷积核,i表示第i层,б为激活函数。IDCNN模型结构如图3所示。

[0056] Di=σ(Hi-1·Di-1)(1)

[0057] 步骤4:

[0058] BiLSTM为序列标记任务生成句子的表示是优于传统RNN模型的,但是也存在相应的限制,主要是由于BiLSTM模型的中间细胞单元只融合了当前字符的部分前向和后向信息,而没有融合完整的句子表示,即中间细胞缺乏全局信息,这是由于BiLSTM模型的内部细胞单元的传输结果所造成的。在命名实体识别NER任务中,若每个字符都能学习到整个句子信息,在解码层对于实体标签的分类效果将会更好。

[0059] 对于BiLSTM在序列标记任务中的局限性这一问题,有些学者提出修改BiLSTM内部细胞单元的传输结构,通常改变模型的内部结构或者设计复杂架构的方法可能会影响模型的推理速度并且在实际应用中实现也会更加复杂。本发明提出在外部对BiLSTM的输出信息调整来补充内部细胞单元的隐藏状态信息。基于此,本发明提出了一个基于Gate权重学习机制的BiLSTM模型(Gate-BiLSTM),模型结构如图4所示。

[0060] BiLSTM的隐藏状态序列为H={H1、H2、……、Hn},完整的句子表示为前向Hn与后向H1的融合,用字母G来表示,如公式(2)所示,其包含了整个句子的完整语义表示;

[0061]

[0062] 整个句子的完整语义信息与当前字符的隐藏状态信息拼接作为当前字符暂时的特征向量Ot,如公式(3)所示;

[0063] Ot=G||Ht(3)

[0064] 在门控权重学习机制中,首先使用两个线性映射从特征向量Ot中选择相关特征,即RH与RG,之后将RH与RG通过两个sigmoid函数得到当前字符所表示的隐藏状态信息的权重iH与整个句子的完整语义信息的权重iG;RH、RG、iH、iG的计算公式如(4)、(5)、(6)、(7)所示;

[0065] RH=WHOt+bH(4)

[0066] RG=WGOt+bG(5)

[0067]

[0068]

[0069] 最后,通过权重iH、iG将当前字符所表示的隐藏状态信息Ht与整个句子的完整语义信息G进行融合作为当前字符最终的特征向量计算公式如(8)所示。

[0070]

[0071] 以上便是门控权重学习机制的整个计算过程,BiLSTM的隐藏状态输出单元经过该层将会获得更完整的句子语义信息。

[0072] 步骤5:

[0073] 将步骤1所提到的数据集MathNer与公开数据集Resume中文简历数据集输入到本发明方法模型中,其中,MathNer数据集训练集有5652条,测试集有629条,Resume数据集的训练集、验证集和测试集分别有3821、463以及477条。

[0074] 两个数据集在传统模型和本发明方法模型上的实验结果如图5所示,其中Resume数据集的实验结果如表(2)所示,MathNer数据集的实验结果如表(3)所示:

[0075] 表2Resume数据集在传统模型和本发明方法模型上的实验结果

[0076]

[0077] 表3MathNer数据集在传统模型和本发明方法模型上的实验结果

[0078]

[0079] 从表2、表3以及图5中可以看出,通过引入IDCNN以及Gate-BiLSTM,模型不论是在Tag级别还是Entity级别上的F1值均比baseline模型的F1值要高,说明多特征对实体标签的识别是有效的。而Entity级别的F1值低于Tag级别,是因为Entity级别要求更高,需要实体中的每个字符标签都预测正确。

[0080] 本发明提出的模型是基于传统模型的基础上所进行的改进,近年来又出现了很多其它模型结构,将本方法与这些模型的结果也进行了比较,在Entity级别上的具体结果如表4,图6所示:

[0081] 表4本方法与其他模型上的性能表现

[0082]

[0083] 从表4中可以看出,Lattice-LSTM模型的效果均低于其他两个模型的F1值,且低于baseline,可能是因为该模型引入了不合适的词信息,使其特征对标签分类造成了干扰。而MECT模型的识别效果高于baseline但低于本方法,对于MathNer数据集来说,可能是因为该数据集中的字母以及数字会比较多,可提取的汉字结构特征较少,而Resume数据集几乎都是汉字所构成的,模型可充分利用汉字结构特征,但由于MECT模型结构较为复杂,使用的是静态词向量模型,并且两个数据集的训练样本也不够多,使得该模型的识别效果低于本发明方法。

[0084] 为了验证本方法提出的两个模块对模型都是有影响的,进行了消融实验,实验结果如图7所示,每个模块对实体识别效果的提升在两个数据集上都是有效的。IDCNN通过引入局部特征,使得模型会注重卷积核内部区域字符之间的相关性,从而影响实体内部标签的分类,而Gate-BiLSTM引入的全局特征,使得每个字符都包含了整个句子的语义信息,使得模型对整个句子的大部分标签识别效果更好。

[0085] 以上具体实施方法的描述仅用于说明本申请,以便于本领域的技术人员理解本发明,但应该清楚,本发明并不限于上述描述的实施方式。在不脱离本发明宗旨和权利要求所保护的范围情况下,本发明的各种变化和改进均属于本发明的保护范围之内。

Claims

1. A named entity recognition method based on multiple features and gating mechanism, characterized in that: The following steps are involved: Step 1. Dataset construction: obtain the test question resources of middle school mathematics; select some test questions from them, clean and manually annotate the collected test question resources; divide the annotated data set into training set and test set; Step 2. Train the traditional model: Select the BERT+BILSTM+CRF model as the baseline model and train it, then make improvements based on the baseline model; Step 3. Local feature extraction module: In the BERT+BILSTM+CRF model mentioned in step 2, multiple DCNN blocks are introduced to form IDCNN, that is, iterative dilated convolutional neural network to extract local features of sentences; Step 4. Improve the extracted global feature module: The BERT+BILSTM+CRF model adjusts the output information of BiLSTM externally to supplement the hidden state information of the internal cell units; Step 5. Integrate the local feature extraction module and global feature module mentioned in steps 3 and 4 into the BERT+BILSTM+CRF model in step 2, and input the training set in step 1 into the BERT+BILSTM+CRF model for training.

2. The method for named entity recognition based on multiple features and a gating mechanism according to claim 1, characterized in that: Step 1 is as follows: (1) Screen out the test questions of the entire junior high school mathematics subject, and select the test questions corresponding to the four major knowledge points of sets, equations, functions and geometry; (2) escape the tags in the filtered HTML string text, delete the duplicate data and manually check; (3) Define nine entity types: set type, function type, equation type, geometry type, inequality type, numerical type, operation type, relationship type, and knowledge point type; and refine some of these entity types, resulting in a total of 20 types; use LabelStudio to label the dataset; (4) Before starting training, the labeled data set is divided into 90% of the labeled data set as a training set and 10% as a test set.

3. The method for named entity recognition based on multiple features and a gating mechanism according to claim 1, characterized in that: Step 2 specifically includes: The specific structure of the BERT+BILSTM+CRF model based on the traditional model can be summarized into three parts: encoding layer, feature extraction layer and decoding layer. Among them, BERT is used as the encoding layer to generate dynamic word vectors for the text; BILSTM is used as the feature extraction layer to extract the global features of the text; CRF is used as the decoding layer to generate the final entity type label of the text.

4. The method for named entity recognition based on multiple features and a gating mechanism according to claim 1, characterized in that: Step 3 specifically includes: when function expression entity types appear in middle school mathematics test questions, the feature relationship between characters inside the entity is stronger than that between characters outside the entity; a local feature extraction module is introduced to extract phrase-level features; Increasing the expansion width of the convolution kernel refers to the number of kernel intervals. Without increasing the size of the convolution kernel, the receptive field is increased so that the output of each convolution contains a larger range of information. The iterative dilated convolutional neural network IDCNN consists of multiple layers of DCNN with different dilation widths. The feature vector of the current dilated convolution is calculated using the dilated convolution of the previous layer. The calculation result is shown in formula (1). Where D is the feature vector after convolution, H is the convolution kernel, i represents the i-th layer, and σ is the activation function. D i =σ(H i-1 ·D i-1 )(1)。 5. The method for named entity recognition based on multiple features and a gating mechanism according to claim 1, characterized in that: Step 4 is as follows: the logical relationships of middle school math questions are complex, and understanding the semantic information of the entire question is more conducive to grasping the key points of the question. Fusion of sentence-level features and single-character features is more beneficial to the classification of entity labels. The hidden state sequence of BiLSTM is H = {H1, H2, ..., H n }, the complete sentence is represented as a forward and backward The fusion of is represented by the letter G, as shown in formula (2), which contains the complete semantic representation of the entire sentence, where Represents the hidden state information of the complete sentence in the backward direction, Represents the forward complete sentence hidden state information; The complete semantic information of the entire sentence and the hidden state information of the current character are concatenated as the temporary feature vector O of the current character. t , as shown in formula (3); O t =G||H t (3) In the gated weight learning mechanism, two linear mappings are first used to transform the feature vector O t Select relevant features, namely R t H With R t G , then R t H With R t G The weight i of the hidden state information represented by the current character t is obtained through two sigmoid activation functions t H The weight i of the complete semantic information of the entire sentence t G ; R t H , R t G 、i t H 、i t G The calculation formulas are shown in (4), (5), (6), and (7), where W H and W G represents the weight parameter, b H and b G Indicates the bias value; Finally, by weight i t H 、i t G The hidden state information H represented by the current character t It is fused with the complete semantic information G of the entire sentence as the final feature vector of the current character The calculation formula is shown in (8); The hidden state output unit of BiLSTM will obtain more complete sentence semantic information after passing through this layer.

6. The method for named entity recognition based on multiple features and a gating mechanism according to claim 1, characterized in that: Step 5 is as follows: Integrate the feature extraction modules mentioned in Steps 3 and 4 into the traditional model of Buzhou 2, extract and fuse the features at the phrase level and sentence level, and improve the prediction accuracy of entity labels; Firstly, the constructed middle school mathematics named entity recognition dataset is divided into training set and test set according to the ratio of 9:1, and input into the BERT+BILSTM+CRF model for training. The number of iterations in the training phase is set to 50. When the loss function of the training set reaches the minimum, the BERT+BILSTM+CRF model is saved, and the saved BERT+BILSTM+CRF model parameters are used to test the test set samples, and the F1 value of the test set is recorded.