A text named entity extraction method based on channel attention and character features
By introducing channel attention and character features into the named entity recognition model, the existing model solves the problem of expressing word ambiguity and word vector singularity when dealing with unstructured text in specific fields, achieving higher accuracy and accuracy of named entity recognition.
Patent Information
- Application Number
- CN202510317358.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-18
AI Technical Summary
When the existing named entity recognition model deals with unstructured text in a specific domain, there are problems with expression word ambiguity and word vector singularity, resulting in poor training results.
A text named entity extraction method based on channel attention and character features is adopted to construct a named entity recognition model including BERT model, CD module, multi-layer bidirectional long and short-term memory network (MBiLSTM) and conditional random field (CRF). The CD module includes the CharCNN module and the DSENet module, which are used to extract character-level features and enhance entity content and position feature expressions.
It effectively solves the problems of character-level feature extraction and important feature weighting, improves the accuracy and accuracy of named entity recognition, and significantly improves the performance on specific domain data sets.
Smart Images

Figure CN119849501B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of named entity recognition, and in particular to a text named entity extraction method based on channel attention and character features. Background Art
[0002] Text contains a wealth of information, and it is of great significance to mine it. Information mining is the problem of named entity recognition (NER).
[0003] Named Entity Recognition (NER) is a key task in Natural Language Processing (NLP), which involves identifying and classifying certain types of information elements, namely named entities (NE). The main goal of NER is to identify and classify entity information from text, such as names of people, places, organization names, and amounts, and is widely used in many fields such as medicine, railway construction, and geological science. In the context of the information explosion era, NER is not only crucial for information extraction and data analysis, but also plays a fundamental role in many application scenarios such as sentiment analysis, question-answering systems, and machine translation.
[0004] Due to the diversity of data, texts show significant differences in format and content, which are manifested in the diversity of formats, grammatical ambiguity, and implicit semantic information. These texts not only contain multiple types of information, but also due to the flexibility of language, the information expression often has a certain degree of ambiguity. For example, important information may appear in multiple ways and is often embedded in complex sentence structures. In addition, terminology and abbreviations in specific fields also increase the difficulty of entity recognition. The high noise and irregularity of such texts increase the challenge of NER model training and reasoning. Extracting important information from large amounts of text is an urgent problem to be solved in related fields. However, the existing Chinese named entity recognition algorithm model has problems such as word ambiguity and word vector singularity when dealing with related tasks, resulting in poor training results of the algorithm model. Therefore, how to effectively process and parse these unstructured texts for specific fields has become an important challenge in current NER research.
[0005] Some traditional models such as conditional random fields (also known as CRFs) and long short-term memory networks (LSTMs) have limitations in handling long-distance dependencies and complex contexts. These models have difficulty capturing deep semantic information in sentences. Although the combination of bidirectional long short-term memory networks (also known as BiLSTMs) and CRFs helps identify entity boundaries, it still faces challenges when dealing with nested entities, continuous entities, or irregular entities. Although the BERT model can understand context bidirectionally and provide rich word meaning representations, it mainly focuses on word-level information and ignores character-level features, which is not conducive to identifying entities based on character features. BERT's self-attention mechanism assigns attention weights to each token, but this assignment is based on the interaction between tokens rather than the importance of features, which may cause the model to pay too much attention to unimportant features and ignore key features in entity recognition.
[0006] In summary, existing named entity recognition models have defects. Summary of the invention
[0007] In order to solve the technical problem that the existing named entity recognition model has poor recognition effect, an embodiment of the present invention provides a text named entity extraction method based on channel attention and character features.
[0008] The technical solution of the embodiment of the present invention is achieved as follows:
[0009] The embodiment of the present invention provides a method for extracting text named entities based on channel attention and character features, the method comprising: obtaining input text to be recognized; constructing a named entity recognition model, the named entity recognition model comprising a BERT model, a CD module, a multi-layer bidirectional long short-term memory network (also known as MBiLSTM) and a CRF; the CD module comprises a CharCNN module and a DSENet module; wherein the BERT model is used to process the input text to obtain features, grammatical structures and semantic representations with context information; the CharCNN module is used to extract character-level features in the output result of the BERT model; the DSENet module is used to enhance the entity content and entity position feature expressions in the output result of the BERT model; the MBiLSTM is used to capture the long-distance dependency in the fusion result after the outputs of the BERT model, the CharCNN module and the DSENet module are fused; the CRF is used to perform entity annotation on the output sequence of the MBiLSTM; the input text to be recognized is input into the named entity recognition model to obtain the named entities output by the named entity recognition model.
[0010] In one embodiment, the BERT model is specifically used to segment the input text into multiple words, map each word to a high-dimensional space, and obtain word embedding; add an additional embedding to each word to obtain segment embedding; add position information to each word to obtain position embedding; add the word embedding, the segment embedding, and the position embedding and input them into a bidirectional Transformer encoder to obtain features, grammatical structures, and semantic representations with contextual information.
[0011] In one embodiment, the CharCNN module is specifically used to decompose each input word into characters, and learn a low-dimensional embedding vector for the characters to obtain character embedding; after using multiple one-dimensional convolutional layers to perform a convolution operation on the character embedding, a pooling layer is used to perform a maximum pooling process on the output after the convolution operation; the features after the maximum pooling process are connected to obtain a comprehensive character-level representation; the comprehensive character-level representation is converted into a low-dimensional space using a linear layer, and the converted output is used as a character-level feature to complete the extraction of character-level features.
[0012] In one embodiment, the DSENet module is specifically used to extract the input features and divide them into a first weight vector for entity content and a second weight vector for entity position; compress the channel information of the first weight vector and the second weight vector respectively to obtain the corresponding first transformation vector and the second transformation vector; use two fully connected layers and a ReLU activation function to reduce the dimension of the first transformation vector and the second transformation vector respectively, and then use a Softmax activation function to process them to generate corresponding third weight vectors and fourth weight vectors; multiply the third weight vector and the fourth weight vector by the original input features to obtain calibration features that enhance important channel information and suppress unimportant channel information.
[0013] In one embodiment, the MBiLSTM includes a forward LSTM layer and a backward LSTM layer; the forward LSTM layer includes multiple LSTM units; and the backward LSTM layer includes multiple LSTM units.
[0014] In one embodiment, the CRF is specifically used to calculate the probabilities of all possible label sequences for a given input sequence, and select the label sequence with the highest probability as the prediction result.
[0015] The method of this embodiment has the following beneficial effects:
[0016] This embodiment proposes a named entity recognition model that can effectively solve the problems of character-level feature extraction and important feature weighting. In the model of this embodiment, the DSENet module is innovatively designed. This module replaces the traditional channel attention mechanism SENet module by emphasizing the features of entity content and entity position, effectively solving the problem of SENet module wasting computing resources on insignificant channel features. At the same time, this embodiment also uses a multi-layer BiLSTM structure to replace the single-layer BiLSTM to better capture long-distance dependencies in the sequence, which is very helpful for understanding complex sentence structures. In addition, this embodiment also uses the CharCNN module to capture semantic information at the character level, and uses the DSENet module to emphasize useful features, suppress unimportant feature extraction, and realize effective information mining. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of a method for extracting named entities from text based on channel attention and character features according to an embodiment of the present invention;
[0018] Figure 2 Schematic diagram of the structure of the TransBERT model according to an embodiment of the present invention;
[0019] Figure 3 Schematic diagram of the structure of the BERT model according to an embodiment of the present invention;
[0020] Figure 4 It is a structural diagram of the CharCNN module according to an embodiment of the present invention;
[0021] Figure 5 It is a structural diagram of a DSENet module according to an embodiment of the present invention;
[0022] Figure 6 This is a schematic diagram of the structure of the BiLSTM module according to an embodiment of the present invention;
[0023] Figure 7 This is a schematic diagram of the structure of an LSTM unit according to an embodiment of the present invention;
[0024] Figure 8 It is a schematic diagram of comparison results of different models using the Precision evaluation index on a text dataset according to an embodiment of the present invention;
[0025] Fig. 9 It is a schematic diagram of comparison results of different models using the Recall evaluation index on a text dataset according to an embodiment of the present invention;
[0026] Fig.10 It is a schematic diagram of comparison results of different models using the F1-score evaluation index on a text dataset according to an embodiment of the present invention;
[0027] Fig.11 1 is a diagram showing the internal structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] Before introducing the solution of this embodiment, the related work in this field is introduced first.
[0029] 1 Named Entity Recognition
[0030] The rule-based and dictionary-based methods are the earliest methods used in named entity recognition. This type of method usually relies on linguistic experts to construct rule templates. Its feature selection covers statistical information, punctuation, keywords, indicator words and direction words, position words, and central words, etc. It is mainly achieved through pattern and string matching. Most of these systems rely on the establishment of knowledge bases and dictionaries. However, these rules often rely on specific languages, fields, and text styles. The compilation process is time-consuming and difficult to cover all language phenomena. It is easy to make errors and the system has poor portability. The rules between different systems need to be rewritten by linguistic experts, which is also a significant shortcoming of the rule-based method. In addition, this method has problems such as long system construction cycle, poor portability, and the need to establish knowledge bases in different fields as an aid to improve the system recognition ability.
[0031] 2 BERT Pre-trained Model
[0032] In recent years, the rapid development of deep learning technology has provided new solutions for NER tasks. Lexical ambiguity has always been the most challenging part in the field of NER. To address this problem, the prior art proposes a new BERT pre-trained language model, which uses a bidirectional Transformer encoder to train the corpus and generate a bidirectional encoding representation of the text. The trained word vectors are dynamic word vectors, thereby improving the representation ability of word vectors. These models effectively capture the global contextual information of the input text, thereby performing well in NER tasks. Although the prior art has improved performance to a certain extent, there is still room for improvement in the generalization ability and feature expression of the model when processing unstructured text in specific fields.
[0033] 3 Attention Mechanism
[0034] In recent years, the attention mechanism has been widely used in the field of NER. Its core concept is to extract key information from complex input data, focus on the most important part of the current task, and build associations between them. Some researchers have proposed a multi-channel graph attention network MCGAT, which takes into account the relative position relationship between characters and words, and combines the statistical information of word frequency with point-by-point mutual information to further improve the performance of the model. Other researchers have proposed CLGP, which uses four specific extractors to obtain the embedding of characters, glyphs, pinyin, and dictionaries, and further uses a cross-attention-based network for multi-embedding fusion. Although the attention mechanism has achieved results in the field of NER, it does not adequately consider the interdependence between feature channels, thereby ignoring the importance recalibration at the feature level.
[0035] Based on this, this embodiment proposes a named entity recognition model called TransBERT, which aims to solve the problems of character-level feature extraction and important feature weighting. In the model, this embodiment innovatively designs the DSENet module, which replaces the traditional channel attention mechanism SENet module by emphasizing the features of entity content and entity position, effectively solving the problem of SENet module wasting computing resources on insignificant channel features. At the same time, this embodiment also uses a multi-layer BiLSTM structure to replace the single-layer BiLSTM to better capture the long-distance dependencies in the sequence, which is very helpful for understanding complex sentence structures. In addition, this embodiment also uses a character-level convolutional neural network CharCNN module to capture semantic information at the character level, and uses the DSENet module to emphasize useful features, suppress unimportant feature extraction, and realize effective information mining. Experimental results show that the model of this embodiment can show better performance than other models on specific domain data sets, significantly improving the accuracy of named entity recognition.
[0036] The present invention will be described in further detail below with reference to the accompanying drawings and embodiments.
[0037] The embodiment of the present invention provides a text named entity extraction method based on channel attention and character features, such as Figure 1 As shown, the method includes:
[0038] Step 101: Obtain input text to be recognized;
[0039] Step 102: construct a named entity recognition model, the named entity recognition model includes a BERT model, a CD module, an MBiLSTM and a CRF; the CD module includes a CharCNN module and a DSENet module; wherein the BERT model is used to process the input text to obtain features, grammatical structures and semantic representations with contextual information; the CharCNN module is used to extract character-level features in the output results of the BERT model; the DSENet module is used to enhance the entity content and entity position feature expressions in the output results of the BERT model; the MBiLSTM is used to capture the long-distance dependency in the fusion result after the outputs of the BERT model, the CharCNN module and the DSENet module are fused; the CRF is used to perform entity annotation on the output sequence of the MBiLSTM;
[0040] Step 103: input the input text to be recognized into the named entity recognition model to obtain the named entity output by the named entity recognition model.
[0041] Among them, the BERT model is specifically used to divide the input text into multiple words, map each word to a high-dimensional space, and obtain word embedding; add an additional embedding to each word to obtain segment embedding; add position information to each word to obtain position embedding; add the word embedding, the segment embedding and the position embedding and input them into a bidirectional Transformer encoder to obtain features, grammatical structure and semantic representation with contextual information.
[0042] The CharCNN module is specifically used to decompose each input word into characters, and learn low-dimensional embedding vectors for the characters to obtain character embeddings; after using multiple one-dimensional convolutional layers to perform convolution operations on the character embeddings, a pooling layer is used to perform maximum pooling processing on the output after the convolution operation; the features after the maximum pooling processing are connected to obtain a comprehensive character-level representation; the comprehensive character-level representation is converted into a low-dimensional space using a linear layer, and the converted output is used as a character-level feature to complete the extraction of character-level features.
[0043] The DSENet module is specifically used to extract the input features and divide them into a first weight vector for entity content and a second weight vector for entity position; compress the channel information of the first weight vector and the second weight vector respectively to obtain the corresponding first transformation vector and the second transformation vector; use two fully connected layers and a ReLU activation function to reduce the dimension of the first transformation vector and the second transformation vector respectively and then use a Softmax activation function to process them to generate the corresponding third weight vector and the fourth weight vector; multiply the third weight vector and the fourth weight vector by the original input features to obtain calibration features that enhance important channel information and suppress unimportant channel information.
[0044] The MBiLSTM includes a forward LSTM layer and a backward LSTM layer; the forward LSTM layer includes multiple LSTM units; the backward LSTM layer includes multiple LSTM units.
[0045] The CRF is specifically used to calculate the probabilities of all possible label sequences for a given input sequence, and select the label sequence with the highest probability as the prediction result.
[0046] This embodiment proposes a named entity recognition model TransBERT based on channel attention and character features, which can effectively realize information mining.
[0047] The model of this embodiment combines the deep semantic representation ability of BERT, the ability of MBiLSTM to capture sequence dependencies, and the annotation consistency constraint of CRF. At the same time, for the problem that BERT divides words into subword units, which may lead to incorrect segmentation at entity boundaries, and the BERT-BiLSTM-CRF model is insensitive to the importance of useful feature information, the model of this embodiment introduces the CharCNN module and the DSENet module to capture character-level semantic information and emphasize entity content and entity position features and suppress unimportant features, thereby improving the expressiveness and accuracy of model recognition.
[0048] Specifically, see Figure 2 , Figure 2This is a schematic diagram of the structure of the TransBERT model of an embodiment of the invention. In this embodiment, the model uses the BERT model to pre-train the data set to be processed to obtain features, grammatical structures, and semantic representations with rich contextual information. At the same time, the model of this embodiment also includes a DSENet module to emphasize entity content and entity position features, and also includes a CharCNN module to capture character-level features. The DSENet module and the CharCNN module together constitute a CD module. In addition, in order to extract global context features, this embodiment performs feature fusion on the output of the BERT model and the output of the CD module, and inputs the result of feature fusion into MBiLSTM. Finally, CRF is used to perform entity annotation on the output sequence of MBiLSTM to complete the named entity recognition process.
[0049] Next, several modules in the model of this embodiment will be introduced in detail.
[0050] 1. BERT model
[0051] The input of the BERT model can be a sentence or a sentence pair. The input text is first segmented into multiple tokens by a tokenizer, where special flags are added. The [CLS] flag is used to indicate the first position of the sentence. The two input sentences are separated by the [SEP] flag, and the [MASK] flag is used to hide certain words in the sentence. The tokens after token segmentation are mapped to a high-dimensional space to form token embeddings (also called Token Embeddings). These embedding vectors capture the semantic information of the tokens, enabling the BERT model to understand and process a variety of complex language phenomena and semantic relationships. Since the BERT model can process two sentences as input, a method is needed to distinguish between the two sentences. Segment embeddings (also called Segment Embeddings) add an additional embedding vector to each token to indicate which sentence it belongs to. Since the Transformer model itself does not have the ability to process the position information of the token in the sequence, position embeddings (also called Position Embeddings) are needed to provide this information. Each position has a unique embedding vector, which is learned during the training process. The token embeddings, segment embeddings, and position embeddings are added together to get the final input vector for each token. The BERT model uses a bidirectional Transformer as an encoder, and these embeddings are then passed to the Transformer for processing. Each encoder layer contains a self-attention mechanism and a feedforward neural network, allowing the model to capture complex dependencies in the input sequence. The structure of the BERT model is as follows: Figure 3 shown.
[0052] 2. CD module
[0053] The CD module consists of the CharCNN module and the DSENet module. The CharCNN module can also be called the CharCNN character-level convolutional neural network module. The DSENet module can also be called the DSENet channel self-attention mechanism module. The CharCNN module is used to extract character-level features of text, and the DSENet module is used to enhance the expression of entity content and entity location features. The combination of the two can better complete the extraction task.
[0054] (1) CharCNN module
[0055] Although the BERT model can capture the contextual information of vocabulary, it may not be able to effectively handle some morphologically rich vocabulary, especially those entities containing spelling variants, abbreviations or special symbols. By introducing the CharCNN module, the model can learn character-level features, so as to better handle the above situation and improve the recognition ability of complex vocabulary. The CharCNN module in this embodiment is a machine learning method that can be used to process string data. First, BiLSTM has generated a vector representation containing rich contextual information for each word in the sentence. Then, the CharCNN module is used to further mine the character-level features inside the word. Specifically, each word is decomposed into characters, and low-dimensional embedding vectors are learned for these characters to capture the morphological details of the word. Subsequently, these character embeddings are convolved through multiple one-dimensional convolutional layers Conv1d, and each convolutional layer uses convolution kernels of different sizes to extract features of different scales. The output of each convolutional layer is max-pooled MaxPool to retain the most significant feature points. These feature points are connected to form a comprehensive character-level representation. Finally, through the linear layer Linear, these features are converted to a lower dimensional space to prepare for subsequent tasks. Finally, the output of the CharCNN module is an enhanced version of the BERT model representation, which incorporates the contextual information of words and character-level features, thereby providing more accurate prediction capabilities in tasks such as named entity recognition. The structure of the CharCNN module is as follows: Figure 4 shown.
[0056] (2) DSENet module
[0057] Although BiLSTM can capture long-distance dependencies in sequences, in some cases, such dependencies may still not be sufficient. By introducing the DSENet module, long-distance dependencies in sequences can be captured more flexibly, especially for entities with larger spans. The DSENet module reweights the output of the CharCNN module layer by learning the importance weights of feature channels, emphasizing important features and suppressing unimportant features. The original SENet focuses on global channel features, while in this model, the goal of this embodiment is to focus on entity content and entity location information. Therefore, this embodiment optimizes and improves the module accordingly. The structure of the DSENet module is as follows: Figure 5 As shown in the figure, the DSENet module in this embodiment consists of three parts: Squeeze operation, Excitation operation, and Scale operation. First, the input features , the feature is obtained by feature extraction Ftr , separate the features of U into two independent weight vectors U 1 and U 2 , one for entity content and one for entity location. 1 and U 2 After Squeeze operation F sq Transformed into and The vector of feature U 1 and U 2 Perform average pooling. This step is to compress the information of the channel and help the model understand the global importance of entity content and entity position in the entire sequence. 1 and Z 2 The vector is subjected to Excitation operation, which consists of two fully connected layers, a ReLU activation function, and a Softmax activation function. The dimensionality is first reduced and then increased. Finally, the weight vector is generated through the sigmoid function. This step is used to learn the correlation between different channels. Then the Scale operation is performed: the channel attention weight obtained in the previous step is multiplied by the original input feature U to obtain the recalibrated feature This step is used to adjust the eigenvalues of each channel, emphasizing the information of important channels and suppressing the information of unimportant channels.
[0058] 3. BiLSTM module
[0059] BiLSTM is a variant of LSTM (Long-short term memory). It runs two LSTMs on the time series simultaneously, one for processing from front to back and the other for processing from back to front. The outputs of these two LSTMs are combined to better capture the bidirectional semantic dependencies in the sequence. The structure of the BiLSTM module is as follows: Figure 6 As shown in Figure 2. Both the forward and backward LSTM layers are composed of LSTM units. The LSTM unit structure is shown in Figure 2. Figure 7 As shown. The calculation process of LSTM can be summarized as forgetting information in the cell state and remembering new information, so that the information useful for subsequent calculations can be transmitted, while the useless information is discarded, and the hidden layer state is output at each time step. , where forgetting, memory, and output are determined by the hidden layer state at the previous moment and the current input Calculated forget gate , Memory Gate , output gate common control . Calculate the forget gate and select the forgotten information: Input: The hidden layer state at the previous moment , the current input word ; Output: the value of the forget gate As shown in the following formula (1). Calculate the memory gate and select the information to be remembered: Input: The hidden layer state at the previous moment , the current input word ; Output: value of memory gate , temporary cell state As shown in the following formula (2) and (3). Calculate the current cell state: Input: The value of the memory gate , the value of the forget gate , temporary cell state ; Output: current cell state As shown in the following formula (4). Calculate the state of the output gate and the hidden layer at the current moment: Input: The state of the hidden layer at the previous moment , the current input word , the current cell state ; Output: the value of the output gate , the hidden layer state As shown in the following equations (5) and (6). Finally, we can get the hidden layer state sequence with the same length as the sentence: .
[0060] (1)
[0061] (2)
[0062] (3)
[0063] (4)
[0064] (5)
[0065] (6)
[0066] in, represents the value of the forget gate, represents the weight matrix of the forget gate, represents the hidden layer state at the previous moment, Indicates the input word at the current moment, represents the bias term of the forget gate, represents the value of the memory gate, represents the weight matrix of the input gate, represents the bias term of the input gate, represents a temporary cell state, represents the weight matrix of the temporary cell, represents the bias term of the temporary cell, Indicates the current cell state. represents a temporary cell state, represents the value of the output gate, represents the weight matrix of the output gate, represents the bias term of the output gate, represents the sigmoid function, represents the tangent function, Represents the hidden layer state.
[0067] 4. CRF module
[0068] The CRF (Conditional Random Field) module is a probabilistic model for sequence labeling that can enhance the model's prediction ability by learning the mutual influence between entity labels. In the prediction stage, the CRF module calculates the probabilities of all possible label sequences for a given input sequence, and selects the sequence with the highest probability as the prediction result. Specifically, after obtaining the label vectors of each position based on BiLSTM, these label vectors will be passed into the CRF module as emission scores. The CRF module decodes a string of label sequences based on this. Since it is uncertain which label the label at the current position will transfer to, the transfer probability of this label is expressed by the transfer score, so the obtained label sequence is not unique. Among all possible paths, the Viterbi algorithm is used to find the label sequence corresponding to the path with the largest sum of the emission score and the transfer score and the best effect. Then this label sequence is the output of the model. Input sequence , represents the corresponding tag sequence, Indicates that from label y i Move to label y i+1 The transfer score is stored in the transfer matrix and will be learned during training as part of the model parameters. The formula for calculating the emission score can be obtained by the following formula (7): is the sum of the transfer score and the emission score, which is obtained by the following formula (8). Finally, the Viterbi algorithm is used to find the label sequence with the highest total score. , as shown in the following formula (9):
[0069] (7)
[0070] (8)
[0071] (9)
[0072] in, , represents the input sequence, , represents the corresponding tag sequence, is the sum of the transfer fraction and the emission fraction, Indicates that from label y i Move to label y i+1 The transfer score, represents all possible label sequences, Represents a given input sequence , the conditional probability of the label sequence y, Indicates the label y at position i i The emission fraction, represents the label sequence with the highest total score, Represents the input sequence and tag sequence The score function value of .
[0073] In addition, in order to verify the effect of this embodiment, the present application also conducted relevant experiments.
[0074] 1 Dataset and annotation method
[0075] The data set of this embodiment consists of relevant texts in a specific field. The data set is preprocessed, and 1753 annotated texts are randomly divided into a training set, a validation set, and a test set in a ratio of 8:1:1.
[0076] The mainstream tagging methods used in NER are BIO and BIOES. This experiment uses the BIO tagging method, where the B label represents the first character of the named entity, the I label represents the non-initial character of the named entity, and the O label represents an unrecognized entity.
[0077] 2 Experimental environment and parameter settings
[0078] In the experiment, this embodiment uses the Chinese-bert-wwm-ext pre-trained language model, which contains 12 layers, 768 hidden dimensions, and 12 attention heads. The model is trained on an RTX 4060 GPU, using the Windows 11 operating system, Python 3.9 programming language, and deep learning Pytorch 2.0.0 as the framework. In the experiment, the LSTM hidden layer size lstm_hidden is 128, the maximum sentence length max_seq_len is set to 512, the batch processing parameter batch_size is set to 12, the AdamW optimizer is used, the BERT learning rate bert_learning_rate is 3e-5, the CRF learning rate crf_learning_rate is 3e-3, and the dropout is 0.1.
[0079] 3 Experimental Evaluation Metrics
[0080] This embodiment uses typical NER evaluation indicators Precision (P), Recall (R) and F1-score (F1) to evaluate the performance of the model in this embodiment. Precision measures how many samples predicted by the model as positive examples are actually positive examples. Recall measures how many of the actual positive examples are correctly identified by the model. F1-score is the harmonic mean of Precision and Recall, which is used to comprehensively evaluate the performance of the model. P (True Positives) indicates the number of entities correctly identified by the model, F P (False Positives) indicates the number of entities that are incorrectly identified by the model, F N (False Negatives) indicates the number of entities that the model failed to recognize.
[0081] The specific calculation formula is as follows:
[0082] (10)
[0083] (11)
[0084] (12)
[0085] Among them, P represents precision, R represents recall, F1 represents F1 score, T P Indicates the number of entities correctly identified by the model, F P represents the number of entities that the model mistakenly identifies, F NRepresents the number of entities that the model failed to recognize.
[0086] 4 Experimental results and analysis
[0087] 4.1 Comparative Experiment
[0088] In order to evaluate the performance of the model proposed in this embodiment, the model of this embodiment is compared with BERT, BERT-LSTM, BERT-BiLSTM and other six models. Figure 8-Figure 10 It can be seen that the model of this embodiment outperforms other models in all evaluation indicators. This shows that the model of this embodiment can better learn the representation information of the entity layer. Compared with other models, the performance improvement of the TransBERT model is mainly attributed to the following points: First, this model is based on the BERT-BiLSTM-CRF model, and is an improvement on it, combining its advantages in capturing contextual information; second, the CD module in this model combines the advantages of the CharCNN and DSENet modules: CharCNN extracts character-level features and DSENet weights feature importance, thereby improving the accuracy of entity prediction; third, multi-layer BiLSTM can capture deeper contextual information in the sequence compared to single-layer BiLSTM, because entities often depend on the contextual information of the entire sentence.
[0089] 4.2 Ablation Experiment
[0090] (1) The impact of the number of BiLSTM layers on the BERT-MBiLSTM-CRF model
[0091] Since BiLSTM can effectively capture the semantic associations between long sequences and alleviate the phenomenon of gradient vanishing or exploding, we conducted experiments on the number of BiLSTM layers in the BERT-MBiLSTM-CRF model, considering the impact of the number of BiLSTM layers on the overall experimental effect. The relevant experimental data are shown in Table 1.
[0092] Table 1 Comparison of the effect of BiLSTM layers on the BERT-MBiLSTM-CRF model
[0093]
[0094] As can be seen from Table 1, when the number of BiLSTM layers in the BERT-MBiLSTM-CRF model is 2, the accuracy, recall, and value of the BERT-MBiLSTM-CRF model are the best overall. As the number of MBiLSTM layers increases, they show a negative growth. Therefore, in the following experiments, the number of BiLSTM layers of MBiLSTM is 2.
[0095] (2) Effectiveness of CharCNN and DSENet modules
[0096] In order to verify the effectiveness of the CharCNN module and the DSENet module, this embodiment designs the following two variants for comparative analysis: <1> Without CharCNN: Remove the CharCNN module from the TransBERT model and explore the impact of the DSENet module on model performance; <2> Without DSENet: Remove the DSENet module from the TransBERT model and explore the impact of the CharCNN module on the model performance. The experimental results are shown in Table 2. By analyzing the experimental results, the following conclusions can be drawn:
[0097] <1> The DSENet channel attention mechanism significantly affects the performance of the model. By weighting important features, the model can focus more on useful features. Removing the DSENet module leads to a decrease in performance.
[0098] <2> Character-level feature extraction is necessary for natural language processing tasks such as named entity recognition, which can achieve better results by improving the model's understanding of entity information in the text. Removing the CharCNN module resulted in a drop in performance.
[0099] Table 2. Module effectiveness ablation experiment results on text dataset
[0100]
[0101] (3) The impact of the location of the CharCNN module and the DSENet module on the model performance
[0102] In order to verify the influence of the position of CharCNN module and DSENet module on the model effect, this experiment designed the following two variants for comparative analysis: <1> CharCNN-DSENet (CD): explores the effect of the placement of the CharCNN module before the DSENet module on model performance; <2> DSENet-CharCNN (DC): Explore the effect of the position of the CharCNN module after the DSENet module on the model performance. The experimental results are shown in Table 3. The experimental results show that the position of the CharCNN module before the DSENet module has a better performance on the model effect than the position of the CharCNN module after the DSENet module.
[0103] Table 3. Experimental results of module position ablation on text dataset
[0104]
[0105] In summary, this embodiment proposes a hybrid neural network method based on channel attention and character features, and introduces the CharCNN module and the DSENet module to study the extraction of named entity recognition in Chinese text in a specific field. The model can extract rich semantic feature information and improve the accuracy of text entity recognition. BERT captures the deep bidirectional contextual information in the text, and MBiLSTM further captures the long-distance dependencies in the sequence data on the basis of BiLSTM. The CharCNN module is used to process character-level information and can extract features that are helpful for entity recognition from the character sequence. The DSENet module weights the importance of feature channels to emphasize entity content and entity position features, and suppresses unimportant features. The CRF layer imposes constraints on the final predicted label to ensure the rationality of the predicted entity label sequence. The experimental results on a specific field text dataset show that the model can efficiently and accurately extract information from text, and its performance is better than other baseline models. The superiority of each module of TransBERT is verified through ablation experiments.
[0106] In order to implement the method of an embodiment of the present invention, an embodiment of the present invention also provides a text named entity extraction system based on channel attention and character features, including: a processor and a memory for storing a computer program that can be run on the processor; wherein, when the processor is used to run the computer program, it executes the steps of the above-mentioned method.
[0107] The above-mentioned system provided in this embodiment belongs to the same concept as the above-mentioned method embodiment. The specific implementation process thereof is detailed in the method embodiment and will not be repeated here.
[0108] In order to implement the method of the embodiment of the present invention, the embodiment of the present invention also provides a computer program product, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of the above method.
[0109] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiment of the present invention, the embodiment of the present invention further provides an electronic device (computer device). Specifically, in one embodiment, the computer device may be a terminal, and its internal structure diagram may be as follows: Fig.11As shown. The computer device includes a processor A01, a network interface A02, a display screen A04, an input device A05 and a memory (not shown in the figure) connected through a system bus. Among them, the processor A01 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes an internal memory A03 and a non-volatile storage medium A06. The non-volatile storage medium A06 stores an operating system B01 and a computer program B02. The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 in the non-volatile storage medium A06. The network interface A02 of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor A01, the method of any one of the above embodiments is implemented. The display screen A04 of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device A05 of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0110] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0111] The device provided by the embodiment of the present invention includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the method of any one of the above embodiments is implemented.
[0112] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0113] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0114] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0116] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0117] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0118] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0119] It can be understood that the memory of the embodiment of the present invention can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAMbus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memories.
[0120] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0121] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A text named entity extraction method based on channel attention and character features, characterized in that: The method comprises: Get the input text to be recognized; Construct a named entity recognition model, the named entity recognition model includes a BERT model, a CD module, a multi-layer bidirectional long short-term memory network and a conditional random field; the CD module includes a CharCNN module and a DSENet module; wherein the BERT model is used to process the input text to obtain features, grammatical structures and semantic representations with contextual information; the CharCNN module is used to extract character-level features in the output results of the BERT model; the DSENet module is used to enhance the entity content and entity position feature expressions in the output results of the BERT model; the multi-layer bidirectional long short-term memory network is used to capture the long-distance dependency in the fusion result after the outputs of the BERT model, the CharCNN module and the DSENet module are fused; the conditional random field is used to perform entity annotation on the output sequence of the multi-layer bidirectional long short-term memory network; Inputting the input text to be recognized into the named entity recognition model to obtain the named entity output by the named entity recognition model; Among them, the DSENet module is specifically used to extract the character-level features input by the CharCNN module and divide them into a first weight vector for entity content and a second weight vector for entity position; compress the channel information of the first weight vector and the second weight vector respectively to obtain the corresponding first transformation vector and the second transformation vector; use two fully connected layers and a ReLU activation function to reduce the dimension and then increase the dimension of the first transformation vector and the second transformation vector respectively, and then use a Softmax activation function to process them to generate the corresponding third weight vector and the fourth weight vector; multiply the third weight vector and the fourth weight vector by the original input features to obtain calibration features that enhance important channel information and suppress unimportant channel information.
2. The text named entity extraction method based on channel attention and character features according to claim 1 is characterized in that: The BERT model is specifically used to segment the input text into multiple tokens, map each of the tokens to a high-dimensional space, and obtain a token embedding; add an additional embedding to each of the tokens to obtain a segment embedding; Adding position information to each of the word units to obtain position embedding; The word unit embedding, the segment embedding and the position embedding are added and input into a bidirectional Transformer encoder to obtain features, grammatical structures and semantic representations with context information.
3. The text named entity extraction method based on channel attention and character features according to claim 1 is characterized in that: The CharCNN module is specifically used to decompose each input word into characters, and learn low-dimensional embedding vectors for the characters to obtain character embeddings; after performing convolution operations on the character embeddings using multiple one-dimensional convolutional layers, the output after the convolution operation is subjected to maximum pooling processing using a pooling layer; The features after the maximum pooling process are connected to obtain a comprehensive character-level representation; the comprehensive character-level representation is converted into a low-dimensional space using a linear layer, and the converted output is used as a character-level feature to complete the extraction of character-level features.
4. The text named entity extraction method based on channel attention and character features according to claim 1 is characterized in that: The multi-layer bidirectional long short-term memory network includes a forward LSTM layer and a backward LSTM layer; the forward LSTM layer includes multiple LSTM units; the backward LSTM layer includes multiple LSTM units.
5. The text named entity extraction method based on channel attention and character features according to claim 4 is characterized in that: The conditional random field is specifically used to calculate the probabilities of all possible label sequences for a given input sequence, and select the label sequence with the highest probability as the prediction result.
Citation Information
Patent Citations
Chinese named entity recognition algorithm based on multi-information enhancement
CN114154504A
Multi-modal false information detection system based on co-situation theory guidance
CN117851894A