A Named Entity Recognition Method and System Based on Evidence Deep Learning

Through the evidence deep learning method based on the Dilicre model, the problems of model uncertainty and sparse entity recognition in named entity recognition are solved, the accuracy and robustness of the model in an open environment are improved, the recognition ability of sparse entities and unknown entities is enhanced, and the utilization efficiency of training samples is improved.

CN118114673BActive Publication Date: 2025-07-29BEIJING ZHONGJING HUAZHI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410392962.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2025-07-29
Estimated Expiration
2044-04-02

AI Technical Summary

Technical Problem

In the existing named entity recognition technology, insufficient quantification of model uncertainty, difficulty in identifying sparse entities, insufficient recognition of unknown vocabulary and entity outside the domain, resulting in poor recognition of models in open environments and low training samples.

Method used

The evidence deep learning method based on the Dilikre model is adopted to optimize the model training process through the bootstrap loss function of confidence and uncertainty, improve the recognition ability of sparse entities and unknown entities, and improve the uncertainty estimation of the model through uncertainty quality optimization.

Benefits of technology

It improves the recognition accuracy and robustness of the model in an open environment, enhances the recognition ability of sparse entities and unknown entities, reduces dependence on labeled data, and improves training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118114673B_ABST
    Figure CN118114673B_ABST
Patent Text Reader

Abstract

The present invention discloses a named entity recognition method and system based on evidence deep learning. The method includes: obtaining a named entity recognition data set; preprocessing the named entity recognition data set to obtain a processed named entity recognition data set; extracting features from the processed named entity recognition data set to obtain feature data; constructing an evidence deep learning model based on the Dirichlet model; training the evidence deep learning model based on the Dirichlet model according to the feature data and an uncertainty-guided loss function for uncertainty quality optimization to obtain a named entity recognition model; and performing named entity recognition according to the named entity recognition model. The present invention can overcome the limitations in the existing named entity recognition technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a named entity recognition method and system based on evidence deep learning. Background Art

[0002] In the field of natural language processing (NLP), named entity recognition (NER) is a key task aimed at identifying and classifying specific types of entities, such as person names, place names, organization names, etc., from text. Although deep learning techniques, especially neural network-based methods, have made significant progress in the NER task, there are still some problems and drawbacks in the existing technologies:

[0003] 1. Insufficient quantification of model uncertainty: Most NER systems focus on improving model performance, such as recognition accuracy and F1 score, but often ignore the quantification of model uncertainty. In an open environment, the uncertainty of model predictions is crucial for the reliability of the system, as it can reveal the possibility of model prediction errors.

[0004] 2. Sparse entity problem: In a text corpus, entities usually only account for a minority, while non-entity type words (such as the "other" category) occupy the majority. This imbalance may cause the model to overfit non-entity words, thus affecting the recognition performance of entity types.

[0005] 3. Out-of-vocabulary (OOV) and out-of-domain (OOD) entity recognition: In practical applications, an NER system may encounter words that do not appear in the training data or entities outside the domain, which requires the model to have good generalization ability. However, existing EDL methods lack explicit modeling when dealing with OOV and OOD entities, resulting in insufficient recognition performance in these cases.

[0006] 4. Sample efficiency: High-quality uncertainty estimation can help improve the training efficiency of the model by selecting more informative samples to reduce the number of labeled samples required. However, the existing methods are not yet mature enough in this regard. Summary of the Invention

[0007] The object of the present invention is to provide a named entity recognition method and system based on evidence deep learning, which can overcome the limitations in the existing named entity recognition technology, especially the accurate quantification problem of model prediction uncertainty, the recognition problem of sparse entities, the effective recognition problem of out-of-vocabulary and out-of-domain entities, and further improve the utilization efficiency of training samples.

[0008] To achieve the above object, the present invention provides the following solutions:

[0009] A named entity recognition method based on evidence deep learning includes:

[0010] Obtain a named entity recognition dataset;

[0011] Preprocess the named entity recognition dataset to obtain a processed named entity recognition dataset;

[0012] Extract features from the processed named entity recognition dataset to obtain feature data;

[0013] Construct an evidence deep learning model based on the Dirichlet model;

[0014] Train the evidence deep learning model based on the Dirichlet model according to the feature data and an uncertainty-guided loss function for uncertainty quality optimization to obtain a named entity recognition model;

[0015] Perform named entity recognition according to the named entity recognition model.

[0016] Optionally, the evidence deep learning model based on the Dirichlet model includes confidence and uncertainty, where the confidence is the probability of the evidence assigned to each class, the uncertainty is used to provide uncertainty estimation, and the uncertainty calculates the class-level uncertainty of each instance using the belief quality of the reference class to adjust the loss.

[0017] Optionally, the named entity recognition dataset includes CoNLL-2003 and OntoNotes 5.0.

[0018] Optionally, the preprocessing of the named entity recognition dataset to obtain a processed named entity recognition dataset specifically includes:

[0019] Perform word segmentation, tokenization, add special symbols, and label conversion on the named entity recognition dataset to obtain a processed named entity recognition dataset.

[0020] Optionally, the feature extraction from the processed named entity recognition dataset to obtain feature data specifically includes:

[0021] Extract features from the processed named entity recognition dataset using a pre-trained Seq2Seq model to obtain feature data.

[0022] To achieve the above object, the present invention also provides the following solution:

[0023] A named entity recognition system based on evidence deep learning includes:

[0024] A named entity recognition dataset acquisition module for obtaining a named entity recognition dataset;

[0025] A named entity recognition dataset processing module, which is used to preprocess the named entity recognition dataset to obtain a processed named entity recognition dataset;

[0026] A feature data extraction module, which is used to extract features from the processed named entity recognition dataset to obtain feature data;

[0027] A Dirichlet model-based evidence deep learning model construction module, which is used to construct a Dirichlet model-based evidence deep learning model;

[0028] An evidence deep learning model determination module, which is used to train the Dirichlet model-based evidence deep learning model according to the feature data and an uncertainty-guided loss function for uncertainty quality optimization to obtain a named entity recognition model;

[0029] A named entity recognition module, which is used to perform named entity recognition according to the named entity recognition model.

[0030] Optionally, the Dirichlet model-based evidence deep learning model includes confidence and uncertainty, where the confidence is the probability of evidence assigned to each category, the uncertainty is used to provide uncertainty estimation, and the uncertainty uses the belief quality of the reference category to calculate the category-level uncertainty of each instance to adjust the loss.

[0031] To achieve the above object, the present invention also provides the following solution:

[0032] An electronic device includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a named entity recognition method based on evidence deep learning.

[0033] To achieve the above object, the present invention also provides the following solution:

[0034] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements a named entity recognition method based on evidence deep learning. To achieve the above object, the present invention also provides the following solution:

[0035] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0036] The present invention provides a named entity recognition method based on evidence deep learning. The method includes: obtaining a named entity recognition data set; preprocessing the named entity recognition data set to obtain a processed named entity recognition data set; extracting features from the processed named entity recognition data set to obtain feature data; constructing an evidence deep learning model based on the Dirichlet model; training the evidence deep learning model based on the Dirichlet model according to the feature data to obtain a trained evidence deep learning model; determining an uncertainty-guided loss function of the trained evidence deep learning model, performing uncertainty quality optimization to obtain a named entity recognition model; and performing named entity recognition according to the named entity recognition model. The present invention can overcome the limitations in the existing named entity recognition (NER) technology, especially the accurate quantification problem of model prediction uncertainty, the recognition problem of sparse entities, the effective recognition problem of out-of-vocabulary (OOV) and out-of-domain (OOD) entities, thereby improving the utilization efficiency of training samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0038] Figure 1 It is a flowchart of the named entity recognition method based on evidence deep learning of the present invention;

[0039] Figure 2 It is a structural diagram of the named entity recognition system based on evidence deep learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0041] The purpose of the present invention is to provide a named entity recognition method and system based on evidence deep learning, which can overcome the limitations in the existing named entity recognition (NER) technology, especially the accurate quantification problem of model prediction uncertainty, the recognition problem of sparse entities, the effective recognition problem of out-of-vocabulary (OOV) and out-of-domain (OOD) entities, thereby improving the utilization efficiency of training samples.

[0042] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0043] Example 1:

[0044] Figure 1 This is a flowchart of the named entity recognition method based on evidence deep learning according to the present invention. As Figure 1 shown, the present invention provides a named entity recognition method based on evidence deep learning, and the method includes:

[0045] Step 101: Obtain a named entity recognition data set.

[0046] The named entity recognition data set includes CoNLL-2003 and OntoNotes 5.0.

[0047] Step 102: Preprocess the named entity recognition data set to obtain a processed named entity recognition data set.

[0048] This step specifically includes:

[0049] Perform word segmentation, tokenization, add special symbols, and label conversion on the named entity recognition data set to obtain a processed named entity recognition data set.

[0050] Step 103: Extract features from the processed named entity recognition data set to obtain feature data.

[0051] This step specifically includes:

[0052] Extract features from the processed named entity recognition data set using a pre-trained Seq2Seq model to obtain feature data.

[0053] Given a word sequence X = {x (1) ,..., x (n)} and a target sequence Y = {y (1 ),..., y (n)}. To obtain the hidden representation H of X, first preprocess the words in the sentence X according to the input form required by the corresponding named entity recognition method. Then, input the processed input into an encoder module to calculate the hidden representation H = Encoder(X), where d h represents the dimension of the hidden representation.

[0054] Note that the input format of the named entity recognition model may vary depending on the paradigm used. Common named entity recognition paradigms include sequence tagging, span recognition, and sequence transduction models. Different input formats require different models for feature processing. For the Seq2Seq model, the present invention selects a pointer-based model (Yan H, Gui T, Dai J, et al. A unified generative framework for various NER subtasks [J]. arXiv preprint arXiv:2106.01223, 2021.), so there is no need to learn the entire vocabulary.

[0055] Once the present invention obtains the hidden representation, the present invention introduces a layer based on the Dirichlet distribution to generate the final prediction distribution. Specifically, for the i-th sample, the hidden representation h is fed into a fully connected layer to output logits, and then the present invention can convert the logits into Dirichlet parameters α. Finally, only one forward calculation is needed to calculate the uncertainty.

[0056] Step 104: Construct an evidence deep learning model based on the Dirichlet model.

[0057] The evidence deep learning model based on the Dirichlet model includes confidence and uncertainty. The confidence is the probability of the evidence assigned to each category, and the uncertainty is used to provide an uncertainty estimate. The uncertainty calculates the class-level uncertainty of each instance using the belief mass of the reference class to adjust the loss.

[0058] The basic structure of a traditional neural network classifier usually refers to a multi-layer feedforward neural network. This network consists of multiple layers, including an input layer, hidden layers, and an output layer. The following are the basic components of this network structure:

[0059] 1. Input layer: This is the first layer of the neural network and is responsible for receiving external input data.

[0060] 2. Hidden layer: The hidden layer is located between the input layer and the output layer and can be multiple.

[0061] 3. Output layer: This is the last layer of the neural network and is responsible for outputting the final classification result. In a classification task, each node in the output layer usually corresponds to a category, and the activation function of the output layer usually uses the Softmax function, which can convert the output into a probability distribution for easy category selection.

[0062] Traditional neural network classifiers typically use a Softmax layer to provide a point estimate of the classification distribution. In contrast, the Dirichlet-based model (DBM) outputs the parameters of a Dirichlet distribution, which are then used to estimate the classification distribution. Specifically, for the i-th sample x in a C-class classification task (i) , the DBM replaces the Softmax layer of the neural network with an activation function layer (such as Softplus) to ensure that the network outputs non-negative values, which are considered as evidence supporting classification e(i) is the output of the entire network.

[0063] The classification evidence can be used to construct a Dirichlet distribution, which models the distribution of different classes. In other words, specific class information, i.e., the classification evidence, can be obtained through this distribution.

[0064] Then, these evidences are used to construct a Dirichlet distribution, which models the distribution of different classes. To this end, the parameters of the Dirichlet distribution are obtained as follows: α (i) = e (i) + 1, where 1 represents a vector with C elements all equal to 1. Finally, the density function of the Dirichlet distribution is:

[0065]

[0066] where B(α (i) ) is the C-dimensional multinomial beta function, and p c is the probability of each class. To learn the model parameters, given a sample (x (i) , y (i) ), where y (i) is the C-dimensional one-hot label of the sample x (i) , an optimization objective is constructed by combining the cross-entropy classification loss and the KL penalty loss .

[0067]

[0068] where ψ(·) is the digamma function, S(i) represents the Dirichlet strength, λ1 is a balancing factor, Dir(p(i)|1) is a special case equivalent to the uniform distribution, represents the masking parameter, ⊙ represents the Hadamard (element-wise) product, a (i) represents the prediction parameter, and KL represents the KL divergence.

[0069] Once the present invention obtains the Dirichlet distribution for prediction, the present invention can estimate the uncertainty of the prediction in a closed form. To this end, EDL provides two probabilities: confidence and uncertainty. Confidence represents the probability of the evidence assigned to each category, and uncertainty provides an uncertainty estimate. The present invention uses the parameter α of the Dirichlet distribution to represent the category probability distribution, which allows the model to efficiently estimate the prediction uncertainty in a single forward pass. In this way, the model can assign a probability value to each category while maintaining the quantification of uncertainty.

[0070] Due to the sparse entity and OOV / OOD entity problems, directly applying EDL to NER will result in suboptimal uncertainty estimation. The present invention improves the traditional EDL method by incorporating confidence and uncertainty into the network training process. Specifically, the present invention introduces two key modifications: (1) Calculate the importance weight for each sample according to the confidence to re-weight the original classification loss. (2) The present invention introduces another term to increase the uncertainty of mispredicted instances, thus significantly improving the quality of uncertainty estimation and helping to detect OOD entities.

[0071] Specifically, in order to make the training pay more attention to entities and increase the evidence corresponding to the reference category, the present invention uses the belief quality of the reference category to calculate the category-level uncertainty of each instance to adjust the loss. Specifically, for the i-th sample, the present invention uses (1 - b (i) ), where b represents the belief quality, which is calculated as follows:

[0072]

[0073] c represents the category, and S(i) represents the Dirichlet strength. As the category-level uncertainty, this uncertainty acts as an important weight for the entity category during the training process. To this end, the present invention uses an important weight (IW), w (i) = (1 - b (i) ) ⊙ y (i) , to replace the true label y(i) of the reference category in one-hot representation, and the loss is:

[0074]

[0075] A higher confidence quality for the true category indicates a higher level of certainty in the prediction. In this case, the assigned importance weight (IW) will be very small. In this way, the training process can pay more attention to sparse but valuable entities.

[0076] Step 105: Train the Dirichlet model-based evidence deep learning model according to the feature data and the uncertainty-guided loss function for uncertainty quality optimization to obtain a named entity recognition model.

[0077] The input is a preprocessed dataset, and the output is specific entity category labels.

[0078] Training method: After sending the input to the model, the model makes predictions, compares the predicted results with the labels, calculates the loss, and uses this loss for backpropagation to optimize the model.

[0079] Assigning a higher uncertainty to OOV / OOD entities helps to detect OOV / OOD entities. However, there are no real OOV / OOD samples available during training. The present invention proposes to regard difficult samples as OOV / OOD samples that are still mispredicted after sufficient model training, and these samples are usually outliers. In this way, the present invention enables the model to detect OOV / OOD data. Specifically, Uncertainty Quality Optimization (UNM) assigns a higher uncertainty to samples that are prone to errors by adding an uncertainty quality penalty term to the mispredicted samples to represent the lack of evidence. The calculation is as follows:

[0080]

[0081] \(\hat{y}^{(i)}\) is the predicted category, \(y^{(i)}\) is the true category, \(\lambda_2\) is a decreasing coefficient, and the calculation is as follows: \(\lambda_2=\lambda_0\exp\{-(\ln\lambda_0 / T)t\}\), where \(\lambda_2\in[\lambda_0,1]\), \(\lambda_0\ll1\) is a small positive constant, \(t\) is the current training epoch, and \(T\) is the total number of training epochs. As the training epoch \(t\) increases towards \(T\), the factor \(\lambda_2\) will monotonically increase from \(\lambda_0\) to 1.0. This allows the network to initially focus on optimizing classification and gradually shift its focus to optimizing UNM as training progresses. \(\exp\) is the exponential function with the constant \(e\) (the base of the natural logarithm) as the base, and \(\ln\) is the logarithm function with \(e\) as the base.

[0082] \(u\) is the uncertainty quality, and the calculation is as follows:

[0083]

[0084] \(C\) is the total number of categories.

[0085] Step 106: Perform named entity recognition according to the named entity recognition model.

[0086] The technical solution of the present invention has the following remarkable advantages and positive effects compared with the prior art:

[0087] 1. Improve the quality of uncertainty estimation: By introducing an uncertainty-guided loss function, the model can more accurately estimate the uncertainty of the prediction results, thus providing more reliable entity recognition results in an open environment.

[0088] 2. Enhance the processing ability for sparse entities: Through importance-weighted loss, the present invention can better focus on and learn entity categories, and maintain a high recognition accuracy even when the entity appears with a low frequency.

[0089] 3. Improve OOV / OOD entity recognition: The uncertainty quality optimization loss helps the model distinguish and recognize unknown words and out-of-domain entities, improving the robustness of the NER system when facing these challenges.

[0090] 4. Enhance sample efficiency: By optimizing the uncertainty estimation, the present invention helps to select more informative samples for training, thereby reducing the dependence on a large amount of labeled data and improving the training efficiency.

[0091] 5. Multi-paradigm applicability: The flexibility of the present invention enables it to adapt to different NER tasks and datasets, enhancing the generality and practicality of the technology.

[0092] 6. Improve system performance: While maintaining or improving the model performance, the present invention enhances the reliability and generalization ability of the NER system by improving the uncertainty estimation, making the system more stable and efficient in practical applications.

[0093] In summary, the technical solution of the present invention shows significant advantages in improving the accuracy, robustness, and efficiency of the NER system, bringing positive progress to the field of natural language processing.

[0094] Embodiment 2:

[0095] This embodiment provides a named entity recognition method based on the Seq2Seq paradigm implemented using the Pytorch and Transformers frameworks, specifically including the following steps:

[0096] Step 1: Configure the basic environment, install Python (Python 3.8 or higher version is recommended), Pytorch, and the Transformers framework.

[0097] Step 2: Select an appropriate dataset and use a pre-trained Seq2Seq model to extract data features. Commonly used NER datasets include CoNLL-2003, OntoNotes 5.0, etc. Data preprocessing: including word segmentation, tokenization, adding special symbols, and converting labels into a format suitable for the model.

[0098] Step 3: Construct an evidence deep learning (EDL) training process based on the Dirichlet model (DBM).

[0099] Step 4: Construct an uncertainty-guided loss function and perform uncertainty quality optimization.

[0100] Step Five: Train the model and conduct verification and evaluation, with special attention to:

[0101] 1. Model performance and uncertainty quantification: Analyze the impact of uncertainty in the model's predictions.

[0102] 2. Sparse entity recognition: Evaluate the model's ability to identify entities with low occurrence frequencies.

[0103] 3. Out-of-vocabulary (OOV) and out-of-domain (OOD) entity recognition: Test the model's performance in handling vocabulary and entities outside the training set.

[0104] Example Three:

[0105] Figure 2 This is the structural diagram of the named entity recognition system based on evidence deep learning according to the present invention. As Figure 2 shown, a named entity recognition system based on evidence deep learning includes:

[0106] A named entity recognition dataset acquisition module 201 for acquiring a named entity recognition dataset;

[0107] A named entity recognition dataset processing module 202 for preprocessing the named entity recognition dataset to obtain a processed named entity recognition dataset;

[0108] A feature data extraction module 203 for extracting features from the processed named entity recognition dataset to obtain feature data;

[0109] An evidence deep learning model construction module 204 based on the Dirichlet model for constructing an evidence deep learning model based on the Dirichlet model;

[0110] A named entity recognition model determination module 205 for training the evidence deep learning model based on the Dirichlet model according to the feature data and an uncertainty-guided loss function for uncertainty quality optimization to obtain a named entity recognition model;

[0111] A named entity recognition module 206 for performing named entity recognition according to the named entity recognition model.

[0112] The evidence deep learning model based on the Dirichlet model includes confidence and uncertainty. Among them, the confidence is the probability of the evidence assigned to each category, and the uncertainty is used to provide uncertainty estimation. The uncertainty calculates the class-level uncertainty of each instance using the belief mass of the reference category to adjust the loss.

[0113] Example Four:

[0114] This embodiment provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the named entity recognition method based on evidence deep learning in Embodiment 1.

[0115] Optionally, the above-mentioned electronic device may be a server.

[0116] In addition, an embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the named entity recognition method based on evidence deep learning in Embodiment 1.

[0117] Embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0118] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0119] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the functions specified in one Figure 1Steps of the functions specified in one or more processes and / or boxes Figure 1 Steps of the functions specified in one or more boxes

[0121] The embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section

[0122] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention

Claims

1. A named entity recognition method based on evidence deep learning, which is used to identify and classify specific types of entities from text, characterized in that, The method includes: Obtaining a named entity recognition data set; Preprocessing the named entity recognition data set to obtain a processed named entity recognition data set; Performing feature extraction on the processed named entity recognition data set to obtain feature data; Constructing an evidence deep learning model based on the Dirichlet model; Training the evidence deep learning model based on the Dirichlet model according to the feature data and an uncertainty-guided loss function for uncertainty quality optimization to obtain a named entity recognition model; Performing named entity recognition according to the named entity recognition model; For the Seq2Seq model, a pointer-based model is selected, so there is no need to learn the entire vocabulary; For the i-th sample, the hidden representation h is fed into a fully connected layer to output logits, and then the logits are converted into Dirichlet parameters α; finally, only one forward calculation is needed to calculate the uncertainty; The evidence deep learning model based on the Dirichlet model includes confidence and uncertainty, where the confidence is the probability of evidence assigned to each category, the uncertainty is used to provide uncertainty estimation, and the uncertainty uses the belief mass of the reference category to calculate the class-level uncertainty of each instance to adjust the loss; Among them, constructing an evidence deep learning model based on the Dirichlet model; the specific steps are as follows: For the i-th sample x in the C-class classification task (i) , the DBM replaces the Softmax layer of the neural network with an activation function layer to ensure that the network outputs non-negative values, which are considered as evidence to support classification e (i) is the output of the entire network; Classification evidence is based on e (i) Construct a Dirichlet distribution that models the distributions of different classes, and obtain specific class information, i.e., classification evidence, through this distribution; Then, these evidences are used to construct a Dirichlet distribution that models the distributions of different classes; to this end, the parameters of the Dirichlet distribution are obtained as follows: α (i) = e (i) + 1, where 1 represents a vector with C elements all equal to 1; finally, the density function of the Dirichlet distribution is: where B(α( i) ) is the C-dimensional polynomial beta function, and p c is the probability for each class; to learn the model parameters, given a sample (x (i) , y (i) ), where y (i) is the C-dimensional one-hot label of the sample x (i) , the optimization objective is constructed by combining the cross-entropy classification loss and the KL penalty loss ; where ψ(·) is the digamma function, S (i) denotes the Dirichlet strength, λ1 is the balance factor, Dir(p (i) |1) is a special case equivalent to the uniform distribution, denotes the mask parameter, ⊙ denotes the Hadamard product, a (i) denotes the prediction parameter, KL represents the KL divergence; After feeding the input into the model, the model generates predictions, compares the predicted results with the labels, calculates the loss, and uses the loss for backpropagation to optimize the model; Assigning a higher uncertainty to OOV / OOD entities helps to detect OOV / OOD entities. However, since there are no true OOV / OOD samples available during training, difficult samples are regarded as OOV / OOD samples that are still mispredicted after sufficient model training. These samples are usually outliers. In this way, the model can detect OOV / OOD data. Specifically, the Uncertainty Quality Optimization (UNM) assigns a higher uncertainty to samples that are prone to errors by adding an uncertainty quality penalty term to the mispredicted samples to represent the lack of evidence, which is calculated as follows: y^(i) is the predicted category, y(i) is the true category, λ2 is a decreasing coefficient, and the calculation is as follows: λ2 = λ0exp{-(lnλ0 / T)t}, where λ2 ∈ [λ0, 1], λ0 << 1 is a small positive constant, t is the current training epoch, and T is the total number of training epochs; as the training epoch t increases towards T, the factor λ2 will monotonically increase from λ0 to 1.0, which allows the network to initially focus on optimizing classification and gradually shift its focus to optimizing UNM as the training progresses. exp is the exponential function with base constant e, and ln is the logarithmic function with base e; u is the uncertainty quality, and the calculation is as follows: C is the total number of categories.

2. The named entity recognition method based on evidence deep learning according to claim 1, characterized in that, The named entity recognition data set includes CoNLL-2003 and OntoNotes 5.

0.

3. The named entity recognition method based on evidence deep learning according to claim 1, wherein The preprocessing of the named entity recognition data set to obtain a processed named entity recognition data set specifically includes: Performing word segmentation, tokenization, adding special symbols, and label conversion on the named entity recognition data set to obtain a processed named entity recognition data set.

4. A named entity recognition system based on evidence deep learning, which applies the named entity recognition method based on evidence deep learning according to any one of the above claims 1-3, characterized in that, The system includes: A named entity recognition data set acquisition module for obtaining a named entity recognition data set; A named entity recognition data set processing module for preprocessing the named entity recognition data set to obtain a processed named entity recognition data set; A feature data extraction module for performing feature extraction on the processed named entity recognition data set to obtain feature data; A building module for an evidence deep learning model based on the Dirichlet model, which is used to build an evidence deep learning model based on the Dirichlet model; A named entity recognition model determination module, which is used to train the evidence deep learning model based on the Dirichlet model according to the feature data and an uncertainty-guided loss function for uncertainty quality optimization, to obtain a named entity recognition model; A named entity recognition module, which is used to perform named entity recognition according to the named entity recognition model; The evidence deep learning model based on the Dirichlet model includes confidence and uncertainty. Among them, the confidence is the probability of the evidence assigned to each category, and the uncertainty is used to provide uncertainty estimation. The uncertainty uses the belief quality of the reference category to calculate the category-level uncertainty of each instance to adjust the loss.

5. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the named entity recognition method based on evidence deep learning according to any one of claims 1-3.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the named entity recognition method based on evidence deep learning according to any one of claims 1-3.

Citation Information

Patent Citations

  • Naming entity recognition method, language recognition method and system

    CN109388795A

  • Keyword extraction method based on Seq2seq framework

    CN110119765A

  • Name entity recognition with deep learning

    CN113853606A

  • Unmanned aerial vehicle aerial photography target credible identification method based on data and model dual uncertainty perception

    CN116740587A