A Tuning Method for Address Named Entity Recognition Based on Deep Learning Model
By building an entity dictionary and optimizing the mask mechanism in the pre-training stage, combined with a real-time data optimization model, the problem that the address named entity recognition model in the existing technology relies on a large amount of labeled data, achieving more efficient model tuning and recognition effects.
Patent Information
- Application Number
- CN202111443614.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-11-30
AI Technical Summary
In the prior art, in address naming entity recognition, there is a problem that model optimization requires relying on a large amount of labeled data and poor model recognition effect.
By collecting industry corpus to build a solid dictionary, using unlabeled corpus to optimize the masking mechanism in the pre-training stage of the neural network language model, fine-tuning the model for downstream tasks, and optimizing the output model through real-time data.
This realizes model tuning that does not rely on a large amount of labeled data, and improves the accuracy and generalization ability of address naming entity recognition.
Smart Images

Figure CN114169332B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to natural language recognition, and specifically to a tuning method for address named entity recognition based on a deep learning model. Background Art
[0002] The named entity recognition task is a very common task in the field of natural language processing, and the purpose of this task is to identify specific types of entities in natural language text. The application of named entity recognition is very extensive. For example, in the express delivery industry, it is necessary to identify information such as the name, phone number, item, and detailed address of the express delivery sender and receiver; in the news media industry, it is necessary to identify information such as person names, place names, and organization names; in the medical industry, it is necessary to identify information such as the names of patients and doctors, pathological names, symptoms, medication names, and dosage instructions; in the field of bioinformatics, it is necessary to extract information such as proteins and DNA.
[0003] The named entity recognition task is usually modeled as a character-level sequence labeling task, that is, for a string of input character sequences, the named entity recognition model needs to predict the named entity label corresponding to each character. Currently, there are mainly two typical named entity recognition models in the practical application of natural language processing technology.
[0004] The first method is a model based on LSTM (Long Short-Term Memory Neural Network) and CRF (Conditional Random Field). The training process of this model is to perform Embedding (word vector encoding) operation on the input Chinese sequence according to character encoding, and then send it into the neural network of LSTM. In order to consider context information, a bidirectional LSTM model is generally used. In order to constrain the relationship between entities and the law of state transition between entities, a CRF layer neural network is generally used to model it later. The named entity model based on LSTM and CRF is relatively simple in structure, has a small number of parameters, occupies less computing resources, and runs fast, but has the disadvantages of low prediction accuracy and weak generalization ability.
[0005] The second method is based on a deeper pre-trained model. Its essence is to use the pre-trained model to replace the role of LSTM in the first method to represent richer and more complex semantic relationships. The current mainstream pre-trained models are a series of models based on Transformer, such as Bert, GPT, etc.
[0006] For entity recognition using the Bert model, pre-training is first carried out, which includes the Masked Language Model (MLM) and Next Sentence Prediction (NSP); then, for fine-tuning specific tasks, the Bert model is used to extract the vector features of each character in the Chinese sequence for label classification of each character, and then in the CRF stage, the label values are constrained within a reasonable range. The above is the model training process based on Bert and CRF. Through the pre-training method, prior knowledge such as grammar and semantics can be obtained from a large amount of unlabeled data.
[0007] However, in the Bert model, only the features of each character are considered, while the relationships between Chinese entities are ignored. To improve the effect, a large amount of corpus must be relied on for model fine-tuning, and there are also bottlenecks in the effect.
[0008] Both of the above methods have obvious defects in terms of model effect and training process. The extraction of Chinese vector features in both methods is based on the character level, and the relationships and features between Chinese "words" cannot be considered. At the same time, to improve the model accuracy, both rely on a large amount of labeled data, and a large amount of manpower is required to analyze the misrecognized data of the model. Although the first method has a simple model structure and short training time, it is difficult to improve the effect; although the second method uses a pre-trained model and incorporates prior knowledge of Chinese in the model initialization stage, it only considers the relationships at the character level and ignores the information between entities, and there is a slight improvement in the named entity recognition task, but it relies on a large amount of data. Summary of the Invention
[0009] (1) Technical problems to be solved
[0010] In view of the above-mentioned disadvantages of the existing technology, the present invention provides an optimization method for address named entity recognition based on a deep learning model, which can effectively overcome the defects of the existing technology that model optimization requires relying on a large amount of labeled data and the model recognition effect is poor.
[0011] (2) Technical solutions
[0012] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0013] An optimization method for address named entity recognition based on a deep learning model, comprising the following steps:
[0014] S1. Collect industry corpora in related fields and construct an industry entity dictionary;
[0015] S2. Collect online Chinese data, perform manual annotation according to the task objective and generate a template, perform data augmentation on the entity names in the template and the industry entity dictionary, and then perform data expansion;
[0016] S3. Utilize the unannotated industry corpus and entity dictionary to optimize the masking mechanism during the pre-training stage of the neural network language model;
[0017] S4. Fine-tune the neural network language model for the downstream recognition task, and select the neural network language model with the highest test accuracy as the output model;
[0018] S5. Collect online real-time data, save the entities whose prediction results of the output model are lower than the confidence threshold in the log file, and optimize the output model using the log file.
[0019] Preferably, in S1, collect the industry corpus in the relevant field and construct an industry entity dictionary, including:
[0020] S1. Integrate the existing publicly available entity dictionaries in the field to form a "public entity dictionary";
[0021] S2. Through a series of entity matching rules constructed by experts in this field based on experience, use string matching or pattern matching methods, combined with keyword vocabulary, proper nouns or structural rule entity features, perform expert experience matching on the collected public corpus, extract entities, and construct an "expert entity dictionary";
[0022] S3. Integrate the "public entity dictionary" and the "expert entity dictionary" to construct an "experience entity dictionary";
[0023] S4. Through an unsupervised method, statistically analyze the frequency of vocabulary occurrences, recall a large number of pending entities through word frequency, calculate their degrees of freedom and compactness, and screen out entities by setting thresholds to form an "unsupervised entity dictionary";
[0024] S5. Select a small amount of corpus to recall candidate words according to word frequency, screen the candidate words through frequency, integrity, information content and co-occurrence degree, and use the screened candidate words and the cross-words in the "experience entity dictionary" as the positive sample set during training;
[0025] S6. Use negative sampling to randomly sample other words to form a negative sample set, and use the positive sample set and the negative sample set to train the Bert model;
[0026] S7. Use the trained Bert model to score the quality of the entities recalled in all the corpus, and select the effective entities;
[0027] S8. Use the AutoNER model to predict the types of these words to form a "supervised entity dictionary".
[0028] S9. Integrate the "unsupervised entity dictionary" and the "supervised entity dictionary" to construct a "mined entity dictionary".
[0029] Preferably, in S2, data augmentation is performed on the entity names in the template and the industry entity dictionary, including:
[0030] Perform "nuclear fission" sampling on the entity names in the template and the industry entity dictionary: Select a small number of samples from the random samples, and according to the characteristics of these small samples, use the user dictionary to select relevant and similar content for replacement to generate new small samples, and so on for data augmentation.
[0031] Preferably, the data expansion in S2 includes:
[0032] Perform data expansion according to the industry corpus template, and the expansion methods used include the pseudo-label strategy, upsampling, slot replacement, and synonym replacement.
[0033] Preferably, in S3, use the unlabeled industry corpus and the entity dictionary to optimize the masking mechanism in the pre-training stage of the neural network language model, including:
[0034] Optimize the random masking strategy of the neural network language model by word to the target masking strategy by actual "entity".
[0035] Preferably, when there is no actual "entity", the target masking strategy degrades to the random masking strategy.
[0036] Preferably, in S4, perform model fine-tuning on the neural network language model for downstream recognition tasks, including:
[0037] Use the data after data augmentation and data expansion to perform downstream business training on the neural network language model, introduce the CRF layer of the neural network language model to capture the features of the transfer between label categories, and fine-tune the learning rates of the model and the CRF layer.
[0038] Preferably, in S5, save the entities with the output model prediction results lower than the confidence threshold in the log file, including:
[0039] Each word in the Chinese sequence to be recognized by the output model is processed by the SoftMax function to obtain a confidence in the label category. The confidence of each entity in the Chinese sequence to be recognized is the average of the sum of the confidences of each word. At the same time, a confidence threshold is set artificially according to the data distribution characteristics, and the entities with a confidence lower than the confidence threshold are saved in the log file.
[0040] (III) Beneficial effects
[0041] Compared with the prior art, a tuning method for address named entity recognition based on a deep learning model provided by the present invention incorporates prior knowledge of Chinese entities, makes full use of the characteristics of feature distribution data and the model being convenient for iterative optimization, etc. At the same time, during the process of fine-tuning the model, it does not rely on a large amount of labeled data, and it is a technical solution that is convenient for model tuning and iterative optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0043] Figure 1 is a flow diagram of the present invention;
[0044] Figure 2 is a flow diagram of constructing an industry entity dictionary in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0046] A tuning method for address named entity recognition based on a deep learning model, as Figure 1 shown, collects industry corpora in related fields and constructs an industry entity dictionary.
[0047] When performing named entity recognition in many professional fields, there will be some problems, such as: the diversity of entities. Due to the large number of synonyms, abbreviations, etc. in Chinese, the same entity often has many expressions (such as "Industrial and Commercial Bank of China" abbreviated as "ICBC"). Generally, drugs in the pharmaceutical industry alone have three types: trade names, generic names, and chemical names (such as "cold medicine", "Gan Kang", "compound paracetamol tablets", and "N-(4-hydroxyphenyl)acetamide molecule"), with strong uncertainty in meaning, and there are also cases in Chinese where different entity meanings represent one statement and polysemy (such as "apple", which can be "Apple Inc." or "fruit"). Therefore, we first collect industry corpora in the relevant field, specify rules and user dictionaries, and at the same time construct an industry entity dictionary.
[0048] The construction of the industry entity dictionary can be carried out through supervised, unsupervised, and remote supervision methods. The sources of industry corpus can be public datasets within the industry, encyclopedia entries, self-built information databases, user search logs, and unstructured user comments, etc.
[0049] As Figure 2 shown, first, integrate the existing publicly available entity dictionaries in the field to form a "public entity dictionary"; then, a series of rules for entity matching are constructed by experts in this field based on experience. Using string matching or pattern matching methods, combined with entity features such as keywords, proper nouns, or structural rules, perform expert experience matching on the collected public corpus, extract entities, and construct an "expert entity dictionary". Integrate the "public entity dictionary" and the "expert entity dictionary" to construct an "experience entity dictionary"; then, through an unsupervised method, count the frequency of word occurrences, recall a large number of pending entities based on word frequency, calculate their degrees of freedom and tightness, and screen out entities by setting thresholds to form an "unsupervised entity dictionary"; then, select a small amount of corpus to recall candidate words according to word frequency, screen the candidate words through frequency, integrity, information content, and co-occurrence degree, and use the cross-words between the screened candidate words and the "experience entity dictionary" as the positive sample set during training. Use negative sampling to randomly sample other words to form a negative sample set, and use the positive sample set and the negative sample set to train the Bert model.
[0050] Use the trained Bert model to score the quality of the entities recalled in all the corpus, and select the effective entities; finally, use the AutoNER model to predict the types of these words to form a "supervised entity dictionary", and integrate the "unsupervised entity dictionary" and the "supervised entity dictionary" to construct a "mined entity dictionary".
[0051] Collect online Chinese data, perform manual annotation according to the task objectives and generate templates, perform data augmentation on the entity names in the templates and the industry entity dictionary, and then perform data expansion.
[0052] ① Perform data augmentation on the entity names in the templates and the industry entity dictionary, including:
[0053] Perform "nuclear fission" sampling on the entity names in the templates and the industry entity dictionary: select a small number of samples from the random samples, and according to the characteristics of these small samples, use the user dictionary to select relevant and similar content for replacement, so as to generate a small number of new samples, and so on for data augmentation.
[0054] ② Data expansion includes:
[0055] Data augmentation is performed according to the industry corpus template, and the augmentation methods include pseudo-label strategy, upsampling, slot replacement, and synonym replacement.
[0056] Generally, a self-owned training set is used to fine-tune the pre-trained model for specific tasks. However, the data in the self-owned training set may be insufficient, such as problems like unbalanced label category distribution and less training data.
[0057] For the problem of unbalanced label category distribution, downsampling the original data can be selected, and high-quality data with appropriate proportions of each category is selected as the training set. However, this approach is only applicable to the case of extremely large data volume. Otherwise, the training set data after downsampling will become even less, and a good model fine-tuning effect cannot be obtained.
[0058] The process of "nuclear fission" is that thermal neutrons bombard uranium-235 atoms, and the atomic nucleus splits into 2 - 4 neutrons. The neutrons generated by the fission then bombard other uranium-235 atoms, and so on, forming a chain reaction. The idea of "nuclear fission" sampling comes from the nuclear fission process and is a sampling method aimed at finding the expected target bodies that meet the set characteristics from a sparse population.
[0059] In real life, there are often such scenarios, such as people who have participated in a certain meeting, people engaged in a certain professional direction, people of a certain ethnic minority, etc. Such groups are often of low probability, perhaps only one in ten thousand or even lower in a specific area. If the conventional sampling method is used to obtain a sample of such people, tens of thousands of people need to be screened, which is usually unrealistic and the cost is huge. The specific method of "nuclear fission" sampling is: first, some people are randomly selected from the crowd as the survey objects, a small number of samples needed are screened through these surveys, and then the investigation continues based on the clues provided by the samples, and so on, forming a chain reaction to obtain more small samples.
[0060] Using unlabeled industry corpus and entity dictionary, mask mechanism optimization is carried out in the pre-training stage of the neural network language model, specifically including:
[0061] The random mask strategy of the neural network language model based on "characters" is optimized into the target mask strategy based on actual "entities". At the same time, when there is no actual "entity", the target mask strategy degrades to the random mask strategy.
[0062] For deep learning language models, the problem that the "word" features are often not considered. During pre-training, a certain probability is often used to randomly use "characters" as the smallest unit as the mask label. This approach has the smallest element as "characters", but it splits the originally connected "words" into single "characters", so the internal information of Chinese text words is not fully utilized in the pre-training stage.
[0063] In the technical solution of this application, various types of entities are extracted from industry corpora according to rule templates, and the entity names are used as the objects of masked labels, which increases the difficulty of predicting masked words in the pre-training stage, makes full use of the information of "words", and thus can effectively improve the pre-training effect.
[0064] Fine-tune the neural network language model for downstream recognition tasks, and select the neural network language model with the highest test accuracy as the output model.
[0065] Among them, fine-tuning the neural network language model for downstream recognition tasks includes:
[0066] Use the data after data augmentation and data expansion to perform downstream business training on the neural network language model, introduce the CRF layer of the neural network language model to capture the features of the transfer between label categories, and fine-tune the learning rates of the model and the CRF layer.
[0067] Collect online real-time data, save the entities whose prediction results of the output model are lower than the confidence threshold in the log file, and use the log file to optimize the output model.
[0068] Among them, saving the entities whose prediction results of the output model are lower than the confidence threshold in the log file includes:
[0069] Each word in the Chinese sequence to be recognized by the output model is processed by the SoftMax function to obtain a confidence in the label category. The confidence of each entity in the Chinese sequence to be recognized is the average of the sum of the confidences of each word (averaged according to the entity word length). At the same time, a confidence threshold is set artificially according to the data distribution characteristics, and the entities with confidences lower than the confidence threshold are saved in the log file.
[0070] In the technical solution of this application, after the model is deployed online, a data feedback mechanism is introduced. A confidence threshold is set artificially according to the data distribution characteristics, and the entities with confidences lower than this confidence threshold are recorded and collected for model iteration, optimization and update, realizing the closed-loop processing of the entire data stream.
[0071] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A tuning method for address named entity recognition based on a deep learning model, characterized in that: It includes the following steps: S1. Collect industry corpora in related fields and construct an industry entity dictionary; S2. Collect online Chinese data, perform manual annotation according to the task objective and generate templates, perform data augmentation on the entity names in the templates and the industry entity dictionary, and then perform data expansion; S3. Utilize the unannotated industry corpora and entity dictionary to optimize the masking mechanism during the pre-training stage of the neural network language model; S4. Fine-tune the neural network language model for the downstream recognition task, and select the neural network language model with the highest test accuracy as the output model; S5. Collect online real-time data, save the entities whose prediction results of the output model are lower than the confidence threshold in the log file, and optimize the output model using the log file; In S1, collecting industry corpora in related fields and constructing an industry entity dictionary includes: S1. Integrate the existing publicly available entity dictionaries in the field to form a "public entity dictionary"; S2. A series of rules for entity matching are constructed by experts in this field according to experience. Using string matching or pattern matching methods, combined with keyword vocabulary, proper nouns or structural rule entity features, perform expert experience matching on the collected public corpora, extract entities, and construct an "expert entity dictionary"; S3. Integrate the "public entity dictionary" and the "expert entity dictionary" to construct an "experience entity dictionary"; S4. Through an unsupervised method, count the frequency of word occurrences, recall a large number of pending entities through word frequency, calculate their degrees of freedom and compactness, and screen out entities by setting thresholds to form an "unsupervised entity dictionary"; S5. Select a small amount of corpora to recall candidate words according to word frequency, screen the candidate words through frequency, integrity, information content and co-occurrence degree, and use the cross-words between the screened candidate words and the "experience entity dictionary" as the positive sample set during training; S6. Use negative sampling to randomly sample other words to form a negative sample set, and use the positive sample set and negative sample set to train the Bert model; S7. Use the trained Bert model to score the quality of the entities recalled in all corpora, and select the effective entities; S8. Perform type prediction on these words through the AutoNER model to form a "supervised entity dictionary"; S9. Integrate the "unsupervised entity dictionary" and the "supervised entity dictionary" to construct a "mined entity dictionary".
2. The tuning method for address named entity recognition based on a deep learning model according to claim 1, characterized in that: In S2, data augmentation of the entity names in the templates and the industry entity dictionary includes: Perform "nuclear fission" sampling on the entity names in the templates and the industry entity dictionary: Select a small number of samples from the random samples, and according to the characteristics of these small samples, use the user dictionary to select relevant and similar content for replacement, so as to generate a new small number of samples, and so on for data augmentation.
3. The tuning method for address named entity recognition based on a deep learning model according to claim 2, characterized in that: In S2, data expansion includes: Data augmentation is performed according to the industry corpus template, and the augmentation methods adopted include the pseudo-label strategy, upsampling, slot replacement, and synonym replacement.
4. The optimization method for address named entity recognition based on a deep learning model according to claim 1, characterized in that: In S3, the unlabeled industry corpus and entity dictionary are used to optimize the masking mechanism in the pre-training stage of the neural network language model, including: The random masking strategy of the neural network language model by "character" is optimized into the target masking strategy by actual "entity".
5. The optimization method for address named entity recognition based on a deep learning model according to claim 4, characterized in that: When no actual "entity" appears, the target masking strategy degrades to the random masking strategy.
6. The optimization method for address named entity recognition based on a deep learning model according to claim 1, characterized in that: In S4, the neural network language model is fine-tuned for the downstream recognition task, including: Using the data after data enhancement and data augmentation to perform downstream business training on the neural network language model, introducing the CRF layer of the neural network language model to capture the features of the transfer between label categories, and fine-tuning the learning rates of the model and the CRF layer.
7. The optimization method for address named entity recognition based on a deep learning model according to claim 1, characterized in that: In S5, the entities with the output model prediction results lower than the confidence threshold are saved in the log file, including: The output model processes each character in the Chinese sequence to be recognized through the SoftMax function to obtain a confidence about the label category. The confidence of each entity in the Chinese sequence to be recognized is the average of the sum of the confidences of each character. At the same time, the confidence threshold is set artificially according to the data distribution characteristics, and the entities with the confidence lower than the confidence threshold are saved in the log file.
Citation Information
Patent Citations
Dialogue generation method and device
CN106951468A
Security event entity recognition method based on pre-training model
CN113312914A