A training method for named entity recognition model

By automatically extracting and filtering key syntactic components and fusing their word embeddings, the training method of named entity recognition model is improved, solving the problems of insufficient generalization capabilities and high manual labeling costs in existing models, and improving recognition capabilities and interpretability.

CN114881031BActive Publication Date: 2025-05-23NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210428560.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-22
Publication Date
2025-05-23
Estimated Expiration
2042-04-22

AI Technical Summary

Technical Problem

When the existing named entity recognition model deals with entities not included in the annotation data, it lacks generalization ability and low recall rate. In addition, the existing technology has problems such as incomplete entity dictionary coverage, unclear distinction between the importance of additional information and high cost of manual labeling.

Method used

The component analysis tree is constructed through a pre-trained component syntax analyzer, and the key syntax components are automatically extracted. The difference in entity recognition probability is calculated using masking and replacement schemes, the most important key syntax components are selected, and the word embedding of entities and key syntax components is fused through a gating mechanism to improve the training method of named entity recognition models.

Benefits of technology

It improves the recognition ability of the model in actual use scenarios, enhances the connection between key syntactic components and entity types, simplifies the human reading and understanding process, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881031B_ABST
    Figure CN114881031B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method for a named entity recognition model, which uses a pre-trained component syntactic analyzer to construct a component analysis tree of an input text; based on a generation rule, a key syntactic component candidate set is formed through the component analysis tree; by masking different key syntactic components, the two most important key syntactic components in the key syntactic component candidate set are screened out; the entity and the two most important key syntactic components are respectively masked to obtain two word embeddings and a gating mechanism is introduced to fuse the two word embeddings to form a final word embedding representation of each word; the final word embedding representation of each word in the text is used as input, input into a conditional random field for training, and a named entity recognition model is obtained. The present invention enhances the expressive power of the final word embedding; saves the labor cost required for labeling sample data; effectively reduces the influence of the complex semantics of the entire sentence, simplifies the process of human reading and understanding, and has strong interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a training method for a named entity recognition model. Background Art

[0002] Named entity recognition is an important basic work in the field of natural language processing. Its task is to extract specific types of text segments with practical significance in natural language text, namely entities, and classify the entities into one of the pre-defined categories. Common entity types include names, places, organizations, etc. Named entity recognition is a leading work for many natural language understanding tasks, and its effect affects the performance of many downstream tasks.

[0003] With the development of deep learning, the named entity recognition model using neural network structure, especially pre-trained language model, has achieved great improvement in effect. The current mainstream solution requires a large amount of text data with entity categories manually labeled for training. However, the text processed in real application scenarios may contain entities that are not included in the labeled data. Therefore, the model needs to have good generalization ability, be able to integrate the semantic information of the context and the feature information of the vocabulary itself, and have good recognition ability for entities not included in the labeled data.

[0004] The existing technology attempts to solve this problem mainly from the following aspects:

[0005] Reference 1: Towards Improving Neural Named Entity Recognition with Gazetteers (Improving Neural Named Entity Recognition with Gazetteers).

[0006] This paper is based on a large entity dictionary obtained from outside. Through the designed matching rules, a large amount of unlabeled corpus is annotated with entities, and labeled data containing more entities is generated, and the number of samples and entities in the training data are expanded for training improvement.

[0007] Reference 2: Improving Named Entity Recognition with Attentive Ensemble of Syntactic Information (Improving Named Entity Recognition by Integrating Syntactic Information Using Attention Mechanism).

[0008] This paper proposes to make up for the lack of labeled data by encoding the linguistic information of the text into the model. The information that can be used for encoding usually includes part-of-speech information, syntactic components, and dependency relationships. The word embedding of this information is spliced ​​together with the original context word embedding as the final encoding of each word to improve the model's capabilities.

[0009] Reference 3: TriggerNER: Learning with Entity Triggers as Explanations for Named Entity Recognition (TriggerNER: Learning with Entity Triggers as Explanations for Named Entity Recognition).

[0010] The paper proposes the concept of "entity triggers", which are a set of words in a sentence that can help identify and classify entities. Based on a large number of manually annotated entity triggers, the paper strengthens the semantic connection between entity triggers and entities, and improves the model's attention to key words in the context to help identify entities.

[0011] For entities that have not been included in the model training and learning, the existing technical solutions can improve their recognition capabilities to a certain extent, but they all have objective shortcomings. Document 1 uses an external entity dictionary to annotate a large amount of unlabeled corpus, and uses the annotated corpus as a new training sample for learning. However, new entities are constantly emerging, the entity dictionary cannot cover all entities, and the quality of the new training corpus obtained by annotation is not high. The final model is more inclined to recognize entities with similar forms to the entities in the dictionary, resulting in a low recall rate of the model for new entities. Document 2 adds linguistic information such as syntax to the model to make up for the shortcomings of the labeled data, which can better mine the potential information of the existing labeled data, but this method simply splices this part of the additional information with the original context text representation, and cannot distinguish the importance of different types of syntactic information. Reference 3 proposes a new concept of "entity triggers". Although this type of additional supervised learning signal provided by manual annotation is of high quality, it is costly and time-consuming, which greatly limits the application field and scale of "entity triggers". Moreover, manually annotated "entity triggers" are generally non-continuous text segments, which are not exactly the same as the way the model serializes and reads text input for encoding. Therefore, there are deficiencies in the encoding modeling of "entity triggers". Summary of the invention

[0012] In order to overcome a series of shortcomings existing in the above-mentioned background technology, the purpose of the present invention is to provide a training method for a named entity recognition model, automatically extract key syntactic components, enhance their semantic connection with entity categories, and improve word embedding encoding to improve the named entity recognition effect. When automatically extracting key syntactic components, for the candidate set of key syntactic components generated based on the component analysis tree, a masking and replacement scheme is used to calculate the difference between the entity recognition probabilities before and after, so as to select the two that have the greatest impact on entity recognition as key syntactic components. Two word embeddings are obtained by masking the entity and the key syntactic components respectively, and a gating mechanism is introduced to fuse the two word embeddings, thereby strengthening the connection between the key syntactic components and the entity type, thereby improving the recognition ability of the model in actual usage scenarios.

[0013] In order to achieve the above object, a first aspect of the present invention provides a method for training a named entity recognition model, comprising the following steps:

[0014] Step 101: Use the pre-trained component syntactic analyzer to construct a component analysis tree of the input text;

[0015] Step 102: Based on the generation rule, a key syntactic component candidate set is formed through the component analysis tree;

[0016] Step 103: screening out the two most important key syntactic components in the key syntactic component candidate set by masking different key syntactic components;

[0017] Step 104: Mask the entity and the two most important key syntactic components respectively to obtain two word embeddings and introduce a gating mechanism to fuse the two word embeddings to form the final word embedding representation of each word;

[0018] Step 105: The final word embedding representation of each word in the text is used as input and input into the conditional random field for training to obtain a named entity recognition model.

[0019] In some possible implementations, the “forming a candidate set of key syntactic components through the component analysis tree based on the generation rule” in step 102 is specifically as follows:

[0020] In a top-down, left-to-right order, the phrase node closest to the leaf node in the component analysis tree is taken as a candidate key syntactic component and added to the key syntactic component candidate set.

[0021] In some possible implementations, for the phrase node determined to be a prepositional phrase or a verb phrase, if the subnodes of the prepositional phrase or the verb phrase include a node determined to be a noun phrase, it is necessary to separate the verb or preposition from the noun phrase, and treat the verb and the preposition as a phrase alone and add them to the candidate set of key syntactic components;

[0022] In some possible implementations, the entity itself is not a candidate key syntactic component. If a phrase contains an entity, the entity needs to be removed, and the phrase formed excluding the entity is added to the key syntactic component candidate set as a new candidate.

[0023] In some possible implementations, the step 103 of "screening out the two most important key syntactic components in the key syntactic component candidate set by masking different key syntactic components" specifically includes the following steps:

[0024] Step 201: Using the encoder module and the classifier module to learn the initial training data;

[0025] Step 202: Using the encoder module and the classifier module trained in step 201, calculate the probability S that the entity is correctly identified 1 ;

[0026] Step 203: Mask the candidate key syntactic component in the key syntactic component candidate set, and calculate the probability S of the entity being correctly identified after the candidate key syntactic component is masked. 2 ;

[0027] Step 204: Calculate the difference in the probability of correct recognition of each candidate key syntactic component before and after masking, and select the two most important key syntactic components.

[0028] In some possible implementations, the step 202 of "using the encoder module and the classifier module trained in step 201 to calculate the probability S that the entity is correctly identified" 1 The details are as follows: the encoder module is used to obtain the word embedding of each word in the text, the classifier module is used to map the word embedding representation to the entity category representation space, and the predicted probabilities of the correct categories corresponding to each word constituting the entity are summed and averaged as the probability S of the entity being correctly identified. 1 .

[0029] In some possible implementations, the step 203 of "masking the candidate key syntactic component in the key syntactic component candidate set, and calculating the probability S of the entity being correctly identified after the candidate key syntactic component is masked" is as follows: 2The specific steps are as follows: for each candidate set of key syntactic components generated by each training text, mask them in the text, that is, replace each word of each candidate key syntactic component in the text with [mask] to form new text data corresponding to each candidate key syntactic component; for each new text data of a candidate key syntactic component in the original sentence that is masked, use the encoder module and the classifier module that are the same as step 202 to perform named entity recognition on the new text data, and add and average the predicted probabilities of the correct categories corresponding to each word constituting the entity, as the probability S of the entity being correctly identified and classified after masking. 2 .

[0030] In some possible implementations, the step 204 of "calculating the difference in the probability of each candidate key syntactic component being correctly recognized before and after masking, and selecting the two most important key syntactic components" is as follows: The difference in the recognition probability of each candidate syntactic component before and after masking is S 1 -S 2 , the difference in recognition probability is defined as the importance score of the candidate syntactic component for named entity recognition of the sentence; for each element in the candidate set of key syntactic components, its importance score for named entity recognition of the sentence is calculated, and all the importance scores obtained are sorted, and the two candidate key syntactic components with the highest importance scores are selected as the final two key syntactic components.

[0031] In some possible implementations, the step 104 of "masking the entity and the two most important key syntactic components respectively, obtaining two word embeddings and introducing a gating mechanism to fuse the two word embeddings to form the final word embedding representation of each word" specifically includes the following steps:

[0032] Step 301: Mask the entities in the text data, that is, replace the words that make up the entities in the text with [mask], and obtain the word embedding representation h of each word in the text through the encoder module described in step 202 1 ;

[0033] Step 302: Mask the two key syntactic components in the text data, that is, replace the words that make up the two key syntactic components in the text with [mask], and obtain the word embedding representation h of each word in the text through the encoder module described in step 202 2 ;

[0034] Step 303: By introducing a gating mechanism, h 1 and h 2 Perform fusion calculation to obtain the final word embedding representation h of each word in the text data; the calculation formula of the gating mechanism is: g=σ(W 1 ·h 1 +W 2 ·h 2 +b), where σ represents the sigmoid function, W 1 and W 2 represents the trainable parameter matrix, and b represents the bias term; represents a term-by-term multiplication operation, 1 represents a high-dimensional vector whose elements are all 1 and whose dimension is the same as h 1 and h 2 same.

[0035] In some possible implementations, in step 105, “taking the word embedding representation of each word in the text as input, inputting it into the conditional random field for training, and obtaining a named entity recognition model” specifically includes the following steps:

[0036] Step 401: The word embedding representation of each word in the input text is used, and the conditional random field performs forward calculation to obtain the score of each word in the text being predicted as each entity category;

[0037] Step 402: Calculate the scores of the correctly labeled paths and the scores of all possible labeled paths, and define the prediction loss function of the model based on maximum likelihood estimation. The prediction loss function formula is: Among them, P real is the correct labeling sequence, P i are all possible annotation sequences, M is the number of words in the text, and n is the number of predefined entity categories; Step 403, calculate the gradient of the loss function and back propagate, update the parameters in the encoder module and the conditional random field described in step 202, and train the model;

[0038] Step 404: Use the validation set to evaluate the performance of the model. The basis for the evaluation is to calculate the F1 score on the validation set. The calculation formula is: Where P represents precision and R represents recall;

[0039] Step 405: Record the model that achieves the best results on the validation set during the training process and use it as the final trained named entity recognition model.

[0040] A second aspect of the present invention provides an electronic device, the electronic device comprising a processor and a memory:

[0041] The memory is used to store program code and transmit the program code to the processor;

[0042] The processor is used to execute the above-mentioned training method of a named entity recognition model according to the instructions in the program code.

[0043] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the above-mentioned training method for a named entity recognition model.

[0044] The beneficial effects of the present invention include two aspects: technical level and application level:

[0045] From a technical perspective:

[0046] 1. The present invention uses a pre-trained component syntactic analyzer to process the text to obtain a component analysis tree, and generates a set of key syntactic component candidates according to the rules. Through masking and replacement schemes, the predicted probability difference between the previous and next target entities is calculated, thereby selecting the two key syntactic components that have the greatest impact on the named entity recognition effect;

[0047] 2. By masking the entity and key syntactic components respectively, two word embeddings are obtained and a gating mechanism is introduced to fuse the two word embeddings, thus establishing the association between the key syntactic components and the entity type. At the same time, the semantic information of the entity word itself is taken into account, thus enhancing the expressive power of the final word embedding.

[0048] From the application level:

[0049] 1. The present invention can automatically extract key syntactic components in the text, and use them to guide and help identify named entities, which can save the labor cost required to annotate sample data;

[0050] 2. Focus on the most critical grammatical components, effectively reduce the impact of the complex semantics of the entire sentence, simplify the process of human reading and understanding, and have strong interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flowchart of the overall generation steps of the named entity recognition model in the embodiment of the present application;

[0052] Figure 2 A visual representation of a component analysis tree corresponding to an example sentence in an embodiment of the present application;

[0053] Figure 3 This is a flowchart of the key syntactic component screening steps in the embodiments of this application;

[0054] Figure 4 This is a flowchart of the steps of enhancing the word embedding representation in the implementation of this application;

[0055] Figure 5 This is a flow chart of the model parameter training steps in the embodiment of this application;

[0056] Figure 6This is a structural diagram of a training device for a named entity recognition model in an embodiment of the present application.

[0057] In the figure: 50, electronic device; 51, processor; 52, memory. DETAILED DESCRIPTION

[0058] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0059] Taking the sentence "We had a fantastic lunch at Rumble Fish yesterday, where the food is my favorite." as an example, the entity that the model needs to recognize is "Rumble Fish", and its corresponding entity type is place.

[0060] This embodiment provides a method for training a named entity recognition model. Figure 1 As shown, the following steps are included:

[0061] Step 101: Use the pre-trained component syntactic analyzer to construct a component analysis tree of the input text. The component syntactic analyzer is used to analyze the components of the sentence, obtain the syntactic structure of the entire sentence, and complete the syntactic analysis for the purpose of obtaining the phrase structure. The result of the analysis is to convert the input unstructured text into a component analysis tree consisting of terminal and non-terminal nodes, wherein the leaf nodes correspond to each word in the text, and the intermediate nodes correspond to phrase nodes consisting of several words. The component syntactic analyzer used in the embodiment of the present application can be trained based on the existing Penn TreeBank (PTB) data set of the University of Pennsylvania using a neural network method, or directly use the StanfordCoreNLP toolkit, which is free and open source, provided by the Stanford University Natural Language Processing Research Group.

[0062] Before using the component parser, each word (including each punctuation mark) in the sentence to be processed needs to be separated by a space to form the input text. Then it is input into the component parser, which converts the sentence into a node sequence stored in a tree data structure, that is, it represents the corresponding component analysis tree. Each word in the sentence is a leaf node and is marked as the corresponding syntactic component, while the phrase composed of several words is an intermediate node and is marked as the corresponding node type. Finally, the component parser will output an instance of a tree type, which has an attribute representing the root node of the corresponding component analysis tree. Starting from the root node, you can traverse to each node in the tree, where the leaf node represents each word in the sentence and the non-leaf node represents a phrase composed of several words.

[0063] Take the sentence "We had a fantastic lunch at Rumble Fish yesterday, where the food is my favorite." as an example. The result after the space segmentation is "We had a fantastic lunch at Rumble Fish yesterday, where the food is my favorite." is sent as input to the component parser. The pre-trained component parser (if you use the StanfordCoreNLP toolkit directly, select the type that needs to be labeled as "constituency parse") will output a tree type instance. Starting from its root node, you can traverse to access each node in the tree. The visual representation of the component analysis tree corresponding to the above example sentence is as follows: Figure 2 shown.

[0064] Step 102: Based on the generation rules, a set of key syntactic component candidates is formed through the component analysis tree; specifically, in a top-down and left-to-right order, phrase nodes closest to leaf nodes in the component analysis tree are taken as candidate key syntactic components and added to the set of key syntactic component candidates.

[0065] For the phrase nodes that are judged to be prepositional phrases and verb phrases, if the child nodes of the prepositional phrases and verb phrases contain nodes that are judged to be noun phrases, it is necessary to separate the verb or preposition from the noun phrase, and treat the verb and preposition as a separate phrase and add them to the candidate set of key syntactic components.

[0066] The entity itself is not a candidate for a key syntactic component. If a phrase contains an entity, the entity needs to be removed, and the phrase formed except the entity is added to the key syntactic component candidate set as a new candidate.

[0067] Based on the above generation rules and the above component analysis tree, a set of key syntactic component candidates is formed. In this example, the elements in the set include {We, had, a fantastic lunch, at, yesterday, where, the food, is, my favorite}.

[0068] Step 103: By masking different key syntactic components, the two most important key syntactic components in the key syntactic component candidate set are screened out, see Figure 3 As shown, the specific screening process includes the following steps:

[0069] Step 201: Use a BERT-based encoder module and a classifier module implemented by a fully connected layer using a sigmoid activation function to learn the initial training data. BERT is a large-scale pre-trained language model proposed by Google in 2018. It can take the entire sentence as input and generate a corresponding high-dimensional word embedding for each word in the sentence based on the attention mechanism; the fully connected layer using the sigmoid activation function takes the above word embedding as input, first maps the word embedding corresponding to each word to a predefined entity category representation space through a linear transformation matrix, generates a vector with a dimension equal to the number of entity categories, and then inputs it into the sigmoid activation function. Each dimension of the new vector representation output by the function represents the probability that the word is classified as the entity type. The purpose of using the sigmoid function here is to ensure that the probability of each word being classified into each entity category is between [0,1].

[0070] Step 202: Using the encoder module and classifier module in step 201, calculate the probability S that the entity is correctly identified 1 , as follows: Use the BERT-based encoder module to obtain the word embedding of each word in the text, use the classifier module implemented by the fully connected layer with the sigmoid activation function to map the word embedding representation to the entity category representation space, and sum and average the probabilities of the correct categories corresponding to each word that constitutes the entity as the probability S of the entity being correctly identified. 1 . Applied to this example, the BERT-based encoder module generates corresponding word embeddings for each word in the sentence "We had a fantastic lunch at Rumble Fish yesterday, where the food is my favorite." The classifier module uses the above word embeddings as input, maps the word embeddings of each word to the entity category space, and obtains the predicted probability of each category. For the entity "Rumble Fish" to be identified in this example, the probability of its correct identification is defined as the arithmetic mean of the probabilities that the two words "Rumble" and "Fish" are correctly classified. Step 203, mask the candidate elements in the candidate set of key syntactic components, and calculate the probability S that the entity is correctly identified after the candidate key syntactic components are masked. 2, specifically as follows: for each candidate set of key syntactic components generated by the training text, mask them in the text respectively, that is, replace each word constituting the candidate key syntactic component in the text with [mask] to form new text data corresponding to each candidate key syntactic component; for each new text data of a candidate key syntactic component in the original sentence that is masked, use the encoder module and the classifier module that are the same as step 201 to perform named entity recognition on the new text data, and add and average the predicted probabilities of the correct categories corresponding to each word constituting the entity, as the probability S of the entity being correctly identified and classified after masking. 2 In this example, we need to mask each element in the candidate set in turn. Taking "a fantastic lunch" as an example, after masking the candidate element, we can generate the sentence "We had [mask][mask][mask] at Rumble Fish yesterday, where the food is my favorite.".

[0071] The BERT-based encoder module takes the masked sentence as input and generates the corresponding word embedding for each word in it. The classifier module takes the above word embedding as input, maps the word embedding of each word to the entity category space, and obtains the predicted probability of each category. For the entity "Rumble Fish" to be identified in this example, the probability of its correct identification is also defined as the arithmetic mean of the probabilities of the two words "Rumble" and "Fish" being correctly classified.

[0072] Step 204: Calculate the difference in the probability of each candidate key syntactic component being correctly recognized before and after masking, and select the two most important key syntactic components, as follows: The difference in the recognition probability of each candidate syntactic component before and after masking is S 1 -S 2 , the difference in recognition probability is defined as the importance score of the candidate syntactic component for named entity recognition of the sentence; for each element in the candidate set of key syntactic components, its importance score for named entity recognition of the sentence is calculated, and all the importance scores obtained are sorted, and the two candidate key syntactic components with the highest importance scores are selected as the final two key syntactic components. Applied to this example, the difference S between the probability of the entity being correctly recognized before and after masking each candidate component is calculated 1 -S 2, sort them as the importance scores of the corresponding candidate components, and select the two with the highest importance scores as the key syntactic components. In this example, after the above operation, it can be obtained that the elements "afantastic lunch" and "the food" are the key syntactic components that affect the recognition of "Rumble Fish".

[0073] Step 104: Mask the entity and the two most important key syntactic components respectively to obtain two word embeddings and introduce a gating mechanism to fuse the two word embeddings to form the final word embedding representation of each word, see Figure 4 As shown, the specific process of obtaining and calculating word embedding representation includes the following steps: Applied to this example, the entity "Rumble Fish" and the key syntactic components "a fantastic lunch" and "the food" are masked respectively to obtain the word embeddings in the two cases, and a gating mechanism is used to generate the final word embedding representation that combines the features of the two.

[0074] Step 301: Mask the entities in the text data, that is, replace the words that make up the entities in the text with [mask], and obtain the word embedding representation h of each word in the text through the encoder module described in step 202 1 Applied to this example, masking "Rumble Fish" yields the sentence "We had a fantastic lunch at [mask] [mask] yesterday, where the food is my favorite.", and obtaining the word embedding h for each word in this case. 1 .

[0075] Step 302: Mask the two key syntactic components in the text data, that is, replace the words that make up the two key syntactic components in the text with [mask], and obtain the word embedding representation h of each word in the text through the encoder module described in step 202 2 Applied to this example, masking "a fantastic lunch" and "the food" gives the sentence "We had [mask][mask][mask] at Rumble Fish yesterday, where [mask][mask] is my favorite." and getting the word embedding h for each word in this case. 2 .

[0076] Step 303: By introducing a gating mechanism, h 1 and h 2 Perform fusion calculation to obtain the final word embedding representation h of each word in the text data; the calculation formula of the gating mechanism is: g=σ(W 1 ·h 1 +W 2 ·h 2 +b), where σ represents the sigmoid function, W 1 and W 2 represents the trainable parameter matrix, and b represents the bias term; represents a term-by-term multiplication operation, 1 represents a high-dimensional vector whose elements are all 1 and whose dimension is the same as h 1 and h 2 The reason for this is that some entity words have distinct feature information and can be correctly identified without relying on context. By introducing a gating mechanism to fuse the two word embedding representations, it is possible to take into account both the information of the entity words themselves and the contextual information provided by the key syntactic components.

[0077] Step 105: The final word embedding representation of each word in the text is used as input and input into the conditional random field for training to obtain a named entity recognition model, that is, the obtained word embedding representation h is input into the conditional random field (CRF) for decoding to generate a predicted label sequence. The model parameters are updated by calculating the loss function between the predicted label sequence and the true label for back propagation, and the model with the best performance during the training process is recorded.

[0078] See also Figure 5 As shown, the specific training process includes the following steps:

[0079] Step 401: The word embedding representation of each word in the input text is used, and the conditional random field performs forward calculation to obtain the score of each word in the text being predicted as each entity category;

[0080] Step 402: Calculate the scores of the correctly labeled paths and the scores of all possible labeled paths, and define the prediction loss function of the model based on maximum likelihood estimation. The prediction loss function formula is: Among them, P real is the correct tag sequence, P i are all possible tag sequences, M is the number of words in the text, and n is the number of predefined entity categories;

[0081] Step 403: Calculate the gradient of the loss function and back-propagate, update the parameters in the encoder module and conditional random field described in step 202, and train the model;

[0082] Step 404: Use the validation set to evaluate the performance of the model. The basis for the evaluation is to calculate the F1 score on the validation set. The calculation formula is: Where P represents precision and R represents recall;

[0083] Step 405: Record the model that achieves the best results on the validation set during the training process and use it as the final trained named entity recognition model.

[0084] See also Figure 6 As shown, the present invention provides an electronic device 50 for training a named entity recognition model, such as Figure 5 As shown, the electronic device 50 includes a processor 51 and a memory 52 coupled to the processor 51 .

[0085] The memory 52 is used to store program codes and transmit the program codes to the processor 51;

[0086] The processor 51 is used to execute the above-mentioned training method of a named entity recognition model according to the instructions in the program code.

[0087] The memory 52 stores program instructions for implementing the training method of the named entity recognition model of the above embodiment or the named entity recognition method of the above embodiment.

[0088] The processor 51 is used to execute program instructions stored in the memory 52 to train the named entity recognition model.

[0089] The processor 51 may also be referred to as a CPU (Central Processing Unit). The processor 51 may be an integrated circuit chip having the ability to process signals. The processor 51 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0090] This embodiment also provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the above-mentioned training method for a named entity recognition model.

[0091] The storage medium stores program instructions that can implement all the above methods, wherein the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or a terminal device such as a computer, a server, a mobile phone, or a tablet.

[0092] The above description is only an implementation mode of the present invention. It should be pointed out that, for ordinary technicians in this field, improvements can be made without departing from the creative concept of the present invention, but these all belong to the protection scope of the present invention.

Claims

1. A training method for a named entity recognition model, It is characterized in that The following steps are involved: Step 101: Use the pre-trained component syntactic analyzer to construct a component analysis tree of the input text; Step 102: Based on the generation rule, a key syntactic component candidate set is formed through the component analysis tree; Step 103: screening out the two most important key syntactic components in the key syntactic component candidate set by masking different key syntactic components; Step 104: Mask the entity and the two most important key syntactic components respectively to obtain two word embeddings and introduce a gating mechanism to fuse the two word embeddings to form the final word embedding representation of each word; Step 105: taking the final word embedding representation of each word in the text as input, inputting it into the conditional random field for training, and obtaining a named entity recognition model; The step 104 of "masking the entity and the two most important key syntactic components respectively, obtaining two word embeddings and introducing a gating mechanism to fuse the two word embeddings to form the final word embedding representation of each word" specifically includes the following steps: Step 301: Mask the entities in the text data, that is, replace the words that make up the entities in the text with [mask], and obtain the word embedding representation h of each word in the text through the encoder module described in step 202 1 ; Step 302: Mask the two key syntactic components in the text data, that is, replace the words that make up the two key syntactic components in the text with [mask], and obtain the word embedding representation h of each word in the text through the encoder module described in step 202 2 ; Step 303: By introducing a gating mechanism, h 1 and h 2 The final word embedding representation h of each word in the text data is obtained by fusion. The calculation formula of the gating mechanism is: g=σ(W 1 ·h 1 +W 2 ·h 2 +b), where σ represents the sigmoid function, W 1 and W 2 represents the trainable parameter matrix, and b represents the bias term; represents a term-by-term multiplication operation, 1 represents a high-dimensional vector whose elements are all 1 and whose dimension is the same as h 1 and h 2 same; In step 105, "taking the word embedding representation of each word in the text as input, inputting it into the conditional random field for training, and obtaining a named entity recognition model" specifically includes the following steps: Step 401: The word embedding representation of each word in the input text is used, and the conditional random field performs forward calculation to obtain the score of each word in the text being predicted as each entity category; Step 402: Calculate the scores of the correctly labeled paths and the scores of all possible labeled paths, and define the prediction loss function of the model based on maximum likelihood estimation. The prediction loss function formula is: Among them, P real is the correct labeling sequence, P i are all possible annotation sequences, M is the number of words in the text, and n is the number of predefined entity categories; Step 403: Calculate the gradient of the loss function and back-propagate, update the parameters in the encoder module and conditional random field described in step 202, and train the model; Step 404: Use the validation set to evaluate the performance of the model. The basis for the evaluation is to calculate the F1 score on the validation set. The calculation formula is: Where P represents precision and R represents recall; Step 405: Record the model that achieves the best results on the validation set during the training process and use it as the final trained named entity recognition model.

2. A method for training a named entity recognition model according to claim 1, It is characterized in that The "forming a key syntactic component candidate set through the component analysis tree based on the generation rule" in step 102 is specifically as follows: In a top-down, left-to-right order, the phrase node closest to the leaf node in the component analysis tree is taken as a candidate key syntactic component and added to the key syntactic component candidate set.

3. A method for training a named entity recognition model according to claim 2, It is characterized in that For the phrase nodes that are judged to be prepositional phrases and verb phrases, if the child nodes of the prepositional phrases and verb phrases contain nodes that are judged to be noun phrases, it is necessary to separate the verb or preposition from the noun phrase, and treat the verb and preposition as a separate phrase and add them to the candidate set of key syntactic components.

4. A method for training a named entity recognition model according to claim 1 or 2, It is characterized in that The entity itself is not a candidate for a key syntactic component. If a phrase contains an entity, the entity needs to be removed, and the phrase formed except the entity is added to the key syntactic component candidate set as a new candidate.

5. The method for training a named entity recognition model according to claim 1, It is characterized in that The step 103 of "screening out the two most important key syntactic components in the key syntactic component candidate set by masking different key syntactic components" specifically includes the following steps: Step 201: Using the encoder module and the classifier module to learn the initial training data; Step 202: Using the encoder module and the classifier module trained in step 201, calculate the probability S that the entity is correctly identified 1 ; Step 203: Mask the candidate key syntactic component in the key syntactic component candidate set, and calculate the probability S of the entity being correctly identified after the candidate key syntactic component is masked. 2 ; Step 204: Calculate the difference in the probability of correct recognition of each candidate key syntactic component before and after masking, and select the two most important key syntactic components.

6. A method for training a named entity recognition model according to claim 5, It is characterized in that In step 202, the probability S of an entity being correctly identified is calculated by using the encoder module and the classifier module trained in step 201. 1 The details are as follows: the encoder module is used to obtain the word embedding of each word in the text, the classifier module is used to map the word embedding representation to the entity category representation space, and the predicted probabilities of the correct categories corresponding to each word constituting the entity are summed and averaged as the probability S of the entity being correctly identified. 1 .

7. A method for training a named entity recognition model according to claim 5, It is characterized in that In step 203, the key syntactic component candidate in the key syntactic component candidate set is masked, and the probability S of the entity being correctly identified after the key syntactic component candidate is masked is calculated. 2 The specific steps are as follows: for each candidate set of key syntactic components generated by each training text, mask them in the text, that is, replace each word of each candidate key syntactic component in the text with [mask] to form new text data corresponding to each candidate key syntactic component; for each new text data of a candidate key syntactic component in the original sentence that is masked, use the encoder module and the classifier module that are the same as step 202 to perform named entity recognition on the new text data, and add and average the predicted probabilities of the correct categories corresponding to each word constituting the entity, as the probability S of the entity being correctly identified and classified after masking. 2 .

8. The method for training a named entity recognition model according to claim 5, It is characterized in that The step 204 of "calculating the difference in the probability of each candidate key syntactic component being correctly recognized before and after masking, and selecting the two most important key syntactic components" is as follows: The difference in the recognition probability of each candidate syntactic component before and after masking is S 1 -S 2 , the difference in recognition probability is defined as the importance score of the candidate syntactic component for named entity recognition of the sentence; for each element in the candidate set of key syntactic components, its importance score for named entity recognition of the sentence is calculated, and all the importance scores obtained are sorted, and the two candidate key syntactic components with the highest importance scores are selected as the final two key syntactic components.

Citation Information

Patent Citations

  • Named entity detection method and device, electronic equipment and readable storage medium

    CN110399616A

  • Industrial text matching model method and device based on deep learning

    CN114282592A