Named entity recognition method, device and storage medium
By entering text sequences into the Transformer model and comparing similarity with the mean vector, the recognition problem of named entity recognition in the zero sample situation is solved, and the recognition accuracy and model robustness are improved.
Patent Information
- Application Number
- CN202211470178.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-11-23
AI Technical Summary
Existing named entity recognition technology is difficult to identify specific categories in the case of zero samples, and the training set labeling errors and lack of Chinese samples.
A named entity recognition method is proposed. By inputting text sequences into the Transformer model, a word vector without entity type is obtained, and similarity is compared with the mean vector in the registration set to determine the entity type. This method uses mean vectors to overcome the zero-sample problem and improves the robustness of the model through the random discarding method.
It improves the accuracy of named entity recognition, can effectively identify words that cannot be recognized by the model, and improves the robustness of the model, especially in Chinese named entity recognition tasks.
Smart Images

Figure CN115935992B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of natural language processing (NLP), and more specifically, to a method, device and storage medium for named entity recognition. Background Art
[0002] Named Entity Recognition (NER), also known as proper name recognition, refers to the recognition of entities with specific meanings in texts, mainly including names of people, places, organizations, proper nouns, etc. Named entities generally refer to entities with specific meanings or strong referentiality in texts. Named entity recognition is a basic task in NLP and an important basic tool for many NLP tasks such as information extraction, question-answering systems, syntactic analysis, and machine translation.
[0003] In the prior art, it is usually assumed that the training set has relatively complete category information. However, in the actual named entity recognition task, there may be a situation where there is no sample (i.e., zero sample) for a certain category, making it difficult to recognize the category. Moreover, the training set is basically labeled by setting rules. During the labeling process, labeling errors may occur due to mistakes or rule failure, which will interfere with named entity recognition. In addition, in the current research on Chinese named entity recognition, there is a general lack of Chinese samples. Summary of the invention
[0004] The embodiments of the present invention provide a method, device and storage medium for named entity recognition.
[0005] The technical solution of the embodiment of the present invention is as follows:
[0006] A named entity recognition method, comprising:
[0007] Inputting the text sequence into a named entity recognition model to obtain a word vector recognized as a non-entity type by the named entity recognition model;
[0008] Determining a similarity between the word vector without the entity type and a mean vector in the registration set, the mean vector corresponding to a specific entity type, the mean vector being determined by a mean operation of a plurality of word vectors that conform to the specific entity type;
[0009] The specific entity type corresponding to the mean vector whose similarity is greater than a predetermined threshold value is determined as the entity type of the word vector without entity type.
[0010] In an exemplary embodiment, the entity recognition model includes a trained first Transformer model, the first Transformer model including N identical encoders, N droppers corresponding to the N identical encoders, a weighted summer, and a decoder, wherein N is a positive integer of at least 2;
[0011] The encoder is adapted to receive the text sequence in parallel and encode the text sequence into a sentence vector in parallel; the discarder is adapted to discard the random part in the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine the weighted summation result of the sentence vector output by the discarder with the random part discarded; the decoder is adapted to perform named entity recognition on the weighted summation result to obtain the word vector of the entity-free type.
[0012] In an exemplary embodiment, the method further includes a training process of the first Transformer model, wherein the training process includes:
[0013] Obtaining training samples, where words in the training samples are labeled with specific entity types;
[0014] The training sample is input into the first Transformer to train the first Transformer model, wherein the model parameters of the first Transformer model are configured through the training to make a predetermined loss function value lower than a preset threshold.
[0015] In an exemplary embodiment, the method further comprises:
[0016] Inputting the training sample into the trained first Transformer model to output a word vector corresponding to a specific entity type from the first Transformer;
[0017] Determine the mean vector of the word vectors corresponding to a particular entity type;
[0018] Include the mean vector in the registration set.
[0019] In an exemplary embodiment, the step of inputting the training sample into the trained first Transformer model to output a word vector corresponding to a specific entity type from the first Transformer includes:
[0020] The training samples are input into the N identical encoders in the trained first Transformer in parallel; wherein the encoders are adapted to receive the training samples in parallel and encode the training samples into sentence vectors in parallel; the discarder is adapted to discard the random part of the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine the weighted summation result of the sentence vector output by the discarder with the random part discarded; the decoder is adapted to determine the word vector of the word labeled with the specific entity type based on the weighted summation result.
[0021] In an exemplary embodiment, the method further comprises:
[0022] Determine pre-training samples;
[0023] Perturbing the pre-training samples;
[0024] Pre-training a second Transformer model using the perturbed training samples;
[0025] Copying the encoder in the pre-trained second Transformer model N times to obtain the N identical encoders;
[0026] The disturbance comprises at least one of the following:
[0027] The minimum unit in the pre-training sample is replaced by a mask; the minimum unit in the pre-training sample is randomly deleted; and the minimum unit in the pre-training sample is transformed in a disordered order.
[0028] In an exemplary embodiment, the pre-training samples include Chinese corpus constructed by using a Chinese vocabulary.
[0029] A named entity recognition device, comprising:
[0030] An input module, used for inputting a text sequence into a named entity recognition model to obtain a word vector recognized as a non-entity type by the named entity recognition model;
[0031] A first determination module is used to determine the similarity between the word vector without entity type and a mean vector in the registration set, wherein the mean vector corresponds to a specific entity type and is determined by a mean operation of multiple word vectors that conform to the specific entity type;
[0032] The second determination module is used to determine the specific entity type corresponding to the mean vector whose similarity is greater than a predetermined threshold value as the entity type of the word vector without entity type.
[0033] In an exemplary embodiment, the entity recognition model includes a trained first Transformer model, the first Transformer model including N identical encoders, N droppers corresponding to the N identical encoders, a weighted summer, and a decoder, wherein N is a positive integer of at least 2;
[0034] The encoder is adapted to receive the text sequence in parallel and encode the text sequence into a sentence vector in parallel; the discarder is adapted to discard the random part in the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine the weighted summation result of the sentence vector output by the discarder with the random part discarded; the decoder is adapted to perform named entity recognition on the weighted summation result to obtain the word vector of the entity-free type.
[0035] In an exemplary embodiment, the apparatus further comprises:
[0036] A training module is used to execute the training process of the first Transformer model, and the training process includes: obtaining training samples, in which words in the training samples are labeled with specific entity types; inputting the training samples into the first Transformer to train the first Transformer model, wherein the model parameters of the first Transformer model are configured through the training to make a predetermined loss function value lower than a preset threshold.
[0037] In an exemplary embodiment, the apparatus further comprises:
[0038] A registration set determination module is used to input the training samples into the trained first Transformer model to output a word vector corresponding to a specific entity type from the first Transformer; determine a mean vector of the word vector corresponding to the specific entity type; and include the mean vector in the registration set.
[0039] In an exemplary embodiment, the registration set determination module is used to input the training samples in parallel to the N identical encoders in the trained first Transformer; wherein the encoders are adapted to receive the training samples in parallel and encode the training samples into sentence vectors in parallel; the discarder is adapted to discard the random part of the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine the weighted summation result of the sentence vector output by the discarder with the random part discarded; and the decoder is adapted to determine the word vector of the word labeled with the specific entity type based on the weighted summation result.
[0040] In an exemplary embodiment, the apparatus further comprises:
[0041] A pre-training module is used to determine a pre-training sample; perturb the pre-training sample; pre-train a second Transformer model using the perturbed training sample; copy the encoder in the second Transformer model after the pre-training N times to obtain the N identical encoders; wherein the perturbation includes at least one of the following: masking to replace the minimum unit in the pre-training sample; randomly deleting the minimum unit in the pre-training sample; and randomly transforming the minimum unit in the pre-training sample.
[0042] In an exemplary embodiment, the pre-training samples include Chinese corpus constructed by using a Chinese vocabulary.
[0043] An electronic device, comprising:
[0044] processor;
[0045] a memory for storing instructions executable by the processor;
[0046] The processor is used to read the executable instructions from the memory and execute the executable instructions to implement the named entity recognition method as described in any one of the above items.
[0047] A computer-readable storage medium stores computer instructions, which, when executed by a processor, can implement any of the above named entity recognition methods.
[0048] A computer program product comprises computer instructions, wherein when the computer instructions are executed by a processor, the method for named entity recognition as described in any one of the above items is implemented.
[0049] It can be seen from the above technical scheme that in the embodiment of the present invention, a text sequence is input into a named entity recognition model to obtain a word vector that is recognized as a non-entity type by the named entity recognition model; the similarity between the word vector of the non-entity type and the mean vector in the registration set is determined, the mean vector corresponds to a specific entity type, and the mean vector is determined by the mean operation of multiple word vectors that conform to the specific entity type; the specific entity type corresponding to the mean vector whose similarity is greater than a predetermined threshold value is determined as the entity type of the word vector of the non-entity type. It can be seen that the embodiment of the present invention can recognize words that cannot be recognized by the model based on the similarity comparison with the mean vector, thereby improving the recognition accuracy.
[0050] Moreover, the embodiment of the present invention proposes a first Transformer model with a novel structure. The first Transformer model includes N encoders that receive text sequences in parallel, so that the text sequence can be copied N times to overcome the small sample problem, and the discarder and the weighted summer use a random discard method to achieve smoothing of sentence disturbance factors, reduce the interference of disturbance data on model prediction results, and improve model robustness.
[0051] In addition, the embodiments of the present invention use Chinese corpus constructed with a Chinese vocabulary for pre-training to achieve the Chinese semantic representation capability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0053] Figure 1 It is an exemplary flow chart of the named entity recognition method according to an embodiment of the present invention.
[0054] Figure 2 It is an exemplary schematic diagram of the pre-training process of an embodiment of the present invention.
[0055] Figure 3 It is an exemplary schematic diagram of the process of training the first Transformer model and generating a registration set according to an embodiment of the present invention.
[0056] Figure 4 It is an exemplary schematic diagram of the named entity recognition process based on the first Transformer model according to an embodiment of the present invention.
[0057] Figure 5 It is an exemplary structural diagram of a named entity recognition device according to an embodiment of the present invention.
[0058] Figure 6 is an exemplary structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0059] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings.
[0060] For the sake of simplicity and intuition in description, the scheme of the present invention is explained below by describing several representative embodiments. A large number of details in the embodiments are only used to help understand the scheme of the present invention. However, it is obvious that the technical scheme of the present invention may not be limited to these details when it is implemented. In order to avoid unnecessarily obscuring the scheme of the present invention, some embodiments are not described in detail, but only the framework is given. Hereinafter, "including" means "including but not limited to", and "according to..." means "at least according to...", but not limited to only according to...". Due to the language habits of Chinese, when the number of a component is not specifically indicated below, it means that the component can be one or more, or can be understood as at least one. The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the embodiments of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here, for example.
[0061] Taking into account the defect in the prior art that it is difficult to achieve named entity recognition in the zero-sample situation, the implementation mode of the present invention calculates the similarity between the word vector of the non-entity type and the mean vector of multiple word vectors of a specific entity type, which can recognize words that the named entity recognition model cannot recognize, that is, realize the prediction of unknown entity categories and improve the recognition accuracy.
[0062] Moreover, considering the defect of the prior art that the robustness to the training set containing interference factors is weak, the embodiment of the present invention constructs a balanced word vector feature representation through a random discarding method to alleviate the result error caused by the deviation.
[0063] In addition, in view of the defect that most of the existing technologies focus on English entity recognition and there is little research on the small sample task of Chinese named entity recognition, the implementation method of the present invention uses Chinese corpus constructed with a Chinese vocabulary for pre-training to realize the Chinese semantic representation ability of the model.
[0064] Figure 1 It is an exemplary flow chart of the named entity recognition method according to an embodiment of the present invention.
[0065] like Figure 1 As shown, the method includes:
[0066] Step 101: Input the text sequence into a named entity recognition model to obtain word vectors that are recognized as having no entity type by the named entity recognition model.
[0067] Here, the named entity recognition model is a deep learning model with named entity recognition capabilities. After the text sequence is input into the named entity recognition model, the named entity recognition model performs named entity recognition processing, in which some words are recognized as entity types (because samples of recognized entity types are usually provided in the training stage); the remaining words cannot be recognized, that is, they are non-entity types (because samples of corresponding entity types are usually not provided in the training stage), and the vectors of words without entity types are calculated, that is, the word vectors recognized as non-entity types.
[0068] Step 102: Determine the similarity between the word vector without entity type and the mean vector in the registration set, where the mean vector corresponds to a specific entity type and is determined by the mean operation of multiple word vectors that conform to the specific entity type.
[0069] Here, the registration set may include multiple mean vectors, each of which has its own specific entity type (the specific entity type can be annotated by a label). Wherein: each mean vector is determined by the mean operation of multiple word vectors that conform to the specific entity type. Preferably, the multiple word vectors are obtained by inputting training samples annotated with the specific entity type into the named entity recognition model.
[0070] For example, determining the similarity between the word vector without entity type and the mean vector in the registration set may be specifically implemented as follows: calculating the cosine similarity between the word vector without entity type and the mean vector in the registration set.
[0071] Step 103: Determine the specific entity type corresponding to the mean vector whose similarity is greater than a predetermined threshold value as the entity type of the word vector without entity type.
[0072] It can be seen that the embodiments of the present invention can recognize words that cannot be recognized by the model based on the similarity comparison process with the mean vector corresponding to the specific entity type, thereby improving the recognition accuracy.
[0073] In one embodiment, the entity recognition model may include a Transformer model. For example, the entity recognition model includes a trained first Transformer model, the first Transformer model includes N identical encoders, N dropouts corresponding to the N identical encoders, a weighted summer, and a decoder, wherein N is a positive integer of at least 2; the encoder is adapted to receive a text sequence in parallel and encode the text sequence into a sentence vector in parallel; the dropout is adapted to drop a random part of the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine a weighted summation result of the sentence vector output by the dropout with the random part dropped; and the decoder is adapted to perform named entity recognition on the weighted summation result to obtain a word vector without entity type.
[0074] Therefore, the embodiment of the present invention proposes a first Transformer model with a novel structure. The first Transformer model includes N encoders that receive text sequences in parallel, so that the text sequence can be copied N times to overcome the small sample problem, and the discarder and the weighted summer use a random discard method to achieve smooth processing of sentence disturbance factors, reduce the interference of disturbance data on model prediction results, and improve model robustness.
[0075] The following describes how to obtain N identical encoders in the first Transformer model.
[0076] In one embodiment, the method further includes: determining a pre-training sample; perturbing the pre-training sample; pre-training the second Transformer model using the perturbed training sample; copying the encoder in the pre-trained second Transformer model N times to obtain N identical encoders; wherein the perturbation includes at least one of the following: masking to replace the smallest unit in the pre-training sample; randomly deleting the smallest unit in the pre-training sample; shuffling the smallest unit in the pre-training sample, and so on. The overall structure of the second Transformer model has a common Transformer structure, that is, it includes an encoder-decoder structure, for example, the encoder and decoder each have 6 layers. For the encoder part, the entire encoder structure includes 6 layers, and each layer further includes a self-attention layer and a fully connected layer.
[0077] It can be seen that the embodiment of the present invention can conveniently implement the first Transformer model by copying the encoder obtained by pre-training the second Transformer model. Moreover, during the pre-training process of the second Transformer model, the robustness of the encoder and decoder can be further improved by perturbation processing, thereby further improving the robustness of the first Transformer model.
[0078] In one embodiment, the pre-training samples include Chinese corpus constructed by using a Chinese vocabulary.
[0079] Therefore, the embodiments of the present invention use Chinese corpus constructed with a Chinese vocabulary for pre-training to achieve the Chinese semantic representation capability of the model.
[0080] Figure 2 It is an exemplary schematic diagram of the pre-training process of an embodiment of the present invention.
[0081] Assume that the pre-training samples include Chinese corpus constructed by Chinese vocabulary. For example, Figure 2In the example, the pre-training sample contains 8 minimum units (tokens), which are t0, t1, t2, t3, t4, t5, t6, and t7 in sequence.
[0082] First, perturbation is performed on the pre-training sample, and the specific perturbation method includes at least one of the following:
[0083] (1) Mask replaces the smallest unit in the pre-training sample. For example, if mask is used to replace t2, the perturbed pre-training sample is: t0, t1, mask, t3, t4, t5, t6, t7.
[0084] (2) Randomly delete the smallest unit in the pre-training sample. For example, if t3 is randomly deleted, the perturbed pre-training samples are: t0, t1, t2, t4, t5, t6, t7.
[0085] (3) Shuffle the smallest unit in the pre-training samples. For example, if t1 is moved to the end of the sample, the perturbed pre-training samples are: t0, t2, t3, t4, t5, t6, t1.
[0086] Then, the perturbed training samples are input into the second Transformer model, and the second Transformer model is pre-trained to obtain Chinese semantic representation capabilities.
[0087] The encoder in the pre-trained second Transformer model is copied into N copies to obtain the above-mentioned N identical encoders. Moreover, each of the N identical encoders is connected to a respective discarder. Next, a common weighted summer is connected to the N discarders, and the weighted summer is connected to the decoder in the pre-trained second Transformer model to obtain the first Transformer model.
[0088] In one embodiment, the method further includes a training process of a first Transformer model, the training process comprising: obtaining training samples, wherein words in the training samples are labeled with specific entity types; inputting the training samples into the first Transformer to train the first Transformer model, wherein the model parameters of the first Transformer model are configured through training to make a predetermined loss function value lower than a preset threshold.
[0089] Therefore, by training the first Transformer model with the novel structure, the first Transformer model can have the ability of named entity recognition.
[0090] In one embodiment, the method also includes: inputting training samples into a trained first Transformer model to output a word vector corresponding to a specific entity type from the first Transformer; determining a mean vector of the word vector corresponding to the specific entity type; and including the mean vector in a registration set.
[0091] It can be seen that the embodiments of the present invention can also quickly obtain a registration set of word vectors for identifying non-entity types based on the trained first Transformer model.
[0092] In one embodiment, inputting training samples into a trained first Transformer model to output a word vector corresponding to a specific entity type from the first Transformer includes: inputting the training samples into N identical encoders in the trained first Transformer in parallel; wherein the encoders are adapted to receive the training samples in parallel and encode the training samples into sentence vectors in parallel; the discarder is adapted to discard the random part of the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine the weighted summation result of the sentence vector output by the discarder with the random part discarded; and the decoder is adapted to determine the word vector of the word labeled with the specific entity type based on the weighted summation result.
[0093] Therefore, the embodiments of the present invention can determine the word vectors of words labeled with specific entity types based on the trained first Transformer model, thereby assisting in generating a registration set.
[0094] The first Transformer model can be used to generate a registration set for recognizing words that the model cannot recognize.
[0095] Figure 3 It is an exemplary schematic diagram of the process of generating a registration set according to an embodiment of the present invention.
[0096] like Figure 3 As shown, the first Transformer model contains N identical encoders in the second Transformer model (hereinafter referred to as encoders, which have been Figure 2 Each encoder is connected to its own dropper. The N droppers are connected to a weighted summer, which is then connected to a decoder in the second Transformer model (hereinafter referred to as the decoder, which has been Figure 2 pre-trained as shown).
[0097] During the training of the first Transformer model:
[0098] First, the first Transformer model receives training samples, where words in the training samples are labeled with specific entity types. For example, the training samples sequentially contain minimum units k0, k1, k2, k3, k4, k5, k6, k7.
[0099] Then, the first Transformer model copies the training sample N times, where each training sample is input into its own encoder, that is, N encoders receive the training sample in parallel. The N encoders encode their received training samples into sentence vectors. Each discarder discards the random part of the sentence vector encoded by the corresponding encoder. The weighted summer performs weighted summation on the sentence vectors output by the N discarders with the random part discarded to obtain a weighted summation result. The decoder performs named entity recognition on the weighted summation result.
[0100] During the training process of the first Transformer model, the model parameters of the first Transformer model are generally configured to make a predetermined loss function value lower than a preset threshold.
[0101] When the first Transformer model is trained, it can be used to perform named entity recognition. Since the first Transformer model contains N encoders that receive text sequences in parallel, the text sequence can be copied N times to overcome the small sample problem. In addition, the dropper and weighted summer use a random dropout method to smooth the sentence disturbance factors, reduce the interference of the disturbance data on the model prediction results, and improve the model robustness.
[0102] However, considering the zero-sample problem, the first Transformer model after training may still generate word vectors that are identified as non-entity types. A registration set may be generated based on the first Transformer model to identify word vectors for non-entity types.
[0103] In one embodiment, the training samples are input again into the trained first Transformer model to output a word vector corresponding to a specific entity type from the first Transformer model output (usually the output of the previous layer of the classification layer in the decoder); determine the mean vector of the word vector corresponding to the specific entity type; and include the mean vector in the registration set.
[0104] For example, suppose training sample 1 is: "I like to eat apples", where the specific entity type of "apple" is labeled as "fruit"; training sample 2 is: "I like to eat oranges", where the specific entity type of "oranges" is labeled as "fruit"; training sample 3 is: "I like to eat cherry tomatoes", where the specific entity type of "cherry tomatoes" is labeled as "fruit". Then calculate the mean vector of the three word vectors of "apple", "orange", and "cherry tomatoes" (for example, sum the average value), and associate the mean vector with the specific entity type "fruit" and save it in the registration set.
[0105] Figure 4 It is an exemplary schematic diagram of the named entity recognition process based on the first Transformer model according to an embodiment of the present invention.
[0106] exist Figure 4 In the example, the first Transformer model receives a test sample for which named entity recognition is to be performed. For example, the test sample sequentially contains the smallest units M0, M1, M2, M3, M4, M5, M6, and M7.
[0107] During testing:
[0108] The first Transformer model copies the test sample N times, where each test sample is input into its own encoder, that is, N encoders receive the test sample in parallel. The N encoders encode the test samples they receive into sentence vectors. Each discarder discards the random part of the sentence vector encoded by the corresponding encoder. The weighted summer performs weighted summation on the sentence vectors output by the N discarders with the random part discarded to obtain a weighted summation result. The decoder performs named entity recognition on the weighted summation result. Among them, for some minimum units, since samples of the recognized entity type are provided in the training stage, the entity type of this part of the minimum units can be identified. For the remaining minimum units, since samples of the corresponding entity type are not provided in the training stage, the entity type cannot be identified, that is, a word vector without entity type is formed.
[0109] Further, the similarity between the word vector without entity type output by the first Transformer model (usually the output of the classification layer of the decoder) and the mean vector in the registration set is determined. The specific entity type corresponding to the mean vector with a similarity greater than a predetermined threshold is determined as the entity type of the word vector without entity type.
[0110] For example, assume that the first Transformer model outputs the word vector of "banana" as a non-entity type. Calculate the similarity between the word vector of "banana" and the mean vector of all specific entity types in the registration set. It can be found that the similarity between the word vector of "banana" and the mean vector of the specific entity type "fruit" is greater than a predetermined threshold value (for example, 0.5), and "banana" is identified as "fruit". It can be seen that the embodiment of the present invention can only identify the non-entity type output by the first Transformer model, thereby improving the recognition accuracy.
[0111] Figure 5 It is an exemplary structural diagram of a named entity recognition device according to an embodiment of the present invention.
[0112] The named entity recognition device 500 comprises:
[0113] An input module 501 is used to input a text sequence into a named entity recognition model to obtain a word vector that is recognized as a non-entity type by the named entity recognition model;
[0114] A first determination module 502 is used to determine the similarity between the word vector without entity type and the mean vector in the registration set, where the mean vector corresponds to a specific entity type and is determined by the mean operation of multiple word vectors that conform to the specific entity type;
[0115] The second determination module 503 is used to determine the specific entity type corresponding to the mean vector whose similarity is greater than a predetermined threshold value as the entity type of the word vector without entity type.
[0116] In an exemplary embodiment, the entity recognition model includes a trained first Transformer model, which includes N identical encoders, N discarders corresponding to the N identical encoders, a weighted summer and a decoder, wherein N is a positive integer of at least 2; the encoder is adapted to receive a text sequence in parallel and encode the text sequence into a sentence vector in parallel; the discarder is adapted to discard a random part of the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine a weighted summation result of the sentence vector output by the discarder with the random part discarded; and the decoder is adapted to perform named entity recognition on the weighted summation result to obtain a word vector without entity type.
[0117] In an exemplary embodiment, the apparatus 500 further includes:
[0118] The training module 504 is used to execute the training process of the first Transformer model, and the training process includes: obtaining training samples, in which words in the training samples are labeled with specific entity types; inputting the training samples into the first Transformer to train the first Transformer model, wherein the model parameters of the first Transformer model are configured through training to make a predetermined loss function value lower than a preset threshold.
[0119] In an exemplary embodiment, the device 500 also includes a registration set determination module 505, which is used to input training samples into the trained first Transformer model to output word vectors corresponding to a specific entity type from the first Transformer; determine the mean vector of the word vectors corresponding to the specific entity type; and include the mean vector in the registration set.
[0120] In an exemplary embodiment, the registration set determination module 505 is used to input the training samples in parallel into N identical encoders in the trained first Transformer; wherein the encoders are adapted to receive the training samples in parallel and encode the training samples into sentence vectors in parallel; the discarder is adapted to discard the random part of the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine the weighted summation result of the sentence vector output by the discarder with the random part discarded; the decoder is adapted to determine the word vector of the word labeled with a specific entity type based on the weighted summation result.
[0121] In an exemplary embodiment, the apparatus 500 further includes a pre-training module 506 for determining a pre-training sample; perturbing the pre-training sample; pre-training the second Transformer model using the perturbed training sample; replicating the encoder in the pre-trained second Transformer model N times to obtain N identical encoders; wherein the perturbation includes at least one of the following: masking and replacing the minimum unit in the pre-training sample; randomly deleting the minimum unit in the pre-training sample; and randomly transforming the minimum unit in the pre-training sample. In an exemplary embodiment, the pre-training sample includes a Chinese corpus constructed by a Chinese vocabulary.
[0122] In summary, in an embodiment of the present invention, a text sequence is input into a named entity recognition model to obtain a word vector that is recognized by the named entity recognition model as a non-entity type; the similarity between the word vector of the non-entity type and the mean vector in the registration set is determined, the mean vector corresponds to a specific entity type, and the mean vector is determined by the mean operation of multiple word vectors that conform to the specific entity type; the specific entity type corresponding to the mean vector whose similarity is greater than a predetermined threshold value is determined as the entity type of the word vector of the non-entity type. It can be seen that the embodiment of the present invention can recognize words that cannot be recognized by the model based on the similarity comparison with the mean vector, thereby improving the recognition accuracy.
[0123] Moreover, the embodiment of the present invention proposes a first Transformer model with a novel structure. The first Transformer model includes N encoders that receive text sequences in parallel, so that the text sequence can be copied N times to overcome the small sample problem, and the discarder and the weighted summer use a random discard method to achieve smoothing of sentence disturbance factors, reduce the interference of disturbance data on model prediction results, and improve model robustness.
[0124] In addition, the embodiments of the present invention use Chinese corpus constructed with a Chinese vocabulary for pre-training to achieve the Chinese semantic representation capability of the model.
[0125] The embodiment of the present invention also provides a computer-readable medium, the computer-readable storage medium stores instructions, and the instructions can execute the steps in the above named entity recognition method when executed by the processor. The computer-readable medium in the practical application may be included in the device / device / system described in the above embodiment, or it may exist alone without being assembled into the device / device / system. The above-mentioned computer-readable storage medium carries one or more programs, and when the above-mentioned one or more programs are executed, the named entity recognition method described in the above-mentioned embodiments can be implemented. According to the embodiment disclosed in the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above, but it is not used to limit the scope of protection of the present invention. In the embodiment disclosed in the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, a device or a device or used in combination with it.
[0126] like Figure 6As shown, the embodiment of the present invention further provides an electronic device, in which the device for implementing the method of the embodiment of the present invention can be integrated. Figure 6 As shown, it shows an exemplary structural diagram of an electronic device involved in an embodiment of the present invention,
[0127] Specifically, the electronic device may include a processor 601 with one or more processing cores, a memory 602 with one or more computer-readable storage media, and a computer program stored in the memory and executable on the processor. When executing the program in the memory 602, the above-mentioned named entity recognition method may be implemented.
[0128] In practical applications, the electronic device may further include components such as a power supply 603, an input unit 604, and an output unit 605. Those skilled in the art will appreciate that Figure 6 The structure of the electronic device shown in the figure does not constitute a limitation on the electronic device, and may include more or less components than shown in the figure, or combine certain components, or arrange components differently. Wherein: the processor 601 is the control center of the electronic device, and uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions of the server and processes data by running or executing software programs and / or modules stored in the memory 602, and calling data stored in the memory 602, so as to monitor the electronic device as a whole. The memory 602 can be used to store software programs and modules, that is, the above-mentioned computer-readable storage medium. The processor 601 executes various functional applications and data processing by running software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the server, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602 .
[0129] The electronic device also includes a power supply 603 for supplying power to each component, which can be logically connected to the processor 601 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 603 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, etc. The electronic device can also include an input unit 604, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control. The electronic device can also include an output unit 605, which can be used to display information input by the user or information provided to the user and various graphical user interfaces, which can be composed of graphics, text, icons, videos and any combination thereof.
[0130] An embodiment of the present invention further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, the method for named entity recognition as described in any of the above embodiments is implemented.
[0131] The flow chart and block diagram in the accompanying drawings of the present invention show the possible architecture, function and operation of the system, method and computer program product according to the various embodiments disclosed in the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in the order of the standards in different drawings. For example, the boxes represented by two connections can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0132] The principles and implementation methods of the present invention are described herein using specific implementation methods. The description of the above implementation methods is only used to help understand the method and core ideas of the present invention, and is not used to limit the present invention. For those skilled in the art, changes can be made in the specific implementation methods and application scopes according to the ideas, spirits and principles of the present invention, and any modifications, equivalent substitutions, improvements, etc. made therein should be included in the scope of protection of the present invention.
Claims
1. A method for named entity recognition, characterized in that: include: Inputting a text sequence into a named entity recognition model to obtain a word vector that is recognized by the named entity recognition model as a word vector without an entity type, wherein the word vector without an entity type is a word vector that cannot be recognized by the named entity recognition model, and the entity recognition model includes a trained first Transformer model; Determining a similarity between the word vector without the entity type and a mean vector in the registration set, the mean vector corresponding to a specific entity type, the mean vector being determined by a mean operation of a plurality of word vectors that conform to the specific entity type; as well as Determine the specific entity type corresponding to the mean vector whose similarity is greater than a predetermined threshold value as the entity type of the word vector without entity type; The registration set is obtained by the following method: Inputting training samples in which words are labeled with specific entity types into the trained first Transformer model to output word vectors corresponding to the specific entity type from the first Transformer; Determine the mean vector of the word vectors corresponding to a particular entity type; as well as including the mean vector in the registration set to obtain a registration set; The first Transformer model comprises N identical encoders, N droppers corresponding to the N identical encoders, a weighted summer, and a decoder, wherein N is a positive integer of at least 2; The encoder is adapted to receive the text sequence in parallel and encode the text sequence into sentence vectors in parallel; the discarder is adapted to discard the random part in the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine the weighted summation result of the sentence vector output by the discarder with the random part discarded; the decoder is adapted to perform named entity recognition on the weighted summation result to obtain the word vector of the entity-free type.
2. The method for named entity recognition according to claim 1, characterized in that: The method further includes a training process of the first Transformer model, wherein the training process includes: Obtaining training samples, where words in the training samples are labeled with specific entity types; The training sample is input into the first Transformer to train the first Transformer model, wherein the model parameters of the first Transformer model are configured through the training to make a predetermined loss function value lower than a preset threshold.
3. The method for named entity recognition according to claim 2, characterized in that: Inputting a training sample in which a word is labeled with a specific entity type into the trained first Transformer model to output a word vector corresponding to the specific entity type from the first Transformer, including: The training samples are input into the N identical encoders in the trained first Transformer in parallel; wherein the encoders are adapted to receive the training samples in parallel and encode the training samples into sentence vectors in parallel; the discarder is adapted to discard the random part of the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine the weighted summation result of the sentence vector output by the discarder with the random part discarded; the decoder is adapted to determine the word vector of the word labeled with the specific entity type based on the weighted summation result.
4. The method for named entity recognition according to any one of claims 1 to 3, characterized in that: The method further comprises: Determine pre-training samples; Perturbing the pre-training samples; Pre-training a second Transformer model using the perturbed training samples; Copying the encoder in the pre-trained second Transformer model N times to obtain the N identical encoders; The disturbance comprises at least one of the following: The minimum unit in the pre-training sample is replaced by a mask; the minimum unit in the pre-training sample is randomly deleted; and the minimum unit in the pre-training sample is transformed in a disordered order.
5. The method for named entity recognition according to claim 4, characterized in that: The pre-training samples include Chinese corpus constructed by using a Chinese vocabulary.
6. A named entity recognition device, characterized in that: include: An input module, used for inputting a text sequence into a named entity recognition model to obtain a word vector recognized by the named entity recognition model as a word vector without an entity type, wherein the word vector without an entity type is a word vector that cannot be recognized by the named entity recognition model, and the entity recognition model includes a trained first Transformer model; A first determination module is used to determine the similarity between the word vector without entity type and a mean vector in the registration set, wherein the mean vector corresponds to a specific entity type and is determined by a mean operation of multiple word vectors that conform to the specific entity type; as well as A second determination module is used to determine the specific entity type corresponding to the mean vector whose similarity is greater than a predetermined threshold value as the entity type of the word vector without entity type; The registration set is obtained by the following method: Inputting training samples in which words are labeled with specific entity types into the trained first Transformer model to output word vectors corresponding to the specific entity type from the first Transformer; Determine the mean vector of the word vectors corresponding to a particular entity type; as well as including the mean vector in the registration set to obtain a registration set; The first Transformer model comprises N identical encoders, N droppers corresponding to the N identical encoders, a weighted summer, and a decoder, wherein N is a positive integer of at least 2; The encoder is adapted to receive the text sequence in parallel and encode the text sequence into sentence vectors in parallel; the discarder is adapted to discard the random part in the sentence vector encoded by the corresponding encoder; the weighted summer is adapted to determine the weighted summation result of the sentence vector output by the discarder with the random part discarded; the decoder is adapted to perform named entity recognition on the weighted summation result to obtain the word vector of the entity-free type.
7. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the method for named entity recognition according to any one of claims 1 to 5 is implemented.
8. A computer program product, characterized in that The method comprises computer instructions, which implement the named entity recognition method according to any one of claims 1 to 5 when executed by a processor.
Citation Information
Patent Citations
Natural language processing method and device and storage medium
CN113590831A
Judicial named entity recognition method based on natural language processing
CN115238697A