Entity recognition model training method, device, equipment and entity recognition method
By training the initial multitasking model and using its output feature vectors to train the entity recognition model, the problem of word segmentation inaccuracy caused by blurred boundaries in Chinese words is solved, and the accuracy of the entity recognition model is improved.
Patent Information
- Application Number
- CN202111243356.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-25
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2041-10-25
AI Technical Summary
The existing technology cannot effectively solve the problem of blurred boundaries of Chinese characters and words, resulting in inaccurate word segmentation and affecting the accuracy of the entity recognition model.
By using the first training sample to train the initial model, a pre-trained model is obtained, and an initial multi-task model is established based on the pre-trained model for performing word segmentation tasks and entity recognition tasks. Then, the initial multitasking model is trained through the target loss function and the second training sample to obtain the target multitasking model. Finally, the entity recognition model is trained using the word participle representation vector, word vector and position representation vector output from the target multitasking model.
It effectively improves the accurate recognition ability of the entity recognition model for Chinese words and solves the problem of inaccurate word segmentation caused by blurred boundaries of Chinese words.
Smart Images

Figure CN114048737B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent recognition technology, and in particular to an entity recognition model training method, device, equipment and entity recognition method. Background Art
[0002] The entity recognition method of the related art cannot solve the problem of blurred boundaries of Chinese words. The boundary markers of Chinese and English are different. English has obvious spaces and some unique form marks, such as capitalizing the first letter, as boundary markers of English. However, Chinese words do not have obvious segmentation marks like English, which leads to blurred front and back boundaries of Chinese words and is not easy to determine, and inaccurate word segmentation. Because the word segmentation task and the entity recognition task affect each other, and the related art does not take into account the relationship between the word segmentation task and the entity recognition task, which affects the accuracy of the entity recognition model for Chinese words. Summary of the invention
[0003] The present application provides an entity recognition model training method, apparatus, device and computer-readable storage medium to solve the problem that the relationship between word segmentation tasks and entity recognition tasks is not considered in the related art, which affects the accuracy of entity recognition model in identifying Chinese words.
[0004] In the first aspect, the present application provides an entity recognition model training method, which uses a first training sample to train an initial model to obtain a pre-trained model, wherein the pre-trained model is used for natural language processing; an initial multi-task model is established based on the pre-trained model, and the initial multi-task model is used to perform word segmentation tasks and entity recognition tasks; the initial multi-task model is trained using a target loss function and a second training sample to obtain a target multi-task model; the target training sample is input into the target multi-task model to obtain a word segmentation representation vector output by the target multi-task model, as well as a word vector and a position representation vector output by the pre-trained model in the target multi-task model; and an entity recognition model is trained using the word segmentation representation vector, the word vector and the position representation vector to obtain a target model.
[0005] In a second aspect, the present application provides an entity recognition method, which includes: performing entity recognition on a target sample using a target model obtained by the entity recognition model training method of any embodiment of the first aspect.
[0006] In the third aspect, the present application provides an entity recognition model training device, including a first training module, which uses a first training sample to train an initial model to obtain a pre-trained model; a second training module, which establishes an initial multi-task model based on the pre-trained model, and the initial multi-task model is used to perform word segmentation tasks and entity recognition tasks; a third training module, which trains the initial multi-task model through a target loss function and a second training sample to obtain a target multi-task model; a fourth training module, which inputs the target training sample into the target multi-task model to obtain a word segmentation representation vector output by the target multi-task model, as well as a word vector and a position representation vector output by the pre-trained model in the target multi-task model; a fifth training module, which uses the word segmentation representation vector, the word vector and the position representation vector to train the entity recognition model to obtain a target model.
[0007] In a fourth aspect, the present application provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0008] Memory, used to store computer programs;
[0009] The processor is used to implement the steps of the entity recognition model training method of any embodiment of the first aspect when executing the program stored in the memory.
[0010] In a fifth aspect, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of the entity recognition model training method of any embodiment of the first aspect are implemented.
[0011] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:
[0012] The entity recognition model training method provided in the embodiment of the present application is applied to entity recognition, and an initial model is trained using a first training sample to obtain a pre-trained model; an initial multi-task model is established according to the pre-trained model, and the initial multi-task model is used to perform word segmentation tasks and entity recognition tasks. The initial multi-task model establishes a semantically rich language representation method and enhances the representation ability of the language model; the initial multi-task model is trained by a target loss function and a second training sample to obtain a target multi-task model, and the target multi-task model is trained on the word segmentation task and entity recognition task of the initial multi-task model by the target loss function and the second training sample, so that the word segmentation task is more accurate; the target training sample is input into the target multi-task model to obtain a word segmentation representation vector output by the target multi-task model, as well as a word vector and a position representation vector output by the pre-trained model in the target multi-task model; the entity recognition model is trained using the word segmentation representation vector, the word vector and the position representation vector to obtain a target model, and considering the influence of the word segmentation task on the entity recognition task, the word segmentation representation vector is introduced to train the entity recognition model, which effectively improves the accuracy of the entity recognition model in recognizing Chinese words. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0014] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0015] Figure 1 A schematic diagram of a hardware environment for an optional entity recognition model training method provided according to an embodiment of the present application;
[0016] Figure 2 A schematic diagram of an optional entity recognition model training method provided according to an embodiment of the present application;
[0017] Figure 3 A schematic diagram of the structure of an optional target multi-task model provided according to an embodiment of the present application;
[0018] Figure 4 A schematic diagram of the structure of an optional entity recognition model provided according to an embodiment of the present application;
[0019] Figure 5 A schematic diagram of the structure of an optional entity recognition model provided according to an embodiment of the present application;
[0020] Figure 6 A schematic diagram of another optional entity recognition model training method provided according to an embodiment of the present application;
[0021] Figure 7 A schematic diagram of another optional entity recognition model training method provided according to an embodiment of the present application;
[0022] Figure 8 A schematic diagram of another optional entity recognition model training method provided according to an embodiment of the present application;
[0023] Fig. 9 A schematic diagram of another optional entity recognition model training method provided according to an embodiment of the present application;
[0024] Fig.10 A schematic diagram of another optional entity recognition model training method provided according to an embodiment of the present application;
[0025] Fig.11 A schematic diagram of another optional entity recognition model training method provided according to an embodiment of the present application;
[0026] Fig.12 A block diagram of an optional model training device provided according to an embodiment of the present application;
[0027] Fig.13 A schematic diagram of an optional electronic device structure provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0029] The entity recognition method of the related technology cannot solve the problem of blurred boundaries of Chinese words. The boundary markers of Chinese and English are different. English has obvious spaces and some unique formal marks, such as capitalizing the first letter, as boundary markers of English. However, Chinese words do not have obvious segmentation marks like English, which leads to blurred front and back boundaries of Chinese words and is not easy to determine, and the word segmentation is inaccurate. This is because the word segmentation task and the entity recognition task influence each other, and the related technology does not take into account the relationship between the word segmentation task and the entity recognition task, which affects the accuracy of the entity recognition model in identifying Chinese words.
[0030] In order to solve the problems mentioned in the background technology, according to one aspect of an embodiment of the present application, an embodiment of an entity recognition model training method is provided.
[0031] Optionally, in an embodiment of the present application, the above entity recognition model training method can be applied to Figure 1 In the hardware environment composed of the terminal 101 and the server 103 shown in FIG. Figure 1 As shown, the server 103 is connected to the terminal 101 via a network, and can be used to provide services for the terminal or a client installed on the terminal. A database 102 can be set on the server or independently of the server to provide data storage services for the server 103. The above-mentioned network includes but is not limited to: a wide area network, a metropolitan area network or a local area network, and the terminal 101 includes but is not limited to a PC, a mobile phone, a tablet computer, etc.
[0032] An entity recognition model training method in the embodiment of the present application can be executed by the server 103 or the terminal 101, or can be executed by the server 103 and the terminal 101 together, such as Figure 2 As shown, the method may include the following steps:
[0033] Step S201, using the first training sample to train the initial model to obtain a pre-trained model 301, wherein the pre-trained model 301 is used for natural language processing.
[0034] In the related art, the model for solving the entity recognition task is mainly based on statistical methods. Among them, the statistical methods include Hidden Markov Mode (HMM), Maximum Entropy (ME), and Support Vector Machine (SVM). However, in entity recognition, the boundary markers of Chinese and English are different. English has obvious spaces and some unique form marks, such as capitalization of the first letter, etc. as the boundary markers of English words. However, Chinese words do not have obvious segmentation marks like English, which leads to the fuzzy front and back boundaries of Chinese words and is not easy to determine. Because the word segmentation task and the entity recognition task affect each other, and in the related art, the entity recognition model based on the statistical method does not take into account the relationship between the word segmentation task and the entity recognition task, which affects the accuracy of the entity recognition model for entity recognition of Chinese words. The entity recognition model training method provided in the embodiment of the present application can effectively improve the word segmentation ability of the entity recognition model, and then improve the accuracy of the entity recognition model in recognizing Chinese words.
[0035] In the embodiment of the present application, the first training sample can be a text corpus with labels or a text corpus without any labels, which is not specifically limited in the embodiment of the present application. The initial model includes, but is not limited to, the RoBerta model, etc. The pre-training model 301 is used for natural language processing, for example, to recognize names of people, places, and organization names in the first training sample, which is not specifically limited in the embodiment of the present application.
[0036] In the embodiment of the present application, the initial model is a RoBerta model. The initial model is trained with a first training sample to obtain a pre-trained model 301. The pre-trained model 301 may be a Pre_RoBerta model. The pre-trained model 301 is used for natural language processing.
[0037] Step S202, establishing an initial multi-task model based on the pre-trained model 301, and the initial multi-task model is used to perform word segmentation tasks and entity recognition tasks.
[0038] In an embodiment of the present application, the pre-trained model 301 can be a Pre_RoBerta model. An initial multi-task model is established based on the Pre_RoBerta model. The initial multi-task model is used to perform word segmentation tasks and entity recognition tasks. Both the entity recognition task and the word segmentation task are sequence labeling tasks, and there is a strong correlation between the two tasks.
[0039] Step S203, training the initial multi-task model through the target loss function and the second training sample 304 to obtain a target multi-task model, and the target loss function is used to combine the word segmentation task and the entity recognition task.
[0040] In the embodiment of the present application, the first training sample and the second training sample 304 mentioned above can be text corpus with labels or text corpus without any labels. They can be set as needed, and the embodiment of the present application does not make any specific limitation on this.
[0041] The target multi-task model of the embodiment of the present application is trained by the initial multi-task model. First, the pre-trained model 301 is connected with the first task layer 302 and the second task layer 303 to obtain the initial multi-task model. The target loss function and the second training sample 304 are used to train the initial multi-task model to obtain the following: Figure 3 The target multi-task model shown in Figure 3 , which is a schematic diagram of a target multi-task model provided in an embodiment of the present application. The target multi-task model includes: a pre-trained model 301, a trained first task layer 302, and a trained second task layer 303.
[0042] In an embodiment of the present application, training the initial multi-task model using the target loss function and the second training sample 304 includes: using the second training sample 304 to train the initial multi-task model multiple times. In the initial stage of training, the word segmentation task of the initial multi-task model is first trained by updating the second training sample 304. When the loss value of the word segmentation task meets the target loss threshold, the second training sample 304 is updated to train the entity recognition task. Therefore, when training the entity recognition task, the entity recognition task of the initial multi-task model has a certain word segmentation ability, which effectively improves the performance and generalization ability of the target multi-task model.
[0043] It can be understood that in the target multi-task model of the embodiment of the present application, the pre-trained model can be a Pre_RoBert model, the trained first task layer 302 can be a Bi-LSTM_n, and the trained second task layer 303 can be a Bi-LSTM_p. Bi-LSTM_n and Bi-LSTM_p are respectively connected to the Pre_RoBert model, and Bi-LSTM_n is used to perform entity recognition tasks; Bi-LSTM_p is also used to perform word segmentation tasks.
[0044] Step S204, input the target training sample into the target multi-task model to obtain the word segmentation representation vector 403 output by the target multi-task model, as well as the character vector 402 and position representation vector 401 output by the pre-training model 301 in the target multi-task model.
[0045] In the embodiment of the present application, the above-mentioned first training sample, second training sample 304 and target training sample can be text corpora with labels, or text corpora without any labels. Any two or three of them can be the same, or all three can be different. They can be set as needed, and the embodiment of the present application does not make specific limitations on this.
[0046] In an embodiment of the present application, the first task layer of the target multi-task model obtained by the above training can be a Bi-LSTM_n, and the pre-training model 301 of the target multi-task model can be a Pre_RoBert model. After the target training sample is input into the target multi-task model, the word segmentation representation vector 403 output by the Bi-LSTM_n and the character vector 402 output by the Pre_RoBert model are obtained. At the same time, the Pre_RoBert model itself encodes the position of the target training sample as the position representation vector 401.
[0047] Step S205, using the word representation vector 403, the character vector 402 and the position representation vector 401 to train the entity recognition model to obtain the target model.
[0048] In the present application embodiment, Figure 4As shown, the entity recognition model includes a language layer 404. The second training sample 304 is input into Figure 3 In the target multi-task model shown in FIG. 1 , the word segmentation representation vector 403 output by the first task layer 302 of the target multi-task model, the character vector 402 output by the pre-trained model 301, and the position representation vector 401 of the pre-trained model 301 are obtained. The word segmentation representation vector 403, the character vector 402, and the position representation vector 401 are used to train Figure 4 In the entity recognition model shown, the language layer 404 of the entity recognition model acts as a feature extractor to extract features from the input vector and output features. The first target feature is the result output by the entity recognition model.
[0049] like Figure 5 As shown, the entity recognition model of the embodiment of the present application may also include a trained first task layer 302 and a conditional random field layer 405, wherein the language layer 404 is connected to the trained first task layer 302, and the trained first task layer 302 is connected to the conditional random field layer 405. After the language layer 404 of the entity recognition model of the above embodiment extracts features, the features are input into the trained first task layer 302, and the trained first task layer 302 performs entity recognition on the input features. Next, the trained first task layer 302 inputs the features after performing entity recognition into the conditional random field layer 405 to learn the dependencies and constraints between the features.
[0050] In an embodiment of the present application, the language layer 404 may be a Bert model, the trained first task layer 302 may be a Bi-LSTM_n, and the conditional random field layer 405 may be a CRF model. The entity recognition model is trained using the word segmentation representation vector 403, the character vector 402, and the position representation vector 401 to obtain a target model. Bi-LSTM_n is used to perform entity recognition tasks, which can effectively accelerate the training speed of the entity recognition model. In addition, it can also increase the sensitivity of the entity recognition model to the temporal characteristics of the input samples and output new features. The CRF model is introduced after the Bi-LSTM_n. The CRF model is used to learn the features output by the Bi-LSTM_n. The learning of the CRF model includes, but is not limited to, the dependencies and constraints between features.
[0051] It can be understood that the front and back boundaries of Chinese words are fuzzy. The entity recognition model of the above embodiment introduces the word segmentation representation vector 403 for training, which can demarcate the front and back boundaries of Chinese words. This is because the word segmentation task and the entity recognition task influence each other. The embodiment of the present application takes into account the relationship between the entity recognition task and the word segmentation task. Therefore, the entity recognition model of the embodiment of the present application can effectively improve the accuracy of entity recognition of Chinese words.
[0052] like Figure 6 As shown, specifically, in the above step S202, an initial multi-task model is established according to the pre-trained model 301, and the initial multi-task model is used to perform word segmentation tasks and entity recognition tasks, which can be implemented by the following steps S601 and S602:
[0053] Step S601, establishing an initial multi-task model according to the pre-trained model 301, the initial multi-task model is used to perform word segmentation tasks and entity recognition tasks.
[0054] Step S602, the pre-trained model 301 is connected to the trained first task layer 302 and the trained second task layer 303 respectively to obtain an initial multi-task model, the trained first task layer 302 is used to perform the entity recognition task, and the trained second task layer 303 is used to perform the word segmentation task.
[0055] In the embodiment of the present application, the pre-trained model 301 can be a Pre_RoBert model, the trained first task layer 302 can be a Bi-LSTM_n, the trained second task layer 303 can be a Bi-LSTM_p, and the Pre_RoBerta model is connected to the Bi-LSTM_n and Bi-LSTM_p respectively, the Bi-LSTM_n is used for feature extraction to perform entity recognition tasks, and the Bi-LSTM_p is also used for feature extraction to perform word segmentation tasks. It can be understood that the Bi-LSTM_n and Bi-LSTM_p in the embodiment of the present application are both used to extract features, and the parameters between the two are not shared.
[0056] like Figure 7 As shown, specifically, in the above step S203, the initial multi-task model is trained by the target loss function and the second training sample 304 to obtain the target multi-task model, which can be implemented by the following steps S701 and S702:
[0057] Step S701 , training the initial multi-task model through the second training sample 304 .
[0058] In an embodiment of the present application, the second training sample 304 can train the initial multi-task model multiple times. The above-mentioned first training sample and second training sample 304 can be text corpora with annotation labels, or can be text corpora without any annotation labels. They can be set as needed, and the embodiment of the present application does not make specific limitations on this.
[0059] Step S702: When the loss value of the target loss function reaches a threshold, it is determined that the initial multi-task model training is completed, and the target multi-task model is obtained. The target loss function is:
[0060] Among them, Loss represents the loss value of the target loss function, loss1 represents the loss value of the entity recognition task, loss2 represents the loss value of the word segmentation task, Step represents the total number of times the initial multi-task model is trained, and i represents the current number of training times.
[0061] In the embodiment of the present application, the loss value of the target loss function is compared with the threshold to determine whether the initial multi-task model is trained. When the loss value of the target loss function reaches the threshold, the initial multi-task model training is determined to be completed, and the target multi-task model is obtained. The design of the target loss function includes but is not limited to: Among them, Loss represents the loss value of the target loss function, loss1 represents the loss value of the entity recognition task, loss2 represents the loss value of the word segmentation task, Step represents the total number of times the initial multi-task model is trained, and i represents the current number of training times. It is sufficient to determine that the training of the initial multi-task model is completed.
[0062] In the embodiment of the present application, the target loss function uses the transformation of the sine and cosine functions as the weighted weight of the entity recognition task loss function value and the word segmentation task loss function value. Both the entity recognition task and the word segmentation task are sequence labeling tasks, but the solution space of the word segmentation task can be smaller than the entity recognition task. Therefore, in the initial stage of training the initial multi-task model, the word segmentation task loss accounts for the main part of the total loss. As the initial multi-task model is trained, the entity recognition task loss accounts for the main part of the total loss. When the loss value of the target loss function reaches a certain threshold, the training of the initial multi-task model is determined to be completed, and the target multi-task model is obtained. It can be understood that the target multi-task model has a certain word segmentation ability after the above training, so that when performing entity recognition, the performance effect and generalization ability of the entity recognition model are effectively improved.
[0063] In an embodiment of the present application, in the initial stage of training the above-mentioned initial multi-task model, a relatively simple entity recognition task and / or word segmentation task can be used, thereby alleviating the cold start problem caused by updating the second training sample 304 to the training of the initial multi-task model.
[0064] like Figure 8 As shown, specifically, in the above step S203, the initial multi-task model is trained by the target loss function and the second training sample 304, which can be implemented by the following steps S801 and S802:
[0065] Step S801 , training the word segmentation task of the initial multi-task model by updating the second training sample 304 .
[0066] In the embodiment of the present application, the first training sample and the second training sample 304 may be text corpora with labels or text corpora without any labels. They may be set as needed, and the embodiment of the present application does not make any specific limitation on this.
[0067] Step S802 , when the loss value of the word segmentation task meets the target loss threshold, the entity recognition task of the initial multi-task model is trained by updating the second training sample 304 .
[0068] In an embodiment of the present application, the word segmentation task of the initial multi-task model is first trained by updating the second training sample 304. When the loss value of the word segmentation task meets the target loss threshold, the entity recognition task of the initial multi-task model is trained by updating the second training sample 304. When the second training sample 304 is updated for the training entity recognition task, the entity recognition task of the initial multi-task model has a certain word segmentation ability, which effectively improves the performance and generalization ability of the target multi-task model obtained by training the initial multi-task model.
[0069] It is understandable that the embodiment of the present application updates the second training sample 304 multiple times to train the word segmentation task and entity recognition task of the initial multi-task model.
[0070] like Fig. 9 As shown, specifically, in the above step S205, the entity recognition model is trained using the word segmentation representation vector 403, the character vector 402 and the position representation vector 401 to obtain the target model, which can be achieved through the following steps S901 and S902:
[0071] Step S901 , performing a sum operation on the word representation vector 403 , the character vector 402 , and the position representation vector 401 to obtain a target data set.
[0072] Step S902, training the language layer 404 of the entity recognition model with the target data set to obtain a target model, wherein the language layer 404 is used to extract features from the input target data set and output a first target feature.
[0073] In an embodiment of the present application, the language layer 404 of the entity recognition model includes, but is not limited to, a Bert model. The target data set after adding the word representation vector 403, the character vector 402 and the position representation vector 401 is input into the Bert model for training. After the Bert model extracts and represents the input target data set, it outputs the first target feature.
[0074] like Fig.10 Specifically, the entity recognition model in the above embodiment may further include step S1001 and step S1002:
[0075] Step S1001 , connecting the language layer 404 to the trained first task layer 302 , and connecting the trained first task layer 302 to the conditional random field layer 405 .
[0076] Step S1002, input the first target feature to the trained first task layer 302 to obtain the second target feature after performing entity recognition on the first target feature, and input the second target feature to the conditional random field layer 405 to learn the dependency and constraint relationship between the second target features.
[0077] In an embodiment of the present application, the above-mentioned language layer 404 may include a Bert model, a conditional random field layer 405, including but not limited to a CRF model, and after the Bert model of the entity recognition model outputs the first target feature, the first target feature is input into the CRF model of the entity recognition model, and the CRF model is used to learn the dependency and constraint relationship between the first target features.
[0078] In an embodiment of the present application, the language layer 404 can be a Bert model, the conditional random field layer 405 can be a CRF model, and the first task layer 302 that has been trained can be a Bi-LSTM_n. The Bert model is connected to the Bi-LSTM_n, and the Bi-LSTM_n is connected to the CRF model. After the Bert model extracts features from the input target data set, it outputs the first target feature to the Bi-LSTM_n. The Bi-LSTM_n performs entity recognition on the first target feature and obtains the second target feature. The second target feature is input into the CRF model to learn the dependency and constraint relationship between the second target features. In an embodiment of the present application, connecting the Bert model to the Bi-LSTM_n can effectively improve the training speed of the entity recognition model, and can also increase the sensitivity of the entity recognition model to the sample time series features. In addition, after the Bi-LSTM_n, the CRF model is connected to learn the dependency and constraint relationship between the second target features.
[0079] In an embodiment of the present application, an entity recognition method is also provided. The entity recognition method includes but is not limited to performing entity recognition on a target sample by obtaining a target model by implementing an entity recognition model training method provided in any of the aforementioned method embodiments.
[0080] In the embodiment of the present application, the above-mentioned first training sample, second training sample 304 and target training sample can be text corpora with labels, or text corpora without any labels. Any two or three of them can be the same, or all three can be different. They can be set as needed, and the embodiment of the present application does not make specific limitations on this.
[0081] like Fig.11 As shown, in one embodiment of the present application, the training method of the entity recognition model in the above embodiment can also be implemented by the following steps S1101, S1102, S1103, S1104, S1105 and S1106:
[0082] Step S1101, start.
[0083] Step S1102: train the initial model using the first training sample.
[0084] In the embodiment of the present application, the first training sample can be a text corpus that does not require any labeling, or can be a text corpus that has labeling. After the initial model is trained using the first training sample, a pre-training model 301 is obtained.
[0085] Step S1103: training an initial multi-task model consisting of a word segmentation task and an entity recognition task.
[0086] In an embodiment of the present application, the pre-trained model 301 is connected to the trained first task layer 302 and the trained second task layer 303 respectively to obtain an initial multi-task model. The initial multi-task model is used to perform entity recognition tasks and word segmentation tasks, the entity recognition task is the trained first task layer 302, and the word segmentation task is the trained second task layer 303. It can be understood that both tasks are sequence labeling tasks and have a strong correlation. After the pre-trained model 301 outputs a tensor, the trained first task layer 302 and the trained second task layer 303 extract features, and the parameters between the trained first task layer 302 and the trained second task layer 303 are not shared.
[0087] It can be understood that the initial multi-task model in the above embodiment also includes a target loss function and a second training sample 304. The target loss function includes an entity recognition task loss function and a word segmentation task loss function. The target loss function can be used to determine whether the initial multi-task model has been trained. When the loss value of the target loss function reaches a threshold, the training of the initial multi-task model is determined to be completed, and the target multi-task model is obtained. The second training sample 304 is used to train the initial multi-task model multiple times. In the initial stage of training, the word segmentation task loss accounts for the main part. As the initial multi-task model is trained, the entity recognition task loss accounts for the main part. In the initial stage of training, using simpler tasks can alleviate the cold start problem of updating the parameters of the initial multi-task model. When training the entity recognition task, the entity recognition task of the initial multi-task model has a certain word segmentation ability, which effectively improves the performance and generalization ability of the target multi-task model.
[0088] It should be noted that, in the above embodiment, the loss function of the initial multi-task model is: Among them, Loss represents the loss value of the target loss function, loss1 represents the loss value of the entity recognition task, loss2 represents the loss value of the word segmentation task, Step represents the total number of times the initial multi-task model is trained, and i represents the current number of training times.
[0089] Step S1104, introducing the Bi-LSTM of the word segmentation task into the input representation of the entity recognition model.
[0090] Step S1105, training an entity recognition model.
[0091] Step S1106, end.
[0092] In an embodiment of the present application, the second training sample 304 is input into the target multi-task model trained in step S1103 to obtain a word segmentation representation vector 403 and a character vector 402 output by the pre-training model 301. The pre-training model 301 encodes the data position and generates a position representation vector 401. The word segmentation representation vector 403, the character vector 402 and the position representation vector 401 are added and then input into the entity recognition model for training. The entity recognition model of the embodiment of the present application includes a language layer, which can be a Bert model. The Bert model is used as a feature extractor to extract and characterize the input word segmentation representation vector 403, the character vector 402 and the position representation vector 401, thereby enhancing the characterization ability of the entity recognition model, introducing the word segmentation representation vector 403 into the entity recognition model for training, expanding the feature space of the entity recognition model, and improving the prediction accuracy of the entity recognition task.
[0093] The entity recognition model in the above embodiment may also include the first task layer 302 and the conditional random field layer obtained through the training in step S1103. The first task layer 302 may be a Bi-LSTM_n, and the conditional random field layer may be a CRF model. The Bi-LSTM_n is connected with the Bert model to accelerate the training speed of the entity recognition model and increase the sensitivity of the entity recognition model to the time series characteristics of the target sample. By connecting the Bi-LSTM_n with the CRF model, the CRF model can learn the dependencies and constraints between the output features of the Bi-LSTM_n.
[0094] like Fig.12 As shown, the embodiment of the present application also provides an entity recognition model training method device, the entity recognition model training method device includes but is not limited to:
[0095] A first training module 1201 uses a first training sample to train an initial model to obtain a pre-trained model 301, wherein the pre-trained model 301 is used for natural language processing;
[0096] The second training module 1202 establishes an initial multi-task model according to the pre-training model 301, and the initial multi-task model is used to perform word segmentation tasks and entity recognition tasks;
[0097] The third training module 1203 trains the initial multi-task model through the target loss function and the second training sample 304 to obtain a target multi-task model, wherein the target loss function is used to combine the word segmentation task and the entity recognition task;
[0098] The fourth training module 1204 inputs the target training sample into the target multi-task model to obtain the word segmentation representation vector 403 output by the target multi-task model, as well as the character vector 402 and the position representation vector 401 output by the pre-training model 301 in the target multi-task model;
[0099] The fifth training module 1205 uses the word segmentation representation vector 403, the character vector 402 and the position representation vector 401 to train the entity recognition model to obtain a target model.
[0100] It should be noted that the first training module 1201 in this embodiment can be used to execute step S201 in the embodiment of the present application, the second training module 1202 in this embodiment can be used to execute step S202 in the embodiment of the present application, the third training module 1203 in this embodiment can be used to execute step S203 in the embodiment of the present application, the fourth training module 1204 in this embodiment can be used to execute step S204 in the embodiment of the present application, and the fifth training module 1205 in this embodiment can be used to execute step S205 in the embodiment of the present application. It should be noted that the above modules as part of the device can be run on Figure 1 In the hardware environment shown, it can be implemented by software or by hardware.
[0101] Optionally, the first training module 1201 is specifically configured to:
[0102] The RoBerta model is trained using the first training sample to obtain a model with initial parameter values, namely the Pre_RoBerta model. The Pre_RoBerta model can recognize named entities such as names of people, places, and organization names in the sample corpus.
[0103] Optionally, the second training module 1202 is specifically configured to:
[0104] The initial multi-task model is established through the Pre_RoBerta model. The initial multi-task model is used to perform entity recognition tasks and word segmentation tasks. Both tasks are essentially sequence labeling tasks. From the perspective of natural language characteristics, there is a strong correlation between entity recognition tasks and word segmentation tasks.
[0105] Optionally, the entity recognition model training device further includes a third training module 1203, which is used to:
[0106] Through the target loss function, the entity recognition task and the word segmentation task of the initial multi-task model are combined, and the initial multi-task model is trained using the second training sample 304. After the Pre_RoBerta model outputs the tensor, the entity recognition task and the word segmentation task respectively use their own Bi-LSTM layers to extract target features, and the Bi-LSTM parameters between the two are not shared with each other.
[0107] When the initial multi-task model is trained using the second training sample 304, the target loss function value is used to determine whether the initial multi-task model is trained. The target loss function is,
[0108] Among them, Loss represents the loss value of the target loss function, loss1 represents the loss value of the entity recognition task, loss2 represents the loss value of the word segmentation task, Step represents the total number of times the initial multi-task model is trained, and i represents the current number of training times. The target loss function value is the entity recognition task function value and the word segmentation task loss function value. In the initial stage of training the initial multi-task model, because the solution space of the word segmentation task is smaller than the solution space of the entity recognition task, the word segmentation task loss accounts for the main part of the total loss. With the continuous training of the initial multi-task model, the entity recognition task loss accounts for the main part of the total loss. When the target loss function value reaches the threshold, it is determined that the initial multi-task model training is completed, thereby obtaining the target multi-task model.
[0109] Optionally, the fourth training module 1204 is specifically configured to:
[0110] The target training sample is input into the target multi-task model to obtain the word segmentation representation vector 403 output by the Bi-LSTM_p of the target multi-task model. At the same time, the character vector 402 output by the Pre_RoBerta model of the target multi-task model and the position representation vector 401 generated by the position encoding of the data by the Pre_RoBerta model itself are obtained.
[0111] Optionally, the fifth training module 1205 is specifically configured to:
[0112] The word segmentation representation vector 403, character vector 402 and position representation vector 401 output by the fourth training module 1204 are used as input for training the entity recognition model in the fifth training module 1205. After comprehensively obtaining the word segmentation representation vector 403, character vector 402 and position representation vector 401, the entity recognition model uses the Bert model of the entity recognition model as a feature extractor to extract features and characterize the input vector.
[0113] It should be noted that after the Bert model, the Bi-LSTM_n trained by the third training module 1203 can be added to accelerate the speed of training the entity recognition model and enhance the sensitivity of the entity recognition model to the temporal characteristics of the target sample.
[0114] It should also be noted that, after the Bi-LSTM_n in the above embodiment, a CRF model may be added, and the CRF model may be used to learn the dependency and constraint relationship between the second target features output by the Bi-LSTM_n.
[0115] According to another aspect of the embodiment of the present application, the embodiment of the present application further provides an electronic device, such as Fig.13 As shown, it includes a processor 1301, a communication interface 1302, a memory 1303 and a communication bus 1304, wherein the processor 1301, the communication interface 1302, and the memory 1303 communicate with each other through the communication bus 1304. The electronic device includes but is not limited to:
[0116] Memory 1303, used for storing computer programs;
[0117] The processor 1301 is used to execute the program stored in the memory 1303. When the processor 1301 executes the computer program stored in the memory 1303, the processor 1301 is used to execute the above-mentioned entity recognition model training method.
[0118] The memory 1303 is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs and non-transitory computer executable programs, such as the entity recognition model training method described in the embodiment of the present application. The processor 1301 implements the above-mentioned entity recognition model training method by running the non-transitory software program and instructions stored in the memory 1303.
[0119] The memory 1303 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store and execute the above-mentioned entity recognition model training method. In addition, the memory 1303 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1303 may optionally include a memory 1303 remotely arranged relative to the processor 1301, and these remote memories may be connected to the processor 1301 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0120] The non-transient software programs and instructions required to implement the above-mentioned entity recognition model training method are stored in the memory 1303. When executed by one or more processors 1301, the above-mentioned entity recognition model training method is executed, for example, Figure 2 The method steps S201 to S205 described in Figure 6 The method steps S601 and S602 described in Figure 7 The method steps S701 and S702 described in Figure 8 The method steps S801, S802 described in Fig. 9 The method steps S901 and S902 described in Fig.10 The method steps S1001 and S1002 described in Fig.11 Method steps S1101 to S1106 described in .
[0121] The memory and processor in the above electronic device communicate through the communication bus and the communication interface. The communication bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0122] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0123] An embodiment of the present invention further provides a storage medium storing computer executable instructions, wherein the computer executable instructions are used to execute the above-mentioned entity recognition model training method.
[0124] In one embodiment, the storage medium stores computer executable instructions, which are executed by one or more control processors 1301, for example, by a processor 1301 in the above electronic device, so that the one or more processors 1301 can execute the above entity recognition model training method, for example, execute Figure 2 The method steps S201 to S205 described in Figure 6 The method steps S601 and S602 described in Figure 7 The method steps S701 and S702 described in Figure 8 The method steps S801, S802 described in Fig. 9 The method steps S901 and S902 described in Fig.10 The method steps S1001 and S1002 described in Fig.11 Method steps S1101 to S1106 described in .
[0125] The above described embodiments are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0126] It is understood that the embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application, or a combination thereof.
[0127] For software implementation, the technology described herein can be implemented by a unit that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0128] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0129] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0130] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0131] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0132] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0133] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.
[0134] It should be noted that, in this article, relational terms such as "first", "second", "third", "fourth", "fifth", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0135] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for training an entity recognition model, characterized in that: include: Using the first training sample to train an initial model to obtain a pre-trained model, wherein the pre-trained model is used for natural language processing; Establishing an initial multi-task model according to the pre-trained model, wherein the initial multi-task model is used to perform word segmentation tasks and entity recognition tasks; The initial multi-task model is trained by using a target loss function and a second training sample to obtain a target multi-task model, wherein the target loss function is used to combine the word segmentation task and the entity recognition task; Input the target training sample into the target multi-task model to obtain the word segmentation representation vector output by the target multi-task model, as well as the word vector and position representation vector output by the pre-trained model in the target multi-task model; The entity recognition model is trained using the word segmentation representation vector, the character vector and the position representation vector to obtain a target model.
2. The entity recognition model training method according to claim 1, characterized in that: Establishing an initial multi-task model according to the pre-trained model includes: The pre-trained model is connected to the first task layer and the second task layer respectively to obtain an initial multi-task model, wherein the first task layer is used to perform the entity recognition task, and the second task layer is used to perform the word segmentation task.
3. The entity recognition model training method according to claim 2, characterized in that: The initial multi-task model is trained by using the target loss function and the second training sample to obtain the target multi-task model, including: Training the initial multi-task model using the second training sample; When the loss value of the target loss function reaches a threshold, it is determined that the training of the initial multi-task model is completed, and the target multi-task model is obtained; The objective loss function is, Among them, Loss represents the loss value of the target loss function, loss1 represents the loss value of the entity recognition task, loss2 represents the loss value of the word segmentation task, Step represents the total number of times the initial multi-task model is trained, and i represents the current number of training times.
4. The entity recognition model training method according to claim 3, characterized in that: The initial multi-task model is trained by using the target loss function and the second training sample, and further includes: Training the word segmentation task of the initial multi-task model by updating the second training sample; When the loss value of the word segmentation task meets the target loss threshold, the entity recognition task of the initial multi-task model is trained by updating the second training sample.
5. The entity recognition model training method according to any one of claims 2 to 4, characterized in that: The entity recognition model is trained using the word segmentation representation vector, the character vector and the position representation vector to obtain a target model including: Performing a sum operation on the word segmentation representation vector, the character vector and the position representation vector to obtain a target data set; The language layer of the entity recognition model is trained by the target data set to obtain the target model, and the language layer is used to extract features from the input target data set and output a first target feature.
6. The entity recognition model training method according to claim 5, characterized in that: The entity recognition model also includes: Connecting the language layer to the trained first task layer, and connecting the trained first task layer to the conditional random field layer; The first target feature is input into the first task layer that has been trained to obtain a second target feature after performing entity recognition on the first target feature, and the second target feature is input into the conditional random field layer to learn the dependency and constraint relationship between the second target features.
7. An entity recognition method, characterized in that: The entity recognition method comprises: performing entity recognition on a target sample using a target model obtained by the entity recognition model training method according to any one of claims 1 to 6.
8. An entity recognition model training device, characterized in that: The device comprises: A first training module uses the first training sample to train the initial model to obtain a pre-trained model; A second training module is used to establish an initial multi-task model based on the pre-trained model, where the initial multi-task model is used to perform word segmentation tasks and entity recognition tasks; A third training module trains the initial multi-task model through a target loss function and a second training sample to obtain a target multi-task model; A fourth training module, inputting the target training sample into the target multi-task model, obtaining the word segmentation representation vector output by the target multi-task model, and the word vector and position representation vector output by the pre-training model in the target multi-task model; The fifth training module uses the word segmentation representation vector, the character vector and the position representation vector to train an entity recognition model to obtain a target model.
9. An electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein: The processor, the communication interface and the memory communicate with each other through the communication bus, and is characterized in that the memory is used to store computer programs, and the processor is used to execute the programs stored in the memory to implement the steps of the entity recognition model training method described in any one of claims 1 to 6.
10. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, characterized in that: When the computer program is executed by a processor, the steps of the entity recognition model training method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Multi-type entity recognition multi-task deep learning model training method and device
CN108920460A
Named entity recognition method and named entity recognition model training method and device
CN109902307A