A training method, device, terminal and storage medium for a named entity model

Through the dual-model training method, iterative training and momentum update of pseudo-labels are used to use models with the same structure but different optimization algorithms to perform iterative training and momentum update of pseudo-labels, which solves the problem of low accuracy of named entity recognition models and achieves improvement of model recognition performance.

CN116070631BActive Publication Date: 2025-07-25ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211449338.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2025-07-25
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

The recognition accuracy of named entity recognition models in the prior art is low, mainly due to the complex structure of the neural network and the large parameter scale, requiring a large number of manual annotation training corpus, resulting in low marking quality and easy artificial errors.

Method used

The dual-model training method is adopted, and the first model and the second model with the same structure but different optimization algorithms are used as learners. By updating the momentum of the first model and iteratively training the error value of the pseudo-label, the pseudo-label noise problem is alleviated and the model recognition accuracy is improved.

Benefits of technology

It effectively alleviates the problem of pseudo-label noise, improves the recognition accuracy of named solid models, and improves the recognition performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116070631B_ABST
    Figure CN116070631B_ABST
Patent Text Reader

Abstract

The present invention provides a training method, device, terminal and storage medium for a named entity model. The training method of the named entity model includes: obtaining a training data set; performing label prediction on second text data through a first model to obtain second text data with pseudo-labels; performing entity prediction on first text data and second text data respectively through a second model to obtain prediction labels corresponding to the first text data and the second text data respectively; iteratively training the second model based on the error value between the prediction label and the annotation information of the same first text data and the error value between the prediction label and the pseudo-label of the same second text data; performing momentum update on the model parameters of the first model based on the model parameters of the trained second model, and using the trained first model as the named entity model. In this application, the two models correct each other, which can effectively alleviate the noise problem of pseudo-labels and improve the recognition accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing in artificial intelligence, and particularly to a training method, device, terminal and computer-readable storage medium for a named entity model. Background Art

[0002] Named Entity Recognition (NER), also known as "Proper Name Recognition", refers to identifying entities with specific meanings in text, mainly including personal names, place names, organization names, proper nouns, etc. It is a basic task in the field of natural language processing and belongs to the information extraction task. Specifically, it refers to extracting entities with specific meanings (or proper nouns, such as personal names, organization names, addresses, times, etc.) from unstructured text, and the entity is uniquely determined by the text position and entity type.

[0003] Currently, the neural network method based on deep learning is an effective method for solving named entity recognition. However, the neural network structure is complex and the parameter scale is large, requiring a large amount of manually annotated training corpus. Generally speaking, the more training corpus is annotated, the stronger the diversity, and the higher the annotation quality, the stronger the accuracy and generalization of the neural network model. This poses relatively high requirements for the domain knowledge and professionalism of the annotators, and a large number of repetitive annotation tasks are also prone to human errors, making it difficult to ensure the correctness and standardization of the annotation, resulting in low data quality. Summary of the Invention

[0004] The main technical problem to be solved by the present invention is to provide a training method, device, terminal and computer-readable storage medium for a named entity model, so as to solve the problem of low recognition accuracy of the model in the prior art.

[0005] To solve the above technical problems, the first technical solution adopted by the present invention is: to provide a training method for a named entity model, and the training method for the named entity model includes: obtaining a training data set, where the training data set includes an annotated text data set and an unannotated text data set; the annotated text data set includes multiple first text data with annotation information, and the unannotated text data set includes multiple unannotated second text data; training an initial parameter model according to the annotated text data set; performing label prediction on the second text data through a first model to obtain second text data with pseudo-labels; the model parameters of the first model are the same as those of the trained initial parameter model; performing entity prediction on the first text data and the second text data respectively through a second model to obtain prediction labels corresponding to the first text data and the second text data respectively; the model parameters of the second model are the same as those of the trained initial parameter model, and the second model has the same structure as the first model but different optimization algorithms for model parameters; iteratively training the second model based on the error value between the prediction label and the annotation information of the same first text data and the error value between the prediction label and the pseudo-label of the same second text data; performing momentum update on the model parameters of the first model based on the model parameters of the trained second model, and using the trained first model as the named entity model.

[0006] Among them, the model parameter learning algorithm of the second model is the stochastic gradient descent algorithm, and the model parameter learning algorithm of the first model is the momentum optimization method.

[0007] Among them, the initial parameter model includes a text encoding module and an entity recognition module, and the text encoding module is connected to the entity recognition module; training the initial parameter model according to the annotated text data set includes: detecting words in the first text data through the text encoding module to obtain vectors corresponding to each word in the first text data; performing entity prediction on the vectors corresponding to each word through the entity recognition module to obtain prediction information corresponding to each word; iteratively training the initial parameter model based on the error value between the annotation information and the prediction information corresponding to each word in the first text data.

[0008] Among them, performing label prediction on the second text data through the first model to obtain second text data with pseudo-labels includes: performing label prediction on each word in the second text data through the first model to obtain label information corresponding to each word in the second text data; determining the pseudo-label information of the second text data based on the label information corresponding to each word respectively.

[0009] Among them, the tag information includes a speculative tag and a speculative confidence level; based on the tag information respectively corresponding to each word, determining the pseudo-tag information of the second text data, including: determining a candidate entity corresponding to the second text data based on the speculative tag of each word; determining an entity confidence level corresponding to the candidate entity based on the speculative confidence level of the word corresponding to the candidate entity; in response to the entity confidence level of the candidate entity being greater than a preset confidence level, retaining the candidate entity and determining the speculative tag of the word corresponding to the candidate entity as the pseudo-tag information of the second text data.

[0010] Among them, the speculative tag has one of a first identifier and a second identifier; the first identifier represents the first word of the entity; the second identifier represents the other words of the entity; determining a candidate entity corresponding to the second text data based on the speculative tag of each word, including: determining the word corresponding to the first identifier as the starting field; in response to there being a next starting field after the current starting field, determining the word corresponding to the last second identifier among all the words after the current starting field and before the next starting field as the ending field; in response to the current starting field being the last starting field, determining the word corresponding to the last second identifier after the current starting field as the ending field; determining the starting field, the corresponding ending field, and the words therebetween as a candidate entity corresponding to the second text data.

[0011] Among them, iteratively training the second model based on the error value between the predicted tag and the annotation information of the same first text data and the error value between the predicted tag and the pseudo-tag of the same second text data, including: iteratively training the second model based on the error value between the predicted tag and the annotation information of each word in the same first text data and the error value between the speculative tag and the predicted tag of each word corresponding to the retained candidate entity in the same second text data.

[0012] Among them, performing momentum update on the model parameters of the first model based on the model parameters of the trained second model, and using the trained first model as the named entity model, including: determining the model parameters of the second model after the current iterative training based on the model parameters of the second model before the current iterative training and the error value corresponding to the second model during the current iterative training; determining the model parameters of the first model after the current iterative training based on the model parameters of the second model after the current iterative training, the model parameters of the first model before the current iterative training, and a preset momentum coefficient; in response to the number of iterations corresponding to the first model reaching a preset number, using the first model after training completion as the named entity model.

[0013] To solve the above technical problems, the second technical solution adopted by the present invention is: to provide a training device for a named entity model, including: an acquisition module for acquiring a training data set, the training data set including an annotated text data set and an unannotated text data set; the annotated text data set includes a plurality of first text data with annotation information, and the unannotated text data set includes a plurality of unannotated second text data; a first training module for training an initial parameter model according to the annotated text data set; a first prediction module for performing label prediction on the second text data through a first model to obtain second text data with pseudo-labels; the model parameters of the first model are the same as those of the trained initial parameter model; a second prediction module for performing entity prediction on the first text data and the second text data respectively through a second model to obtain prediction labels corresponding to the first text data and the second text data respectively; the model parameters of the second model are the same as those of the trained initial parameter model, and the second model has the same structure as the first model but different optimization algorithms for model parameters; a second training module for iteratively training the second model based on the error value between the prediction label and the annotation information of the same first text data and the error value between the prediction label and the pseudo-label of the same second text data; an update module for performing momentum update on the model parameters of the first model based on the model parameters of the trained second model, and using the trained first model as the named entity model.

[0014] To solve the above technical problems, the third technical solution adopted by the present invention is: to provide a terminal, the terminal includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor is used to execute program data to implement the steps in the above-named entity model training method.

[0015] To solve the above technical problems, the fourth technical solution adopted by the present invention is: to provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps in the above-named entity model training method.

[0016] The beneficial effects of the present invention are as follows: Different from the prior art, a training method, device, terminal, and storage medium for a named entity model are provided. The training method for the named entity model includes: obtaining a training data set, which includes an annotated text data set and an unannotated text data set; the annotated text data set includes multiple first text data with annotation information, and the unannotated text data set includes multiple unannotated second text data; training an initial parameter model according to the annotated text data set; performing label prediction on the second text data through a first model to obtain second text data with pseudo-labels; the model parameters of the first model are the same as those of the trained initial parameter model; performing entity prediction on the first text data and the second text data respectively through a second model to obtain prediction labels corresponding to the first text data and the second text data respectively; the model parameters of the second model are the same as those of the trained initial parameter model, and the second model has the same structure as the first model but different model parameter optimization algorithms; iteratively training the second model based on the error value between the prediction label and the annotation information of the same first text data and the error value between the prediction label and the pseudo-label of the same second text data; performing momentum update on the model parameters of the first model based on the model parameters of the trained second model, and using the trained first model as the named entity model. In this application, the first model and the second model with the same structure but different model parameter optimization algorithms are used as learners respectively, and the two models correct each other, which can effectively alleviate the noise problem of pseudo-labels, thereby improving the accuracy of pseudo-labels. Based on the pseudo-labels with high accuracy, the recognition accuracy of the model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0018] Figure 1 is a flowchart of the training method for the named entity model provided by the present invention;

[0019] Figure 2 is a flowchart of a specific embodiment of the training method for the named entity model provided by the present invention;

[0020] Figure 3 is a flowchart of an embodiment of the training method for the named entity model provided by the present invention;

[0021] Figure 4 is Figure 1 a flowchart of a specific embodiment of step S2 in the training method for the named entity model provided

[0022] Figure 5 is Figure 1 A schematic flowchart of a specific embodiment of step S3 in the training method of the provided named entity model;

[0023] Figure 6 is Figure 5 A schematic flowchart of a specific embodiment of step S32 in the training method of the provided named entity model;

[0024] Figure 7 is Figure 1 A schematic flowchart of a specific embodiment of step S6 in the training method of the provided named entity model;

[0025] Figure 8 A schematic framework diagram of an embodiment of the training device for the named entity model provided by the present invention;

[0026] Figure 9 A schematic framework diagram of an embodiment of the terminal provided by the present application;

[0027] Figure 10 A schematic framework diagram of an embodiment of the computer-readable storage medium provided by the present application. Specific Embodiments

[0028] The following describes the solutions of the embodiments of the present application in detail with reference to the accompanying drawings of the specification.

[0029] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0030] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.

[0031] To enable those skilled in the art to better understand the technical solutions of the present invention, the following further describes in detail a training method for a named entity model provided by the present invention in combination with the drawings and specific embodiments.

[0032] Please refer to Figure 1 、 2 、3, Figure 1 A schematic flowchart of the training method for the named entity model provided by the present invention; Figure 2 A schematic flowchart of a specific embodiment of the training method for the named entity model provided by the present invention; Figure 3It is a schematic flowchart of an embodiment of the training method of the named entity model provided by the present invention.

[0033] In this embodiment, a training method of a named entity model is provided. The training method of the named entity model includes the following steps.

[0034] S1: Obtain a training data set.

[0035] Specifically, the training data set includes an annotated text data set and an unannotated text data set; the annotated text data set includes a plurality of first text data with annotation information, and the unannotated text data set includes a plurality of unannotated second text data.

[0036] In this embodiment, text data in a vertical domain is obtained on the Internet or other channels. Based on the field to which the named entity model is to be applied, text data in the corresponding professional field is selected.

[0037] After obtaining the text data, the text data is filtered to delete illegal characters, whitespace characters, etc. Among them, illegal characters refer to characters that do not normally appear in the text. For example, control characters in ASCII code, these characters will not be displayed in the text, but will indeed occupy a character position and will affect the word segmentation algorithm.

[0038] Some of the filtered text data is manually annotated to obtain first text data with annotation information, and the unannotated text data is used as second text data. Among them, all the first text data with annotation data constitutes the annotated text data set. All the unannotated second text data constitutes the unannotated text data set.

[0039] In a specific embodiment, the named entity model needs to be applied to the scenario of judicial cases, where data needs to be extracted and entities such as time, location, name, WeChat, phone number, and items in the case need to be labeled. For example, the first text data can be "The complainant Chen Mouhai came to the police station by himself. After understanding, the WeChat ID (WXID_1521****336) of Chen Mouhai's mobile phone was stolen by someone", and the annotation information of the first text data is "O O O B-name I-name I-name O O O O O O O O O O O O O B-name I-name I-name O O O O O O O B-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat I-WeChat O O O O O". Among them, the first character and other characters of the entity are distinguished by the prefixes "B-" and "I-" to facilitate determining the number of entities in the text data. Specifically, "B-name" is the annotation of the first character of the name entity, and "I-name" is the annotation of other characters of the name entity. According to the annotation information, the number of entities and the entity set of the first text data can be determined. For example, the number of entities in the first text data is three, and the entity set is {"Chen Mouhai", "Chen Mouhai", "WXID_1521****336"}.

[0040] The second text data can be "On December 3, 2021, Xu Mouxing and Lin Mouhu entered Wang's house and stole Wang's table lamp." The second text data does not have annotation information.

[0041] S2: Train the initial parameter model according to the annotated text data set.

[0042] Specifically, the initial parameter model includes a text encoding module and an entity recognition module, and the text encoding module is connected to the entity recognition module. Among them, the text encoding module is a pre-trained bidirectional encoding model (Bidirectional Encoder Representation from Transformers, BERT), and the pre-trained bidirectional encoding model is a pre-trained language representation model. The entity recognition module is a Conditional Random Field (CRF), and CRF is a sequential labeling algorithm. The initial parameter model is constructed by the pre-trained bidirectional encoding model and the conditional random field.

[0043] In a specific embodiment, the initial parameter model is trained with the first text data with annotated text to obtain an initial parameter model with good convergence effect.

[0044] Please refer to Figure 4 , Figure 4 is Figure 1Flow diagram of a specific embodiment of step S2 in the provided training method of the named entity model.

[0045] S21: Detect words in the first text data through a text encoding module to obtain vectors corresponding to each word in the first text data.

[0046] Specifically, input the first text data into the text encoding module, and the text encoding module performs text detection on the text sequence composed of the first text data to obtain vectors corresponding to each word in the first text data. In this embodiment, the text encoding module detects each word text in the first text data to obtain a sequence of word vectors corresponding to the word text sequence.

[0047] S22: Perform entity prediction on the vectors corresponding to each word through an entity recognition module to obtain prediction information corresponding to each word.

[0048] Specifically, input the sequence of word vectors corresponding to the first text data into the entity recognition module, and the entity recognition module performs entity recognition on the words corresponding to the word vectors, thereby obtaining the predicted category and category confidence corresponding to each word. Among them, the prediction information includes the predicted category and category confidence.

[0049] S23: Iteratively train the initial parameter model based on the error value between the annotation information and the prediction information corresponding to each word in the first text data.

[0050] Specifically, based on the error value between the annotation information and the prediction information of each word in the same first text data, determine the error value between the annotation information and the prediction information of the first text data, and then train the initial parameter model according to the error value between the annotation information and the prediction information of the first text data.

[0051] In an alternative embodiment, the result of the initial parameter model is backpropagated, and the weights of the initial parameter model are corrected according to the error value between the annotation information and the prediction information corresponding to the same first text data, so as to realize the training of the initial parameter model.

[0052] Input the first text data into the initial parameter model, and the initial parameter model analyzes the first text data. When the error value between the annotation information and the prediction information corresponding to the same first text data is less than a preset threshold, the preset threshold can be set by oneself, such as 1%, 5%, etc., then stop training the initial parameter model.

[0053] In another embodiment, in response to the number of iterations of the initial parameter model reaching a preset number, the training of the initial parameter model can be stopped.

[0054] In one embodiment, a first model and a second model are constructed based on a text encoding module and an entity recognition module. That is to say, both the first model and the second model include a text encoding module and an entity recognition module connected to each other. Among them, the model structures of the first model and the second model are the same as the model structure of the initial parameter model. The first model and the second model are initialized based on the model parameters of the trained initial parameter model. That is to say, the model parameters of the first model and the model parameters of the second model are both consistent with the model parameters of the initial parameter model. That is, the model parameters of the first model are denoted as θ m , and the model parameters of the second model are denoted as θ s . However, the optimization algorithm for the model parameters of the first model is different from the optimization algorithm for the model parameters of the second model. In this embodiment, the model parameter learning algorithm of the second model is the stochastic gradient descent algorithm. Therefore, the second model can also be called a gradient update model. The model parameter learning algorithm of the first model is the momentum optimization method. Therefore, the first model can also be called a momentum update model.

[0055] S3: Use the first model to perform label prediction on the second text data to obtain the second text data with pseudo-labels.

[0056] Specifically, performing entity prediction on the second text data through the first model specifically includes the following steps.

[0057] Please refer to Figure 5 , Figure 5 which is Figure 1 a schematic flowchart of a specific embodiment of step S3 in the training method of the named entity model provided.

[0058] S31: Use the first model to perform label prediction on each word in the second text data to obtain the label information corresponding to each word in the second text data.

[0059] Specifically, since the model parameters of the first model are trained based on the first text data with annotations and have reference value. Therefore, it is possible to predict whether each word in the second text data is an entity based on the first model.

[0060] In one embodiment, the second text data is input into the first model, and the text encoding module in the first model performs text detection on the second text data to obtain the vector corresponding to each word in the second text data. In this embodiment, the text encoding module detects each word text in the second text data to obtain the word vector sequence corresponding to the word text sequence.

[0061] Input the word vector sequence corresponding to the second text data into the entity recognition module in the first model. The entity recognition module performs entity recognition on the words corresponding to the word vectors, and then obtains the label information corresponding to each word.

[0062] In a specific embodiment, the label information includes a speculative label and a speculative confidence level. Among them, the value range of the speculative confidence level is [0, 1]. The speculative confidence level represents the credibility of the speculative label corresponding to the word. If the speculative confidence level is relatively low, it indicates that the possibility of the speculative label corresponding to the word being incorrect is relatively high. Therefore, the speculative confidence level with a higher value can be used as the pseudo-label of the second text data to improve the accuracy of the pseudo-label of the second text data.

[0063] In a specific embodiment, the second text data can be "On December 3, 2021, Xu Mouxing and Lin Mouhu entered Wang's house and stole Wang's table lamp.", then the label information of each word in the second text data detected by the first model is represented in the form of "(character, speculative label, speculative confidence level)". For example, the label information of each word in the second text data is (2, B-time, 0.95), (0, I-time, 0.94), (2, I-time, 0.93), (1, I-time, 0.96), (year, I-time, 0.95), (1, I-time, 0.97), (2, I-time, 0.95), (month, I-time, 0.91), (3, I-time, 0.9), (number, I-time, 0.91), (,, O, 1.0), (Xu, B-name, 0.95), (Mou, I-name, 0.95), (Xing, I-name, 0.95), (partner, O, 1.0), (with, O, 1.0), (Lin, B-name, 0.95), (Mou, I-name, 0.95), (Hu, I-name, 0.95), (enter, O, 1.0), (into, O, 1.0), (Wang, B-location, 0.4), (Mou, I-location, 0.5), (house, I-location, 0.7), (,, O, 1.0), (steal, O, 1.0), (theft, O, 1.0), (, O, 1.0), (Wang, B-name, 0.93), (Mou, I-name, 0.94), (of, O, 1.0), (table, B-item, 0.6), (lamp, I-item, 0.7), (., O, 1.0).

[0064] S32: Determine the pseudo-label information of the second text data based on the label information corresponding to each word.

[0065] Please refer to Figure 6 , Figure 6 is Figure 5 a schematic flowchart of a specific embodiment of step S32 in the training method of the named entity model provided.

[0066] S321: Determine the candidate entities corresponding to the second text data based on the speculative tags of each word.

[0067] Specifically, the speculative tag has one of a first identifier and a second identifier; the first identifier represents the first word of the entity; the second identifier represents the other words of the entity. In this embodiment, the first identifier is the prefix "B-", and the second identifier is the prefix "I-". In other embodiments, other prefixes may also be used.

[0068] Determine the word corresponding to the first identifier as the starting field; in response to there being another starting field after the current starting field, determine the word corresponding to the last second identifier among all the words after the current starting field and before the next starting field as the ending field; in response to the current starting field being the last starting field, determine the word corresponding to the last second identifier after the current starting field as the ending field; determine the starting field, the corresponding ending field, and the words therebetween as a candidate entity corresponding to the second text data.

[0069] In a specific embodiment, the data form of the tag information corresponding to each word in the second text data is similar to the data form of the annotation information corresponding to each word in the first text data. That is, the tag information in the second text data also distinguishes the first character and other characters of the entity through the prefixes "B-" and "I-", so as to facilitate determining the number of entities in the text data. Therefore, it can be determined that the number of entities in the above embodiment of the second text data is 6 according to the number of the prefix "B-" in the tag information corresponding to each word in the second text data, and the candidate entities corresponding to the second text data are respectively "December 3, 2021", "Xu Mouxing", "Lin Mouhu", "Wang Moujia", "Wang Mou", and "table lamp".

[0070] S322: Determine the entity confidence corresponding to the candidate entity based on the speculative confidence of the words corresponding to the candidate entity.

[0071] Specifically, calculate the average value of the speculative confidences corresponding to all the words included in each entity, and determine the average confidence as the entity confidence corresponding to the entity.

[0072] In a specific embodiment, the entity confidences corresponding to the candidate entities of the second text data are respectively: (December 3, 2021, time, 0.94), (Xu Mouxing, name, 0.95), (Lin Mouhu, name, 0.95), (Wang Moujia, name, 0.53), (Wang Mou, name, 0.94), (table lamp, item, 0.65).

[0073] S323: If the entity confidence of the candidate entity is greater than the preset confidence, retain the candidate entity and determine the speculative label of the word corresponding to the candidate entity as the pseudo-label information of the second text data.

[0074] Specifically, to improve the accuracy of the pseudo-labels of the second text data, the speculative labels of the candidate entities with high entity confidence and the speculative labels of the words that have not been determined as entities with relatively high speculative confidence are used as the pseudo-labels of the second text data.

[0075] In one embodiment, the speculative label of the word corresponding to the candidate entity whose entity confidence is greater than the preset confidence is determined as the pseudo-label data of the second text data.

[0076] In one specific embodiment, the preset confidence can be set to 0.9. Since the entity confidences of the candidate entities "table lamp" and "Wang's house" corresponding to the second text data are lower than the preset confidence, the speculative labels of the words corresponding to "December 3, 2021", "Xu Xing", "Lin Lake", and "Wang" are determined as the pseudo-label data of the second text data. Further, the speculative labels of the words that have not been determined as candidate entities and whose speculative confidence is greater than the preset confidence are also determined as the pseudo-label data of the second text data.

[0077] S4: Use the second model to perform entity prediction on the first text data and the second text data respectively to obtain the prediction labels corresponding to the first text data and the second text data respectively.

[0078] Specifically, input the first text data with annotation information and the second text data with pseudo-label data into the second model, and use the second model to predict each corresponding word in the first text data and the second text data to obtain the prediction labels of each word.

[0079] S5: Iteratively train the second model based on the error value between the prediction label and the annotation information of the same first text data and the error value between the prediction label and the pseudo-label of the same second text data.

[0080] Specifically, iteratively train the second model based on the error value between the prediction label and the annotation information of each word in the same first text data and the error value between the speculative label and the prediction label of each word corresponding to the retained candidate entities in the same second text data.

[0081] In a specific embodiment, a second model is trained based on the error value between the predicted label corresponding to each word in the first text data and the annotation information. The second model is trained based on the error value between the predicted label of the word with the pseudo label in the second text data and the pseudo label data. That is, the second model is trained based on the error value between the predicted labels of the words corresponding to "December 3, 2021", "Xu Mouxing", "Lin Mouhu", and "Wang Mou" in the second text data and the predicted labels of the words that have not been confirmed as candidate entities and the pseudo label data.

[0082] S6: Perform momentum update on the model parameters of the first model based on the model parameters of the trained second model, and use the trained first model as the named entity model.

[0083] Please refer to Figure 7 , Figure 7 is Figure 1 a schematic flowchart of a specific embodiment of step S6 in the provided training method of the named entity model.

[0084] S61: Determine the model parameters of the second model after the current iterative training based on the model parameters of the second model before the current iterative training and the error value corresponding to the second model during the current iterative training.

[0085] Specifically, based on the loss value of the entity recognition module in the second model during the current iterative training and the corresponding gradient during the current iterative training Update the model parameters once using the gradient descent method to obtain the model parameters of the second model after the current iterative training.

[0086] In a specific embodiment, the model parameters of the second model after the current iterative training are obtained based on the following formula.

[0087]

[0088] In the formula: represents the model parameters of the second model after the current iterative training; represents the model parameters of the second model before the current iterative training; η represents the learning rate.

[0089] Among them, η is generally set to a relatively small value, such as it can be selected within the range of 1e-5 to 1e-3, so as to prevent the second model from oscillating and diverging.

[0090] S62: Determine the model parameters of the first model after the current iterative training based on the model parameters of the second model after the current iterative training, the model parameters of the first model before the current iterative training, and the preset momentum coefficient.

[0091] Specifically, update the model parameters of the first model according to the model parameters of the second model after the current iterative training.

[0092] In one embodiment, the model parameters of the first model after the current iterative training are obtained based on the following formula.

[0093] θ m t ←αθ m t-1 +(1-α)θ s t (Formula 2)

[0094] In the formula: is the model parameter of the first model after the current iterative training; α is the momentum coefficient; is the model parameter of the first model before the current iterative training; is the model parameter of the second model after the current iterative training.

[0095] Among them, the value range of α is [0, 1], and generally a relatively large value is set, such as 0.999, to prevent the parameters from changing too significantly before and after the update of the first model, resulting in unstable pseudo-labels.

[0096] In one embodiment, the model parameters of the first model are iteratively updated based on the method in the above steps S3 to S6.

[0097] S63: In response to the iteration count corresponding to the first model reaching the preset count, the first model after training is used as the named entity model.

[0098] Specifically, when the iteration count corresponding to the first model reaches the preset count, or the number of entities of the predicted entity in the second text data is in a stable state, that is, the number of entities corresponding to the same second text data obtained by multiple detections remains consistent, the training of the first model and the second model is stopped, and the trained first model is used as the model entity model.

[0099] Since the first model can be regarded as an integration of multiple second models (gradient update models) during the iteration process, generally the performance will be better. Therefore, after the training is completed, the first model (momentum update model) is output as the result, that is, only the trained first model is used in the inference stage, which does not affect the inference performance.

[0100] In this embodiment, during the iteration, the parameters are updated based on both the first model and the second model, and the model performance and the accuracy of the pseudo-labels are gradually improved through iteration. The stability of the pseudo-label data can be guaranteed during the process.

[0101] The training method of the named entity model provided in this embodiment includes: obtaining a training data set, where the training data set includes an annotated text data set and an unannotated text data set; the annotated text data set includes multiple first text data with annotation information, and the unannotated text data set includes multiple unannotated second text data; training an initial parameter model according to the annotated text data set; predicting labels for the second text data through a first model to obtain second text data with pseudo-labels; the model parameters of the first model are the same as those of the trained initial parameter model; predicting entities for the first text data and the second text data respectively through a second model to obtain predicted labels corresponding to the first text data and the second text data respectively; the model parameters of the second model are the same as those of the trained initial parameter model, and the second model has the same structure as the first model but different optimization algorithms for model parameters; iteratively training the second model based on the error value between the predicted label and the annotation information of the same first text data and the error value between the predicted label and the pseudo-label of the same second text data; performing momentum update on the model parameters of the first model based on the model parameters of the trained second model, and using the trained first model as the named entity model. In this application, by using the first model and the second model with the same structure but different optimization algorithms for model parameters as learners respectively, the two models correct each other, which can effectively alleviate the noise problem of pseudo-labels, thereby improving the accuracy of pseudo-labels. Based on the pseudo-labels with high accuracy, the recognition accuracy of the model can be improved.

[0102] Refer to Figure 8 , Figure 8 FIG. is a schematic framework diagram of an embodiment of a training device for a named entity model provided by the present invention. This embodiment provides a training device 60 for a named entity model. The training device 60 for a named entity model includes an acquisition module 61, a first training module 62, a first prediction module 63, a second prediction module 64, a second training module 65, and an update module 66.

[0103] The acquisition module 61 is used to obtain a training data set, where the training data set includes an annotated text data set and an unannotated text data set; the annotated text data set includes multiple first text data with annotation information, and the unannotated text data set includes multiple unannotated second text data.

[0104] The first training module 62 is used to train an initial parameter model according to the annotated text data set.

[0105] The first prediction module 63 is used to predict labels for the second text data through a first model to obtain second text data with pseudo-labels; the model parameters of the first model are the same as those of the trained initial parameter model.

[0106] The second prediction module 64 is configured to perform entity prediction on the first text data and the second text data respectively through a second model, and obtain prediction labels corresponding to the first text data and the second text data respectively; the model parameters of the second model are the same as those of the initially trained parameter model, and the second model has the same structure as the first model, but different optimization algorithms for the model parameters.

[0107] The second training module 65 is configured to iteratively train the second model based on the error value between the prediction label and the annotation information of the same first text data and the error value between the prediction label and the pseudo-label of the same second text data.

[0108] The update module 66 is configured to perform momentum update on the model parameters of the first model based on the model parameters of the trained second model, and use the trained first model as the named entity model.

[0109] In the training device of the named entity model provided in this embodiment, the first model and the second model with the same structure but different optimization algorithms for the model parameters are respectively used as learners, and the two models correct each other, which can effectively alleviate the noise problem of pseudo-labels, and then improve the accuracy of pseudo-labels. Based on the pseudo-labels with high accuracy, the recognition accuracy of the model can be improved.

[0110] Please refer to Figure 9 , Figure 9 FIG. is a schematic framework diagram of an embodiment of a terminal provided in the present application. The terminal 80 includes a memory 81 and a processor 82 which are coupled to each other. The processor 82 is configured to execute program instructions stored in the memory 81 to implement the steps of any of the above-mentioned embodiments of the training method of the named entity model. In a specific implementation scenario, the terminal 80 may include, but is not limited to: a microcomputer, a server. In addition, the terminal 80 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.

[0111] Specifically, the processor 82 is used to control itself and the memory 81 to implement the steps of the method embodiments of any of the above-named entity models. The processor 82 may also be referred to as a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip with the ability to process signals. The processor 82 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 82 may be implemented jointly by integrated circuit chips.

[0112] Please refer to Figure 10 , Figure 10 which is a schematic framework diagram of an embodiment of the computer-readable storage medium provided by this application. The computer-readable storage medium 90 stores program instructions 901 that can be run by a processor, and the program instructions 901 are used to implement the steps of the method embodiments of any of the above-named entity models.

[0113] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the method embodiments above. The specific implementation can refer to the description of the method embodiments above. For the sake of brevity, it will not be elaborated here.

[0114] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, they will not be elaborated in this article.

[0115] In several embodiments provided by this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0116] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0117] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0118] If the technical solution of the present application involves personal information, before the product applying the technical solution of the present application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of the present application involves sensitive personal information, before the product applying the technical solution of the present application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is regarded as consenting to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0119] The above are only the embodiments of the present invention, and do not limit the patent protection scope of the present invention accordingly. All equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present invention.

Claims

1. A training method for a named entity model, characterized in that, The training method of the named entity model includes: Obtain a training data set, where the training data set includes an annotated text data set and an unannotated text data set; the annotated text data set includes multiple first text data with annotation information, and the unannotated text data set includes multiple unannotated second text data; Train an initial parameter model according to the annotated text data set; the initial parameter model includes a text encoding module and an entity recognition module, and the text encoding module is connected to the entity recognition module; Perform label prediction on the second text data through a first model to obtain the second text data with pseudo-labels; the model parameters of the first model are the same as those of the trained initial parameter model; Perform entity prediction on the first text data and the second text data respectively through a second model to obtain the predicted labels corresponding to the first text data and the second text data respectively; the model parameters of the second model are the same as those of the trained initial parameter model, the second model has the same structure as the first model, and the optimization algorithms of the model parameters are different; the model parameter learning algorithm of the second model is the stochastic gradient descent algorithm, and the model parameter learning algorithm of the first model is the momentum optimization method; Iteratively train the second model based on the error value between the predicted label and the annotation information of the same first text data and the error value between the predicted label and the pseudo-label of the same second text data; Perform momentum update on the model parameters of the first model based on the model parameters of the trained second model, and use the trained first model as the named entity model.

2. The training method of the named entity model according to claim 1, wherein The training of the initial parameter model according to the annotated text data set includes: Detect words in the first text data through the text encoding module to obtain vectors corresponding to each word in the first text data; Perform entity prediction on the vectors corresponding to each word through the entity recognition module to obtain prediction information corresponding to each word; Iteratively train the initial parameter model based on the error value between the annotation information and the prediction information corresponding to each word in the first text data.

3. The training method of the named entity model according to claim 1, wherein The performing label prediction on the second text data through a first model to obtain the second text data with pseudo-labels includes: Perform label prediction on each word in the second text data through the first model to obtain label information corresponding to each word in the second text data; Determine the pseudo-label information of the second text data based on the label information corresponding to each word respectively.

4. The training method of the named entity model according to claim 3, characterized in that The label information includes a speculated label and a speculated confidence level; The determining the pseudo-label information of the second text data based on the label information corresponding to each word respectively includes: Determine candidate entities corresponding to the second text data based on the speculative tags of each of the words; Determine the entity confidence corresponding to the candidate entity based on the speculative confidence of the words corresponding to the candidate entity; In response to the entity confidence of the candidate entity being greater than a preset confidence, retain the candidate entity and determine the speculative tag of the word corresponding to the candidate entity as the pseudo-tag information of the second text data.

5. The training method of the named entity model according to claim 4, wherein The speculative tag has one of a first identifier and a second identifier; the first identifier represents the first word of the entity; The second identifier represents other words of the entity; The determining of the candidate entities corresponding to the second text data based on the speculative tags of each of the words includes: Determine the word corresponding to the first identifier as the starting field; In response to there being a next starting field after the current starting field, determine the word corresponding to the last second identifier among all the words after the current starting field and before the next starting field as the ending field; In response to the current starting field being the last starting field, determine the word corresponding to the last second identifier after the current starting field as the ending field; Determine the starting field, the corresponding ending field, and the words therebetween as one candidate entity corresponding to the second text data.

6. The training method of the named entity model according to claim 4, wherein The iterative training of the second model based on the error value between the prediction tag and the annotation information of the same first text data and the error value between the prediction tag and the pseudo-tag of the same second text data includes: Iteratively train the second model based on the error value between the prediction tag corresponding to each word in the same first text data and the annotation information, and the error value between the speculative tag and the prediction tag of each word corresponding to the candidate entity retained in the same second text data.

7. The training method of the named entity model according to claim 1, wherein The momentum update of the model parameters of the first model based on the model parameters of the trained second model, and taking the trained first model as the named entity model includes: Calculate the parameter gradient and update based on the model parameters of the second model before the current iterative training and the error value corresponding to the second model during the current iterative training, to obtain the model parameters of the second model after the current iterative training; Determine the model parameters of the first model after the current iterative training based on the model parameters of the second model after the current iterative training, the model parameters of the first model before the current iterative training, and a preset momentum coefficient; In response to the number of iterations corresponding to the first model reaching a preset number, take the trained first model as the named entity model.

8. A training device for a named entity model, characterized in that The training apparatus of the named entity model includes: An acquisition module, configured to acquire a training data set, where the training data set includes an annotated text data set and an unannotated text data set; the annotated text data set includes a plurality of first text data with annotation information, and the unannotated text data set includes a plurality of unannotated second text data; A first training module, configured to train an initial parameter model according to the annotated text data set; the initial parameter model includes a text encoding module and an entity recognition module, and the text encoding module is connected to the entity recognition module; A first prediction module, configured to perform label prediction on the second text data through a first model to obtain the second text data with pseudo-labels; the model parameters of the first model are the same as those of the trained initial parameter model; A second prediction module, configured to perform entity prediction on the first text data and the second text data respectively through a second model to obtain prediction labels corresponding to the first text data and the second text data respectively; the model parameters of the second model are the same as those of the trained initial parameter model, and the second model has the same structure as the first model and different optimization algorithms for the model parameters; the model parameter learning algorithm of the second model is the stochastic gradient descent algorithm, and the model parameter learning algorithm of the first model is the momentum optimization method; A second training module, configured to iteratively train the second model based on the error value between the prediction label of the same first text data and the annotation information and the error value between the prediction label of the same second text data and the pseudo-label; An update module, configured to perform momentum update on the model parameters of the first model based on the model parameters of the trained second model, and use the trained first model as the named entity model.

9. A terminal, characterized in that, The terminal includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor is configured to execute program data to implement the steps in the training method of the named entity model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps in the training method of the named entity model according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Named entity identification method and system based on active learning

    CN113919358A

  • Named entity identification method and device, electronic equipment and storage medium

    CN115099237A