Method, device and electronic equipment for training entity marking model

By training the entity marking model, generating pseudo labels and re-noting, and finally re-training the model, the low prediction accuracy problem under the influence of noise data in the prior art is solved, and higher prediction accuracy and broader label error denoising capabilities are achieved.

CN113971183BActive Publication Date: 2025-05-06ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010710014.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-22
Publication Date
2025-05-06
Estimated Expiration
2040-07-22

AI Technical Summary

Technical Problem

When existing entity marking models process noise data in remote supervision construct data, the prediction accuracy is low, and commonly used assumptions are defective, resulting in poor final prediction results.

Method used

By using training data containing noise labels to train the entity marking model, generate pseudo labels, and re-label them, and finally re-train the model using the re-labeled data.

Benefits of technology

This method can effectively reduce the negative impact of noise data on the model, improve the prediction accuracy of the model, and is scalable, and can handle a wider range of label errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971183B_ABST
    Figure CN113971183B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a training method, device and electronic device for an entity labeling model, the method comprising: obtaining an original training sample set, and using the original training sample set to train the entity labeling model to establish a first entity labeling model; using the first entity labeling model to predict the label distribution of entities in the training sample; using the original label distribution of the entities in the training sample and the correct information contained in the predicted label distribution, re-labeling the entities in the training sample, and re-training the entity labeling model based on the re-labeled training sample to establish a second entity labeling model. Through the embodiment of the present application, the prediction accuracy of the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of entity marking technology, and in particular to a training method, device and electronic equipment for an entity marking model. Background Art

[0002] The entity labeling task refers to labeling an entity with a corresponding label based on a pre-set label set given the entity and the context in which the entity is located. Entity labeling is an important subtask of natural language processing information extraction. For example, the entity "Liu" in the sentence "Liu is a Hong Kong singer" can be labeled as "person / singer / Hong Kong singer". The labeled entity labels can be applied to many downstream tasks, such as named entity recognition, relationship extraction, knowledge graph expansion, etc.

[0003] In practical applications, entity labeling is usually considered a classification problem based on a pre-set set of labels. With the rise of deep neural networks, deep learning representation models represented by Bi-LSTM (Long Short-Term Memory, Bi-LSTM refers to the combination of forward LSTM and backward LSTM) have achieved good results in entity labeling tasks.

[0004] Existing entity labeling models are generally trained based on datasets constructed with remote supervision. The remote supervision method links the entities to be labeled with the entities in the knowledge graph, and uses all the label sets corresponding to the entities in the knowledge graph as the label sets for remote supervision labeling. Although the remote supervision labeling method is very efficient and can theoretically construct an infinite amount of labeled data, the data based on remote supervision labeling usually contains a large amount of incorrectly labeled data (noise data), which is bound to affect the accuracy of the model. Therefore, how to better reduce the negative impact of noise data on the model is of great significance to improving the prediction accuracy of the model.

[0005] At present, the methods for processing noisy data in remote supervision construction data in entity labeling tasks generally make the following assumptions: training data containing only one label must be correct, and data sets containing multiple labels contain noise. Based on this assumption, one type of method is to process these two types of data sets separately, for example, treating data with multiple labels as regular data. In addition, there are also some schemes to re-label data sets with multiple labels, that is, to select a label as the correct label in the multi-label set. However, the main problem with the above scheme is that the pre-proposed assumptions are flawed. Since the data in the knowledge graph may be wrong, or there may be problems in the linking, the training data with only one label is not necessarily correct, and the label set of the data set containing multiple labels may not necessarily contain the correct label. Therefore, the above scheme is still flawed in the final prediction effect.

[0006] Therefore, how to further improve the prediction accuracy of the model has become a technical problem that needs to be solved by technical personnel in this field. Summary of the invention

[0007] The present application provides a method, device and electronic equipment for training an entity labeling model, which can improve the prediction accuracy of the model.

[0008] This application provides the following solutions:

[0009] A training method for an entity labeling model, comprising:

[0010] Using training data containing noise labels to train an entity labeling model, and using the trained entity labeling model to predict label distribution of entities in the training data;

[0011] Obtaining a pseudo label of the entity according to the original label distribution of the entity included in the training data and the label distribution predicted by the entity labeling model;

[0012] The entities in the training data are relabeled using the pseudo labels, and the entity labeling model is trained using the relabeled training data to obtain a trained entity labeling model.

[0013] A training method for an entity labeling model, comprising:

[0014] Obtain an original training sample set, and use the original training sample set to train an entity labeling model to establish a first entity labeling model; wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity;

[0015] Using the first entity labeling model to predict label distribution of entities in the training samples;

[0016] Re-labeling the entities in the training samples using the original label distribution of the entities in the training samples and the correct information contained in the predicted label distribution;

[0017] The entity labeling model is retrained according to the re-labeled training samples to establish a second entity labeling model.

[0018] A training sample processing method, comprising:

[0019] Obtain an original training sample set, and use the original training sample set to train an entity labeling model to establish an entity labeling model; wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, and the original label includes a noise label; the entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information;

[0020] Using the entity labeling model to predict label distribution of entities in the training samples;

[0021] According to the original label distribution of the entities in the training samples and the predicted label distribution, the entities in the training samples are relabeled to generate a new training sample set.

[0022] A method for extending an entity tag knowledge graph, comprising:

[0023] Obtain an original training sample set, wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, wherein the original label information is obtained based on the original knowledge graph, including a noise label;

[0024] The entity labeling model is trained using the original training sample set to establish an entity labeling model. The entity labeling model is used to determine a corresponding label for an entity from a candidate label set according to an input entity and associated text information;

[0025] Using the entity labeling model to predict label distribution of entities in the training samples;

[0026] Re-labeling the entities in the training sample according to the original label distribution of the entities in the training sample and the predicted label distribution;

[0027] If the re-labeled label does not belong to the original label of the corresponding entity, the corresponding relationship between the entity and the re-labeled label is added to the knowledge graph.

[0028] A training device for an entity marking model, comprising:

[0029] A first training unit is used to train an entity labeling model using training data containing noise labels, and use the trained entity labeling model to predict label distribution of entities in the training data;

[0030] A pseudo label obtaining unit, configured to obtain a pseudo label of the entity according to the original label distribution of the entity included in the training data and the label distribution predicted by the entity labeling model;

[0031] The second training unit is used to re-label the entities in the training data using the pseudo labels, and train the entity labeling model using the re-labeled training data to obtain a trained entity labeling model.

[0032] A training device for an entity marking model, comprising:

[0033] A first training unit is used to obtain an original training sample set, and use the original training sample set to train the entity labeling model to establish a first entity labeling model; wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity;

[0034] A prediction unit, configured to use the first entity labeling model to predict label distribution of entities in the training sample;

[0035] A re-labeling unit, configured to re-label the entities in the training sample using the original label distribution of the entities in the training sample and the correct information contained in the predicted label distribution;

[0036] The second training unit is used to retrain the entity labeling model according to the re-labeled training samples to establish a second entity labeling model.

[0037] A training sample processing device, comprising:

[0038] A training sample acquisition unit is used to acquire an original training sample set, and use the original training sample set to train an entity labeling model to establish an entity labeling model; wherein the original training sample set includes a plurality of training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, and the original label includes a noise label; the entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information;

[0039] A prediction unit, configured to use the entity labeling model to predict label distribution of entities in the training sample;

[0040] The re-labeling unit is used to re-label the entities in the training samples according to the original label distribution of the entities in the training samples and the predicted label distribution to generate a new training sample set.

[0041] A device for expanding an entity tag knowledge graph, comprising:

[0042] A training sample acquisition unit is used to acquire an original training sample set, wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, wherein the original label information is obtained based on the original knowledge graph, including a noise label;

[0043] A training unit, used to train the entity labeling model using the original training sample set, and establish the entity labeling model. The entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information;

[0044] A prediction unit, configured to use the entity labeling model to predict label distribution of entities in the training sample;

[0045] A re-labeling unit, configured to re-label the entities in the training sample according to the original label distribution of the entities in the training sample and the predicted label distribution;

[0046] An adding unit is used to add the corresponding relationship between the entity and the re-labeled label to the knowledge graph if the re-labeled label does not belong to the original label of the corresponding entity.

[0047] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0048] Through the embodiments of the present application, the original labels of the training samples can be first constructed by remote supervision, which include noise labels. In this state, they can be directly used to pre-train the entity labeling model. Although errors may occur when the entity labeling model trained in this way is used for prediction, it actually contains a large amount of correct information. At the same time, the original labels of the training samples actually also contain a large amount of correct information. Therefore, the embodiments of the present application can re-label the entities in the training samples on the basis of making full use of the correct information in the above two, and then use the re-labeled training samples to train the entity labeling model. Since the process of re-labeling the entities in the training samples in the embodiments of the present application does not rely on any assumptions, and can integrate the prediction results of the pre-trained entity labeling model and the correct information contained in the original noise labels, therefore, it is possible to denoise a larger range and more common label errors, and has scalability. In addition, since the pre-trained entity labeling model does not completely simulate the training sample set, it is possible that for a certain entity, when predicting the label in combination with its context, other labels other than the original noise label may also have a certain probability of being correct in the prediction result. On this basis, through the calculation in the embodiment of the present application, it is possible to improve the correct probability of other labels other than the original noise label, and even eventually re-label them as the correct label to the corresponding entity. Therefore, through the scheme of the embodiment of the present application, when re-labeling the entities in the training samples, it is no longer limited to selecting one from the original noise labels. Therefore, for situations where there is a large error in the original noise label, or even the original noise label does not contain a correct label, there is an opportunity to re-label the specific training sample with a truly correct label, thereby improving the quality of the training sample. After re-training the model, the prediction accuracy of the model can be improved.

[0049] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0051] Figure 1 It is a schematic diagram of the system architecture provided by the embodiment of the present application;

[0052] Figure 2 is a flow chart of the first method provided in an embodiment of the present application;

[0053] Figure 3 is a flow chart of the second method provided in an embodiment of the present application;

[0054] Figure 4 is a flowchart of the third method provided in an embodiment of the present application;

[0055] Figure 5 is a flowchart of the fourth method provided in an embodiment of the present application;

[0056] Figure 6 is a schematic diagram of a first device provided in an embodiment of the present application;

[0057] Figure 7 is a schematic diagram of a second device provided in an embodiment of the present application;

[0058] Figure 8 is a schematic diagram of a third device provided in an embodiment of the present application;

[0059] Fig. 9 is a schematic diagram of a fourth device provided in an embodiment of the present application;

[0060] Fig.10 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of this application.

[0062] In order to facilitate the understanding of the technical solution provided by the embodiment of the present application, the following is a brief introduction to the process of constructing a data set based on remote supervision and training an entity labeling model. First, a relatively small knowledge graph can be provided, in which (entity, label) pairs can be saved, and the same entity may correspond to multiple labels. Then more training samples can be constructed based on the small knowledge graph for training the entity labeling model. For example, for the entity "cainiao", the following information may exist in the knowledge graph (cainiao, company, network company), (cainiao, service, logistics service provider), etc. When constructing more samples based on the small knowledge graph, assuming that one of the samples is "I want to go to the Cainiao post station to pick up a courier", then word segmentation can be performed first, and entity recognition can be performed from it. After identifying the entity "cainiao", it can be determined based on the above knowledge graph that the corresponding label may be an Internet company or a logistics service provider, etc. In other words, an entity may correspond to multiple labels, which may include correct labels that can reflect the current sample context scenario, and may also include incorrect labels (called noise labels). After that, the existing processing method is to select one of the above multiple labels as the correct label. For example, for the samples in the above example, for the entity "Cainiao", you can select one of the labels from the Internet company or the logistics service provider, re-label the entity, and then use the re-labeled samples to train the model.

[0063] However, in practical applications, the entity labels obtained by linking the knowledge graph have large errors. For example, the labels of some entities in the knowledge graph may be inaccurate, or errors may occur during linking, or the included scenarios are not comprehensive enough, etc., resulting in that the labels of the entities obtained through the knowledge graph cannot correctly reflect the scenarios in the current sample context. For example, a sample is "I am still a rookie in some aspects". By linking the aforementioned knowledge graph, the label determined for the entity "rookie" still includes an Internet company or a logistics service provider. At this time, according to the scheme in the prior art, only one of the labels can be selected from the Internet company or the logistics service provider as the label of the entity "rookie" in the above sample. However, it is obvious that in this sample, the entity "rookie" is neither an Internet company nor a logistics service provider, but a popular Internet term used to express newcomers in a certain field. The scheme in the prior art cannot label the entity with the truly correct label of "Internet popular term".

[0064] Based on the above situation, the embodiment of the present application provides a corresponding solution. In this solution, the original label of the training sample can be first constructed by remote supervision, which contains a noise label. In this state, it can be directly used to pre-train the entity labeling model. When the entity labeling model trained in this way is used for prediction, although errors may occur, it actually contains a large amount of correct information. At the same time, the original label of the training sample actually contains a large amount of correct information. Therefore, the embodiment of the present application can re-label the entities in the training sample on the basis of making full use of the correct information in the above two, and then use the re-labeled training sample to train the entity labeling model. Since the process of re-labeling the entities in the training sample in the embodiment of the present application does not rely on any assumptions, and can integrate the prediction results of the pre-trained entity labeling model and the correct information contained in the original noise label, therefore, it is possible to denoise a larger range and more common label errors, and has scalability. In addition, since the pre-trained entity labeling model does not completely simulate the training sample set, it is possible that for a certain entity, when predicting the label in combination with its context, other labels other than the original noise label may also have a certain probability of being correct in the prediction result. On this basis, through the calculation in the embodiment of the present application, it is possible to improve the correct probability of other labels other than the original noise label, and even eventually re-label them as the correct label to the corresponding entity. Therefore, through the scheme of the embodiment of the present application, when re-labeling the entities in the training samples, it is no longer limited to selecting one from the original noise labels. Therefore, for situations where there is a large error in the original noise label, or even the original noise label does not contain the correct label, there is an opportunity to re-label the specific training sample with a truly correct label.

[0065] Among them, in order to be able to utilize the information of both the predicted label distribution (predicted by the pre-trained entity labeling model for training samples) and the original noise label distribution, there can be a variety of specific implementation methods. For example, in an optional way, the concept of pseudo-label distribution is proposed in an embodiment of the present application. Specifically, the pseudo-label distribution can be initialized according to the original label distribution corresponding to the sample, and then, the objective function about the pseudo-label distribution, the original label distribution and the predicted label distribution is constructed to utilize the model prediction results and the information in the original noise label. Moreover, through multiple rounds of iterations, the parameter matrix of the model and the pseudo-label distribution can be updated so that the cost value of the objective function gradually approaches the target (for example, minimum). Afterwards, the updated pseudo-label distribution can be used to re-label the entities in the sample. Among them, there can be a variety of ways to construct the objective function. For example, in an optional implementation method, such as Figure 1As shown, three constraints can be designed, including the KL distance constraint (mainly used to limit the pseudo-label distribution from deviating too much from the label distribution predicted by the model, so as to better utilize the model prediction results), the deviation constraint (mainly used to limit the pseudo-label distribution from deviating too much from the original label distribution, so as to better utilize the original label information), and the one-hot constraint (mainly used to limit the number of correct labels to one). The weighted sum of the above three sub-functions is minimized as the goal to iterate, and the pseudo-label distribution and other parameters in the model are updated in each iteration. Finally, the updated pseudo-label distribution can be obtained, and the entities in the training samples can be re-labeled.

[0066] The specific technical solutions provided in the embodiments of the present application are described in detail below.

[0067] Embodiment 1

[0068] First, this embodiment provides a method for training an entity labeling model. Figure 2 , the method may specifically include:

[0069] S201: Obtain an original training sample set, and use the original training sample set to train an entity labeling model to establish a first entity labeling model; wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity;

[0070] The original training sample set may be data constructed by remote supervision. Since the entity labeling model needs to be trained, each training sample may include an entity, text associated with the entity, and the original label information corresponding to the entity. The specific original label information can be obtained by connecting with the knowledge graph. Such original labels usually include noise labels, that is, labels with incorrect labels. In an embodiment of the present application, after obtaining the above-mentioned original label information, it can be directly used to pre-train the entity labeling model.

[0071] There may be multiple entity labeling models to be trained. For example, one method may be the NFETC model. The NFETC model is a supervised fine-grained entity labeling model that can be trained as a baseline model in the entire method. This entity labeling model is used to determine the corresponding label for the entity from a candidate label set based on the input entity and the associated text information. Specifically, for an input text c = {w1, w2, …, w n} and entity m={w m1 , w m2 ,…w ml} (an entity is a continuous character segment in a text sentence), the entity labeling model NFETC aims to label the entity m with an entity label selected from a candidate label set. The label set can be pre-set, in which multiple possible labels can be listed. For example, a label set may include a total of 200 labels (for example, (company, technology company), (star, singer), etc.), then the role of the entity labeling model is to select one of the correct labels for the entity in a certain input sentence from these 200 labels.

[0072] Although there may be some noise information in the original labels corresponding to the entities in the training samples, the probability of containing the correct labels is still relatively high. Therefore, this part of information can be fully utilized when re-labeling the training samples. Specifically, in order to utilize this part of information, the original label distribution of the entities in the training samples can be determined first. Among them, the original labels corresponding to the entities in the training samples (that is, the labels obtained from the knowledge graph through remote supervision) may be one or more, and these original labels usually also belong to the aforementioned label set. For the convenience of calculation, the original labels of the entities in each training sample can be expressed in the form of original label distribution. Specifically, this original label distribution can exist in the form of a vector, and the number of dimensions of the vector can be equal to the number of labels in the label set, for example, it can be a 200-dimensional vector. Among them, for an entity of a specific sample, the dimension where the original label is located is 1, and the other dimensions are 0. For example, an entity in a sample corresponds to three original labels, which are located in the 10th, 13th, and 14th dimensions of the vector, respectively. In the original label distribution of the entity, the 10th, 13th, and 14th dimensions of the vector are 1, and the other dimensions are 0. In addition, in the specific implementation, in order to make the numbers on each dimension in the original label distribution express the probability, the above label distribution can also be normalized. For example, for the above example, in the normalized vector, the 10th, 13th, and 14th dimensions are 0.33, and the other dimensions are 0, and so on.

[0073] S202: Using the first entity labeling model to predict label distribution of entities in the training sample;

[0074] After the first entity labeling model is obtained through training, the first entity labeling model can also be used directly to predict the label distribution of entities in each training sample. Since the first entity labeling model obtained through training does not completely fit the labeling information in the training sample during the prediction process, but will be combined with the contextual information of the specific entity for comprehensive analysis and calculation, and finally obtain the prediction result, it also contains a large amount of correct information. In the embodiment of the present application, when re-labeling the training sample, in addition to using the correct information contained in the original label, it can also be combined with the correct information contained in the prediction result. Among them, when making a specific prediction, each piece of data in the training sample (including only the sentence and the entity itself, excluding the labeling information) can be directly input into the entity labeling model, and the entity labeling model can output the prediction result. The specific prediction result can also be represented in the form of a vector, and the dimension of the vector can also be the same as the number of labels in the label set. The value on each dimension in the vector represents the correct probability of the label on the corresponding dimension for the current entity.

[0075] It should be noted that when the first entity labeling model is used for prediction, the training sample data is not completely fitted. For example, if the number of sentences containing a certain entity in the training sample set is relatively small, the model will learn less information from the sample's annotation information during the training process. At this time, when predicting this entity, the context information of the entity may be used more for prediction. In turn, it is possible that other labels other than the original noise labels obtain a certain probability of correctness. For example, for an entity in the training sample, its original noise labels are distributed in the 10th, 13th, and 14th dimensions, with respective probabilities of 0.33; however, when the first entity labeling model is used to predict the entity, through the analysis and calculation of the context information of the entity, in the specific prediction results, the label on the 20th dimension also has a certain probability of correctness, and its probability may even be higher than the original probability on the 10th, 13th, and 14th dimensions, and so on.

[0076] S203: Re-labeling the entities in the training sample using the original label distribution of the entities in the training sample and the correct information contained in the predicted label distribution;

[0077] In an embodiment of the present application, the entities in the training samples can be re-labeled by combining the original label distribution corresponding to the entities in the training samples and the correct information contained in the prediction results given by the pre-trained model, and then the model can be re-trained using the re-labeled training samples to improve the prediction accuracy of the model.

[0078] There are multiple ways to utilize the original label distribution and the predicted label distribution information in the re-labeling process. For example, in one way, the pseudo-label distribution corresponding to the entity in the training sample can be initialized according to the original label distribution, and then the objective function about the pseudo-label distribution, the original label distribution and the predicted label distribution is constructed, and the pseudo-label distribution is updated through multiple rounds of iterations. Then, the training sample is re-labeled according to the updated pseudo-label distribution.

[0079] There can also be multiple ways to construct the objective function. For example, in one way, the constructed objective function can ensure that the difference between the pseudo-label distribution and the predicted label distribution and the original label distribution during the update process remains within a preset range. To this end, the specific objective function can include a KL distance constraint. Specifically, the KL distance constraint can be expressed according to the cross entropy between the predicted label distribution and the pseudo-label distribution, so that the offset of the pseudo-label distribution relative to the predicted label distribution is controlled within the first target range. Among them, the KL distance is a function used to measure the similarity between distributions. By setting the KL distance constraint, the pseudo-label distribution can be made to not deviate too much from the predicted label distribution during the update process. Through this constraint, the proportion of correct labels contained in the final prediction information can be greater than the proportion of correct information contained in the original training set (training set constructed based on remote supervision). Especially for simple samples, the influence of noise labels can be basically overcome.

[0080] Of course, if the objective function is constructed only by the above-mentioned KL distance constraint, the accuracy of the prediction results of the model that is finally re-labeled and retrained has a lower limit, that is, it is higher than the proportion of correct information contained in the original training set. Among them, for relatively simple samples, the accuracy after retraining may be significantly improved, while for more complex samples, it may only be slightly higher than the proportion of correct information contained in the original training set. In addition, if only the KL distance constraint is set, the pseudo-label distribution may be too close to the label distribution predicted by the model. Therefore, in practical applications, in order to further improve the accuracy and to balance the offset between the pseudo-label distribution and the label distribution in the prediction results, other constraints can be added to the objective function.

[0081] For example, in one way, an offset constraint can also be added, and the offset constraint is expressed by the cross entropy between the updated pseudo-label distribution and the initialized pseudo-label distribution, so that the offset between the updated pseudo-label distribution and the initialized pseudo-label distribution is within the second target range. That is, through this constraint, it is possible to avoid excessive offset of the pseudo-label distribution relative to the original label distribution during the iteration process. The purpose of doing so is to make full use of the correct information contained in the original label distribution.

[0082] Through the above-mentioned KL distance constraint and offset constraint, the label distribution information predicted by the first entity labeling model and the original label distribution information can both participate in the calculation, so that the correct information contained in both can be fully utilized, which is conducive to improving the accuracy of re-labeling.

[0083] In addition, in practical applications, if only KL distance constraints and offset constraints are set, multiple labels may appear in the updated pseudo-label distribution, each corresponding to a probability value, which is used to indicate the probability that the corresponding label is correct. At this time, the training sample can be directly re-labeled with this probability value, or, in a more preferred manner, a one-hot constraint can be added to the objective function so that only one label on one dimension in the updated pseudo-label distribution meets the target condition. In other words, by adding a one-hot constraint, there can be only one label on one dimension in the final updated pseudo-label distribution, and the labels on other dimensions are all 0, so that a training sample is only labeled with one label, and the probability that the label belongs to the correct label is relatively high. Therefore, when using such re-labeled sample data for model training, the prediction accuracy of the model can be improved.

[0084] It should be noted that, since the embodiment of the present application utilizes the prediction results of the pre-trained entity labeling model and the original label distribution when re-labeling the entities in the training samples, and in the label distribution in the prediction results, since the context information of the entity can be combined for comprehensive calculation, other labels other than the original noise labels may also have a certain probability of being correct, thus appearing in the predicted label distribution. In the process of re-labeling using this information, the embodiment of the present application may improve the correct probability of other labels other than the original noise labels, and then when re-labeling, the other labels may be labeled as correct labels to the entities in the corresponding training samples. It can be seen that through the embodiment of the present application, the specific re-labeled labels can include labels other than the original labels, that is, it is no longer limited to re-labeling after selecting labels from the original noise labels, so that a wider range and more common label errors can be denoised. In addition, since the re-labeling process of the embodiment of the present application does not require assumptions or rely on any additional labeling information, it has good scalability.

[0085] S204: Retrain the entity labeling model according to the re-labeled training samples to establish a second entity labeling model.

[0086] After completing the re-labeling of the training samples, the entity labeling model can be re-trained based on the re-labeled training samples to establish a second entity labeling model. Since the re-labeled training samples have higher accuracy, in an optional manner, they can also correspond to only one label with a relatively high accuracy rate. Therefore, after using such samples for model training, a higher prediction accuracy rate can be obtained for the model.

[0087] In summary, through the embodiments of the present application, the original labels of the training samples can be first constructed by remote supervision, which contain noise labels. In this state, they can be directly used to pre-train the entity labeling model. Although errors may occur when the entity labeling model trained in this way is used for prediction, it actually contains a large amount of correct information. At the same time, the original labels of the training samples actually also contain a large amount of correct information. Therefore, the embodiments of the present application can re-label the entities in the training samples on the basis of making full use of the correct information in the above two, and then use the re-labeled training samples to train the entity labeling model. Since the process of re-labeling the entities in the training samples in the embodiments of the present application does not rely on any assumptions, and can integrate the prediction results of the pre-trained entity labeling model and the correct information contained in the original noise labels, therefore, it is possible to denoise a larger range and more common label errors, and has scalability. In addition, since the pre-trained entity labeling model does not completely simulate the training sample set, it is possible that for a certain entity, when predicting the label in combination with its context, other labels other than the original noise label may also have a certain probability of being correct in the prediction result. On this basis, through the calculation in the embodiment of the present application, it is possible to improve the correct probability of other labels other than the original noise label, and even eventually re-label them as the correct label to the corresponding entity. Therefore, through the scheme of the embodiment of the present application, when re-labeling the entities in the training samples, it is no longer limited to selecting one from the original noise labels. Therefore, for situations where there is a large error in the original noise label, or even the original noise label does not contain a correct label, there is an opportunity to re-label the specific training sample with a truly correct label, thereby improving the quality of the training sample. After re-training the model, the prediction accuracy of the model can be improved.

[0088] Embodiment 2

[0089] In the above-mentioned embodiment 1, after the training samples are re-labeled, the re-labeled samples can be used to retrain the model to improve the prediction accuracy of the model. In other application scenarios, the re-labeled samples can also be used in other scenarios, for example, in training relationship extraction models, etc. Therefore, in this embodiment 2, a training sample processing method is provided, see Figure 3 , the method may specifically include:

[0090] S301: Obtain an original training sample set, and use the original training sample set to train an entity labeling model to establish an entity labeling model; wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, and the original label includes a noise label; the entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information;

[0091] S302: Using the entity labeling model to predict label distribution of entities in the training sample;

[0092] S303: Re-labeling the entities in the training samples according to the original label distribution of the entities in the training samples and the predicted label distribution to generate a new training sample set.

[0093] Embodiment 3

[0094] As mentioned above, in the embodiment of the present application, for a training sample, the re-labeled label may be a label other than the original label. In this case, the re-labeled label can also be used to expand the original knowledge graph. Specifically, this embodiment three also provides a method for expanding the entity label knowledge graph, see Figure 4 , the method may specifically include:

[0095] S401: Acquire an original training sample set, wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, wherein the original label information is obtained based on the original knowledge graph, including a noise label;

[0096] S402: Using the original training sample set to train an entity labeling model, and establishing an entity labeling model, wherein the entity labeling model is used to determine a corresponding label for an entity from a candidate label set according to an input entity and associated text information;

[0097] S403: Using the entity labeling model to predict label distribution of entities in the training samples;

[0098] S404: re-labeling the entities in the training sample according to the original label distribution of the entities in the training sample and the predicted label distribution;

[0099] S405: If the re-labeled label does not belong to the original label of the corresponding entity, the corresponding relationship between the entity and the re-labeled label is added to the knowledge graph.

[0100] In addition to the applications described in the above-mentioned embodiments 2 and 3, other application scenarios may also be included in the specific implementation. For example, in a commodity object online sales system, the name of the commodity object may be relatively long. When processing such a name text, it may be necessary to perform entity recognition and labeling from it. At this time, the solution provided in the embodiment of the present application can be used to train the entity labeling model to improve the accuracy of entity labeling, etc.

[0101] Embodiment 4

[0102] This fourth embodiment provides a training method for an entity labeling model from another perspective. Figure 5 , the method may specifically include:

[0103] S501: training an entity labeling model using training data containing noise labels, and using the trained entity labeling model to predict label distribution of entities in the training data;

[0104] S502: Obtain a pseudo label of the entity according to the original label distribution of the entity included in the training data and the label distribution predicted by the entity labeling model;

[0105] In specific implementation, the pseudo-label distribution can be first initialized according to the original label distribution, and then the pseudo-label distribution is updated. During the updating process, the difference between the pseudo-label distribution and the predicted label distribution and the original label distribution is kept within a preset range. After the update is completed, the label with a probability that meets the conditions in the updated pseudo-label distribution is selected as the pseudo-label.

[0106] Among them, in specific implementation, the objective function can be constructed according to the KL distance constraint, and the pseudo-label distribution is updated through multiple rounds of iterations; the KL distance constraint is expressed by the cross entropy between the predicted label distribution and the pseudo-label distribution, so that during the iteration process, the offset of the pseudo-label distribution relative to the predicted label distribution is within the first target range. Alternatively, the objective function is constructed according to the offset constraint, and the pseudo-label distribution is updated through multiple rounds of iterations; the offset constraint is expressed by the cross entropy between the pseudo-label distribution and the original label distribution, so that during the iteration process, the offset of the pseudo-label distribution relative to the original label distribution is within the second target range. In addition, the objective function can also be constructed according to the one-hot constraint, and the pseudo-label distribution can be updated through multiple rounds of iterations, so that only the labels on a single dimension in the updated pseudo-label distribution meet the target conditions.

[0107] S503: Re-labeling entities in the training data using the pseudo labels, and training the entity labeling model using the re-labeled training data to obtain a trained entity labeling model.

[0108] Among them, for the parts not described in detail in Examples 2, 3, and 4, please refer to the description in the aforementioned Example 1, and no further details will be given here.

[0109] Corresponding to the first embodiment, the present application embodiment also provides a training device for an entity marking model, see Figure 6 , the device may specifically include:

[0110] The first training unit 601 is used to obtain an original training sample set, and use the original training sample set to train the entity labeling model to establish a first entity labeling model; wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity;

[0111] A prediction unit 602 is used to predict label distribution of entities in the training sample using the first entity labeling model;

[0112] A re-labeling unit 603, configured to re-label the entities in the training sample by using the original label distribution of the entities in the training sample and the correct information contained in the predicted label distribution;

[0113] The second training unit 604 is used to retrain the entity labeling model according to the re-labeled training samples to establish a second entity labeling model.

[0114] Among them, the first entity labeling model is used to determine the corresponding label for the entity from the candidate label set according to the input entity and the associated text information; the label distribution is: the distribution of the label in the multidimensional vector corresponding to the label set.

[0115] The re-labeling unit may specifically include:

[0116] A pseudo-label initialization subunit, used to initialize the pseudo-label distribution corresponding to the entity in the training sample according to the original label distribution;

[0117] A pseudo label updating subunit, used to construct an objective function about the pseudo label distribution, the original label distribution and the predicted label distribution, and update the pseudo label distribution through multiple rounds of iterations;

[0118] The re-labeling subunit is used to re-label the training samples according to the updated pseudo-label distribution.

[0119] Specifically, during the updating process, the difference between the pseudo label distribution and the predicted label distribution and the original label distribution is kept within a preset range.

[0120] Among them, the objective function includes a KL distance constraint, and the KL distance constraint is expressed by the cross entropy between the predicted label distribution and the pseudo label distribution, so that during the iteration process, the offset of the pseudo label distribution relative to the predicted label distribution is within a first target range.

[0121] Alternatively, the objective function includes an offset constraint, which is expressed by a cross entropy between the pseudo label distribution and the original label distribution, so that during the iteration process, the offset of the pseudo label distribution relative to the original label distribution is within a second target range.

[0122] In addition, the objective function may also include a one-hot constraint so that only labels on a single dimension in the updated pseudo-label distribution meet the target condition.

[0123] The re-labeled labels include labels other than the original labels in the label set.

[0124] Corresponding to the second embodiment, the present application embodiment also provides a training sample processing device, see Figure 7 , the device may specifically include:

[0125] The training sample acquisition unit 701 is used to acquire an original training sample set, and use the original training sample set to train the entity labeling model to establish the entity labeling model; wherein the original training sample set includes a plurality of training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, and the original label includes a noise label; the entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information;

[0126] A prediction unit 702, configured to use the entity labeling model to predict label distribution of entities in the training sample;

[0127] The re-labeling unit 703 is used to re-label the entities in the training samples according to the original label distribution of the entities in the training samples and the predicted label distribution to generate a new training sample set.

[0128] Corresponding to the third embodiment, the present application embodiment also provides a device for expanding the entity tag knowledge graph, see Figure 8 , the device may include:

[0129] A training sample acquisition unit 801 is used to acquire an original training sample set, wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, wherein the original label information is obtained based on the original knowledge graph, including a noise label;

[0130] A training unit 802 is used to train the entity labeling model using the original training sample set to establish an entity labeling model. The entity labeling model is used to determine a corresponding label for an entity from a candidate label set according to an input entity and associated text information;

[0131] A prediction unit 803 is used to predict label distribution of entities in the training sample using the entity labeling model;

[0132] A re-labeling unit 804, configured to re-label the entities in the training sample according to the original label distribution of the entities in the training sample and the predicted label distribution;

[0133] The adding unit 805 is used to add the corresponding relationship between the entity and the re-labeled label to the knowledge graph if the re-labeled label does not belong to the original label of the corresponding entity.

[0134] Corresponding to the fourth embodiment, the present application embodiment also provides a training device for an entity marking model, see Fig. 9 , the device may specifically include:

[0135] A first training unit 901 is used to train an entity labeling model using training data containing noise labels, and use the trained entity labeling model to predict label distribution of entities in the training data;

[0136] A pseudo label obtaining unit 902 is used to obtain a pseudo label of the entity according to the original label distribution of the entity included in the training data and the label distribution predicted by the entity labeling model;

[0137] The second training unit 903 is used to re-label the entities in the training data using the pseudo labels, and train the entity labeling model using the re-labeled training data to obtain a trained entity labeling model.

[0138] The second training unit may specifically include:

[0139] An initialization subunit, used to initialize the pseudo label distribution according to the original label distribution;

[0140] An updating subunit, configured to update the pseudo label distribution, wherein the difference between the pseudo label distribution and the predicted label distribution and the original label distribution is kept within a preset range during the updating process;

[0141] The pseudo-label determination subunit is used to select a label whose probability meets the conditions in the updated pseudo-label distribution as a pseudo-label.

[0142] The updating subunit may be specifically used for:

[0143] An objective function is constructed according to a KL distance constraint, and the pseudo-label distribution is updated through multiple rounds of iterations; the KL distance constraint is expressed by the cross entropy between the predicted label distribution and the pseudo-label distribution, so that during the iteration process, the offset of the pseudo-label distribution relative to the predicted label distribution is within a first target range.

[0144] Alternatively, an objective function is constructed according to an offset constraint, and the pseudo-label distribution is updated through multiple rounds of iterations; the offset constraint is expressed by the cross entropy between the pseudo-label distribution and the original label distribution, so that during the iteration process, the offset of the pseudo-label distribution relative to the original label distribution is within a second target range.

[0145] In addition, an objective function may be constructed according to the one-hot constraint, and the pseudo-label distribution may be updated through multiple rounds of iterations so that only labels on a single dimension in the updated pseudo-label distribution meet the target condition.

[0146] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.

[0147] And a computer system comprising:

[0148] one or more processors; and

[0149] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.

[0150] in, Fig.10The electronic device architecture is shown as an example, which may include a processor 1010, a video display adapter 1011, a disk drive 1012, an input / output interface 1013, a network interface 1014, and a memory 1020. The processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, and the memory 1020 may be communicatively connected via a communication bus 1030.

[0151] Among them, the processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided in this application.

[0152] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store an operating system 1021 for controlling the operation of the electronic device 1000, and a basic input and output system (BIOS) for controlling the low-level operation of the electronic device 1000. In addition, a web browser 1023, a data storage management system 1024, and a model training processing system 1025, etc. can also be stored. The above-mentioned model training processing system 1025 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided in the present application is implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0153] The input / output interface 1013 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0154] The network interface 1014 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0155] The bus 1030 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, and the memory 1020).

[0156] It should be noted that, although the above device only shows a processor 1010, a video display adapter 1011, a disk drive 1012, an input / output interface 1013, a network interface 1014, a memory 1020, a bus 1030, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.

[0157] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0158] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0159] The training method, device and electronic device of the entity marking model provided by this application are introduced in detail above. The principle and implementation method of this application are explained by using specific examples in this article. The description of the above embodiments is only used to help understand the method and core idea of ​​this application. At the same time, for those skilled in the art, according to the idea of ​​this application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A training method for an entity labeling model, characterized in that: include: Using training data containing noise labels to train an entity labeling model, and using the trained entity labeling model to predict label distribution of entities in the training data, wherein the training data includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity; Obtaining a pseudo label of the entity according to the original label distribution of the entity included in the training data and the label distribution predicted by the entity labeling model; Re-labeling entities in the training data using the pseudo labels, and training the entity labeling model using the re-labeled training data to obtain a trained entity labeling model; Among them, the pseudo-label of the entity is obtained according to the original label distribution of the entity included in the training data and the label distribution predicted by the entity labeling model, including: initializing the pseudo-label distribution according to the original label distribution; updating the pseudo-label distribution using the target constraint to select the pseudo-label in the updated pseudo-label distribution.

2. The method according to claim 1, characterized in that Updating the pseudo-label distribution using the target constraint to select the pseudo-label in the updated pseudo-label distribution includes: The pseudo label distribution is updated using the target constraint, and during the updating process, the difference between the pseudo label distribution and the predicted label distribution and the original label distribution is kept within a preset range; Select the label whose probability meets the conditions in the updated pseudo-label distribution as the pseudo-label.

3. The method according to claim 2, characterized in that The updating of the pseudo label distribution by using the target constraint includes: An objective function is constructed according to a KL distance constraint, and the pseudo-label distribution is updated through multiple rounds of iterations; the KL distance constraint is expressed by the cross entropy between the predicted label distribution and the pseudo-label distribution, so that during the iteration process, the offset of the pseudo-label distribution relative to the predicted label distribution is within a first target range.

4. The method according to claim 2, characterized in that: The updating of the pseudo label distribution by using the target constraint includes: The objective function is constructed according to the offset constraint, and the pseudo label distribution is updated through multiple rounds of iterations; the offset constraint is expressed by the cross entropy between the pseudo label distribution and the original label distribution, so that in the iterative process, the The offset of the pseudo label distribution relative to the original label distribution is within the second target range.

5. The method according to claim 2, characterized in that: The updating of the pseudo label distribution includes: An objective function is constructed according to the one-hot constraint, and the pseudo-label distribution is updated through multiple rounds of iterations so that only labels on a single dimension in the updated pseudo-label distribution meet the target condition.

6. A method for training an entity labeling model, characterized in that: include: Obtain an original training sample set, and use the original training sample set to train an entity labeling model to establish a first entity labeling model; wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity; Using the first entity labeling model to predict label distribution of entities in the training samples; Re-labeling the entities in the training samples using the original label distribution of the entities in the training samples and the correct information contained in the predicted label distribution; Retrain the entity labeling model according to the re-labeled training samples to establish a second entity labeling model; Among them, the entities in the training samples are re-labeled using the correct information contained in the original label distribution and the predicted label distribution, including: initializing the pseudo-label distribution corresponding to the entities in the training samples according to the original label distribution; updating the pseudo-label distribution using the objective function to re-label according to the updated pseudo-label distribution.

7. The method according to claim 6, characterized in that The first entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information; The label distribution refers to the distribution of labels in the multidimensional vector corresponding to the label set.

8. The method according to claim 6, characterized in that The updating of the pseudo label distribution by using the objective function to re-label according to the updated pseudo label distribution includes: Constructing the objective function about the pseudo label distribution, the original label distribution and the predicted label distribution, and updating the pseudo label distribution through multiple rounds of iterations; The training samples are relabeled according to the updated pseudo-label distribution.

9. The method according to claim 8, characterized in that During the updating process, the difference between the pseudo label distribution, the predicted label distribution and the original label distribution is kept within a preset range.

10. The method according to claim 9, characterized in that The objective function includes a KL distance constraint, which is expressed by a cross entropy between the predicted label distribution and the pseudo label distribution, so that during an iteration, an offset of the pseudo label distribution relative to the predicted label distribution is within a first target range.

11. The method according to claim 9, characterized in that Also includes: The objective function includes an offset constraint, which is expressed by a cross entropy between the pseudo label distribution and the original label distribution, so that during the iteration process, the offset of the pseudo label distribution relative to the original label distribution is within a second target range.

12. The method according to claim 8, characterized in that Also includes: The objective function includes a one-hot constraint so that only labels on a single dimension in the updated pseudo-label distribution meet the target condition.

13. The method according to any one of claims 6 to 12, characterized in that: The relabeled labels include labels other than the original labels in the label set.

14. A training sample processing method, characterized in that: include: Obtain an original training sample set, and use the original training sample set to train an entity labeling model to establish an entity labeling model; wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, and the original label includes a noise label; the entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information; Using the entity labeling model to predict label distribution of entities in the training samples; Re-labeling the entities in the training samples according to the original label distribution of the entities in the training samples and the predicted label distribution to generate a new training sample set; Among them, according to the original label distribution of the entities in the training sample and the predicted label distribution, the entities in the training sample are re-labeled, including: initializing the pseudo-label distribution corresponding to the entities in the training sample according to the original label distribution; updating the pseudo-label distribution using the target constraint to re-label according to the updated pseudo-label distribution.

15. A method for expanding an entity tag knowledge graph, characterized in that: include: Obtain an original training sample set, wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, wherein the original label information is obtained based on the original knowledge graph, including a noise label; The entity labeling model is trained using the original training sample set to establish an entity labeling model; the entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information; Using the entity labeling model to predict label distribution of entities in the training samples; Re-labeling the entities in the training sample according to the original label distribution of the entities in the training sample and the predicted label distribution; If the re-labeled label does not belong to the original label of the corresponding entity, adding the corresponding relationship between the entity and the re-labeled label to the knowledge graph; Among them, according to the original label distribution of the entities in the training sample and the predicted label distribution, the entities in the training sample are re-labeled, including: initializing the pseudo-label distribution corresponding to the entities in the training sample according to the original label distribution; updating the pseudo-label distribution using the target constraint to re-label according to the updated pseudo-label distribution.

16. A training device for an entity marking model, characterized in that: include: A first training unit is used to train an entity labeling model using training data containing noise labels, and use the trained entity labeling model to predict label distribution of entities in the training data, wherein the training data includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity; A pseudo label obtaining unit, configured to obtain a pseudo label of the entity according to the original label distribution of the entity included in the training data and the label distribution predicted by the entity labeling model; A second training unit is used to re-label the entities in the training data using the pseudo labels, and train the entity labeling model using the re-labeled training data to obtain a trained entity labeling model; The second training unit is further used for: initializing the pseudo label distribution according to the original label distribution; and updating the pseudo label distribution by using the target constraint to select the pseudo label in the updated pseudo label distribution.

17. A training device for an entity marking model, characterized in that: include: A first training unit is used to obtain an original training sample set, and use the original training sample set to train the entity labeling model to establish a first entity labeling model; wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity; A prediction unit, configured to use the first entity labeling model to predict label distribution of entities in the training sample; A re-labeling unit, configured to re-label the entities in the training sample using the original label distribution of the entities in the training sample and the correct information contained in the predicted label distribution; A second training unit is used to retrain the entity labeling model according to the re-labeled training samples to establish a second entity labeling model; The re-labeling unit is further used to: initialize the pseudo-label distribution corresponding to the entities in the training sample according to the original label distribution; and update the pseudo-label distribution using an objective function to re-label according to the updated pseudo-label distribution.

18. A training sample processing device, characterized in that: include: A training sample acquisition unit is used to acquire an original training sample set, and use the original training sample set to train an entity labeling model to establish an entity labeling model; wherein the original training sample set includes a plurality of training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, and the original label includes a noise label; the entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information; A prediction unit, configured to use the entity labeling model to predict label distribution of entities in the training sample; A re-labeling unit, configured to re-label the entities in the training samples according to the original label distribution of the entities in the training samples and the predicted label distribution, so as to generate a new training sample set; The re-labeling unit is further used to: initialize the pseudo-label distribution corresponding to the entities in the training sample according to the original label distribution; and update the pseudo-label distribution using the target constraint to re-label according to the updated pseudo-label distribution.

19. A device for expanding an entity tag knowledge graph, characterized in that: include: A training sample acquisition unit is used to acquire an original training sample set, wherein the original training sample set includes multiple training samples, each training sample includes an entity, text associated with the entity, and original label information corresponding to the entity, wherein the original label information is obtained based on the original knowledge graph, including a noise label; A training unit, used to train the entity labeling model using the original training sample set to establish the entity labeling model; the entity labeling model is used to determine a corresponding label for the entity from a candidate label set according to the input entity and the associated text information; A prediction unit, configured to use the entity labeling model to predict label distribution of entities in the training sample; A re-labeling unit, configured to re-label the entities in the training sample according to the original label distribution of the entities in the training sample and the predicted label distribution; An adding unit, configured to add a corresponding relationship between the entity and the re-labeled label to the knowledge graph if the re-labeled label does not belong to the original label of the corresponding entity; The re-labeling unit is further used to: initialize the pseudo-label distribution corresponding to the entities in the training sample according to the original label distribution; and update the pseudo-label distribution using the target constraint to re-label according to the updated pseudo-label distribution.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 15 are implemented.

21. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of claims 1 to 15.

Citation Information

Patent Citations

  • Semi-supervised Chinese named entity recognition method based on deep learning

    CN108959252A

  • Fraud recognition model training method, fraud recognition method and device

    CN109598331A