A processing method, device and electronic device

By analyzing the data set for errors in the machine learning model, determining the characteristics associated with the correct annotation results, and performing batch corrections, the problem of insufficient accuracy of artificial intelligence processing results in the existing technology is solved, and realizing immediate and accurate error corrections are achieved.

CN114548317BActive Publication Date: 2025-06-27LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210208150.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-03
Publication Date
2025-06-27
Estimated Expiration
2042-03-03

AI Technical Summary

Technical Problem

The existing technology has shortcomings in improving the accuracy of artificial intelligence processing results, especially in the error label correction of machine learning models. The existing methods cannot achieve immediate correction and the correction accuracy is difficult to guarantee.

Method used

By obtaining data sets with machine learning model annotation type, we determine the correction information of specific data in the data set, and analyze the features associated with the correct annotation result, and then batch correct the annotation types of other data containing the feature.

Benefits of technology

It realizes batch instant correction of the same type of error labeling, ensures the accuracy of the correction operation, and avoids the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548317B_ABST
    Figure CN114548317B_ABST
Patent Text Reader

Abstract

The present application discloses a processing method, apparatus and electronic device. The method includes: obtaining a first data set, where the data in the first data set is labeled with a label type obtained by processing based on a machine learning model; determining correction information for first data in the first data set, the correction information indicating that the label type of the first data is corrected from a first type to a second type; determining a first feature in the first data that has an association relationship with the second type; and adjusting the label type of other data in the first data set that includes the first feature from the first type to the second type. This solution can automatically process and analyze the features in the sample data that have an association relationship with the correct annotation result when manually revising an incorrectly labeled sample, and then uniformly batch correct other sample data that includes the feature and is incorrectly labeled based on the association relationship between the feature and the correct annotation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology. More specifically, it relates to a processing method, apparatus, and electronic device. Background Art

[0002] Artificial intelligence technology has currently been maturely applied in various fields, but there is still a great deal of room for improvement in enhancing the accuracy of the processing results of artificial intelligence.

[0003] For example, the prediction or annotation results automatically obtained through a machine learning model are often not completely correct. Therefore, manual inspection is required in practical applications. Based on the current inspection scheme, if a large number of corrections are desired, one method is to retrain the model and then use the retrained model to make predictions and annotations again. However, this method cannot correct errors immediately. Another method is to rigidly specify similar errors based on rules, such as rules where samples have certain identical contexts, etc. However, it is difficult to guarantee the correction accuracy of this method. Summary of the Invention

[0004] In view of this, this application provides the following technical solutions:

[0005] A processing method includes:

[0006] Obtain a first data set, where the data in the first data set has an annotation type, and the annotation type is obtained based on the processing of a machine learning model;

[0007] Determine the correction information of the first data in the first data set, where the correction information indicates that the annotation type of the first data is corrected from a first type to a second type;

[0008] Determine a first feature in the first data that has an association relationship with the second type;

[0009] Adjust the annotation type of other data in the first data set that includes the first feature from the first type to the second type.

[0010] Optionally, after determining the first feature in the first data that has an association relationship with the second type, it further includes:

[0011] Correct the machine learning model based on the association relationship between the second type and the first feature.

[0012] Optionally, the determining the first feature in the first data that has an association relationship with the second type includes:

[0013] Determine the first feature on which the machine learning model depends to process and output the second type based on interpretable artificial intelligence technology.

[0014] Optionally, determining the first features on which the machine learning model processes and outputs the second type based on the interpretability artificial intelligence technology includes:

[0015] Obtaining an approximate sample set of the first data based on the first data;

[0016] Obtaining the label and confidence of each sample data in the approximate sample set through the machine learning model;

[0017] Determining the similarity weight between each sample data and the first data;

[0018] Determining the associated data between each feature in the first data and the annotation type based on the label, confidence, and similarity weight corresponding to the sample data;

[0019] Determining the first features on which the machine learning model processes and outputs the second type from each feature of the first data.

[0020] Optionally, determining the similarity weight between each sample data and the first data includes:

[0021] Performing a first process on each sample data in the approximate sample set, where the first process is used to transform the text data into machine-recognizable identification data;

[0022] Determining the cosine distance between the corresponding sample data and the first data based on the identification data as the similarity weight.

[0023] Optionally, performing the first process on each sample data in the approximate sample set includes:

[0024] Performing a first process on each sample data in the approximate sample set through a bag-of-words feature, a pre-trained model, or a word vector model.

[0025] Optionally, determining the associated data between each feature in the first data and the annotation type based on the label, confidence, and similarity weight corresponding to the sample data includes:

[0026] Constructing a weighted regression model of the approximate sample set based on the label, confidence, and similarity weight corresponding to each sample data, and outputting the associated data between each feature in the first data and the annotation type.

[0027] Optionally, after determining the correction information of the first data in the first data set, it further includes:

[0028] Determining the second features in the first data that have an associated relationship with the first type;

[0029] Then, adjusting the annotation type of other data in the first data set that contains the first feature from the first type to the second type includes:

[0030] Adjusting the annotation type of other data in the first data set that contains both the first feature and the second feature from the first type to the second type.

[0031] This application also discloses a processing device, including:

[0032] A data set acquisition module, configured to acquire a first data set, where the data in the first data set has an annotation type, and the annotation type is obtained based on processing by a machine learning model;

[0033] A correction determination module, configured to determine correction information of first data in the first data set, where the correction information indicates correcting the annotation type of the first data from a first type to a second type;

[0034] A feature determination module, configured to determine a first feature in the first data that has an association relationship with the second type;

[0035] A correction adjustment module, configured to adjust the annotation type of other data in the first data set that contains the first feature from the first type to the second type.

[0036] Furthermore, this application also discloses an electronic device, including:

[0037] A processor;

[0038] A memory, configured to store executable program instructions of the processor;

[0039] Wherein, the executable program instructions include: acquiring a first data set, where the data in the first data set has an annotation type, and the annotation type is obtained based on processing by a machine learning model; determining correction information of first data in the first data set, where the correction information indicates correcting the annotation type of the first data from a first type to a second type; determining a first feature in the first data that has an association relationship with the second type; adjusting the annotation type of other data in the first data set that contains the first feature from the first type to the second type.

[0040] As can be seen from the above technical solutions, embodiments of the present application disclose a processing method, apparatus, and electronic device. The method includes: obtaining a first data set, where the data in the first data set is labeled with a label type obtained by processing based on a machine learning model; determining correction information for first data in the first data set, where the correction information indicates that the label type of the first data is corrected from a first type to a second type; determining a first feature in the first data that has an association relationship with the second type; and adjusting the label type of other data in the first data set that includes the first feature from the first type to the second type. This solution can automatically analyze and process the features in the sample data that have an association relationship with the correct labeling result when manually revising a mislabeled case, and then uniformly batch correct other mislabeled sample data that includes the feature based on the association relationship between the feature and the correct labeling result; since the correction principle is implemented based on the logical relationship between the feature and the labeling result itself, batch instant correction of the same type of mislabeling can be achieved, and the accuracy of the correction operation can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.

[0042] Figure 1 It is a flowchart of a processing method disclosed in an embodiment of the present application;

[0043] Figure 2 It is a flowchart of determining the first feature in the first data disclosed in an embodiment of the present application;

[0044] Figure 3 It is a flowchart of determining the similarity weight between sample data and the first data disclosed in an embodiment of the present application;

[0045] Figure 4 It is a schematic flowchart of using interpretable artificial intelligence technology to correct the labeling result disclosed in an embodiment of the present application;

[0046] Figure 5 It is a schematic diagram of the implementation of the working process of the LIME algorithm disclosed in an embodiment of the present application;

[0047] Figure 6 It is a schematic structural diagram of a processing apparatus disclosed in an embodiment of the present application;

[0048] Figure 7 It is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] For reference and clarity, the descriptions, abbreviations, or acronyms of technical terms used hereinafter are summarized as follows:

[0050] LIME: Local Interpretable Model-agnostic Explanations, a machine learning model interpretation tool.

[0051] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0052] The embodiments of the present application can be applied to electronic devices. The present application does not limit the product form of the electronic device, which may include but is not limited to smart phones, tablet computers, wearable devices, personal computers (PCs), netbooks, etc., and can be selected according to application requirements.

[0053] Figure 1 It is a flowchart of a processing method disclosed in an embodiment of the present application. Refer to Figure 1 As shown, the processing method may include:

[0054] Step 101: Obtain a first data set, where the data in the first data set has an annotation type, and the annotation type is obtained based on the processing of a machine learning model.

[0055] The first data set includes a plurality of sample data, and these sample data have been processed by the annotation of the machine learning model and have an annotation type. Since there is still a gap between the artificial intelligence processing ability of the machine learning model and that of real humans, there will be a certain error rate in the annotation type of the sample data in the first data set; the processing method disclosed in the embodiments of the present application is to provide a solution that can efficiently correct the mislabeled data of the machine learning model in real time.

[0056] Step 102: Determine the correction information of the first data in the first data set, where the correction information indicates that the annotation type of the first data is corrected from a first type to a second type.

[0057] Among them, the first type is the original annotation type of the first data by the machine learning model, and the second type is the correct annotation type corrected by the user.

[0058] Since there is a certain error rate in the processing results of the machine learning model, after the machine learning model finishes annotating the sample data in the first dataset, the user will manually verify the annotation results. After discovering the incorrect annotation type during the verification, the annotation result will be manually modified. For example, the annotation "negative evaluation" will be corrected to "positive evaluation".

[0059] Step 103: Determine the first feature in the first data that has an association with the second type.

[0060] Specifically, the first feature on which the machine learning model depends for processing and outputting the second type can be determined based on explainable artificial intelligence technology. Explainable artificial intelligence technology can be but is not limited to the LIME algorithm.

[0061] After the user manually corrects an incorrect annotation type, the processing method described in this embodiment will analyze the first data that was originally misannotated and the first feature that has an association with the correct annotation result, that is, the second type. It should be noted that the first feature here can include one or more features. For example, in a machine learning model for annotating gender based on name, for the name "Zhang Ting", the annotation result is "female", and the feature "Ting" has an association with the annotation result "female". In this example, there is one first feature, which is the name "Ting". Another example is in a machine learning model for annotating gender based on name and height. One sample data is "Chen Mingming, 189 cm", and another sample data is "Chen Mingming, 163 cm". Since the name "Mingming" could be female or male, it is difficult to give a definite annotation result relying solely on the name. In this case, by combining the height information in the sample data, an accurate annotation result can be obtained. For the sample data "Chen Mingming, 189 cm", the annotation result is "male", and for the sample data "Chen Mingming, 163 cm", the annotation result is "female". In this example, there are two first features, one is the name "Mingming", and the other is the height "189 cm" or "163 cm".

[0062] In the solution of this application, the relationship between the first data and the correct annotation result is not determined based on specific rules. For example, rules with certain same contexts for the first data often have great limitations and are difficult to cover other situations with slight differences. In the embodiment of this application, based on the features of the first data itself, the influence of each feature on the correct annotation result is analyzed and processed, so as to obtain the logical association relationship of feature → annotation result. That is, essentially, it is judged under what circumstances the machine learning model will give what kind of annotation result. The logical association relationship obtained here is not interfered by other data outside the sample data, that is, the first data, which helps to determine other annotation errors of the same type subsequently.

[0063] Step 104: Adjust the annotation type of other data in the first dataset that contains the first feature from the first type to the second type.

[0064] After determining the first feature in the first data that has an association with the second type, the logical association relationship between the feature and the annotation result is determined. In this way, based on this logical association relationship, other sample data including the first feature can be determined. If the annotation type of other sample data including the first feature is also the first type, it needs to be corrected to the second type.

[0065] It should be noted that the execution of Step 104 is automatically completed in batches, and the process does not require user participation. It only needs to be based on the previously determined association relationship between the first feature and the annotation result of the second type, so as to dynamically and real-time improve the output result of the machine learning model.

[0066] The processing method described in this embodiment is directed to the dataset initially annotated by the machine learning model. In the case of manually revising a mislabeled one, it can automatically process and analyze the features in the sample data that have an association with the correct annotation result, and then uniformly batch correct other sample data that contains the feature and is mislabeled based on the association relationship between the feature and the correct annotation result; since the correction principle is implemented based on the logical relationship between the feature and the annotation result itself, it can achieve batch instant correction of the same type of mislabeling and ensure the accuracy of the correction operation.

[0067] In some other implementations, after determining the first feature in the first data that has an association with the second type, it may further include the step of correcting the machine learning model based on the association relationship between the second type and the first feature.

[0068] Correcting the machine learning model based on the association relationship between the second type and the first feature ensures that the corrected machine learning model will not output incorrect annotation results when encountering sample data containing the first feature in the future, thus effectively improving the prediction accuracy of the machine learning model and reducing the manual workload at the same time.

[0069] Figure 2 This is a flowchart for determining the first feature in the first data disclosed in the embodiments of the present application. Refer to Figure 2 As shown, in the above embodiments, the determining the first feature on which the machine learning model processes and outputs the second type based on the explainable artificial intelligence technology may include:

[0070] Step 201: Obtain an approximate sample set of the first data based on the first data.

[0071] The principle of generating an approximate sample set can be to randomly replace some words in the sample data. For example, for the sample data "novel in style, but poor in workmanship", the generated approximate samples can be "nice in style", "poor in quality", "nice in style but poor in quality", etc. These approximate samples constitute the approximate sample set.

[0072] Step 202: Obtain the label and confidence of each sample data in the approximate sample set through the machine learning model.

[0073] The machine learning model therein is the machine learning model that labels the results of the first data set. Using the same machine learning model to label the results of the sample data in the approximate sample set ensures the consistency of the processing algorithm. The labeling results therein include the label and confidence.

[0074] Step 203: Determine the similarity weight between each sample data and the first data.

[0075] The sample data are all approximate samples of the first data. Different approximate samples have different degrees of approximation to the first data, that is, different similarity weights; and these approximate samples are used to analyze and determine the association relationship between each feature in the first data and the labeling type. The similarity weight will also affect the analysis result to a certain extent. Therefore, it is necessary to determine the similarity weight between each sample data and the first data. Specifically, how to determine the similarity weight will be introduced in detail in the following content and will not be elaborated here.

[0076] Step 204: Determine the association data between each feature in the first data and the labeling type based on the label, confidence, and similarity weight corresponding to the sample data.

[0077] Specifically, a weighted regression model of the approximate sample set can be constructed based on the label, confidence, and similarity weight corresponding to each sample data, and the association data between each feature in the first data and the labeling type are output.

[0078] Step 205: Determine the first feature on which the machine learning model depends to process and output the second type from each feature of the first data.

[0079] After determining the association data between each feature in the first data and the labeling type, the first feature on which the machine learning model depends to process and output the second type can be found from it.

[0080] The above content completely introduces the process of determining the first feature from the first data, so as to facilitate those skilled in the art to better understand and implement the technical solution of the present application.

[0081] Figure 3The flowchart for determining the similarity weight between sample data and first data disclosed in the embodiments of the present application. In combination with Figure 3 as shown, determining the similarity weight between each of the sample data and the first data may include:

[0082] Step 301: Perform a first process on each sample data in the approximate sample set, where the first process is used to transform text data into machine-recognizable identification data.

[0083] Specifically, each sample data in the approximate sample set can be processed by a bag-of-words feature, a pre-trained model, a word vector model, etc., to realize the conversion from text to numbers, so as to obtain machine-recognizable identification data.

[0084] Step 302: Determine the cosine distance between the corresponding sample data and the first data based on the identification data as the similarity weight.

[0085] After transforming the text data into machine-recognizable identification data, the cosine distance between the sample data and the first data, that is, the similarity weight, can be conveniently calculated.

[0086] To better understand the content of the present application, a process introduction with a category annotation task as an example will be given below.

[0087] 1. Before manual annotation, pre-label the data in dataset C through a machine learning model.

[0088] 2. During manual annotation, check the pre-labeled data and revise the label of the wrong sample X from Positive to the correct label Negative.

[0089] 3. Use the interpretable artificial intelligence algorithm LIME to obtain which features of the wrong sample X cause the model to misclassify to the label Positive. Figure 4 The flowchart for correcting the annotation result using interpretable artificial intelligence technology disclosed in the embodiments of the present application. Figure 5 The schematic diagram for implementing the working process of the LIME algorithm disclosed in the embodiments of the present application. It can be combined with Figure 4 and Figure 5 to understand the implementation process of this example.

[0090] A. For the wrong sample "The screen is large, but it is not convenient to carry", generate an approximate sample set, such as "The screen is large and very good", "The large screen is good for watching videos", "It is not convenient to carry", etc.

[0091] B. Use the original machine learning model to label these approximate samples with labels and probability confidence levels.

[0092] C. Tokenize the approximate sample set to obtain bag-of-words features;

[0093] D. Based on the bag-of-words features, calculate the cosine distance between the approximate samples and the original incorrect samples, which is regarded as the similarity weight.

[0094] E. Based on the approximate sample set, the similarity weight, and the probability confidence, perform a weighted regression model on these approximate samples. The output of the regression model is the features and feature weights, obtaining the influence of different features on the classifier (machine learning model). For example, "large screen" supports the "Positive" category, while the feature "inconvenient" is against the "Positive" category.

[0095] 4. For the subsequent samples in dataset C, if they also contain these features and the pre-labeled label is Positive, then automatically batch-revise them to the label Negative. For example, the subsequent samples to be revised such as "inconvenient to carry because the screen is too large", "inconvenient to carry, large screen", etc. are automatically revised to the "Negative" category because they contain "inconvenient" which is against the "Positive" category.

[0096] Based on the above, in other implementations, after determining the correction information of the first data in the first dataset, it may further include: determining the second feature in the first data that has an association relationship with the first type.

[0097] Then, the adjustment of the annotation type of the other data in the first dataset that contains the first feature from the first type to the second type may include: adjusting the annotation type of the other data in the first dataset that simultaneously contains the first feature (such as "hand" shown in Figure 4 and the second feature (such as "large screen" shown in Figure 4 from the first type to the second type.

[0098] The processing method described in the embodiments of this application can automatically process and analyze the features in the sample data that have an association relationship with the correct annotation result in the case of manually revising an incorrect annotation, and then uniformly batch-correct other sample data that contains this feature and is incorrectly annotated based on the association relationship between this feature and the correct annotation result. The implementation process does not require re-training the machine learning model, nor does it require understanding the internal structure and principle of the machine learning model, and can be implemented for the results of any machine learning model, having good applicability.

[0099] For each of the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0100] In the embodiments of the present application disclosed above, the method is described in detail. The method of the present application can be implemented by means of various forms of devices. Therefore, the present application also discloses a device, and specific embodiments are given below for detailed description.

[0101] Figure 6 It is a schematic structural diagram of a processing device disclosed in an embodiment of the present application. Refer to Figure 6 As shown, the processing device 60 may include:

[0102] A data set acquisition module 601, configured to acquire a first data set, where the data in the first data set has an annotation type, and the annotation type is obtained based on processing by a machine learning model.

[0103] A correction determination module 602, configured to determine correction information of the first data in the first data set, where the correction information indicates that the annotation type of the first data is corrected from a first type to a second type.

[0104] A feature determination module 603, configured to determine a first feature in the first data that has an association relationship with the second type.

[0105] A correction adjustment module 604, configured to adjust the annotation type of other data in the first data set that includes the first feature from the first type to the second type.

[0106] For the data set initially annotated by the machine learning model in this embodiment, in the case of manually revising a mislabeled one, the processing device of this embodiment can automatically process and analyze the features in the sample data that have an association relationship with the correct annotation result, and then uniformly batch correct other sample data that includes the feature and is mislabeled based on the association relationship between the feature and the correct annotation result; since the correction principle is implemented based on the logical relationship between the feature and the annotation result itself, batch instant correction of the same type of mislabeling can be achieved, and the accuracy of the correction operation can be ensured.

[0107] In one implementation, the processing device may further include: a model correction module, configured to correct the machine learning model based on the association relationship between the second type and the first feature after the feature determination module determines the first feature in the first data that has an association relationship with the second type.

[0108] In one implementation, the feature determination module may specifically be configured to: determine, based on interpretable artificial intelligence technology, a first feature on which the machine learning model depends for processing and outputting the second type.

[0109] In one implementation, the feature determination module may include: an approximate sample obtaining module, configured to obtain an approximate sample set of the first data based on the first data; an approximate sample annotation module, configured to obtain, through the machine learning model, a label and a confidence level of each sample data in the approximate sample set; a weight determination module, configured to determine a similarity weight between each sample data and the first data; a correlation relationship determination module, configured to determine, based on the label, the confidence level, and the similarity weight corresponding to the sample data, correlation data between each feature in the first data and the annotation type; and a feature determination sub-module, configured to determine, from each feature of the first data, a first feature on which the machine learning model depends for processing and outputting the second type.

[0110] In one implementation, the weight determination module may include: a first processing module, configured to perform a first processing on each sample data in the approximate sample set, where the first processing is used to convert text data into machine-recognizable identification data; and a weight determination sub-module, configured to determine a cosine distance between the corresponding sample data and the first data based on the identification data as the similarity weight.

[0111] In one implementation, the first processing module may specifically be configured to: perform a first processing on each sample data in the approximate sample set through a bag-of-words feature, a pre-trained model, or a word vector model.

[0112] In one implementation, the correlation relationship determination module may specifically be configured to: construct a weighted regression model of the approximate sample set based on the label, the confidence level, and the similarity weight corresponding to each sample data, and output correlation data between each feature in the first data and the annotation type.

[0113] In one implementation, the correction determination module is further configured to: determine a second feature in the first data that has a correlation relationship with the first type. Then, the correction adjustment module is specifically configured to: adjust the annotation type of other data in the first data set that simultaneously includes the first feature and the second feature from the first type to the second type.

[0114] Any one of the processing devices described in the above embodiments includes a processor and a memory. The data set acquisition module, correction determination module, feature determination module, correction adjustment module, model correction module, approximate sample acquisition module, weight determination module, etc. in the above embodiments are all stored in the memory as program modules, and the processor executes the above program modules stored in the memory to realize corresponding functions.

[0115] The processor includes a kernel, which retrieves the corresponding program module from the memory. One or more kernels can be set, and the processing of the access data can be realized by adjusting the kernel parameters.

[0116] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0117] In an exemplary embodiment, a computer-readable storage medium is also provided, which can be directly loaded into the internal memory of a computer and contains software code. After the computer loads and executes the computer program, the steps shown in any embodiment of the above processing method can be implemented.

[0118] In an exemplary embodiment, a computer program product is also provided, which can be directly loaded into the internal memory of a computer, wherein the computer program contains software codes, and after being loaded and executed by a computer, the computer program can implement the steps shown in any embodiment of the processing method described above.

[0119] Furthermore, an embodiment of the present application provides an electronic device. Figure 7 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application. Figure 7 As shown, the electronic device includes at least one processor 701, and at least one memory 702 and a bus 703 connected to the processor; wherein the memory is used to store executable instructions of the processor; the processor and the memory communicate with each other through the bus; the processor is used to call the executable program instructions in the memory to execute the above-mentioned processing method.

[0120] Among them, the executable program instructions stored in the memory include: obtaining a first data set, the data in the first data set has a label type, and the label type is obtained based on machine learning model processing; determining correction information of the first data in the first data set, and the correction information indicates that the label type of the first data is corrected from the first type to the second type; determining a first feature in the first data that has an association with the second type; adjusting the label type of other data in the first data set that contains the first feature from the first type to the second type.

[0121] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0122] It should also be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0123] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, in software modules executed by a processor, or in a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0124] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A processing method, comprising: Obtaining a first data set, the first data set including text data, and data in the first data set being with an annotation type, the annotation type being obtained based on processing by a machine learning model; Determining correction information of first data in the first data set, the correction information indicating that the annotation type of the first data is corrected from a first type to a second type; Determining a first feature in the first data that has an association relationship with the second type; Adjusting the annotation type of other data in the first data set that includes the first feature from the first type to the second type.

2. The processing method according to claim 1, after determining the first feature in the first data that has an association relationship with the second type, further comprising: Correcting the machine learning model based on the association relationship between the second type and the first feature.

3. The processing method according to claim 1, the determining the first feature in the first data that has an association relationship with the second type includes: Determining, based on an explainable artificial intelligence technique, the first feature on which the machine learning model depends for processing and outputting the second type, wherein the explainable artificial intelligence technique includes: the LIME algorithm.

4. The processing method according to claim 3, the determining, based on the explainable artificial intelligence technique, the first feature on which the machine learning model depends for processing and outputting the second type includes: Obtaining an approximate sample set of the first data based on the first data; Obtaining the label and confidence of each sample data in the approximate sample set through the machine learning model; Determining the similarity weight between each sample data and the first data; Determining the association data between each feature in the first data and the annotation type based on the label, confidence, and similarity weight corresponding to the sample data; Determining, from each feature of the first data, the first feature on which the machine learning model depends for processing and outputting the second type.

5. The processing method according to claim 4, the determining the similarity weight between each sample data and the first data includes: Performing a first processing on each sample data in the approximate sample set, the first processing being used to transform the text data into machine-recognizable identification data; Determining the cosine distance between the corresponding sample data and the first data based on the identification data as the similarity weight.

6. The processing method according to claim 5, the performing a first processing on each sample data in the approximate sample set includes: Performing a first processing on each sample data in the approximate sample set through a bag-of-words feature, a pre-trained model, or a word vector model.

7. The processing method according to claim 4, the determining the association data between each feature in the first data and the annotation type based on the label, confidence, and similarity weight corresponding to the sample data includes: Constructing a weighted regression model of the approximate sample set based on the label, confidence, and similarity weight corresponding to each sample data, and outputting the association data between each feature in the first data and the annotation type.

8. The processing method according to claim 1, after determining the correction information of the first data in the first data set, further comprising: Determining a second feature in the first data that has an association relationship with the first type; Then, the adjusting the annotation type of other data in the first data set that includes the first feature from the first type to the second type includes: Adjusting the annotation type of other data in the first data set that includes both the first feature and the second feature from the first type to the second type.

9. A processing device, comprising: A data set obtaining module, configured to obtain a first data set, where the first data set includes text data, and data in the first data set is provided with an annotation type, and the annotation type is obtained based on processing by a machine learning model; A correction determining module, configured to determine correction information of the first data in the first data set, where the correction information indicates that the annotation type of the first data is corrected from a first type to a second type; A feature determining module, configured to determine a first feature in the first data that has an association relationship with the second type; A correction adjusting module, configured to adjust the annotation type of other data in the first data set that includes the first feature from the first type to the second type.

10. An electronic device, comprising: A processor; A memory, configured to store executable program instructions of the processor; Wherein, the executable program instructions include: obtaining a first data set, where the first data set includes text data, and data in the first data set is provided with an annotation type, and the annotation type is obtained based on processing by a machine learning model; determining correction information of the first data in the first data set, where the correction information indicates that the annotation type of the first data is corrected from a first type to a second type; determining a first feature in the first data that has an association relationship with the second type; adjusting the annotation type of other data in the first data set that includes the first feature from the first type to the second type.

Citation Information

Patent Citations

  • Label correction method and device for sample picture, equipment and storage medium

    CN111382798A

  • Information collection device and information collection method

    JP2018206189A