A method, device, terminal and storage medium for annotating a data set

Through the method of classifying identification allocation, feature learning and matching adjustment of the data in the data set, automatic annotation of the data set is realized, solving the problems of high cost and low efficiency of manual annotation in the prior art, and improving the accuracy and objectivity of the annotation results.

CN110263851BActive Publication Date: 2025-06-10CLOUDMINDS SHANGHAI ROBOTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910533056.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-19
Publication Date
2025-06-10
Estimated Expiration
2039-06-19

AI Technical Summary

Technical Problem

In the prior art, the training of deep learning models relies on manually labeled data sets, resulting in high cost and low efficiency of data set labeling, and there are subjectivity and errors in manual labeling, which limits the promotion of deep learning in practical applications.

Method used

By assigning preset classification identifiers to the data in the data set, performing feature learning and testing, and adjusting the classification identifiers using the matching degree until the matching degree is greater than the threshold, determining the annotation of the data set.

Benefits of technology

Automatic labeling of data sets is realized, labor costs are reduced, human error is avoided, and the accuracy and objectivity of labeling results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110263851B_ABST
    Figure CN110263851B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention relates to the field of data processing, and discloses a method, device, terminal, and storable medium for annotating a data set. The method for annotating a data set in the present invention includes: an allocation step of respectively allocating a preset classification identifier to each piece of data in the data set to obtain a first classification result; a feature learning step of performing feature learning on each preset classification by using the data set after the classification identifier is allocated to obtain the data features of each classification; a testing step of classifying each piece of data in the data set by using the data features of each classification to obtain a second classification result; when the matching degree between the second classification result and the first classification result is less than or equal to a first threshold, re-executing the allocation step to the testing step until the matching degree between the second classification result and the first classification result is greater than the first threshold; determining the annotation of each piece of data in the data set according to the allocated classification identifier. The above solution enables automatic annotation of data with high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of data processing, and in particular to a method, device, terminal and storable medium for labeling a data set. Background Art

[0002] In recent years, deep learning classification methods have achieved significant breakthroughs in classification effects, completely surpassing the level that traditional methods can achieve. With the continuous introduction of various deep learning network models such as Resnet (residual network), deep learning methods continue to refresh the limits of data classification accuracy, making deep learning methods the most popular and reliable classification method.

[0003] Deep learning mainly conducts forward and reverse transmission in the model through a huge number of training sets, and automatically improves the parameters of the model through continuous reciprocation, so that the model can finally achieve the ideal classification effect. Therefore, in addition to being affected by the model structure, the model effect obtained by training also depends greatly on the representativeness of the training set to the category and the accuracy of its corresponding label. In order to ensure the accuracy of the label, the current training set labels are all labeled manually, that is, the category of each data is labeled through human cognition. This method certainly guarantees the accuracy of the data set to a certain extent, but it also has great defects. Because for some more complex classification tasks, the number of data sets required is often at the level of hundreds of thousands or even millions, and manual labeling of these data will consume a lot of manpower and time. For example, the imagenet image classification competition, which has a huge influence, does not provide data for labeling by a single organization, but relies on the Mturk crowdsourcing platform. At the same time, due to the subjectivity of manual labeling, in order to ensure the objectivity and accuracy of the labeling results, it is often necessary to screen the labeling results or supervise the labeling process, which further increases the cost of manual labeling. Therefore, the training of the model mostly relies on a few fixed data sets, and can only classify the categories contained in these data sets. However, in real production, it is often necessary to classify different categories according to different environments and needs, and the categories are diverse and complex. Obviously, a few fixed data sets cannot cover all of these categories. Therefore, in real scenarios, it is often necessary to build your own data set according to your own needs to achieve classification of specific categories. However, building a data set by yourself will consume a lot of manpower to label and verify the data after labeling. Therefore, the reliance on manual labeling has greatly limited the comprehensive promotion of deep learning in practical applications. Summary of the invention

[0004] The purpose of the embodiments of the present invention is to provide a method, device, terminal and storage medium for labeling a data set, so that data can be automatically labeled with high accuracy.

[0005] To solve the above technical problems, an embodiment of the present invention provides a method for annotating a data set, including: an allocation step of respectively allocating a preset classification identifier to each data in the data set to obtain a first classification result, where the data set is a data set of images; a feature learning step of using the data set after allocating the classification identifier to perform feature learning on each preset classification to obtain data features of each classification; a testing step of classifying each data in the data set by using the data features to obtain a second classification result; when the matching degree between the second classification result and the first classification result is less than or equal to a first threshold, re-executing the allocation step to the testing step until the matching degree between the second classification result and the first classification result is greater than the first threshold; determining the annotation of each data in the data set according to the classification identifier allocated when the matching degree between the second classification result and the first classification result is greater than the first threshold.

[0006] An embodiment of the present invention further provides an apparatus for annotating a data set, including: an allocation module for respectively allocating a preset classification identifier to each data in the data set to obtain a first classification result, where the data set is a data set of images; a feature learning module for using the data set after allocating the classification identifier to perform feature learning on each preset classification to obtain data features of each classification; a testing module for classifying each data in the data set by using the data features to obtain a second classification result; a comparison module for, when the matching degree between the second classification result and the first classification result is less than or equal to a first threshold, re-executing the allocation step to the testing step until the matching degree between the second classification result and the first classification result is greater than the first threshold; an annotation determination module for determining the annotation of each data in the data set according to the classification identifier allocated when the matching degree between the second classification result and the first classification result is greater than the first threshold.

[0007] An embodiment of the present invention further provides a terminal, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for annotating a data set as described above.

[0008] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the method for annotating a data set as described above.

[0009] In the embodiments of the present invention, compared with the prior art, by assigning classification identifiers of uncertain correctness to the data set, and using such classification results to determine the data characteristics under each classification, and then according to the characteristic that the characteristics of the data under a classification identifier are basically the same, reversely verify whether the previously assigned classification identifier is accurate, and when it is inaccurate, replace the assigned identifier until the assigned classification identifier conforms to the consistent characteristics of the data in the belonging classification, it is possible to automatically confirm the annotation of the data set, and the result is accurate and reliable, avoiding the way that existing data sets are all manually annotated, not only greatly reducing the labor cost, but also avoiding human errors, making the annotation result more objectively accurate.

[0010] As a further improvement, the feature learning step specifically uses the method of model training on the data set after the classification identifier is assigned to perform feature learning on each preset classification. The above solution clearly performs feature learning through the method of model training, so that the feature learning is not restricted by the predetermined features, and thus more suitable feature representations can be found, making the data features more accurate and appropriate.

[0011] As a further improvement, the assignment step includes: manually assigning a part of the data in the data set and automatically assigning another part of the data in the data set; when the assignment step is re-executed, re-assigning the classification identifier to the data in the data set that is automatically assigned the classification identifier. The above solution clearly states that when assigning the identifier, part of the data is manually assigned, and when re-assigning, only the data that is automatically assigned the classification identifier is re-assigned, making the identifier results of this part of the data accurate, reducing the amount of data with uncertain identifiers, and accelerating the finding of the appropriate classification identifier assignment result. At the same time, only a part of the data is manually assigned, which reduces the labor cost while greatly accelerating the acquisition of the annotation result.

[0012] As a further improvement, the amount of data manually assigned is less than the amount of data automatically assigned. It is clear that the amount of manual assignment is less than the amount of automatic assignment, which controls the manual workload while greatly accelerating the speed of finding the appropriate classification result.

[0013] As a further improvement, after the comparison step and before the annotation determination step, it includes:

[0014] Category expansion step: when the matching degree between the second classification result and the first classification result is greater than the first threshold and the number of classification identifiers corresponding to the data in the dataset is less than the second threshold, perform category expansion on each classification identifier, assign the expanded classification identifiers to the data in the dataset, and re - execute the feature learning step to the comparison step until the matching degree between the second classification result and the first classification result is greater than the first threshold, and the number of the expanded classification identifiers is greater than or equal to the second threshold; the annotation determination step: determine the annotation of each data in the dataset according to the classification identifiers assigned when the matching degree between the second classification result and the first classification result is greater than the first threshold and the number of the expanded classification identifiers is greater than or equal to the second threshold. The above - mentioned solution clearly states that when presetting classification identifiers, fewer categories can be set first. Since the number of identifiers is small, the speed of finding the accurate classification result can be accelerated. Then, intra - class expansion can be gradually performed to increase classifications. Through the above - mentioned process from coarser to finer, the overall speed of finding the accurate classification result can be effectively accelerated.

[0015] As a further improvement, before re - executing the assignment step, it includes: calculating the classification result matching degrees under each classification identifier respectively; determining the classification identifiers with matching degrees less than or equal to the third threshold, denoted as the first - type classification identifiers; re - assigning classification identifiers to each data in the dataset, specifically: re - assigning classification identifiers to each data belonging to the first - type classification identifiers in the dataset. In this embodiment, by only re - assigning the identifiers with relatively poor classification results during re - assignment, the amount of re - assignment is reduced, which is beneficial to accelerating the speed of obtaining appropriate classification results.

[0016] As a further improvement, the second classification result includes: the confidence levels of the classification results of each data; when re - executing the assignment step, re - assign classification identifiers to each data with a relatively low confidence level of the classification result in the dataset. In this embodiment, during testing, the determination of confidence level is combined. Those with high confidence levels are considered basically accurate and can no longer be re - assigned, so that the amount of data for re - assignment is reduced, and the speed of finding the accurate classification result is faster. Description of the Drawings

[0017] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute a scale limitation.

[0018] Figure 1 is the flowchart of the method for annotating a dataset according to the first embodiment of the present invention;

[0019] Figure 2 is the flowchart of the method for annotating a dataset according to the second embodiment of the present invention;

[0020] Figure 3 It is a flowchart for reassigning classification identifiers in the annotation method of the dataset according to the third embodiment of the present invention;

[0021] Figure 4 It is a flowchart of the annotation method of the dataset according to the fourth embodiment of the present invention;

[0022] Figure 5 It is a flowchart of the annotation method of the dataset according to the fifth embodiment of the present invention;

[0023] Figure 6 It is a schematic diagram of the annotation device of the dataset according to the sixth embodiment of the present invention;

[0024] Figure 7 It is a schematic diagram of the terminal according to the seventh embodiment of the present invention. Specific Embodiments

[0025] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will elaborate on each embodiment of the present invention with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present invention, many technical details are provided for the readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation to the specific implementation of the present invention. The various embodiments can be combined and cross-referenced with each other on the premise of no contradiction.

[0026] The first embodiment of the present invention relates to a method for annotating a dataset. It is applied to a terminal, which can be a computer, a server, etc., and the terminal can perform model training for deep learning. The annotation process is as Figure 1 shown as follows:

[0027] Step 101, the allocation step.

[0028] Specifically, preset classification identifiers are respectively assigned to each data in the dataset to obtain a first classification result, and the first classification result may include the classification identifiers corresponding to each data in the dataset.

[0029] In this embodiment, the allocation process is as follows: assume that the number of classification categories is N, then for any data d i ∈ D in the dataset D, a classification identifier l i is automatically and randomly assigned, where l i∈ {0, 1, …, N - 1}, that is to say, the preset classification identifier can be an identification number. When the terminal initially assigns the classification identifier, it cannot know the actual accurate classification results of each data. Therefore, in the first assignment, a random assignment method can be adopted.

[0030] Step 102, the feature learning step.

[0031] Specifically, use the data set after the classification identifier is assigned to perform feature learning on each preset classification to obtain the data features of each classification. More specifically, several features can be preset during feature learning, and the feature values of the preset features are extracted from each data, and the feature values corresponding to each classification are learned.

[0032] In one example, in this step, the feature learning of each preset classification can be specifically performed by using the method of training the model on the data set after the classification identifier is assigned. Specifically, the data set D can be used as the training set, and the classification identifier l i assigned in step 101 is used as the groundtruth (the calibrated real data in machine learning) to train the model, and finally the model M is obtained.

[0033] Step 103, the testing step.

[0034] Specifically, use the data features of each classification to classify each data in the data set to obtain the second classification result. The data features learned in step 102 can be used to test the data set D, that is, for any data d i in D, forward propagation is performed to obtain the predicted value p i , and the predicted value p i represents the classification corresponding to d i , and the second classification result can include each tested data and its corresponding classification.

[0035] Continuing to explain, during testing, all the data in the data set D can be tested, or a part of the data can be sampled from the data set D for testing to reduce the data processing volume during the testing process. For example, if there are 10,000 data in the data set D, 100 data are sampled at a ratio of 1% and tested using the features learned in step 102 to obtain the classifications corresponding to these 100 data, which is recorded as the second classification result.

[0036] In one example, if in step 102, the feature learning is performed by using the model training method and the model M is obtained, then during testing, the model M can be used to test each data in the data set D and obtain the predicted value. Through model testing, it is beneficial to improve the automation rate of the process and speed up the testing speed.

[0037] Step 104: Determine whether the matching degree between the second classification result and the first classification result is less than or equal to the first threshold. If so, return to execute Step 101. If not, continue to execute Step 105.

[0038] Specifically, determine whether the classification is correct by the classification of a certain data in the second classification result and the first classification result. For example, for data d i The classification in the first classification result is l i , and the classification in the second classification result is p i , compare p i with l i . If p i == l i , it means the classification is correct; otherwise, it means the classification is incorrect. According to this rule, the classification accuracy rate of the data participating in the test can be calculated, and this classification accuracy rate is the matching degree r between the second classification result and the first classification result. Compare the matching degree with the first threshold σ. If r > σ, the test is passed; otherwise, the test fails and it is necessary to return to Step 101. That is to say, when the matching degree between the second classification result and the first classification result is less than the first threshold, the allocation step is re-executed, and the classification identifiers of each data are re-allocated during re-allocation.

[0039] Continue to explain that after re-allocation, the classification identifier allocated to at least one data is different from the previously allocated one, so as to obtain a new first classification result.

[0040] It can be seen that after re-allocation of the classification identifier, feature learning and testing will be performed again, and continue to judge the relationship between the matching degree of the front and back classification results and the first threshold. If the matching degree is still low and the test cannot be passed, then continue to re-allocate the classification identifier until the features that can be learned according to the allocated classification identifier are appropriate. That is to say, when using the learned features for testing again, the matching degree can be higher than the first threshold, and then continue with the subsequent steps. It can be seen that when the matching degree between the second classification result and the first classification result is less than or equal to the first threshold, the allocation step to the testing step is re-executed until the matching degree between the second classification result and the first classification result is greater than the first threshold.

[0041] Step 105: Label determination step.

[0042] Specifically, in this step, the annotation of each data in the dataset is determined according to the classification identifier assigned when the matching degree between the second classification result and the first classification result is greater than or equal to the first threshold. More specifically, when reassigning, an iterative method of the assignment result can be adopted. When reassigning, the newly assigned classification identifier is used to replace the classification identifier assigned last time. Therefore, when determining the classification identifier assigned when the matching degree between the second classification result and the first classification result is greater than or equal to the first threshold, the currently corresponding classification identifier of each data is the classification identifier assigned when the matching degree between the second classification result and the first classification result is greater than or equal to the first threshold.

[0043] Continuing the description, this classification identifier can be directly used as the annotation of each data. That is to say, for any data d i ∈ D in the dataset D, its annotated classification is l i .

[0044] Furthermore, the annotation in this embodiment can be the identifier number of the classification identifier, such as 0, 1, 2, ……, N - 1, or it can be a specific classification name. The specific classification name can be confirmed after the classification result is confirmed, which will not be elaborated here.

[0045] It can be seen that in this embodiment, by assigning classification identifiers of uncertain correctness to the dataset, and using such classification results to determine the data characteristics under each classification, and then according to the characteristic that the data characteristics under a classification identifier are basically the same, it is verified in reverse whether the previously assigned classification identifier is accurate. When it is inaccurate, the assigned identifier is replaced until the assigned classification identifier conforms to the consistent data characteristics in the classification, so as to automatically confirm the annotation of the dataset, and the result is accurate and reliable. This avoids the way that existing datasets are all manually annotated, which not only greatly reduces the labor cost, but also avoids human errors, making the annotation result more objectively accurate. In addition, in this embodiment, feature learning is carried out through model training, so that feature learning is not restricted by predefined features, and thus more appropriate feature representations can be found, making the data characteristics more accurate and appropriate.

[0046] The second embodiment of the present invention relates to a method for annotating a dataset. The second embodiment is substantially the same as the first embodiment, and the main difference is that: in the first embodiment, in the assignment step, each data is assigned in an automatic assignment manner, while in the second embodiment of the present invention, some data in the dataset are manually assigned and some data are assigned in an automatic assignment manner. Since the classification of the manually assigned data is accurate, the data with uncertain classification is effectively reduced, thus accelerating the speed of finding the accurate classification result.

[0047] The flowchart of the method for annotating the dataset in this embodiment is as Figure 2 shown, specifically as follows:

[0048] Step 201, the allocation step.

[0049] Specifically, in this embodiment, a part of the data in the dataset can be manually allocated, and the other part of the data can be automatically allocated. For example, the dataset D in the above example is divided into dataset A and dataset B. For the data d i ∈A, its labeli i ∈{0, 1,..., N - 1} is given manually, and for the data d i ∈B, its labell i ∈{0, 1,..., N - 1} is automatically allocated.

[0050] Continuing the description, the amount of data for manual allocation can be less than the amount of data for automatic allocation. For example, there are 10,000 data in the dataset D in total, 3,000 data in dataset A, and the remaining 7,000 data in dataset B. The manually allocated part is manually labeled, and the automatically allocated part still uses the random allocation method. Compared with the prior art, the amount of manually allocated data is relatively small, but the overall speed of obtaining accurate classification results for the dataset can be effectively accelerated.

[0051] Steps 202 and 203 in this embodiment are similar to steps 102 and 103 in the first embodiment, and will not be elaborated here.

[0052] Step 204, determine whether the matching degree between the second classification result and the first classification result is less than or equal to the first threshold; if so, return to execute step 201; if not, continue to execute step 205.

[0053] Specifically, the process of calculating the matching degree between the second classification result and the first classification result in this step is similar to that in the first embodiment. After that, if the matching degree between the second classification result and the first classification result is less than or equal to the first threshold, then return to execute the allocation step again. That is to say, when the matching degree is too low, reallocation will be performed. When reallocating, this step can only reallocate the classification identifiers for the data in dataset B. For example, only reallocate the data in dataset B that was automatically allocated classification identifiers in the above example. Among them, reallocation can still be performed in an automatic random manner.

[0054] Step 205 in this embodiment is similar to step 105 in the first embodiment, and will not be elaborated here.

[0055] The inventors of this application found that during the annotation process, the accuracy of automatic annotation is uncertain, and during reallocation, the accuracy still cannot be determined.

[0056] When clearly allocating identifiers in this embodiment, some data are manually allocated. When reallocating, only the data with automatically allocated classification identifiers are reallocated, making the identifier results of this part of the data accurate, reducing the amount of data with uncertain identifiers, and accelerating the finding of appropriate classification identifier allocation results. At the same time, only part of the data is manually allocated, which not only reduces the labor cost but also greatly speeds up the acquisition of the annotation results. In addition, in this embodiment, the amount of manual allocation is less than that of automatic allocation, which not only controls the manual workload but also greatly speeds up the finding of appropriate classification results.

[0057] The third embodiment of the present invention relates to a method for annotating a data set. The third embodiment is a further improvement based on the first embodiment. The main improvement lies in that: in the first embodiment, during testing, the classification of each data is obtained, while in this embodiment, in addition to classifying each data using data features, the confidence level of the classification result of each data is also obtained. In this embodiment, during testing, the determination of the confidence level is combined. Those with a high confidence level are considered to be basically accurate and can no longer be reallocated, so as to reduce the amount of data to be reallocated and make the speed of finding accurate classification results faster.

[0058] Specifically, the method for annotating the data set in this embodiment is as Figure 1 shown as follows:

[0059] Steps 101 and 102 are similar to those in the first embodiment and will not be elaborated here.

[0060] Step 103, the testing step.

[0061] Specifically, the second classification result obtained by testing in the first embodiment may include each tested data and its corresponding classification. In this embodiment, when classifying each data in the data set using data features, in addition to obtaining the predicted value p i , the confidence level q i corresponding to each predicted value p i is also obtained.

[0062] Continuing to explain, if the confidence level is relatively high, it is considered that the previous classification result is credible and no reallocation is required. Otherwise, it is considered not credible and reallocation is required. Specifically, if the confidence level q i of its test value is greater than the threshold ε, we consider that the model has well extracted the features of this data and has well corresponded this feature with its classification identifier l i , so the label value of this data remains unchanged; if the confidence level q i of its test value is less than the threshold ε, we consider that the correspondence between the features extracted by the model and its classification identifier l i is not clear enough. Therefore, a new classification identifier l i is reallocated to di ', where l i ' ∈ {0, 1, …, N - 1} and l i ' ≠ l i .

[0063] After that, when re - executing the assignment step, re - assign classification labels to each data with a relatively low confidence in the classification result in the dataset. Specifically, in one example, a corresponding threshold ε can be set during confidence determination. A relatively low confidence can be that q is less than or equal to ε. That is to say, when re - assigning, it is necessary to compare q and ε for the data to be re - assigned. Specifically, as Figure 3 shown, for a certain data d, if q > ε, then the confidence in the classification result of this data is high, and the pre - assigned classification label is credible and does not need to be re - assigned. If q is less than or equal to ε, then the confidence in the classification result of this data is low, and the pre - assigned classification label is not credible and needs to be re - assigned.

[0064] Step 104 and step 105 are similar to those in the first embodiment and will not be elaborated here.

[0065] It can be seen that in this embodiment, during testing, the determination of confidence is combined. Those with high confidence are considered to be basically accurate and can no longer be re - assigned, so as to reduce the amount of data to be re - assigned and make the speed of finding the accurate classification result faster.

[0066] It should also be further noted that this embodiment can also be used in conjunction with the annotation method in the second embodiment. Specifically, when re - assigning the data with automatically assigned classification labels, it is also possible to first judge the confidence level of the data to be re - assigned. If the confidence q i is greater than the threshold ε, then we consider that the model has well extracted the features of this data and has well corresponded this feature with its label l i , so keep the label value of this data unchanged; if the confidence q i of its test value is less than the threshold ε, then we consider that the correspondence between the features extracted by the model and its label l i is not clear enough. Therefore, re - assign a label l′ i to d i , where l′ i ≠ l i , and only re - assign the data with relatively low confidence, further reducing the amount of data to be re - assigned.

[0067] The fourth embodiment of the present invention relates to a method for annotating a data set. The fourth embodiment is a further improvement based on the first embodiment. The main improvement lies in that: in the first embodiment, after it is determined that the matching degree between the second classification result and the first classification result is high, it is considered that the appropriate classification result has been found. However, in this embodiment, it is also necessary to determine in combination with the number of classifications, so that the number of classifications can be gradually confirmed from less to more, which can effectively accelerate the convergence speed of the algorithm.

[0068] Specifically, the flowchart of the method for annotating a data set in this embodiment is as Figure 4 shown as follows:

[0069] Steps 401 to 403 are similar to steps 101 to 103 in the first embodiment, and will not be elaborated here.

[0070] Step 404, determine whether the matching degree between the second classification result and the first classification result is less than or equal to the first threshold; if so, return to execute step 401; if not, continue to execute step 405.

[0071] Step 405, determine whether the number of classification identifiers is less than the second threshold; if so, execute step 406; if not, execute step 408.

[0072] In this step, it is specifically determined whether the number of classification identifiers corresponding to the data in the data set is less than the second threshold when the matching degree between the second classification result and the first classification result is greater than the first threshold. If it is less than the second threshold, it is considered that the number of classification identifiers has not reached the requirement, so the classification identifiers can be continuously expanded. If it is greater than or equal to the second threshold, it is considered that the number of classification identifiers has reached the requirement, so there is no need to continue to expand the classification identifiers, and enter the annotation confirmation step.

[0073] Step 406, perform category expansion on each classification identifier.

[0074] Specifically, in the expansion, in-class expansion can be performed on each category. In one example, the following calculation is performed on the current number of classifications n: n' = min(2 * n, N), where N is the number of classification identifiers required. If n' == 2 * n, that is, the number of expanded categories is twice the original number of categories, then the data in each category is binary-classified, so that the number of label categories increases to n'; if n' < 2 * n, then randomly select n' - n categories from the n categories for binary-classification operations, so that the number of classification identifiers increases to n'. Finally, let n = n t , and continue with the subsequent in-class reallocation steps.

[0075] In one example, if N > 2 * n, then first expand the categories to 2 * n, and then perform multiple category expansions until n t= N.

[0076] Step 407, within-class reallocation.

[0077] In this step, specifically, extended classification identifiers are assigned to the data in the dataset. During reallocation, reallocation is performed according to class extension. For example, for the dataset C under the classification identifier n, if n is extended to n1 and n2, then the data in dataset C can be assigned the classification identifier n1 or n2. If there is no extended class, then the data under that class does not need to be reallocated either. After the reallocation of the data that needs to be reallocated is completed, return to step 402, and perform steps 402 to 404 again until the matching degree between the second classification result and the first classification result is greater than the first threshold, and the number of extended classification identifiers is greater than or equal to the second threshold.

[0078] Step 408 is similar to step 105 in the first embodiment, and will not be elaborated again here.

[0079] The inventors of the present application found that there may be requirements for the number of categories in the annotation of the dataset. If a large number of categories are directly used to classify each data, it may lead to too slow or difficult convergence of the algorithm. During the annotation process, the number of classifications can be gradually increased to speed up the convergence. Therefore, in this embodiment, when presetting the classification identifier, a smaller number of categories can be set first. Since the number of identifiers is small, the speed of finding the accurate classification result can be accelerated. Then, within-class extension can be gradually performed to increase the classification. Through the above process from coarse to fine, the overall speed of finding the accurate classification result can be effectively accelerated.

[0080] The fifth embodiment of the present invention relates to a method for annotating a dataset. The fifth embodiment is a further improvement based on the first embodiment. The main improvement lies in that: before reallocation, the annotation accuracy is determined for each category respectively, the classification identifiers with higher accuracy are fixed, and the remaining classification identifiers are reallocated, thereby continuously reducing the amount of data during reallocation and effectively accelerating the convergence speed of the algorithm.

[0081] The method for annotating the dataset in this embodiment is as Figure 5 shown, specifically as follows:

[0082] Steps 501 to 503 are similar to steps 101 to 103 in the first embodiment, and will not be elaborated here.

[0083] Step 504, determine whether the matching degree between the second classification result and the first classification result is less than or equal to the first threshold; if so, return to execute step 505; if not, continue to execute step 507.

[0084] Step 505, calculate the classification result matching degree under each classification identifier respectively.

[0085] Step 506: Determine the classification identifiers with a matching degree less than or equal to the third threshold, denoted as the first type of classification identifiers.

[0086] Specifically, in steps 505 to 506, the classification accuracy rate r of each category is calculated respectively j , where j ∈ {0, 1,..., N - 1}. If the accuracy rate r j of a certain category j is greater than the third threshold τ, then this category belongs to the reliable category X (i.e., the classification identifier that does not need to be reallocated), otherwise, it belongs to the unreliable category Y (i.e., the first type of classification identifier that needs to be reallocated). Among them, the third threshold τ and the first threshold σ can be set as needed, and can be set to the same value or different values.

[0087] Then return to step 501. When reallocating the classification identifiers for each data in the dataset, reallocate the classification identifiers for each data belonging to the first type of classification identifiers in the dataset.

[0088] After that, step 507 is similar to step 105 in the first embodiment, and will not be elaborated here.

[0089] It can be seen that in this embodiment, by only reallocating the identifiers with poor classification results during reallocation, the amount of reallocation is reduced, which is beneficial to accelerating the speed of obtaining appropriate classification results.

[0090] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process are all within the protection scope of this patent.

[0091] The sixth embodiment of the present invention relates to a labeling device for a dataset, as Figure 6 shown, including:

[0092] An allocation module, configured to respectively allocate preset classification identifiers to each data in the dataset to obtain a first classification result.

[0093] A feature learning module, configured to perform feature learning on each preset classification by using the dataset after allocating classification identifiers to obtain the data features of each classification.

[0094] A testing module, configured to classify each data in the dataset by using the data features of each classification to obtain a second classification result.

[0095] A comparison module, configured to re - execute the allocation step to the testing step until the matching degree between the second classification result and the first classification result is greater than the first threshold when the matching degree between the second classification result and the first classification result is less than or equal to the first threshold.

[0096] A labeling determination module, configured to determine the labels of the data in the dataset according to the classification identifiers assigned when the matching degree between the second classification result and the first classification result is greater than the first threshold.

[0097] Furthermore, in one example, the feature learning module can specifically perform feature learning on each preset classification by training a model on the dataset after the classification identifiers are assigned.

[0098] In one example, the allocation module can specifically manually allocate some of the data in the dataset and automatically allocate the other part of the data in the dataset; correspondingly, when the allocation module re - allocates, it re - allocates the classification identifiers of the data in the dataset that are automatically allocated classification identifiers. Specifically, the amount of data manually allocated can be less than the amount of data automatically allocated.

[0099] In one example, the labeling device of the dataset further includes: a category expansion module, which is respectively connected to the comparison module and the labeling determination module. When the number of classification identifiers corresponding to the data in the dataset is less than the second threshold when the matching degree between the second classification result and the first classification result is greater than the first threshold, it expands the categories of each classification identifier, allocates the expanded classification identifiers to the data in the dataset, and re - executes the feature learning step to the comparison step until the matching degree between the second classification result and the first classification result is greater than the first threshold, and the number of the expanded classification identifiers is greater than or equal to the second threshold. Correspondingly, the labeling determination module can determine the labels of the data in the dataset according to the classification identifiers assigned when the matching degree between the second classification result and the first classification result is greater than the first threshold and the number of the expanded classification identifiers is greater than or equal to the second threshold.

[0100] In one example, the labeling device of the dataset further includes:

[0101] A calculation module, configured to calculate the matching degree of the classification results under each classification identifier respectively.

[0102] A determination module, configured to determine the classification identifiers whose matching degree is less than or equal to the third threshold, denoted as the first - type classification identifiers.

[0103] Correspondingly, when the allocation module re - allocates the classification identifiers of the data in the dataset, it re - allocates the classification identifiers of each data belonging to the first - type classification identifiers in the dataset.

[0104] In one example, the second classification result obtained by the testing module may further include: the confidence of the classification result of each data;

[0105] Correspondingly, when the allocation module re-executes the allocation step, it can re-allocate classification identifiers to the data with relatively low confidence in the classification results in the dataset.

[0106] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details mentioned in the above method embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.

[0107] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of the present invention, units that are not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0108] The seventh embodiment of the present invention relates to a terminal, as Figure 7 shown, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for annotating a dataset as described in the first to fifth embodiments above.

[0109] Wherein, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0110] Wherein, the processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when performing operations.

[0111] The eighth embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method embodiments described above are implemented.

[0112] That is, those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0113] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made to them in form and details without departing from the spirit and scope of the present invention.

Claims

1. A method for annotating a dataset, characterized in that, the method is applied to model training in deep learning, and the method includes: An assignment step of respectively assigning a preset classification identifier to each data in the dataset to obtain a first classification result, wherein the dataset is a dataset of images; A feature learning step of performing feature learning on each preset classification by using the dataset after the classification identifiers are assigned to obtain data features of each classification; A testing step of classifying each data in the dataset by using the data features of each classification to obtain a second classification result; A comparison step of, when the matching degree between the second classification result and the first classification result is less than or equal to a first threshold, re-executing the assignment step to the testing step until the matching degree between the second classification result and the first classification result is greater than the first threshold; An annotation determination step of determining the annotation of each data in the dataset according to the classification identifier assigned when the matching degree between the second classification result and the first classification result is greater than the first threshold, and the annotated dataset is used for model training in a deep learning manner; After the comparison step and before the annotation determination step, it includes: A category expansion step of, when the number of classification identifiers corresponding to the data in the dataset is less than a second threshold when the matching degree between the second classification result and the first classification result is greater than the first threshold, expanding the category of each classification identifier, assigning the expanded classification identifier to the data in the dataset, and re-executing the feature learning step to the comparison step until the matching degree between the second classification result and the first classification result is greater than the first threshold and the number of the expanded classification identifiers is greater than or equal to the second threshold; The annotation determination step of determining the annotation of each data in the dataset according to the classification identifier assigned when the matching degree between the second classification result and the first classification result is greater than the first threshold and the number of the expanded classification identifiers is greater than or equal to the second threshold.

2. The method for annotating a dataset according to claim 1, characterized in that, the feature learning step specifically performs feature learning on each preset classification by using a method of training a model on the dataset after the classification identifiers are assigned.

3. The method for annotating a dataset according to claim 1, characterized in that, the assignment step includes: Manually assigning a part of the data in the dataset and automatically assigning another part of the data in the dataset; When re-executing the assignment step, re-assigning the classification identifier to the data in the dataset that is automatically assigned the classification identifier.

4. The method for annotating a dataset according to claim 3, characterized in that, the amount of data manually assigned is less than the amount of data automatically assigned.

5. The method for annotating a dataset according to claim 1, characterized in that, before re-executing the assignment step, it includes: Calculating the classification result matching degree under each classification identifier respectively; Determining the classification identifier with a matching degree less than or equal to a third threshold, denoted as the first type of classification identifier; Reassign classification identifiers to each piece of data in the dataset, specifically: reassign classification identifiers to each piece of data belonging to the first type of classification identifier in the dataset.

6. The method for annotating a dataset according to any one of claims 1 to 5, wherein, the second classification result includes: the confidence level of the classification result of each piece of data; when re - executing the assignment step, reassign classification identifiers to each piece of data with a relatively low confidence level in the classification result in the dataset.

7. An apparatus for annotating a dataset, wherein, the apparatus is applied to the model training of deep learning, and the apparatus includes: an assignment module, configured to respectively assign preset classification identifiers to each piece of data in the dataset to obtain a first classification result, wherein the dataset is a dataset of images; a feature learning module, configured to perform feature learning on each preset classification using the dataset after the classification identifiers are assigned to obtain the data features of each classification; a testing module, configured to classify each piece of data in the dataset using the data features of each classification to obtain a second classification result; a comparison module, configured to, when the matching degree between the second classification result and the first classification result is less than or equal to a first threshold, re - execute the assignment step to the testing step until the matching degree between the second classification result and the first classification result is greater than the first threshold; an annotation determination module, configured to determine the annotation of each piece of data in the dataset according to the classification identifiers assigned when the matching degree between the second classification result and the first classification result is greater than the first threshold, and the annotated dataset is used for model training in a deep - learning manner; a category expansion module, respectively connected to the comparison module and the annotation determination module. When the number of classification identifiers corresponding to the data in the dataset is less than a second threshold when the matching degree between the second classification result and the first classification result is greater than the first threshold, perform category expansion on each classification identifier, assign the expanded classification identifiers to the data in the dataset, and re - execute the feature learning step to the comparison step until the matching degree between the second classification result and the first classification result is greater than the first threshold and the number of the expanded classification identifiers is greater than or equal to the second threshold; the annotation determination module determines the annotation of each piece of data in the dataset according to the classification identifiers assigned when the matching degree between the second classification result and the first classification result is greater than the first threshold and the number of the expanded classification identifiers is greater than or equal to the second threshold.

8. A terminal, wherein, it includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for annotating a dataset according to any one of claims 1 to 6.

9. A computer - readable storage medium storing a computer program, wherein, the computer program, when executed by a processor, implements the method for annotating a dataset according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Clothes identification method and system based on weakly annotated images

    CN107506793A

  • A method and apparatus for generating a training set

    CN109241997A