Data labeling method and device, electronic equipment and computer readable storage medium

By using labeling quality inspection data and real-time monitoring of labeling officers during the data labeling process, the problem of difficult data labeling quality in the existing technology is solved, and efficient and accurate data labeling is achieved.

CN120068020APending Publication Date: 2025-05-30NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411946714.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology conducts comprehensive acceptance after data annotation is completed, which is inefficient and prone to errors, resulting in low quality of data annotation.

Method used

By obtaining the preset ratio of the data to be marked as the marking quality inspection data, the data task package containing the marking quality inspection data is allocated to the marking officer, and when the marking officer error reaches the preset threshold, the remaining data tasks are assigned to other marking officers.

Benefits of technology

Real-time monitoring of the data labeling quality of the labeling officer is realized, and the labeling officer with low labeling level is discovered and stopped in a timely manner, preventing error accumulation, and improving the quality and efficiency of the data labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068020A_ABST
    Figure CN120068020A_ABST
Patent Text Reader

Abstract

The invention discloses a data annotation method and device, electronic equipment and a computer readable storage medium. The method comprises the steps that multiple pieces of data to be annotated are acquired, and data of a preset proportion is extracted from the data to be annotated to serve as annotation quality inspection data; the multiple pieces of data to be labeled are divided into a first number of data task packages, the data task packages are distributed to a second number of labeling personnel for data labeling tasks, and each data task package comprises at least one piece of labeled quality inspection data; in the data labeling process of a first labeling person, a labeling result, submitted by the first labeling person, of any labeled quality inspection data is obtained, and the first labeling person is any labeling person; and when the labeling result indicates that the number of labeling errors of the first labeling person reaches a preset number threshold value, distributing the remaining unlabeled data in the data task packet which is currently subjected to the data labeling operation by the first labeling person to at least one second labeling person. According to the invention, the data annotation quality of the annotator can be monitored in real time, and the data annotation quality and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data annotation, and specifically relates to a data annotation method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] Data annotation refers to manually or semi-automatically marking, classifying, annotating, or correcting data so that computer systems can understand and process this data. However, the professional capabilities of data annotators vary, and their work attitudes are also very different. To obtain accurate and reliable annotated data, it is necessary to identify efficient and conscientious annotators to assign work, thereby ensuring the quality of the annotated data.

[0003] In related technologies, usually after data annotation is completed, it is necessary to conduct data acceptance on the annotated data, calculate the annotation accuracy rate of each annotator, and eliminate the annotated data of the annotators with too low annotation accuracy rate to ensure the quality of the annotated data. However, this method must conduct data acceptance after all data annotation is completed. In the case of a large amount of data, the efficiency of comprehensive acceptance is low and it is prone to errors, resulting in low quality of the annotated data after acceptance. Summary of the Invention

[0004] The present application provides a data annotation method, apparatus, electronic device, and computer-readable storage medium to improve the quality of data annotation.

[0005] In a first aspect, an embodiment of the present application provides a data annotation method, and the method includes:

[0006] Obtain multiple pieces of data to be annotated and extract a preset proportion of the data as annotation quality inspection data from them. At least one of the following is included in the multiple pieces of data to be annotated: data that has been pre-annotated by a machine learning model and data that has not been pre-annotated by a machine learning model;

[0007] Obtain the standard annotation results corresponding to each of the annotation quality inspection data;

[0008] Divide the multiple pieces of data to be annotated into a first number of data task packages, and allocate the first number of data task packages to a second number of annotators for data annotation tasks. At least one piece of annotation quality inspection data is included in the data task package;

[0009] During the process of the first annotator performing data annotation, obtain the annotation result submitted by the first annotator for any one of the annotation quality inspection data, where the first annotator is any one of the annotators;

[0010] When the annotation result indicates that the number of annotation errors of the first annotator reaches a preset quantity threshold, the remaining unannotated data in the data task package for which the first annotator is currently performing data annotation operations is assigned to at least one second annotator.

[0011] In a second aspect, an embodiment of the present application provides a data annotation device, which includes:

[0012] A first acquisition module, configured to acquire multiple pieces of data to be annotated and extract a preset proportion of the data as annotation quality inspection data, where the multiple pieces of data to be annotated include at least one of the following: data that has been pre-annotated by a machine learning model and data that has not been pre-annotated by a machine learning model;

[0013] A second acquisition module, configured to acquire the standard annotation result corresponding to each piece of the annotation quality inspection data;

[0014] A first allocation module, configured to divide the multiple pieces of data to be annotated into a first number of data task packages, and allocate the first number of data task packages to a second number of annotators for data annotation tasks, where each data task package includes at least one piece of annotation quality inspection data;

[0015] A processing module, configured to, during the process of the first annotator performing data annotation, acquire the annotation result submitted by the first annotator for any piece of annotation quality inspection data, where the first annotator is any one of the annotators;

[0016] A second allocation module, configured to, when the annotation result indicates that the number of annotation errors of the first annotator reaches a preset quantity threshold, allocate the remaining unannotated data in the data task package for which the first annotator is currently performing data annotation operations to at least one second annotator.

[0017] In a third aspect, an embodiment of the present application provides an electronic device, which includes:

[0018] A memory and a processor, the memory and the processor being coupled;

[0019] The memory is used to store one or more computer instructions;

[0020] The processor is used to execute the one or more computer instructions to implement the data annotation method according to any one of the first aspects described above.

[0021] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which one or more computer instructions are stored, and characterized in that the instructions are executed by a processor to implement the data annotation method according to any one of the first aspects described above.

[0022] Fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the data annotation method according to any one of the above first aspects when executed by a processor.

[0023] Compared with the prior art, the present application has the following advantages:

[0024] For the data annotation method provided by the present application, multiple pieces of data to be annotated are obtained, and a preset proportion of the data is extracted as annotation quality inspection data. The multiple pieces of data to be annotated include at least one of the following: data that has been pre-annotated by a machine learning model and data that has not been pre-annotated by a machine learning model. The standard annotation results corresponding to each piece of annotation quality inspection data are obtained. The multiple pieces of data to be annotated are divided into a first number of data task packages, and the first number of data task packages are assigned to a second number of annotators for data annotation tasks. Each data task package includes at least one piece of annotation quality inspection data. During the process of the first annotator performing data annotation, the annotation result submitted by the first annotator for any piece of annotation quality inspection data is obtained, and the first annotator is any one of the annotators. When the annotation result indicates that the number of annotation errors of the first annotator reaches a preset quantity threshold, the remaining unannotated data in the data task package that the first annotator is currently performing data annotation on is assigned to at least one second annotator.

[0025] Compared with the prior art, in the present application, by assigning a data task package containing a certain number of annotation quality inspection data to each annotator, and during the process of each first annotator performing data annotation, the annotation result submitted by the first annotator for any piece of annotation quality inspection data is obtained. In this way, during the process of the annotator performing data annotation, the data annotation quality of each annotator can be monitored in real time to a certain extent. When the annotation result indicates that the number of annotation errors of the annotator reaches a preset quantity threshold, the remaining unannotated data in the data task package that the first annotator is currently performing data annotation on is assigned to at least one second annotator. In this way, it can be timely discovered and the annotation work of the annotator with low annotation level can be stopped, which can effectively prevent error accumulation, and improve the overall data annotation quality and efficiency. At the same time, the present application does not need to wait until all data annotation is completed to accept the annotation quality, but can synchronously accept the annotation quality of each annotator during the process of the annotator performing data annotation, thereby reducing the difficulty of the acceptance work, shortening the annotation task duration and the acceptance duration, and further reducing the cost. Description of the Drawings

[0026] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0027] Figure 1Schematic flowchart of the data annotation method provided by one embodiment of the present application;

[0028] Figure 2 Schematic structural diagram of the data annotation device provided by one embodiment of the present application;

[0029] Figure 3 Schematic hardware structure diagram of the electronic device provided by one embodiment of the present application.

[0030] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be given later. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0031] To make the objectives, advantages, and features of the present application clearer, the present application will be clearly and completely described below in conjunction with the accompanying drawings and specific implementation manners. In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0032] It should be noted that in the description of the present application, terms such as "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance, as well as a specific order or sequence. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. In addition, in the description of the present application, unless otherwise specified, the term "plurality" means two or more. The term "and / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0033] To facilitate an understanding of the technical solution of the present application, the relevant concepts involved in the present application will be introduced first.

[0034] Data annotation refers to manually or semi-automatically marking, classifying, annotating, or correcting data so that computer systems can understand and process this data. This data processing method is usually applied in the fields of machine learning and artificial intelligence to create training datasets, train models, or evaluate model performance. However, the professional capabilities and work attitudes of data annotators vary greatly. To obtain accurate and reliable annotated data, it is necessary to identify efficient and conscientious annotators to assign work, thereby ensuring the quality of the annotated data.

[0035] Next, the prior art related to this application and the problems existing in the prior art will be described:

[0036] In the related art, usually after the data is annotated, the quality inspector manually screens and accepts the data, calculates the annotation accuracy of each annotator, and eliminates the annotated data of the annotators with too low annotation accuracy, so as to ensure the quality of the annotated data.

[0037] However, the above prior art still has the following problems:

[0038] Problem 1: The above solution must conduct data acceptance after all the data is annotated. When the quality inspection workload is large and the quality inspector is under great work pressure, comprehensive manual acceptance is prone to errors, resulting in low quality of the annotated data after acceptance and prolonging the duration of the annotation task, wasting costs.

[0039] Problem 2: The above solution lacks a real-time data annotation quality monitoring and warning system and cannot track the data annotation quality of each annotator in real time during the process of the data annotation task. That is to say, it is impossible to timely discover and terminate the annotation operations of "low-quality" annotators during the data annotation task and reassign the annotation task to other annotators, resulting in uneven data annotation quality, increasing the difficulty of quality inspection and acceptance work, and ultimately prolonging the duration of the annotation task and wasting costs.

[0040] To solve the above existing problems, this application provides a data annotation method, a data annotation device corresponding to this method, an electronic device capable of implementing this data annotation method, and a computer-readable storage medium. The following provides embodiments to describe the above method, device, electronic device, and computer-readable storage medium in detail.

[0041] To make the objectives and technical solutions of this application clearer and more intuitive, the following will, in conjunction with the accompanying drawings and embodiments, provide a detailed description of the method provided in the embodiments of this application. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. It can be understood that the following several embodiments can exist independently, and when there is no conflict between the embodiments provided in this application, the following embodiments and the features in the embodiments can be combined with each other. For the same or similar content, it will not be repeated in different embodiments. In addition, the step timings in the following method embodiments are only examples and are not strictly limited. In some cases, the steps shown or described can be executed in a different order than this.

[0042] This application provides a data annotation method, apparatus, electronic device, and computer-readable storage medium. Specifically, the data annotation method in one embodiment of this application can be executed by a computer device, where the computer device can be a terminal or a server, etc. The terminal can be a terminal device such as a smart phone, a tablet computer, a laptop computer, a touch screen, etc. The terminal can also include a client, and the client can be a game application client, a browser client carrying a game program, or an instant messaging client, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms.

[0043] Next, in conjunction with Figure 1 , a data annotation method provided in one embodiment of this application will be described. Figure 1 is a schematic flowchart of the data annotation method provided in one embodiment of this application.

[0044] As Figure 1 shown, the data annotation method includes steps S10 - S50:

[0045] S10. Obtain multiple pieces of data to be annotated and extract a preset proportion of the data as annotation quality inspection data. The multiple pieces of data to be annotated include at least one of the following: data that has been pre-annotated by a machine learning model and data that has not been pre-annotated by a machine learning model.

[0046] S20. Obtain the standard annotation results corresponding to each piece of annotation quality inspection data.

[0047] S30. Divide the multiple pieces of data to be annotated into a first number of data task packages and allocate the first number of data task packages to a second number of annotators for data annotation tasks. Each data task package includes at least one piece of annotation quality inspection data.

[0048] S40. During the process of data annotation by the first annotator, obtain the annotation result submitted by the first annotator for any annotation quality inspection data, where the first annotator is any one of the annotators.

[0049] S50. When the annotation result indicates that the number of annotation errors of the first annotator reaches the preset quantity threshold, allocate the remaining unannotated data in the data task package for which the first annotator is currently performing data annotation operations to at least one second annotator.

[0050] Next, steps S10 - S50 will be described in detail.

[0051] As mentioned above, the data to be annotated refers to unannotated data, and at least one of the following is included in the multiple pieces of data to be annotated: data that has been pre - annotated by a machine learning model and data that has not been pre - annotated by a machine learning model. That is to say, all of the multiple pieces of data to be annotated can be data that has been pre - annotated by a machine learning model, or all can be data that has not been pre - annotated by a machine learning model, or can include both data that has been pre - annotated by a machine learning model and data that has not been pre - annotated by a machine learning model. If a piece of data to be annotated is data that has been pre - annotated by a machine learning model, then this piece of data also includes the annotation result pre - annotated by the machine learning model. When the data to be annotated includes data that has been pre - annotated by a machine learning model, the purpose of subsequent re - annotation is to review and correct the data pre - annotated by the machine learning model and the corresponding annotation results.

[0052] As mentioned above, the types of data to be annotated include one or more of the following: text, numbers, pictures, tabular data. By way of example only, the embodiments of the present application do not make any limitations on the data types.

[0053] In the embodiments of the present application, multiple pieces of data to be annotated are obtained, and these multiple pieces of data to be annotated need to be subsequently annotated by data annotators (abbreviated as annotators). Data annotation can be understood as annotating the data to be annotated, such as tagging, classification, annotation, correction, etc. Among them, the annotator can be a machine learning model or a real person, by way of example only. Among them, a machine learning model is, for example, a model trained through a large amount of training data, which is used to receive input data and output the predicted label corresponding to the input data to tag the input data, and a real person is, for example, a professionally trained annotator, a user who accepts data annotation tasks on a crowdsourcing platform, etc.

[0054] In the embodiments of the present application, a preset proportion of data is extracted from the multiple pieces of data to be labeled as labeled quality inspection data, and the standard labeling results corresponding to each piece of labeled quality inspection data are obtained. Among them, the standard labeling result corresponding to the labeled quality inspection data refers to the correct labeling result of this piece of labeled quality inspection data. That is to say, a small part of data is extracted from a large amount of data to be labeled as labeled quality inspection data. During the subsequent data labeling process by labelers, by checking the data labeling quality of this small part of standard quality inspection data, the data labeling level of the labelers can be known. The preset proportion can be, for example, 10%, 20%, etc.

[0055] Exemplarily, taking text data as an example, category labeling is performed on the text data. For example, text data 1 is "This is an article about artificial intelligence. Artificial intelligence is a branch of computer science...", and the standard labeling result corresponding to this text data 1 is technology; text data 2 is "The snowflakes in winter gently fall, covering the earth and forming a white blanket. The trees are dressed in silver...", and the standard labeling result corresponding to this text data 2 is natural landscape.

[0056] In the embodiments of the present application, after obtaining the standard labeling results corresponding to each piece of labeled quality inspection data, the multiple pieces of data to be labeled are divided into a first number of data task packages, and each data task package includes at least one piece of labeled quality inspection data. For example, among 100 pieces of data to be labeled, 20 pieces of labeled quality inspection data are included. The 100 pieces of data to be labeled are divided into 5 data task packages, and each data task package includes at least 3 pieces of labeled quality inspection data.

[0057] The first number of data task packages are assigned to a second number of labelers for data labeling tasks, so that each labeler's obtained data task package includes a certain number of labeled quality inspection data. Among them, the second number is less than or equal to the first number. Specifically, when assigning data task packages to the second number of labelers, one data task package can be assigned to each person at the beginning. After a labeler completes the data labeling tasks for all the data in the currently owned data task package, then the next data task package is assigned to this labeler.

[0058] In the embodiments of the present application, after each labeler obtains the data task package, the labeler starts to label the data to be labeled in the data task package. During the data labeling process of the first labeler, the first labeler submits the labeling result of the data immediately after completing the labeling of each piece of data to be labeled, and obtains the labeling result submitted by the first labeler for any piece of labeled quality inspection data. Among them, the first labeler is any labeler. That is to say, during the data labeling process of each labeler, the above steps are all executed.

[0059] Optionally, after obtaining the annotation result submitted by the first annotator for any annotation quality inspection data, the annotation result is compared with the standard annotation result corresponding to the annotation quality inspection data. When the comparison result is inconsistent, the number of annotation errors of the first annotator is accumulated, and the initial value of the number of annotation errors is zero.

[0060] Since each annotator's data task package includes a certain number of annotation quality inspection data, when an annotator completes the annotation of one piece of annotation quality inspection data each time, the annotation result given by the annotator will be compared with the standard annotation result corresponding to the annotation quality inspection data. If the comparison result is consistent, it indicates that the annotation result given by the annotator for the annotation quality inspection data is correct; on the contrary, if the comparison result is inconsistent, it indicates that the annotation result given by the annotator for the annotation quality inspection data is wrong. Once the comparison result is detected to be inconsistent, the number of annotation errors of the annotator is accumulated. In this way, the data annotation quality of the annotator can be monitored in real time during the data annotation process.

[0061] Exemplarily, taking text data as an example, category annotation is performed on the text data. For example, text data 1 is "This is an article about artificial intelligence. Artificial intelligence is a branch of computer science...", and the standard annotation result corresponding to this text data 1 is technology. The annotation result given by annotator 01 is humanities, and the comparison result is inconsistent. Therefore, the number of annotation errors of this annotator 01 is accumulated, that is, the number of annotation errors = 0 + 1 = 1; text data 2 is "The snowflakes in winter gently fall, covering the earth and forming a white blanket. The trees are dressed in silver...", and the standard annotation result corresponding to this text data 2 is natural landscape. The annotation result given by annotator 01 is plants, and the comparison result is inconsistent. Therefore, the number of annotation errors of this annotator 01 is accumulated, that is, the number of annotation errors = 1 + 1 = 2.

[0062] In the embodiment of the present application, when the annotation result indicates that the number of annotation errors of the first annotator reaches the preset number threshold, the remaining unannotated data in the data task package currently being annotated by the first annotator is assigned to at least one second annotator. Among them, the second annotator is different from the first annotator.

[0063] It can be understood that when the annotation result indicates that the number of annotation errors of the first annotator reaches the preset number threshold, this means that the first annotator has many annotation errors and low annotation level. In order to avoid uneven data annotation quality in the end, in the present application, the annotation work of the annotator with low annotation level is stopped in time, and the remaining unannotated data in the data task of the annotator with low annotation level is assigned to other annotators for annotation. Among them, the remaining unannotated data can be assigned to one or more second annotators with higher annotation levels.

[0064] The data annotation method provided by the embodiment of the present application obtains multiple pieces of data to be annotated and extracts a preset proportion of the data as annotation quality inspection data. Among the multiple pieces of data to be annotated, at least one of the following is included: data that has been pre-annotated by a machine learning model and data that has not been pre-annotated by a machine learning model. Obtain the standard annotation results corresponding to each annotation quality inspection data. Divide the multiple pieces of data to be annotated into a first number of data task packages, and allocate the first number of data task packages to a second number of annotators for data annotation tasks. Each data task package includes at least one piece of annotation quality inspection data. During the process of the first annotator performing data annotation, obtain the annotation result submitted by the first annotator for any piece of annotation quality inspection data, where the first annotator is any one of the annotators. When the annotation result indicates that the number of annotation errors of the first annotator reaches a preset quantity threshold, allocate the remaining unannotated data in the data task package currently being used by the first annotator for data annotation to at least one second annotator.

[0065] Compared with the prior art, in the present application, by allocating a data task package containing a certain number of annotation quality inspection data to each annotator, during the process of each first annotator performing data annotation, obtain the annotation result submitted by the first annotator for any piece of annotation quality inspection data. In this way, during the process of the annotator performing data annotation, the data annotation quality of each annotator can be monitored in real time to a certain extent. When the annotation result indicates that the number of annotation errors of the annotator reaches a preset quantity threshold, allocate the remaining unannotated data in the data task package currently being used by the first annotator for data annotation to at least one second annotator. In this way, it is possible to timely discover and stop the annotation work of the annotator with low annotation level, effectively prevent error accumulation, and improve the overall data annotation quality and efficiency. At the same time, in the present application, it is not necessary to wait until all data annotation is completed to accept the annotation quality. Instead, during the process of the annotator performing data annotation, the annotation quality of each annotator can be synchronously accepted, thereby reducing the difficulty of the acceptance work, shortening the annotation task duration and the acceptance duration, and further reducing the cost.

[0066] Based on the above embodiments, the data annotation method provided by the embodiment of the present application is further described below.

[0067] An optional implementation manner. The possible implementation manners of the above step 20 "obtain the standard annotation results corresponding to each annotation quality inspection data" include the following steps S201 - S203:

[0068] S201. Obtain the user portraits of multiple annotators. The user portraits include professional capabilities and work attitudes.

[0069] S202. Determine at least one model annotator from multiple annotators according to the professional capabilities and work attitudes in the user portraits.

[0070] S203. Assign the labeled quality inspection data to at least one exemplary labeler for data labeling, and obtain the standard labeling results corresponding to each piece of labeled quality inspection data submitted by the exemplary labeler.

[0071] As described above, the user profile of the labeler is used to comprehensively understand the data labeling ability and attitude of the labeler. The user profile includes specific scores or descriptions of the professional ability (i.e., labeling ability) and work attitude of the labeler. According to the professional ability and work attitude in the user profile, select the second number of labelers with strong professional ability and serious work attitude from multiple labelers as exemplary labelers. Assign the labeled quality inspection data to at least one exemplary labeler for data labeling, and obtain the standard labeling results corresponding to each piece of labeled quality inspection data submitted by the exemplary labeler.

[0072] It should be noted that the multiple labelers here may or may not include the second number of labelers who are assigned data task packages in step S30. That is to say, this application can select exemplary labelers from labelers other than the second number of labelers who are assigned data task packages in step S30. Only by way of example, this application does not impose any restrictions on this.

[0073] In the embodiment of this application, at least one exemplary labeler is determined from multiple labelers according to the professional ability and work attitude in the user profile, so that the labeling ability and work attitude of the labeler can be comprehensively evaluated, and the most excellent exemplary labeler can be selected to be responsible for labeling the labeled quality inspection data to obtain the standard labeling results corresponding to the labeled quality inspection data, which can ensure the correctness of the standard labeling results.

[0074] An alternative implementation, the user profile further includes at least one of the following: age, education level, professional background, the correct number and accuracy rate of labeling for each type of data, so that the overall situation of the labeler can also be obtained according to other evaluation criteria in the user profile.

[0075] An alternative implementation, the data labeling method provided by the embodiment of this application further includes step S60:

[0076] S60. During the process of the first labeler performing the data labeling task, update the user profile and labeling level label of the first labeler according to the number of labeling errors of the first labeler.

[0077] In the embodiment of this application, during the process of the first labeler performing the data labeling task, the user profile and labeling level label of the first labeler are updated in real time according to the number of labeling errors of the first labeler, so as to ensure that the user profile of the labeler is up-to-date and can more accurately reflect the overall situation of the labeler.

[0078] As described above, the annotation level label can be determined according to the number of annotation errors of the annotator. For example, if the number of annotation errors is 0-1, the annotation level label is excellent; if the number of annotation errors is 2-5, the annotation level label is good; if the number of annotation errors is 5-8, the annotation level label is excellent; if the number of annotation errors is more than 8, the annotation level label is poor. The above is only an example, and the embodiments of the present application do not make any limitations thereto.

[0079] In the embodiments of the present application, by dynamically updating the user profile and the annotation level label, the data annotation task assignment can be dynamically adjusted to ensure the workload balance of each annotator and avoid excessive burden on individual annotators. At the same time, by dynamically updating the user profile, the annotator can timely understand their own deficiencies, make improvements, and improve their annotation skills and quality awareness.

[0080] An optional implementation manner is that, in steps S10-S50, or on the basis of step S60, before "assigning the remaining unannotated data in the data task package currently being used by the first annotator for data annotation to at least one second annotator" in step S50, the data annotation method provided by the embodiments of the present application further includes steps S501-S502:

[0081] S501. Obtain the user profiles of the second number of annotators.

[0082] S502. Determine at least one second annotator from the second number of annotators according to the professional ability and work attitude in the user profile.

[0083] In the embodiments of the present application, the user profiles of the second number of annotators are obtained. According to the professional ability and work attitude in the user profile, at least one second annotator is determined from the second number of annotators to replace the first annotator and perform data annotation work on the remaining unannotated data in the data task package currently being used by the first annotator for data annotation. According to the professional ability and work attitude in the user profile, at least one annotator with strong professional ability and serious work attitude is selected from the second number of annotators other than the first annotator as the second annotator. This can further ensure that at least one selected second annotator can complete the annotation work on the remaining unannotated data of the first annotator with high quality, and further improve the data annotation quality.

[0084] An optional implementation manner is that, before step S30 "dividing multiple pieces of data to be annotated into the first number of data task packages", the data annotation method provided by the embodiments of the present application further includes step A1:

[0085] A1. Randomly configure the order of the annotation quality inspection data among the multiple pieces of data to be annotated.

[0086] In the embodiments of the present application, before dividing multiple pieces of data to be labeled into the first number of data task packages, the order of the labeled quality inspection data among the multiple pieces of data to be labeled is randomly configured, that is, randomly shuffled. This can ensure that the labeled quality inspection data is evenly distributed into each data task package, avoiding the situation where the number of labeled quality inspection data in some data task packages is too large while the number of labeled quality inspection data in some data task packages is extremely small or even zero.

[0087] In addition, before the step S50 of "assigning the first number of data task packages to the second number of labelers for data labeling tasks", the data labeling method provided by the embodiments of the present application further includes step A2:

[0088] A2. For each data task package, randomly configure the order of the labeled quality inspection data in the data task package.

[0089] In the embodiments of the present application, the order of the labeled quality inspection data in the data task package is shuffled, which can prevent the labeler from discovering the order of the labeled quality inspection data in the data task package. Furthermore, it can avoid the speculative behavior that the labeler only seriously answers the labeled quality inspection data and randomly labels other data due to discovering or being unaware of the order of the labeled quality inspection data in the data task package. The order of the labeled quality inspection data in the data task package is prevented from being discovered by the labeler, which can further urge the quality inspector to seriously label each piece.

[0090] In an optional implementation manner, when the number of labeling errors of the first labeler reaches the preset number threshold, the data labeled by the first labeler and the corresponding data labeling results are deleted. On this basis, the embodiments of the present application may further include step S501:

[0091] S501. Assign the data in the data task package where the first labeler is currently performing data labeling operations to at least one second labeler.

[0092] In the implementation of this application, it is possible to determine whether to delete the data marked by the first annotator and the corresponding data annotation results when the number of annotation errors of the first annotator reaches a preset quantity threshold according to the current data annotation project budget. When the current data annotation project budget is sufficient, the data marked by the first annotator and the corresponding data annotation results are deleted, and the data in the data task package for which the first annotator is currently performing data annotation operations is assigned to at least one second annotator. On the one hand, by deleting the data marked by the first annotator and the corresponding data annotation results, it is possible to avoid the low-quality data annotation results of this annotator with low annotation level affecting the overall data annotation quality. On the other hand, by assigning the data in the data task package for which the first annotator is currently performing data annotation operations to at least one second annotator, the second annotator with a higher data annotation level can re-annotate the data marked by the first annotator with low data annotation level. In this way, while improving the data annotation quality, it also avoids the problem that some data are not annotated due to the deletion of the data marked by the first annotator and the corresponding data annotation results.

[0093] The data annotation device provided by this application will be described below. The data annotation device described below can be mutually corresponded and referred to the data annotation method described above.

[0094] Figure 2 It is a schematic structural diagram of the data annotation device provided by one embodiment of this application. As Figure 2 shown, the data annotation device 200 includes: a first acquisition module 201, a second acquisition module 202, a first allocation module 203, a processing module 204, and a second allocation module 205.

[0095] The first acquisition module is used to acquire multiple pieces of data to be annotated and extract a preset proportion of the data as annotation quality inspection data. At least one of the following is included in the multiple pieces of data to be annotated: data that has been pre-annotated by a machine learning model and data that has not been pre-annotated by a machine learning model;

[0096] The second acquisition module is used to acquire the standard annotation results corresponding to each of the annotation quality inspection data;

[0097] The first allocation module is used to divide the multiple pieces of data to be annotated into a first number of data task packages and allocate the first number of data task packages to the second number of annotators for data annotation tasks. At least one piece of annotation quality inspection data is included in the data task package;

[0098] The processing module is used to acquire the annotation result submitted by the first annotator for any annotation quality inspection data during the process of the first annotator performing data annotation. The first annotator is any one of the annotators;

[0099] A second allocation module, configured to, when the annotation result indicates that the number of annotation errors of the first annotator reaches a preset quantity threshold, allocate the remaining unannotated data in the data task package currently being used by the first annotator for data annotation tasks to at least one second annotator.

[0100] An alternative implementation, the second acquisition module is specifically configured to:

[0101] Acquire user portraits of multiple annotators, where the user portraits include professional capabilities and work attitudes;

[0102] Determine at least one exemplary annotator from the multiple annotators according to the professional capabilities and work attitudes in the user portraits;

[0103] Allocate the annotation quality inspection data to the at least one exemplary annotator for data annotation, and obtain the standard annotation results corresponding to each of the annotation quality inspection data submitted by the exemplary annotator.

[0104] An alternative implementation, the device further includes an update module, and the update module is specifically configured to:

[0105] During the process of the first annotator performing data annotation tasks, update the user portrait and annotation level label of the first annotator according to the number of annotation errors of the first annotator.

[0106] An alternative implementation, the second allocation module is further configured to:

[0107] Acquire the user portraits of the second number of annotators;

[0108] Determine at least one second annotator from the second number of annotators according to the professional capabilities and work attitudes in the user portraits.

[0109] An alternative implementation, the user portrait further includes at least one of the following: age, education level, professional background, the number of correctly annotated data of each type, and accuracy rate.

[0110] An alternative implementation, the first allocation module is further configured to:

[0111] Randomly configure the order of the annotation quality inspection data among the multiple pieces of data to be annotated;

[0112] For each of the data task packages, randomly configure the order of the annotation quality inspection data in the data task package.

[0113] An alternative implementation, the device further includes a deletion module, and the deletion module is specifically configured to:

[0114] When the number of annotation errors of the first annotator reaches a preset number threshold, the data annotated by the first annotator and the corresponding data annotation results are deleted;

[0115] The first allocation module is further configured to:

[0116] Allocate the data in the data task package for which the first annotator is currently performing data annotation tasks to at least one second annotator.

[0117] The processing module is further configured to:

[0118] Compare the annotation result with the standard annotation result corresponding to the annotation quality inspection data, and accumulate the number of annotation errors of the first annotator when the comparison result is inconsistent. The initial value of the number of annotation errors is zero.

[0119] The data annotation device provided in this embodiment can be used to execute the technical solutions of the above data annotation method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0120] Figure 3 It is a schematic hardware structure diagram of an electronic device provided in one embodiment of the present application. As Figure 3 shown, the electronic device 300 in this embodiment includes: a processor 301 and a memory 302; where

[0121] The memory 302 is used to store computer execution instructions;

[0122] The processor 301 is configured to execute the computer execution instructions stored in the memory to implement each step executed by the data annotation method in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiment.

[0123] Optionally, the memory 302 can be either independent or integrated with the processor 301.

[0124] When the memory 302 is independently provided, the electronic device further includes a bus 303 for connecting the memory 302 and the processor 301.

[0125] One embodiment of the present application further provides a computer-readable storage medium, in which computer execution instructions are stored. When the processor executes the computer execution instructions, the technical solutions corresponding to the data annotation methods in any of the above embodiments executed by the electronic device are implemented.

[0126] One embodiment of the present application further provides a computer program product, which includes: a computer program stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and the execution of the computer program by the at least one processor causes the electronic device to execute the technical solution corresponding to the data annotation method in any of the above embodiments.

[0127] Although the present application is disclosed above in preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be determined by the scope defined by the claims of the present application.

[0128] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical or other forms.

[0129] The integrated modules implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present application.

[0130] It should be understood that the above processor can be a central processing module (English: Central Processing Unit, abbreviated as: CPU), and can also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0131] The memory may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc.

[0132] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the attached drawings of this application are not limited to only one bus or one type of bus.

[0133] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0134] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disk, or optical disc and other media that can store program codes.

[0135] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data labeling method, characterized in that: The method comprises: Acquire multiple pieces of data to be labeled and extract a preset proportion of data therefrom as labeling quality inspection data, wherein the multiple pieces of data to be labeled include at least one of the following: data that has been pre-labeled by the machine learning model and data that has not been pre-labeled by the machine learning model; Obtaining standard annotation results corresponding to each of the annotation quality inspection data; Dividing the plurality of data to be annotated into a first number of data task packages, and allocating the first number of data task packages to the second number of annotators to perform data annotation tasks, wherein the data task packages include at least one annotated quality inspection data; During the process of data labeling by the first labeler, obtaining a labeling result of any labeling quality inspection data submitted by the first labeler, wherein the first labeler is any labeler; When the labeling result indicates that the number of labeling errors of the first labeler reaches a preset number threshold, the remaining unlabeled data in the data task package of the data labeling job currently being performed by the first labeler is allocated to at least one second labeler.

2. The method according to claim 1, characterized in that The obtaining of the standard annotation results corresponding to each of the annotation quality inspection data includes: Obtaining user profiles of multiple labelers, wherein the user profiles include professional ability and work attitude; Determining at least one exemplary annotator from the plurality of annotators according to the professional ability and work attitude in the user portrait; The annotated quality inspection data are assigned to the at least one model annotator for data annotation, and a standard annotation result corresponding to each annotated quality inspection data submitted by the model annotator is obtained.

3. The method according to claim 1, characterized in that The method further comprises: During the process of the first labeler performing the data labeling task, the user profile and labeling level label of the first labeler are updated according to the number of labeling errors made by the first labeler.

4. The method according to claim 1 or 3, characterized in that: Before allocating the remaining unlabeled data in the data task package currently being used for the data labeling operation by the first labeler to at least one second labeler, the method further includes: Obtain user portraits of the second number of labelers; At least one second annotator is determined from the second number of annotators according to the professional ability and work attitude in the user portrait.

5. The method according to claim 2, characterized in that: The user portrait also includes at least one of the following: age, education background, professional background, and the number and accuracy of correct annotations for each type of data.

6. The method according to claim 1, characterized in that Before dividing the plurality of pieces of data to be labeled into a first number of data task packages, the method further includes: Randomly configuring the order of the labeled quality inspection data in the multiple pieces of data to be labeled; The allocating the first number of data task packages to the second number of labelers to perform data labeling tasks, the method further includes: For each of the data task packages, the order of the labeled quality inspection data in the data task package is randomly configured.

7. The method according to claim 1, characterized in that The method further comprises: When the number of labeling errors of the first labeler reaches a preset number threshold, deleting the data labeled by the first labeler and the corresponding data labeling results; Allocating remaining unlabeled data in the data task package currently being labeled by the first labeler to at least one second labeler includes: The data in the data task package of the data labeling operation currently being performed by the first labeler is distributed to at least one second labeler.

8. The method according to claim 1, characterized in that After obtaining the labeling result submitted by the first labeler for any labeling quality inspection data, the method further includes: The labeling result is compared with the standard labeling result corresponding to the labeling quality inspection data. When the comparison result is inconsistent, the number of labeling errors of the first labeler is accumulated, and the initial value of the number of labeling errors is zero.

9. A data labeling device, characterized in that: The device comprises: A first acquisition module is used to acquire multiple pieces of data to be labeled and extract a preset proportion of data therefrom as labeling quality inspection data, wherein the multiple pieces of data to be labeled include at least one of the following: data that has been pre-labeled by the machine learning model and data that has not been pre-labeled by the machine learning model; A second acquisition module is used to obtain the standard annotation results corresponding to each of the annotation quality inspection data; A first allocation module, configured to divide the plurality of data to be annotated into a first number of data task packages, and allocate the first number of data task packages to the second number of annotators to perform data annotation tasks, wherein the data task packages include at least one annotated quality inspection data; A processing module, used for obtaining a labeling result of any labeling quality inspection data submitted by a first labeler during a process of labeling data by the first labeler, wherein the first labeler is any labeler; A second allocation module is configured to allocate remaining unlabeled data in a data task package currently being used for data labeling by the first labeler to at least one second labeler when the labeling result indicates that the number of labeling errors by the first labeler reaches a preset number threshold.

10. An electronic device, characterized in that: The electronic device comprises: Processor; and The memory is used to store a data processing program. After the electronic device is powered on and runs the program through the processor, the data labeling method according to any one of claims 1 to 8 is executed.

11. A computer-readable storage medium, characterized in that: A data processing program is stored, and the program is run by a processor to execute the data labeling method as described in any one of claims 1-8.

Citation Information

Cited By

  • Annotation support device, method, and program

    JP7792174B1

  • Annotation quality assurance support device, method, and program

    JP7792175B1

  • Annotation support device, method, and program

    JP7793244B1