Data annotation method and device and electronic equipment
By dividing into multiple labeling rounds in the data annotation method and calculating the placement priority based on the convergence difficulty of the data, the problem of insufficient quantity of high-quality data when the labeling cost is limited in the prior art is solved, and more efficient labeling data acquisition is achieved.
Patent Information
- Application Number
- CN202411805177.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-02
AI Technical Summary
In the case where the labeling cost is limited, the quantity of high-quality labeling data is affected, and the convergence difficulty of the data to be labeled is not considered.
By dividing the annotation task into multiple annotation rounds, the placement priority is calculated based on the current annotation result of the unqualified data, and based on this priority, part of the unqualified data is put into the next annotation round for annotation.
Under the limited annotation cost, the possibility of unmet data meeting the standards is increased and the number of high-quality annotation data is increased.
Smart Images

Figure CN119917853A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data annotation method, device, electronic device, and computer-readable storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, the demand for high-quality human-annotated data is becoming increasingly strong. In traditional annotation schemes, data is annotated based on a fixed number of people. For some data with inconsistent annotation results or excessive human subjective influence, its credibility is often low and cannot be used for tasks such as model training. In the subsequent annotation scheme, a method for annotating data based on dynamic number of people is provided. In this method, the user sets the target confidence of the annotation task and the upper limit of the number of annotations, and distributes the data to be annotated to the annotators. If the confidence of the annotation result reaches the target confidence or the number of annotations reaches the upper limit of the number of annotations, the annotation is terminated, otherwise it will be distributed to the next annotator for re-annotation. This method undoubtedly increases the credibility of recyclable data (that is, data whose confidence of the annotation result reaches the target confidence). However, in the data distribution process of this method, all data that do not meet the standard will be distributed to the next annotator with the same probability, without considering the convergence difficulty of each data, resulting in the number of high-quality annotated data being affected when the annotation cost is limited.
[0003] Therefore, there is an urgent need for a data labeling method that can obtain more high-quality labeled data under limited labeling costs, so as to solve the technical problem in the prior art that the amount of high-quality labeled data is affected when the labeling cost is limited. Summary of the invention
[0004] The present application provides a data labeling method, device, electronic device, and computer-readable storage medium to solve the technical problem that the existing data labeling method does not consider the convergence difficulty of the data to be labeled, resulting in the number of high-quality labeled data being affected when the labeling cost is limited.
[0005] In a first aspect, an embodiment of the present application provides a data labeling method, the method comprising: obtaining a plurality of non-compliant data, the non-compliant data being data to be labeled whose confidence of a current labeling result does not reach a preset confidence threshold after the labeling operation in a current labeling round; calculating a first number of labeling persons corresponding to each non-compliant data according to the current labeling result of each non-compliant data, the first number of labeling persons being used to characterize the minimum number of labeling persons required to make the confidence of the current labeling result of the non-compliant data reach the preset confidence threshold; sorting the plurality of non-compliant data according to the first number of labeling persons corresponding to each non-compliant data, wherein the smaller the first number of labeling persons, the higher the priority of the non-compliant data for delivery; obtaining a first number of the non-compliant data with a high delivery priority according to the delivery priority sorting result of the plurality of non-compliant data; and putting the first number of the non-compliant data into the next labeling round for labeling operation.
[0006] In a second aspect, an embodiment of the present application provides a data labeling device, the device comprising: a first data acquisition unit, a labeling person-time calculation unit, a delivery priority sorting unit, a second data acquisition unit, and a data delivery unit; the first data acquisition unit is used to acquire a plurality of non-standard data, the non-standard data being the data to be labeled whose confidence of the current labeling result does not reach a preset confidence threshold after the labeling operation of the current labeling round; the labeling person-time calculation unit is used to calculate the first labeling person-time corresponding to each non-standard data according to the current labeling result of each non-standard data, the first labeling person-time is used to characterize the number of labeling people used to label the data. The confidence level of the current annotation result of the non-compliant data reaches the minimum number of annotation persons required for the preset confidence threshold; the delivery priority sorting unit is used to sort the delivery priorities of the multiple non-compliant data according to the first number of annotation persons corresponding to each of the non-compliant data, wherein the delivery priority of the non-compliant data with a smaller first number of annotation persons is higher; the second data acquisition unit is used to acquire a first number of the non-compliant data with a high delivery priority according to the delivery priority sorting result of the multiple non-compliant data; the data delivery unit is used to deliver the first number of the non-compliant data to the next marking round for marking operation.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to implement the above method.
[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having one or more computer instructions stored thereon, which, when executed by a processor, executes the above method.
[0009] Compared with the prior art, the data labeling method provided by the present application includes: obtaining multiple non-compliant data, wherein the non-compliant data are the data to be labeled whose confidence of the current labeling result does not reach a preset confidence threshold after the labeling operation of the current labeling round; calculating the first labeling person-time corresponding to each non-compliant data according to the current labeling result of each non-compliant data, wherein the first labeling person-time is used to characterize the minimum labeling person-time required to make the confidence of the current labeling result of the non-compliant data reach the preset confidence threshold; sorting the multiple non-compliant data according to the first labeling person-time corresponding to each non-compliant data, wherein the smaller the first labeling person-time, the higher the priority of the non-compliant data for delivery; obtaining a first number of non-compliant data with high delivery priority according to the delivery priority sorting results of the multiple non-compliant data; and putting the first number of non-compliant data into the next labeling round for labeling operation. First, this method divides the labeling task into multiple labeling rounds. After the current labeling round is completed, the priority of the non-compliant data will be determined based on the current labeling results of the non-compliant data, and part of the non-compliant data will be put into the next labeling round for labeling based on the priority of delivery, rather than putting all the non-compliant data into the next labeling round with the same probability. Under limited labeling costs, the delivery of non-compliant data in this way undoubtedly increases the possibility of compliance of non-compliant data. Second, this method calculates the minimum number of labeling people required for the confidence of the current labeling results of non-compliant data to reach a preset confidence threshold, and uses the minimum number of labeling people as the basis for determining the priority of delivery of non-compliant data, so that the evaluation of delivery priority is more in line with the principle of data compliance (increasing the number of labeling people), and further increases the possibility of compliance of non-compliant data put into the next compliance round. In summary, the data labeling method provided by the present application is a method that can obtain more high-quality labeled data at a limited labeling cost, which solves the technical problem that the number of high-quality labeled data is affected when the labeling cost is limited in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is an application system diagram of the data annotation method provided in the embodiment of the present application;
[0011] Figure 2 is a flow chart of the data labeling method provided in the first embodiment of the present application;
[0012] Figure 3 is a structural schematic diagram of a data labeling device provided in the second embodiment of the present application;
[0013] Figure 4It is a schematic diagram of the structure of an electronic device provided in the third embodiment of the present application. DETAILED DESCRIPTION
[0014] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0015] With the rapid development of artificial intelligence technology, the demand for high-quality human-labeled data is becoming increasingly urgent. For example, a large amount of high-quality labeled data is needed for model training.
[0016] In traditional labeling schemes, data is labeled based on a fixed number of people, and then the final labeling results are obtained based on majority voting. For some data with inconsistent labeling results, the credibility of the labeling results is low, and they are often discarded in later tasks such as model training. For some labeling tasks where human subjective influence is too large, such as comparing the beauty of faces, comparing the similarity of faces, and evaluating the quality of text creation, it is difficult to obtain unified labeling results in the labeling of a fixed number of people, making it impossible to recover sufficient effective labeling data.
[0017] In the subsequent labeling scheme, a method for labeling data based on dynamic number of people is provided. In this labeling method, the user sets the target confidence of the labeling task and the upper limit of the number of labeling people, distributes the data to be labeled to the labelers, and uses statistical methods to calculate the confidence of the labeling results based on the labeling results given by the labelers. Then, based on the confidence of the labeling results and the preset target confidence, it is judged whether the data to be labeled meets the standard. Specifically, if the confidence reaches the target confidence or the number of labeling people reaches the upper limit of the number of labeling people, the labeling is terminated, otherwise it will be distributed to the next labeler for re-labeling. This method undoubtedly increases the credibility of recyclable data (that is, data whose confidence of the labeling results reaches the target confidence).
[0018] However, in the data distribution process of this method, all data that do not meet the standards will be distributed to the next labeler with the same probability, without considering the difficulty of convergence of each data. Specifically, the data that is difficult to converge after a certain number of annotations is still distributed to subsequent labelers for re-annotation. Under the condition of limited annotation cost, continuing to label these data that are difficult to converge will undoubtedly consume a certain amount of annotation cost, so that some data that are easier to converge will eventually fail to converge due to lack of annotation cost, reducing the overall number of high-quality annotated data. For example, assuming that data A can converge after 15 annotations, and data B can converge after 50 annotations, the current annotation cost can only support 20 annotations. In this method, data A and data B will be distributed to subsequent labelers with the same probability, that is, data A and data B will be labeled by 10 people each, and the annotation cost will be consumed. However, after data A and data B are labeled by 10 people each, neither data A nor data B converges, and neither is high-quality annotated data.
[0019] Therefore, the existing data labeling methods have a technical problem that the amount of high-quality labeled data is affected when the labeling cost is limited because they do not consider the convergence difficulty of the data to be labeled.
[0020] In view of this, the present application provides a data labeling method. First, the method divides the labeling task into multiple labeling rounds. After the current labeling round is completed, the priority of the non-compliant data will be determined based on the current labeling results of the non-compliant data, and part of the non-compliant data will be put into the next labeling round for labeling based on the priority of delivery, rather than putting all the non-compliant data into the next labeling round with the same probability. Under limited labeling costs, the delivery of non-compliant data in this way undoubtedly increases the possibility of compliance of non-compliant data. Second, the method calculates the minimum number of labeling people required for the confidence of the current labeling results of non-compliant data to reach a preset confidence threshold, and uses the minimum number of labeling people as the basis for determining the priority of delivery of non-compliant data, so that the evaluation of delivery priority is more in line with the principle of data compliance (increasing the number of labeling people), and further increases the possibility of compliance of non-compliant data put into the next compliance round. In summary, the data labeling method provided by the present application is a method that can obtain more high-quality labeled data at a limited labeling cost.
[0021] The data annotation method, device, electronic device, and computer-readable storage medium described in this application are further described in detail below in conjunction with specific embodiments and drawings.
[0022] Figure 1 is an application system diagram of the data annotation method provided in the embodiment of the present application. Figure 1As shown, the system includes a user terminal 101 and a server terminal 102. The user terminal 101 can be any device such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant (PDA), etc. The server terminal 102 can be a service module inside the user terminal 101, or a service device electrically connected to the user terminal 101, or a cloud server connected to multiple user terminals 101 in communication. The data labeling method provided by the present application is deployed on the server terminal 102, and in response to receiving the data to be labeled and the data labeling instruction sent by the user terminal 101, the data to be labeled is labeled based on the method.
[0023] The first embodiment of the present application provides a data annotation method, which is deployed in Figure 1 The server 102 shown is used to provide data labeling services to users.
[0024] The labeling task provided by the user often includes multiple data to be labeled. In the method provided in this embodiment, the labeling task will be divided into multiple labeling rounds for multiple rounds of labeling. Specifically, in addition to labeling all the data to be labeled in the initial labeling round, in the subsequent labeling rounds, some non-compliant data will be labeled according to the calculated delivery priority sequence to ensure that the non-compliant data that is easier to meet the standards can meet the standards as much as possible, thereby increasing the number of qualified data (i.e., high-quality data).
[0025] Figure 2 is a flow chart of the data labeling method provided in this embodiment. Figure 2 The data annotation method provided in this embodiment is described in detail. The embodiments described below are used to explain the technical solution of this application and are not intended to be used as limitations for actual use.
[0026] like Figure 2 As shown, the data labeling method provided in this embodiment includes the following steps S210 to S250:
[0027] Step S210, obtaining a plurality of non-standard data, wherein the non-standard data is the data to be labeled whose confidence of the current labeling result does not reach a preset confidence threshold after the labeling operation of the current labeling round.
[0028] The current labeling round may be understood as the labeling round currently being performed when executing the labeling task, and may be the initial labeling round currently being performed, or may be another labeling round other than the initial labeling round currently being performed.
[0029] The non-standard data can be understood as the data to be labeled that still does not meet the high quality standard after being put into the current labeling round for labeling operation. The high quality standard is measured by confidence. Specifically, a confidence threshold is preset. If the confidence of the current labeling result of the data to be labeled does not reach the preset confidence threshold, it means that the data to be labeled does not meet the high quality standard and is non-standard data; on the contrary, when the data to be labeled has undergone the labeling operation of the current labeling round, the confidence of the current labeling result of the data to be labeled reaches the preset confidence threshold, which means that the data to be labeled has reached the high quality standard and is qualified data.
[0030] The current annotation result can be understood as the overall annotation record obtained after the non-compliant data has been annotated in the current annotation round, combined with the annotation value of the non-compliant data in the current annotation round and the annotation value of the non-compliant data in the historical annotation round (the annotation round before the current annotation round). Exemplarily, the non-compliant data A in the current annotation round is marked as favorable, and the non-compliant data A has been annotated for 5 rounds before the current annotation round (that is, there are 5 historical annotation rounds), and the annotation values of the non-compliant data A in the historical annotation rounds are favorable, unfavorable, unfavorable, favorable, and favorable, respectively. Then, the current annotation result is 4 favorable times out of 6 annotations, or 2 unfavorable times out of 6 annotations.
[0031] After determining the current annotation result of the non-compliant data, the confidence corresponding to the current annotation result can be calculated based on the current annotation result. For example, the confidence of approval is calculated based on 4 approvals among 6 annotations, or the confidence of disapproval is calculated based on 2 disapprovals among 6 annotations. It should be noted that, under normal circumstances, the confidence of the annotation result with a greater tendency is calculated based on the current annotation result. For example, among 6 annotations, 4 people are in favor and 2 people are against, and the result with a greater tendency is approval, and the confidence of approval is usually calculated.
[0032] After obtaining the confidence of the current annotation result, by comparing the confidence of the current annotation result with the preset confidence threshold, it is determined whether the confidence of the current annotation result does not reach the preset confidence threshold. Specifically, in response to the confidence of the current annotation result being less than the preset confidence threshold, it is determined that the confidence of the current annotation result does not reach the preset confidence threshold; in response to the confidence of the current annotation result being greater than or equal to the preset confidence threshold, it is determined that the confidence of the current annotation result reaches the preset confidence threshold.
[0033] The methods for calculating confidence are relatively mature, such as statistical methods, methods based on neural network models, etc., and no specific restrictions are made here.
[0034] In this embodiment, a statistical method is provided for calculating confidence for binomial distributed data (i.e., the labeling result of the data is 0 (opposition) or 1 (approval)). Specifically, a confidence level is first set, and a confidence interval under the confidence level is calculated based on the current labeling result. If the confidence interval falls within [0, 0.5) or (0.5, 1] (i.e., the confidence interval is within the interval [0, 0.5) indicating opposition, or within the interval (0.5, 1] indicating approval), it means that the confidence level is less than or equal to the confidence level corresponding to the current labeling result (it is known that under the same result distribution, the size of the confidence level and the confidence interval are positively correlated, i.e., the higher the confidence level, the larger the confidence interval). Then a larger confidence level can be set to calculate whether the confidence interval under the larger confidence level falls within If it is within [0,0.5) or (0.5,1], then you can continue to set a larger confidence level for recalculation. If not, it means that the confidence level set for the first time is the confidence level corresponding to the current annotation result. On the contrary, if the confidence interval does not fall within [0,0.5) or (0.5,1] (that is, the confidence interval spans the interval [0,0.5) indicating opposition and the interval (0.5,1] indicating approval), it means that the confidence level is greater than the confidence level corresponding to the current annotation result. Then you can set a smaller confidence level for recalculation to determine whether the smaller confidence level is the confidence level corresponding to the current annotation result.
[0035] Specifically, the confidence interval under a certain confidence level can be calculated based on the following expression:
[0036]
[0037] Among them, CI represents the confidence interval, p represents the approval / disapproval ratio, n represents the total number of annotations, and Z α / 2 It represents the critical value corresponding to the confidence level under the standard normal distribution, which can be obtained according to the standard normal distribution table.
[0038] For example, the current annotation result is that 100 people participated in the data annotation, of which 60 people annotated in favor and 40 people annotated in opposition. Calculating the confidence corresponding to the current annotation result (since the annotation result is more inclined to be in favor, in this example, only the confidence corresponding to the approval is calculated) may include the following steps:
[0039] First, set the confidence level to 0.94 and calculate the confidence interval at 0.94 confidence level based on the above expression:
[0040]
[0041] The calculated confidence interval at a confidence level of 0.94 is approximately (0.5073, 0.6927), which falls within the interval (0.5, 1]. Next, a larger confidence level is set.
[0042] Second, set the confidence level to 0.95 and calculate the confidence interval at 0.95 confidence level based on the above expression:
[0043]
[0044] The calculated confidence interval at a confidence level of 0.95 is approximately (0.504, 0.696), which falls within the interval (0.5, 1]. Next, a larger confidence level is set.
[0045] Third, set the confidence level to 0.96 and calculate the confidence interval at 0.96 confidence level based on the above expression:
[0046]
[0047] The calculated confidence interval at a confidence level of 0.96 is approximately (0.4996, 0.7004), which is outside the interval (0.5, 1], indicating that the confidence level corresponding to the current annotation result is 0.95.
[0048] Step S220, based on the current labeling result of each non-compliant data, calculate the first labeling person-times corresponding to each non-compliant data, where the first labeling person-times are used to represent the minimum labeling person-times required to make the confidence of the current labeling result of the non-compliant data reach a preset confidence threshold.
[0049] In this embodiment, after obtaining the non-compliant data, the non-compliant data will not be directly put into the next marking round. Instead, the priority of each non-compliant data will be calculated, so that a part of the non-compliant data with a high priority will be put into the next compliant round for priority marking.
[0050] In this embodiment, the method for determining the priority of delivery of non-compliant data is to determine the minimum number of annotations required when the non-compliant data meets the standards (i.e., the confidence of the current annotation results of the non-compliant data reaches a preset confidence threshold) based on the current annotation results of the non-compliant data, and use the minimum number of annotations as the basis for determining the delivery priority. In this embodiment, the minimum number of annotations is defined as the first number of annotations. The current annotation results of each non-compliant data are different, and the minimum number of annotations required is different, and accordingly, the delivery priority is also different.
[0051] In an optional implementation, according to the current labeling result of each non-standard data, the first labeling person-times corresponding to each non-standard data is calculated, which may specifically include the following steps S221 to S223:
[0052] Step S221, determining the tendency value corresponding to the current labeling result of the non-standard data according to the current labeling result of the non-standard data.
[0053] Step S222, using the tendency value as the annotation value, adopting a binary search algorithm, searching within a preset annotation person-time interval, and obtaining the minimum number of annotation persons required to make the confidence of the current annotation result of the non-compliant data reach a preset confidence threshold.
[0054] Step S223: taking the minimum number of annotations as the first number of annotations corresponding to the data that does not meet the standard.
[0055] The tendency value can be understood as the annotation value that the current annotation result is closer to, for example, if it is closer to approval, the tendency value is 1, which indicates approval or approval, and if it is closer to disapproval, the tendency value is 0, which indicates disapproval or disapproval. For example, assuming that 10 people participated in data annotation, 8 of them agreed (selected the value 1) and 2 disagreed (selected the value 2), then the approval ratio is 80% and the disapproval ratio is 20%. The current annotation result is closer to approval, and the tendency value corresponding to the current annotation result is 1, which indicates approval.
[0056] When calculating the minimum number of annotations, it is assumed that subsequent annotation personnel will give the same annotation value as the propensity value, that is, the propensity value is used as the annotation value. Then set the annotation interval, which is a preset numerical interval. Usually, the minimum value of this interval can be 1 or the minimum number of people the user expects to participate in the annotation (for example, the user expects at least 10 people to participate in the annotation of each data to be annotated). The maximum value of this interval can often be set larger to ensure that each data that does not meet the standard can find an annotation number that meets the preset confidence threshold within this interval. Exemplarily, the annotation interval can be set to [1,1000]. Then, through the binary search algorithm, search in the annotation area for the number of annotations that are required to make the confidence of the current annotation result reach the preset confidence threshold, that is, the minimum number of annotations.
[0057] For example, 100 people participated in data annotation, of which 60 people annotated in favor and 40 people annotated against. The calculated confidence of the current annotation result is 0.95, and the preset confidence threshold is 0.99. Assume that all subsequent annotators give an affirmative annotation value, and set the annotation interval to [lower, upper], where lower = 1, upper = 1000, update the annotation interval to the initial search interval, and use the binary search algorithm to search in the initial search interval [1, 1000] how many more annotations are needed to make the current annotation result have a confidence of 0.99. First, calculate the confidence of the current marking result when 500 (mid = (1 + 1000) / 2 ≈ 500) additional annotations are in favor (that is, 600 (100 + 500 = 600) people participated in data annotation, 560 (60 + 500 = 560) people annotated in favor, and 40 people annotated against). If the confidence is greater than the preset confidence threshold, the next search interval is [lower, mid-1], that is, [1,499]; if the confidence is less than or equal to the preset confidence threshold, the next search interval is [mid+1, upper], that is, [501,1000]. Repeat the above steps until the minimum number of annotations that meet the preset confidence threshold is found, that is, the minimum number of annotations.
[0058] Step S230, sorting the delivery priorities of the plurality of non-compliant data according to the first number of annotations corresponding to each non-compliant data, wherein the non-compliant data with a smaller first number of annotations has a higher delivery priority.
[0059] The smaller the number of first annotations of non-compliant data, the fewer the number of annotations required for the non-compliant data to meet the standard, and the easier it is for the non-compliant data to meet the standard. In this embodiment, data that is easier to meet the standard is determined as data with a higher delivery priority and is delivered first in the next annotation round.
[0060] In an optional implementation, the multiple non-compliant data are prioritized for delivery according to the first number of annotations corresponding to each non-compliant data, which may specifically include the following steps S231 to S232:
[0061] Step S231, according to the first number of annotations corresponding to each non-compliant data, multiple non-compliant data are arranged in ascending order.
[0062] Step S232, arranging the results in ascending order as the priority sorting result of the multiple data that do not meet the standards.
[0063] The ascending order means that the more times the non-compliant data has been first labeled, the lower its order will be. For example, there are 5 non-compliant data q1, q2, q3, q4, q5, and their corresponding first labeled times are 10, 8, 6, 4, and 2, respectively. According to the first labeled times corresponding to these 5 non-compliant data, the 5 non-compliant data are arranged in ascending order, and the ascending order result is {q5, q4, q3, q2, q1}, which is the result of the priority sorting of the 5 non-compliant data. Among them, the non-compliant data q5 has the highest priority, and the non-compliant data q1 has the lowest priority. The non-compliant data q5 is easier to meet the standard, and should be given priority in the next round of meeting the standard.
[0064] Of course, you can also directly use the confidence of the current annotation results of the non-compliant data as the basis for determining the delivery priority. However, among multiple non-compliant data, there is no equal positive relationship between confidence and the number of annotations. Therefore, directly using confidence as the basis for determining delivery priority is not as accurate as using the minimum number of annotations as the basis for determination.
[0065] Step S240 , obtaining a first number of non-compliant data with high delivery priority according to the delivery priority sorting results of the plurality of non-compliant data.
[0066] The first number refers to a number of non-compliant data that are ranked at the top in the delivery priority sorting result, that is, a part of the non-compliant data that is easier to meet the standards among all the non-compliant data.
[0067] In this embodiment, the first number is not a preset value, but a value determined based on the number of non-compliant data. Based on this, in an optional implementation, before the step of obtaining a first number of non-compliant data with a high delivery priority according to the delivery priority sorting results of multiple non-compliant data, the method provided in this embodiment may also include the following steps: determining a first number of non-compliant data to be pre-invested in the next labeling round according to the number of multiple non-compliant data and the first preset number; wherein the first preset number is used to represent the number of non-compliant data to be labeled that is preset for the labeling task and is allowed to be non-compliant.
[0068] Understandably, for labeling tasks, considering the labeling cost, it is basically impossible to achieve that all the data to be labeled meet the standards. Therefore, the proportion of non-standard data is expected in advance before executing the labeling task. For example, if there are 1,000 data to be labeled in the labeling task, and 10% of the data is allowed to be non-standard, then the number of non-standard data to be labeled is 100.
[0069] In the method provided in this embodiment, the number of non-compliant data to be put into the next labeling round among all non-compliant data is determined based on the preset number of non-compliant data to be labeled that is allowed to be non-compliant, that is, the first data is determined. In this embodiment, the number of non-compliant data to be labeled that is allowed to be non-compliant for the labeling task is defined as the first preset number.
[0070] In a specific implementation, the first amount of non-compliant data to be pre-invested in the next marking round may be determined based on the relationship between the amount of non-compliant data and the first preset amount, which may specifically include the following situations:
[0071] Case 1: in response to the number of the plurality of non-standard data being greater than twice the first preset number, half of the number of the plurality of non-standard data is used as the first number.
[0072] Case 2: In response to the number of the plurality of non-compliant data being less than or equal to twice the first preset number and greater than or equal to the first preset number, the first preset number is used as the first number.
[0073] Case 3: In response to the number of the plurality of non-standard data being less than a first preset number, the number of the plurality of non-standard numbers is taken as the first number.
[0074] Specifically, the first quantity can be expressed by the following expression:
[0075] First quantity =
[0076] min(max(number of non-compliant data / 2, first preset number), number of non-compliant data)
[0077] Exemplarily, there are 1000 data to be labeled in the labeling task, and 10% of the data are allowed to be substandard, that is, the number of substandard data to be labeled is allowed to be 100, and the first preset number is 100. When the number of substandard data is greater than 200 (assuming that the number of substandard data is 1000), the first number is 500 (min(max(1000 / 2,100),1000)=500); when the number of substandard data is less than or equal to 200 and greater than 100 (assuming that the number of substandard data is 150), the first number is 100 (min(max(150 / 2,100),150)=100); when the number of substandard data is less than 100 (assuming that the number of substandard data is 50), the first number is 50 (min(max(50 / 2,100),50)=50).
[0078] To sum up, when the number of non-compliant data is relatively large, half of the non-compliant data will be put into the next marking round; when the number of non-compliant data is moderate, the first preset number of non-compliant data will be put into the next marking round; when the number of non-compliant data is relatively small, all non-compliant data will be put into the next marking round.
[0079] Step S250: putting the first amount of non-compliant data into the next labeling round for labeling operation.
[0080] After determining the number of non-compliant data to be put into the next round of marking, the first number of non-compliant data with high delivery priority can be put into the next round of marking for re-marking according to the delivery priority sorting results of multiple non-compliant data. For example, if there are 1,000 non-compliant data, according to the delivery priority sorting results, the top 500 non-compliant data are put into the next round of marking for re-marking; if there are 150 non-compliant data, according to the delivery priority sorting results, the top 100 non-compliant data are put into the next round of marking for re-marking; if there are 50 non-compliant data, all 50 non-compliant data are put into the next round of marking for re-marking.
[0081] When the amount of non-compliant data is large, compared with putting all non-compliant data into the next labeling round with the same probability, the method provided in this embodiment gives priority to putting part of the non-compliant data with a high priority into the next labeling round, which can increase the possibility of the non-compliant data meeting the standard. Therefore, under the same labeling cost, the method provided in this embodiment can increase the amount of compliant data (i.e., high-quality data).
[0082] In the method provided in this embodiment, the labeling round of the labeling task is related to the labeling budget preset for the labeling task. Specifically, when the preset labeling budget is used up or is not enough to support the labeling cost of the next labeling round, the labeling will be ended.
[0083] Based on this, in an optional implementation, before the step of obtaining multiple non-compliant data, the method provided in this embodiment may also include: determining the remaining labeling budget value after the current labeling round based on a preset total labeling budget value, and the labeling cost values consumed by the current labeling round and historical labeling rounds; wherein the historical labeling round is the labeling round before the current labeling round.
[0084] When issuing a labeling task, the user often provides a preset budget value for the labeling task to control the labeling cost of the labeling task. In this embodiment, the budget value preset by the user is defined as the total labeling budget value. When the total labeling budget value is used up or is not enough to support the labeling cost of the next labeling round, the labeling will end. Based on this, the current remaining labeling budget value will be determined according to the preset total labeling budget value and the currently consumed labeling cost value. Exemplarily, the total labeling budget value is B, and the currently consumed labeling cost value is A, then the current remaining labeling budget value is BA.
[0085] In this embodiment, the currently consumed annotation cost value includes the annotation cost value consumed by the current annotation round and the annotation cost value consumed by each annotation round before the current annotation round. In this embodiment, the annotation round before the current annotation round is defined as a historical annotation round. Exemplarily, the total annotation budget value is B, the annotation cost value consumed by the current annotation round is A1, the current annotation round is the fifth round, and the annotation cost values consumed by the first annotation round, the second annotation round, the third annotation round, and the fourth annotation round are A5, A4, A3, and A2 respectively. Then, A1+A2+A3+A4+A5 is the annotation cost value consumed by the current annotation round and the historical annotation rounds, and the remaining annotation budget value after the current annotation round is B-(A1+A2+A3+A4+A5).
[0086] When the remaining annotation budget value is 0, that is, the total annotation budget value has been consumed, the annotation operation is terminated, the next round of annotation is no longer performed, and subsequent operations are stopped. When the remaining annotation budget value is greater than 0, that is, the total annotation budget value has not been consumed, the subsequent operations are continued. Specifically, in response to the remaining annotation budget value after the current annotation round being greater than zero, a plurality of non-compliant data are obtained. In response to the remaining annotation budget value after the current annotation round being equal to zero, the annotation operation is terminated.
[0087] Of course, if all the data to be labeled meet the standards after the current labeling round ends, that is, if there is no data that does not meet the standards, the labeling operation will also end.
[0088] Through the above steps, it is determined whether the remaining annotation budget value after the current annotation round is zero. In a specific implementation, it is also necessary to determine whether the remaining annotation budget after the current annotation round can still support the cost of the next annotation round.
[0089] Based on this, in an optional implementation, before the step of putting the first number of non-compliant data into the next marking round for marking operation, the method provided in this embodiment also includes: determining the marking budget value corresponding to the next marking round, and determining whether the marking operation of the next marking round can still be performed based on the remaining marking budget value after the current marking round and the marking budget value corresponding to the next marking round.
[0090] In a specific implementation, determining the annotation budget value corresponding to the next annotation round may specifically include the following steps S11 to S13:
[0091] Step S11, based on the first number of annotation people corresponding to each non-compliant data in the first number of non-compliant data, determine the second number of annotation people corresponding to the non-compliant data in the next annotation round, and the second number of annotation people is used to represent the average number of annotation people for the first number of non-compliant data in the next annotation round.
[0092] Step S12, determining the actual single annotation cost value corresponding to the data to be annotated according to the annotation cost value consumed in the current annotation round and the historical annotation rounds, and the actual number of annotation persons for the data to be annotated in the current annotation round and the historical annotation rounds.
[0093] Step S13, determining a labeling budget value corresponding to the next labeling round according to the actual single labeling cost value, the second labeling person-times, and the first quantity.
[0094] In this embodiment, the number of non-compliant data to be pre-released for the next labeling round (i.e., the first number), the average number of labeling personnel for the first number of non-compliant data to be pre-released for the next labeling round, and the single labeling cost are determined, and the labeling budget value for releasing the first number of non-compliant data into the next labeling round can be calculated.
[0095] In this embodiment, the average number of annotations of the first number of non-standard data pre-placed in the next annotation round is defined as the second number of annotations, and the second number of annotations is determined based on the first number of annotations. The first number of annotations represents the minimum number of annotations for each non-standard data, and the average number of annotations for each non-standard data is averaged to obtain the average number of annotations for the non-standard data, that is, the second number of annotations.
[0096] For example, the number of non-compliant data to be pre-posted into the next round of annotation is M", and the number of first annotations corresponding to the non-compliant data obtained by calculation is n. i The first annotation times corresponding to the M" non-compliant data are expressed from small to large as:
[0097] N * ={n1,n2,……,n M′′}, where n1≤n2≤…≤n M′′
[0098] The second annotation number of people N" can be expressed as:
[0099]
[0100] In this embodiment, the single annotation cost value is calculated based on the annotation cost value consumed in the current annotation round and the historical annotation round, and the actual number of annotation personnel for the data to be annotated in the current annotation round and the historical annotation round. Since the single annotation cost value is calculated based on the actual consumed annotation cost value and the actual number of annotation personnel, it is a real single annotation cost value. In this embodiment, the real single annotation cost value is defined as the actual single annotation cost value.
[0101] For example, the annotation cost value consumed in the current annotation round is A1, the current annotation round is the fifth round, the annotation cost values consumed in the first annotation round, the second annotation round, the third annotation round, and the fourth annotation round are A5, A4, A3, and A2 respectively, then the annotation cost value consumed in the current annotation round and the historical annotation rounds is A1+A2+A3+A4+A5, assuming that the actual number of annotation workers for the data to be annotated in the first annotation round, the second annotation round, the third annotation round, the fourth annotation round, and the current annotation round is C1, C2, C3, C4, and C5 respectively, then the actual number of annotation workers for the data to be annotated in the current annotation round and the historical annotation rounds is C1+C2+C3+C4+C5. The actual single annotation cost value P is (A1+A2+A3+A4+A5) / (C1+C2+C3+C4+C5).
[0102] After determining the first quantity of non-compliant data to be pre-put into the next labeling round, the second number of labeling personnel corresponding to the non-compliant data, and the actual single labeling cost value, the first quantity, the second number of labeling personnel, and the actual single labeling cost value can be multiplied to obtain the labeling budget value corresponding to the next labeling round.
[0103] For example, the first amount of non-standard data to be pre-placed in the next annotation round is M", the second number of annotation people corresponding to the non-standard data is N", and the actual single annotation cost value is P. Then, the annotation budget value X corresponding to the next annotation round can be expressed as:
[0104] X=M″×N″×P
[0105] In another optional implementation, the labeling budget value corresponding to the next labeling round may be determined according to the actual single labeling cost value and the first labeling person-time corresponding to each non-compliant data in the first number of non-compliant data.
[0106] For example, the first annotation times corresponding to M" non-compliant data are expressed from small to large as follows:
[0107] N * ={n1,n2,……,n M″}, where n1≤n2≤…≤n M″
[0108] The actual single annotation cost value is P, then the annotation budget value X corresponding to the next annotation round can be expressed as:
[0109]
[0110] In an optional implementation, after determining the remaining marking budget value after the current marking round and the marking budget value corresponding to the next marking round, it can be determined by comparison whether to start the marking operation of the next marking round. Based on this, the method provided in this embodiment may also include the following steps: according to the marking budget value corresponding to the next marking round and the remaining marking budget value after the current marking round, determine whether to put the first number of non-standard data into the next marking round for marking operation. Specifically, in response to the marking budget value corresponding to the next marking round being less than or equal to the remaining marking budget value after the current marking round, the first number of non-standard data is put into the next marking round for marking operation. In response to the marking budget value corresponding to the next marking round being greater than the remaining marking budget value after the current marking round, the marking operation is ended.
[0111] Exemplarily, the total annotation budget is B, the annotation cost consumed by the current annotation round and the historical annotation round is A, the remaining annotation budget after the current annotation round is BA, and the annotation budget corresponding to the next annotation round obtained by calculation is X. When X≤(BA), it means that the remaining annotation budget after the current annotation round can support the annotation operation of the next annotation round, and the first number of substandard data can be put into the next annotation round for annotation operation. When X>(BA), it means that the remaining annotation budget after the current annotation round cannot support the annotation operation of the next annotation round, and the annotation operation can be terminated, and the current annotation result after the current annotation round is used as the annotation result corresponding to the data to be annotated.
[0112] In the method provided in this embodiment, after determining to put the first amount of non-standard data into the next labeling round, it is also necessary to determine the maximum number of labelers for the non-standard data in the next labeling round, so as to perform constrained labeling (i.e., the maximum number of labelers) in the next labeling round. In this embodiment, the maximum number of labelers for the determined non-standard data in the next labeling round is defined as the third number of labelers.
[0113] Based on this, in an optional implementation, before the step of putting the first number of non-compliant data into the next marking round for marking operation, the method provided in this embodiment may also include: determining the third number of marking people corresponding to the non-compliant data in the next marking round based on the first number of marking people corresponding to each non-compliant data in the first number of non-compliant data, and the third number of marking people is used to represent the maximum number of marking people for the first number of non-compliant data in the next marking round.
[0114] In this embodiment, the third number of annotations is still determined based on the first number of annotations. Specifically, the largest first number of annotations among the first number of annotations corresponding to the first number of non-compliant data is used as the maximum number of annotations for the first number of non-compliant data in the next annotation round, that is, the third number of annotations.
[0115] For example, the number of non-compliant data to be pre-posted into the next round of annotation is M", and the number of first annotations corresponding to the non-compliant data obtained by calculation is n. i The first annotation times corresponding to the M" non-compliant data are expressed from small to large as:
[0116] N * ={n1,n2,……,n M″}, where n1≤n2≤…≤n M″
[0117] The third notation number of people N' can be expressed as:
[0118] N′=n M″
[0119] After determining that the remaining annotation budget value after the current annotation round is not zero (and there is non-standard data), and can support the annotation operation of the next annotation round, the first amount of non-standard data can be put into the next annotation round for annotation operation, which can specifically include the following steps S21 to S22:
[0120] Step S21, the non-compliant data is distributed to the labeling personnel, and the labeling results of the non-compliant data are determined according to the labeling values of the non-compliant data by the labeling personnel and the current labeling results of the non-compliant data.
[0121] Step S22, in response to the confidence of the labeling result of the non-compliant data reaching a preset confidence threshold, or the number of labeling people for the non-compliant data reaching the third number of labeling people, the labeling operation of the non-compliant data in the next labeling round is terminated.
[0122] Specifically, a number of non-standard data can be randomly selected from the first number of non-standard data and distributed to the labeling personnel 1, and then a number of non-standard data can be randomly selected from the remaining non-standard data and distributed to the labeling personnel 2, ..., to ensure that in one distribution cycle, only one labeling personnel will label one non-standard data. After the labeling value marked by the labeling personnel is recovered, the labeling result of the non-standard data after the labeling of the labeling personnel can be calculated according to the current labeling result and the labeling value of the non-standard data. If the confidence of the labeling result reaches the preset confidence threshold, it means that the non-standard data has reached the standard, and the labeling operation of the non-standard data can be stopped. If the confidence of the labeling result still does not reach the preset confidence threshold, it is necessary to continue to determine whether the number of labeling people of the non-standard data has reached the third labeling person (i.e., the maximum number of labeling people). If it has not reached, the non-standard data can continue to be distributed to the next labeling personnel. If it has reached, the labeling operation of the non-standard data in the next labeling round can be stopped.
[0123] It should be noted that when the data that does not meet the standard is put into the next marking round for marking operation, the next marking round is the current marking round, that is, when the marking operation of the next marking round begins, the current marking round is transformed into the historical marking round, and the next marking round is transformed into the current marking round. After completing a round of marking operation, if there is still data that does not meet the standard, the method provided in this embodiment can be re-executed to perform a new marking operation of the marking round.
[0124] In the method provided in this embodiment, since the convergence difficulty of each data to be labeled is not clear in the initial labeling round, all the data to be labeled in the labeling task will be put into the initial labeling round for labeling operation. In an optional implementation, if the current labeling round is the initial labeling round, before obtaining multiple non-standard data, the labeling operation of the current labeling round may specifically include the following steps S31 to S34:
[0125] Step S31, according to the number of data to be labeled included in the labeling task, the labeling budget value preset for the initial labeling round, and the single labeling cost value preset for the data to be labeled, determine the fourth labeling person-time corresponding to the data to be labeled in the initial labeling round, and the fourth labeling person-time is used to represent the maximum number of labeling people for the data to be labeled in the initial labeling round.
[0126] Step S32: the data to be annotated are distributed to annotators, and the annotating results of the data to be annotated are determined according to the annotating values of the data to be annotated by the annotators.
[0127] Step S33, in response to the confidence of the annotation result of the data to be annotated reaching a preset confidence threshold, or the number of annotations of the data to be annotated by the number of people reaching the fourth number of annotations, the annotation operation of the data to be annotated is ended.
[0128] Step S34: the data to be annotated whose annotated times reach the fourth annotated times and whose confidence level does not reach the preset confidence threshold is regarded as non-compliant data.
[0129] The annotation budget value preset for the initial annotation round may be preset by the user when issuing the annotation task, or may be calculated based on the minimum number of annotation persons preset by the user for the data to be annotated and the preset single annotation cost value, and there is no specific limitation. Usually, the annotation budget value preset for the initial annotation round is less than the preset total annotation budget value. When the annotation budget value preset for the initial annotation round is equal to the total annotation budget value, it means that the user only wants to perform the annotation operation for one annotation round.
[0130] Before performing the labeling operation of the initial labeling round, the maximum number of labeling people for the data to be labeled in the initial labeling round can be determined according to the number of data to be labeled, the labeling budget value corresponding to the initial labeling round, and the single labeling cost value. In this embodiment, the maximum number of labeling people for the data to be labeled in the initial labeling round is defined as the fourth labeling person. For example, the number of data to be labeled is M, the labeling budget value corresponding to the initial labeling round is A, and the single labeling cost value is c. The fourth labeling person N can be expressed as:
[0131] N = A / (M × c)
[0132] Optionally, after the fourth number of annotations is determined, several data to be annotated can be randomly selected from all the data to be annotated and distributed to annotation personnel 1, and then several data to be annotated can be randomly selected from the remaining data to be annotated and distributed to annotation personnel 2, ..., to ensure that in one distribution cycle, only one annotation personnel annotates one data to be annotated. After the annotation value given by the annotation personnel is recovered, the annotation result of the data to be annotated after the annotation of the annotation personnel can be calculated according to the annotation value. If the confidence of the annotation result reaches the preset confidence threshold, it means that the data to be annotated has reached the standard, and the annotation operation of the data to be annotated can be stopped, and the data to be annotated is the data that has reached the standard. If the confidence of the annotation result does not reach the preset confidence threshold, it is necessary to continue to determine whether the number of annotations of the data to be annotated has reached the fourth number of annotations. If it has not reached, the data to be annotated can continue to be distributed to the next annotation personnel. If it has reached, the annotation operation of the data to be annotated can be stopped.
[0133] The data to be annotated whose number of annotations has reached the fourth number of annotations and whose confidence has not reached the preset confidence threshold is considered as non-compliant data, and a new round of annotation operations can be continued according to the method provided in this embodiment.
[0134] This embodiment provides an optional data labeling method, which may specifically include the following steps S301 to S316:
[0135] Step S301, initializing the labeling budget value Cost=A, and the remaining labeling budget value Budget=BA, where B is the total labeling budget value.
[0136] Step S302, initialize the data pool Q = {q1, q2, ..., q M}, where q i is the data to be labeled, M is the number of data to be labeled, i = {1, 2, ..., M}.
[0137] Step S303, initializing the labeling round=1.
[0138] Step S304, initializing the maximum number of annotations by N=A / (M×c), where c is the single annotation cost value provided by the user.
[0139] Step S305, initializing the actual number of labeling people LabelCount=0.
[0140] Step S306: randomly take out X data to be labeled from the data pool Q and distribute them to the labeling personnel.
[0141] Step S307, collect the labeling values given by the labeling personnel, and update the actual labeling times LabelCount=LabelCount+X.
[0142] Step S308, determine whether the data to be labeled meets the standard. If it does not meet the standard and the number of labeling people for the data to be labeled does not reach the maximum number of labeling people N, the data to be labeled is returned to the data pool Q.
[0143] Step S309, repeating steps S306 to S308 until the data pool Q is empty.
[0144] Step S310: In response to the data to be annotated not fully meeting the requirements and the remaining annotation budget value Budget>0, proceed to the next step; otherwise, the delivery is terminated.
[0145] Step S311 , updating the data pool Q, which includes M′ unqualified to-be-annotated data (ie, unqualified data), where M′≤M.
[0146] Step S312, calculating the actual single labeling cost value P=Cost / LabelCount.
[0147] Step S313, calculating the pre-delivery quantity M"=min(max(M' / 2, first preset quantity), M') for the next marking round.
[0148] Step S314, calculating the annotation budget value X=M"×N"×P corresponding to the next annotation round, where N" is the average number of annotation persons.
[0149] Step S315, in response to the annotation budget value X corresponding to the next annotation round being ≤ the remaining annotation budget value Budget, perform the following operations:
[0150] Step S315-1, update the marking round = marking round + 1;
[0151] Step S315-2, updating the annotation budget value Cost=Cost+X and the remaining annotation budget value Budget=Budget-X, and the maximum number of annotation persons N=N', where N' is the maximum number of annotation persons corresponding to the next annotation round;
[0152] Step S315-3, jump to step S306.
[0153] The above first embodiment provides an optional data labeling method. First, the method divides the labeling task into multiple labeling rounds. After the current labeling round is completed, the release priority of the non-compliant data will be determined based on the current labeling results of the non-compliant data, and part of the non-compliant data will be put into the next labeling round for labeling based on the release priority, rather than putting all the non-compliant data into the next labeling round with the same probability. Under limited labeling costs, putting non-compliant data in this way undoubtedly increases the possibility of non-compliant data meeting the standards. Second, the method calculates the minimum number of labeling people required for the confidence of the current labeling results of the non-compliant data to reach the preset confidence threshold, and uses the minimum number of labeling people as the basis for determining the release priority of the non-compliant data, so that the evaluation of the release priority is more in line with the data compliance principle (increasing the number of labeling people), and further increases the possibility of non-compliant data entering the next compliance round. In summary, the data labeling method provided by the present application is a method that can obtain more high-quality labeled data at a limited labeling cost.
[0154] It should be noted that the examples in the first embodiment are only for explaining the method described in this application and are not intended to be limiting for actual use. The data annotation method provided in this application includes but is not limited to the method described in the first embodiment.
[0155] The second embodiment of the present application provides a data labeling device. Figure 3 It is a structural diagram of the data labeling device provided in this embodiment.
[0156] like Figure 3As shown, the data labeling device provided in this embodiment includes: a first data acquisition unit 301, a labeling person-times calculation unit 302, a delivery priority sorting unit 303, a second data acquisition unit 304, and a data delivery unit 305.
[0157] The first data acquisition unit 301 is used to acquire a plurality of non-standard data, where the non-standard data is data to be labeled whose confidence level of the current labeling result does not reach a preset confidence threshold after the labeling operation of the current labeling round.
[0158] Optionally, before the step of obtaining a plurality of non-standard data, it is also used to:
[0159] Determine the remaining annotation budget value after the current annotation round according to the preset total annotation budget value and the annotation cost values consumed in the current annotation round and the historical annotation rounds; wherein the historical annotation rounds are the annotation rounds before the current annotation round;
[0160] The obtaining of multiple non-standard data includes:
[0161] In response to the remaining annotation budget value after the current annotation round being greater than zero, the plurality of non-standard data are acquired.
[0162] Optionally, also used for:
[0163] In response to the remaining annotation budget value after the current annotation round being equal to zero, and / or the non-standard data not existing, the annotation operation is terminated.
[0164] Optionally, if the current marking round is an initial marking round; before acquiring a plurality of non-standard data, the marking operation of the current marking round includes:
[0165] Determine, according to the number of to-be-annotated data included in the annotation task, a preset annotation budget value for the initial annotation round, and a preset single annotation cost value for the to-be-annotated data, a fourth number of annotation persons corresponding to the to-be-annotated data in the initial annotation round, wherein the fourth number of annotation persons is used to represent a maximum number of annotation persons for the to-be-annotated data in the initial annotation round;
[0166] The data to be labeled is distributed to a labeler, and a labeling result of the data to be labeled is determined according to the labeling value of the data to be labeled by the labeler;
[0167] In response to the confidence of the labeling result of the data to be labeled reaching the preset confidence threshold, or the number of labeling people for the data to be labeled reaching the fourth number of labeling people, ending the labeling operation on the data to be labeled;
[0168] The data to be labeled whose number of labeling persons reaches the fourth number of labeling persons and whose confidence does not reach the preset confidence threshold is taken as the non-compliant data.
[0169] The labeling number calculation unit 302 is used to calculate the first number of labeling people corresponding to each non-compliant data based on the current labeling result of each non-compliant data. The first number of labeling people is used to represent the minimum number of labeling people required to make the confidence of the current labeling result of the non-compliant data reach the preset confidence threshold.
[0170] Optionally, the calculating, according to the current labeling result of each piece of non-compliant data, the first labeling person-times corresponding to each piece of non-compliant data includes:
[0171] Determining, according to the current labeling result of the non-standard data, a tendency value corresponding to the current labeling result of the non-standard data;
[0172] Using the tendency value as the annotation value, a binary search algorithm is used to search within a preset annotation number interval to obtain the minimum number of annotations required to make the confidence of the current annotation result of the non-standard data reach the preset confidence threshold;
[0173] The minimum number of annotated persons is used as the first number of annotated persons corresponding to the non-compliant data.
[0174] The delivery priority sorting unit 303 is used to sort the delivery priorities of the multiple non-compliant data according to the first number of annotations corresponding to each non-compliant data, wherein the non-compliant data with a smaller number of first annotations has a higher delivery priority.
[0175] Optionally, the prioritizing the multiple non-compliant data according to the first number of annotations corresponding to each non-compliant data includes:
[0176] Arrange the plurality of non-compliant data in ascending order according to the first number of annotation persons corresponding to each non-compliant data;
[0177] The ascending order result is used as the delivery priority sorting result of the multiple non-compliant data.
[0178] The second data acquisition unit 304 is configured to acquire a first number of the non-compliant data with high delivery priorities according to delivery priority sorting results of the plurality of non-compliant data.
[0179] Optionally, before the step of acquiring a first number of the non-compliant data with a high delivery priority according to the delivery priority sorting results of the plurality of non-compliant data, the method is further configured to:
[0180] Determining the first number of the non-compliant data to be pre-invested in the next marking round according to the number of the plurality of non-compliant data and a first preset number;
[0181] The first preset number is used to represent the number of unqualified to-be-annotated data that is preset for the labeling task.
[0182] Optionally, determining the first number of the non-compliant data to be pre-invested in the next marking round according to the number of the plurality of non-compliant data and a first preset number includes:
[0183] In response to the number of the plurality of non-compliant data being greater than twice the first preset number, taking half of the number of the plurality of non-compliant data as the first number;
[0184] In response to the number of the plurality of non-standard data being less than or equal to twice the first preset number and greater than or equal to the first preset number, using the first preset number as the first number;
[0185] In response to the number of the plurality of non-standard data being less than the first preset number, the number of the plurality of non-standard data is taken as the first number.
[0186] The data delivery unit 305 is used to deliver the first quantity of non-compliant data to the next marking round for marking operation.
[0187] Optionally, before the step of putting the first number of non-standard data into the next labeling round for labeling operation, it is also used to:
[0188] According to the first number of annotation people corresponding to each non-compliant data in the first number of non-compliant data, determine the third number of annotation people corresponding to the non-compliant data in the next annotation round, and the third number of annotation people is used to represent the maximum number of annotation people for the first number of non-compliant data in the next annotation round.
[0189] Optionally, the step of putting the first quantity of non-standard data into a next labeling round for labeling includes:
[0190] The non-compliant data is distributed to a labeling staff, and a labeling result of the non-compliant data is determined according to the labeling value of the non-compliant data by the labeling staff and the current labeling result of the non-compliant data;
[0191] In response to the confidence of the labeling result of the non-compliant data reaching the preset confidence threshold, or the number of labeling people for the non-compliant data reaching the third number of labeling people, the labeling operation on the non-compliant data in the next labeling round is terminated.
[0192] Optionally, before the step of putting the first number of non-standard data into the next labeling round for labeling operation, it is also used to:
[0193] Determine, according to the first number of annotation persons corresponding to each non-compliant data in the first number of non-compliant data, a second number of annotation persons corresponding to the non-compliant data in a next annotation round, wherein the second number of annotation persons is used to represent an average number of annotation persons for the first number of non-compliant data in the next annotation round;
[0194] Determine the actual single annotation cost value corresponding to the data to be annotated according to the annotation cost value consumed in the current annotation round and the historical annotation rounds, and the actual number of annotation persons for the data to be annotated in the current annotation round and the historical annotation rounds;
[0195] A labeling budget value corresponding to a next labeling round is determined according to the actual single labeling cost value, the second labeling person-times, and the first quantity.
[0196] Optionally, also used for:
[0197] According to the labeling budget value corresponding to the next labeling round and the remaining labeling budget value after the current labeling round, determining whether to put the first number of non-standard data into the next labeling round for labeling operation;
[0198] The step of putting the first number of non-standard data into the next labeling round for labeling includes:
[0199] In response to a labeling budget value corresponding to a next labeling round being less than or equal to a remaining labeling budget value after a current labeling round, the first quantity of the non-standard data is put into a next labeling round for labeling operation.
[0200] Optionally, also used for:
[0201] In response to the annotation budget value corresponding to the next annotation round being greater than the remaining annotation budget value after the current annotation round, the annotation operation is ended.
[0202] A third embodiment of the present application provides an electronic device, Figure 4 It is a schematic diagram of the structure of the electronic device provided in this embodiment.
[0203] like Figure 4 As shown, the electronic device provided in this embodiment includes: a memory 401, a processor 402;
[0204] The memory 401 is used to store computer instructions for executing the data labeling method;
[0205] The processor 402 is configured to execute computer instructions stored in the memory 401 to perform the following operations:
[0206] Acquire a plurality of non-standard data, wherein the non-standard data is data to be annotated whose confidence of the current annotation result does not reach a preset confidence threshold after the annotation operation of the current annotation round;
[0207] Calculate, according to the current labeling result of each of the non-standard data, the first labeling person-times corresponding to each of the non-standard data, where the first labeling person-times is used to represent the minimum labeling person-times required to make the confidence of the current labeling result of the non-standard data reach the preset confidence threshold;
[0208] sorting the plurality of non-compliant data according to the first number of annotations corresponding to each non-compliant data, wherein the non-compliant data with a smaller number of first annotations has a higher priority for delivery;
[0209] According to the delivery priority sorting results of the plurality of non-compliant data, a first number of the non-compliant data with high delivery priority are acquired;
[0210] The first quantity of the non-standard data is put into the next labeling round for labeling operation.
[0211] Optionally, the calculating, according to the current labeling result of each piece of non-compliant data, the first labeling person-times corresponding to each piece of non-compliant data includes:
[0212] Determining, according to the current labeling result of the non-standard data, a tendency value corresponding to the current labeling result of the non-standard data;
[0213] Using the tendency value as the annotation value, a binary search algorithm is used to search within a preset annotation number interval to obtain the minimum number of annotations required to make the confidence of the current annotation result of the non-standard data reach the preset confidence threshold;
[0214] The minimum number of annotated persons is used as the first number of annotated persons corresponding to the non-compliant data.
[0215] Optionally, the prioritizing the multiple non-compliant data according to the first number of annotations corresponding to each non-compliant data includes:
[0216] Arrange the plurality of non-compliant data in ascending order according to the first number of annotation persons corresponding to each non-compliant data;
[0217] The ascending order result is used as the delivery priority sorting result of the multiple non-compliant data.
[0218] Optionally, before the step of acquiring a first number of the non-compliant data with high delivery priority according to the delivery priority sorting results of the plurality of non-compliant data, the following is further performed:
[0219] Determining the first number of the non-compliant data to be pre-invested in the next marking round according to the number of the plurality of non-compliant data and a first preset number;
[0220] The first preset number is used to represent the number of unqualified to-be-annotated data that is preset for the labeling task.
[0221] Optionally, determining the first number of the non-compliant data to be pre-invested in the next marking round according to the number of the plurality of non-compliant data and a first preset number includes:
[0222] In response to the number of the plurality of non-compliant data being greater than twice the first preset number, taking half of the number of the plurality of non-compliant data as the first number;
[0223] In response to the number of the plurality of non-standard data being less than or equal to twice the first preset number and greater than or equal to the first preset number, using the first preset number as the first number;
[0224] In response to the number of the plurality of non-standard data being less than the first preset number, the number of the plurality of non-standard data is taken as the first number.
[0225] Optionally, before the step of obtaining a plurality of non-standard data, the following steps are further performed:
[0226] Determine the remaining annotation budget value after the current annotation round according to the preset total annotation budget value and the annotation cost values consumed in the current annotation round and the historical annotation rounds; wherein the historical annotation rounds are the annotation rounds before the current annotation round;
[0227] The obtaining of multiple non-standard data includes:
[0228] In response to the remaining annotation budget value after the current annotation round being greater than zero, the plurality of non-standard data are acquired.
[0229] Optionally, also execute:
[0230] In response to the remaining annotation budget value after the current annotation round being equal to zero, and / or the non-standard data not existing, the annotation operation is terminated.
[0231] Optionally, before the step of putting the first quantity of the non-standard data into the next labeling round for labeling, the following is further performed:
[0232] Determine, according to the first number of annotation persons corresponding to each non-compliant data in the first number of non-compliant data, a second number of annotation persons corresponding to the non-compliant data in a next annotation round, wherein the second number of annotation persons is used to represent an average number of annotation persons for the first number of non-compliant data in the next annotation round;
[0233] Determine the actual single annotation cost value corresponding to the data to be annotated according to the annotation cost value consumed in the current annotation round and the historical annotation rounds, and the actual number of annotation persons for the data to be annotated in the current annotation round and the historical annotation rounds;
[0234] A labeling budget value corresponding to a next labeling round is determined according to the actual single labeling cost value, the second labeling person-times, and the first quantity.
[0235] Optionally, also execute:
[0236] According to the labeling budget value corresponding to the next labeling round and the remaining labeling budget value after the current labeling round, determining whether to put the first number of non-standard data into the next labeling round for labeling operation;
[0237] The step of putting the first number of non-standard data into the next labeling round for labeling includes:
[0238] In response to a labeling budget value corresponding to a next labeling round being less than or equal to a remaining labeling budget value after a current labeling round, the first quantity of the non-standard data is put into a next labeling round for labeling operation.
[0239] Optionally, also execute:
[0240] In response to the annotation budget value corresponding to the next annotation round being greater than the remaining annotation budget value after the current annotation round, the annotation operation is ended.
[0241] Optionally, before the step of putting the first quantity of the non-standard data into the next labeling round for labeling, the following is further performed:
[0242] According to the first number of annotation people corresponding to each non-compliant data in the first number of non-compliant data, determine the third number of annotation people corresponding to the non-compliant data in the next annotation round, and the third number of annotation people is used to represent the maximum number of annotation people for the first number of non-compliant data in the next annotation round.
[0243] Optionally, the step of putting the first quantity of non-standard data into a next labeling round for labeling includes:
[0244] The non-compliant data is distributed to a labeling staff, and a labeling result of the non-compliant data is determined according to the labeling value of the non-compliant data by the labeling staff and the current labeling result of the non-compliant data;
[0245] In response to the confidence of the labeling result of the non-compliant data reaching the preset confidence threshold, or the number of labeling people for the non-compliant data reaching the third number of labeling people, the labeling operation on the non-compliant data in the next labeling round is terminated.
[0246] Optionally, if the current marking round is an initial marking round; before acquiring a plurality of non-standard data, the marking operation of the current marking round includes:
[0247] Determine, according to the number of to-be-annotated data included in the annotation task, a preset annotation budget value for the initial annotation round, and a preset single annotation cost value for the to-be-annotated data, a fourth number of annotation persons corresponding to the to-be-annotated data in the initial annotation round, wherein the fourth number of annotation persons is used to represent a maximum number of annotation persons for the to-be-annotated data in the initial annotation round;
[0248] The data to be labeled is distributed to a labeler, and a labeling result of the data to be labeled is determined according to the labeling value of the data to be labeled by the labeler;
[0249] In response to the confidence of the labeling result of the data to be labeled reaching the preset confidence threshold, or the number of labeling people for the data to be labeled reaching the fourth number of labeling people, ending the labeling operation on the data to be labeled;
[0250] The data to be labeled whose number of labeling persons reaches the fourth number of labeling persons and whose confidence does not reach the preset confidence threshold is taken as the non-compliant data.
[0251] A fourth embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium includes computer instructions. When the computer instructions are executed by a processor, they are used to implement the methods described in the embodiments of the present application.
[0252] It should be noted that the relational terms such as "first" and "second" in this article are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "include", "have", "include" and "includes" and other similar forms are the same in meaning, and the end of any one or more items after any of the above words is open-ended, and any of the above nouns does not mean that the one or more items have been exhaustively listed or are limited to the one or more items listed.
[0253] As used herein, unless expressly stated otherwise, the term "or" includes all possible combinations, except those that are not feasible. For example, if it is expressed that a database may include A or B, then unless otherwise specifically stated or not feasible, it may include database A, or B, or A and B. As a second example, if it is expressed that a database may include A, B, or C, then unless otherwise specifically stated or not feasible, the database may include database A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.
[0254] It is worth noting that the above embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. When the software is executed by a processor, it can execute the above-mentioned disclosed method. The computing unit and other functional units described in the present disclosure can be implemented by hardware or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above-mentioned multiple modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.
[0255] In the above detailed description, the embodiments have been described with reference to many specific details, which may vary from implementation to implementation. Certain adaptations and modifications may be made to the embodiments. For those skilled in the art, other embodiments may be readily apparent from the specific embodiments disclosed in the present application. This specification and examples are for exemplary purposes only, and the true scope and essence of the present application are described by the claims. The sequence of steps shown in the diagrams is also for the purpose of explanation only and is not intended to be limited to any particular steps or sequence. Therefore, those skilled in the art will appreciate that these steps may be performed in different orders when implementing the same method.
[0256] In the drawings and detailed description of the present application, exemplary embodiments are disclosed. However, many variations and modifications may be made to these embodiments. Accordingly, although specific terms are used, these terms are only general and descriptive and not for the purpose of limitation.
Claims
1. A data labeling method, characterized in that: The method comprises: Acquire a plurality of non-standard data, wherein the non-standard data is data to be annotated whose confidence of the current annotation result does not reach a preset confidence threshold after the annotation operation of the current annotation round; Calculate, according to the current labeling result of each of the non-standard data, the first labeling person-times corresponding to each of the non-standard data, where the first labeling person-times is used to represent the minimum labeling person-times required to make the confidence of the current labeling result of the non-standard data reach the preset confidence threshold; sorting the plurality of non-compliant data according to the first number of annotations corresponding to each non-compliant data, wherein the non-compliant data with a smaller number of first annotations has a higher priority for delivery; According to the delivery priority sorting results of the plurality of non-compliant data, a first number of the non-compliant data with high delivery priority are acquired; The first quantity of the non-standard data is put into the next labeling round for labeling operation.
2. The method according to claim 1, characterized in that The calculating, according to the current labeling result of each of the non-standard data, the first labeling person-times corresponding to each of the non-standard data includes: Determining, according to the current labeling result of the non-standard data, a tendency value corresponding to the current labeling result of the non-standard data; Using the tendency value as the annotation value, a binary search algorithm is used to search within a preset annotation number interval to obtain the minimum number of annotations required to make the confidence of the current annotation result of the non-standard data reach the preset confidence threshold; The minimum number of annotated persons is used as the first number of annotated persons corresponding to the non-compliant data.
3. The method according to claim 1, characterized in that The step of sorting the plurality of non-compliant data according to the first number of times of annotation corresponding to each non-compliant data, includes: Arrange the plurality of non-compliant data in ascending order according to the first number of annotation persons corresponding to each non-compliant data; The ascending order result is used as the delivery priority sorting result of the multiple non-compliant data.
4. The method according to claim 1, characterized in that: Before the step of acquiring a first number of the non-compliant data with a high delivery priority according to the delivery priority sorting results of the plurality of non-compliant data, the method further includes: Determining the first number of the non-compliant data to be pre-invested in the next marking round according to the number of the plurality of non-compliant data and a first preset number; The first preset number is used to represent the number of unqualified to-be-annotated data that is preset for the labeling task.
5. The method according to claim 4, characterized in that The determining, according to the number of the plurality of non-compliant data and a first preset number, the first number of the non-compliant data to be pre-invested in the next marking round includes: In response to the number of the plurality of non-compliant data being greater than twice the first preset number, taking half of the number of the plurality of non-compliant data as the first number; In response to the number of the plurality of non-standard data being less than or equal to twice the first preset number and greater than or equal to the first preset number, using the first preset number as the first number; In response to the number of the plurality of non-standard data being less than the first preset number, the number of the plurality of non-standard data is taken as the first number.
6. The method according to claim 1, characterized in that Before the step of obtaining a plurality of non-standard data, the method further includes: Determine the remaining annotation budget value after the current annotation round according to the preset total annotation budget value and the annotation cost values consumed in the current annotation round and the historical annotation rounds; wherein the historical annotation rounds are the annotation rounds before the current annotation round; The obtaining of multiple non-standard data includes: In response to the remaining annotation budget value after the current annotation round being greater than zero, the plurality of non-standard data are acquired.
7. The method according to claim 6, characterized in that The method further comprises: In response to the remaining annotation budget value after the current annotation round being equal to zero, and / or the non-standard data not existing, the annotation operation is terminated.
8. The method according to claim 6, characterized in that Before the step of putting the first number of non-standard data into the next labeling round for labeling, the method further includes: Determine, according to the first number of annotation persons corresponding to each of the non-compliant data in the first number of the non-compliant data, a second number of annotation persons corresponding to the non-compliant data in a next annotation round, wherein the second number of annotation persons is used to represent an average number of annotation persons for the first number of the non-compliant data in the next annotation round; Determine the actual single annotation cost value corresponding to the data to be annotated according to the annotation cost value consumed in the current annotation round and the historical annotation rounds, and the actual number of annotation persons for the data to be annotated in the current annotation round and the historical annotation rounds; A labeling budget value corresponding to a next labeling round is determined according to the actual single labeling cost value, the second labeling person-times, and the first quantity.
9. The method according to claim 8, characterized in that The method further comprises: According to the labeling budget value corresponding to the next labeling round and the remaining labeling budget value after the current labeling round, determining whether to put the first number of non-standard data into the next labeling round for labeling operation; The step of putting the first number of non-standard data into the next labeling round for labeling includes: In response to the fact that the annotation budget value corresponding to the next annotation round is less than or equal to the remaining annotation budget value after the current annotation round, the first quantity of the non-standard data is put into the next annotation round for annotation operation.
10. The method according to claim 9, characterized in that The method further comprises: In response to the annotation budget value corresponding to the next annotation round being greater than the remaining annotation budget value after the current annotation round, the annotation operation is ended.
11. The method according to claim 1, characterized in that: Before the step of putting the first number of non-standard data into the next labeling round for labeling, the method further includes: According to the first number of annotation people corresponding to each non-compliant data in the first number of non-compliant data, determine the third number of annotation people corresponding to the non-compliant data in the next annotation round, and the third number of annotation people is used to represent the maximum number of annotation people for the first number of non-compliant data in the next annotation round.
12. The method according to claim 11, characterized in that The step of putting the first number of non-standard data into the next labeling round for labeling includes: The non-compliant data is distributed to a labeling staff, and a labeling result of the non-compliant data is determined according to the labeling value of the non-compliant data by the labeling staff and the current labeling result of the non-compliant data; In response to the confidence of the labeling result of the non-compliant data reaching the preset confidence threshold, or the number of labeling people for the non-compliant data reaching the third number of labeling people, the labeling operation on the non-compliant data in the next labeling round is terminated.
13. The method according to claim 1, characterized in that If the current marking round is an initial marking round; before obtaining a plurality of non-standard data, the marking operation of the current marking round includes: Determine, according to the number of to-be-annotated data included in the annotation task, a preset annotation budget value for the initial annotation round, and a preset single annotation cost value for the to-be-annotated data, a fourth number of annotation persons corresponding to the to-be-annotated data in the initial annotation round, wherein the fourth number of annotation persons is used to represent a maximum number of annotation persons for the to-be-annotated data in the initial annotation round; The data to be labeled is distributed to a labeler, and a labeling result of the data to be labeled is determined according to the labeling value of the data to be labeled by the labeler; In response to the confidence of the labeling result of the data to be labeled reaching the preset confidence threshold, or the number of labeling people for the data to be labeled reaching the fourth number of labeling people, ending the labeling operation on the data to be labeled; The data to be labeled whose number of labeling persons reaches the fourth number of labeling persons and whose confidence does not reach the preset confidence threshold is taken as the non-compliant data.
14. A data labeling device, characterized in that: The device comprises: a first data acquisition unit, a marked person-time calculation unit, a delivery priority sorting unit, a second data acquisition unit, and a data delivery unit; The first data acquisition unit is used to acquire a plurality of non-standard data, wherein the non-standard data are data to be labeled whose confidence of the current labeling result does not reach a preset confidence threshold after the labeling operation of the current labeling round; The annotation person-time calculation unit is used to calculate the first annotation person-time corresponding to each non-standard data according to the current annotation result of each non-standard data, wherein the first annotation person-time is used to represent the minimum annotation person-time required for the confidence of the current annotation result of the non-standard data to reach the preset confidence threshold; The delivery priority sorting unit is used to sort the delivery priorities of the plurality of non-compliant data according to the first number of annotations corresponding to each of the non-compliant data, wherein the non-compliant data with a smaller number of first annotations has a higher delivery priority; The second data acquisition unit is used to acquire a first number of the non-compliant data with high delivery priority according to the delivery priority sorting result of the plurality of non-compliant data; The data delivery unit is used to deliver the first quantity of non-standard data into the next marking round for marking operation.
15. An electronic device, characterized in that: include: Memory, processor; The memory is used to store one or more computer instructions; The processor is used to execute the one or more computer instructions to implement the method according to any one of claims 1-13.
16. A computer-readable storage medium having one or more computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the method according to any one of claims 1 to 13 is performed.