Data labeling method and device, equipment, storage medium and program product
Patent Information
- Application Number
- CN202511417922.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-09-30
AI Technical Summary
[0005]本公开提供一种数据标注方法、装置、设备、存储介质及程序产品,以至少解决现有音频情感数据的标注效率较低,数据标注的准确度较低的问题
在本公开的一些实施例中,获取标注任务、标注任务对应的标签集合、全量待标注数据和多个原始AI模型;根据标注任务、标签集合、全量待标注数据和第一AI模型进行标注人员能力验证操作,确定符合标注条件的标注人员名单,通过基于标注任务和标签集合筛选适配的第一AI模型对标注人员进行能力验证;其中,第一AI模型是多个原始AI模型中能够覆盖标签集合中的基础标签子集的AI模型;根据标注任务、标签集合和第二AI模型,利用第一AI模型进行加权众投操作,得到第一标注数据集;其中,第二AI模型是多个原始AI模型中不能覆盖标签集合的AI模型;根据第一标注数据集的数据质量,从第一标注数据集中选择出数据质量小于质量阈值的标注数据集进行修正,得到修正后的第二标注数据集,利用覆盖不全的第二AI模型结合第一AI模型进行加权众投与质量的迭代修正,提高数据标注的效率,提高数据标注的准确度。
Smart Images

Figure CN121456455B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a data annotation method, apparatus, device, storage medium, and program product. Background Technology
[0002] With the deepening development of artificial intelligence technology in fields such as speech processing, affective computing, and human-computer interaction, subjective data annotation tasks are playing an increasingly crucial role in constructing high-quality training datasets. Among them, audio affective annotation, as a typical subjective annotation task, aims to assign emotion category labels (such as joy, sadness, anger, etc.) to speech data, and its annotation results directly affect the performance of models such as affective recognition and affective speech synthesis.
[0003] Current methods for annotating audio emotion data mainly rely on manual annotation, which makes the results susceptible to human influence.
[0004] Currently, the annotation efficiency and accuracy of audio emotion data are low. Summary of the Invention
[0005] This disclosure provides a data annotation method, apparatus, device, storage medium, and program product to at least solve the problems of low annotation efficiency and low accuracy of existing audio emotion data.
[0006] The technical solution disclosed herein is as follows: This disclosure provides a data annotation method, including: Obtain the annotation task, the label set corresponding to the annotation task, the full amount of data to be annotated, and multiple original AI models; Based on the annotation task, the tag set, the full amount of data to be annotated, and the first AI model, an annotation personnel capability verification operation is performed to determine a list of annotation personnel who meet the annotation conditions; wherein, the first AI model is an AI model among multiple original AI models that can cover a subset of basic tags in the tag set; Based on the annotation task, the label set, and the second AI model, a weighted crowdsourcing operation is performed using the first AI model to obtain a first labeled dataset; wherein, the second AI model is an AI model among the multiple original AI models that cannot cover the label set; Based on the data quality of the first labeled dataset, labeled datasets with data quality less than a quality threshold are selected from the first labeled dataset and corrected to obtain the corrected second labeled dataset.
[0007] Optionally, the step of performing an annotation personnel capability verification operation based on the annotation task, the tag set, the full amount of data to be annotated, and the first AI model to determine the list of annotation personnel who meet the annotation conditions includes: Based on the tag set, determine whether there exists an AI model among the multiple original AI models that can cover the tag set; If there is no AI model among the multiple original AI models that can cover the set of labels, then the first AI model that can cover the basic subset of labels is selected from the multiple original AI models. The first AI model is used to infer the full set of unlabeled data to obtain a set of probability distribution vectors for each data point in the basic label space. Based on the set of probability distribution vectors, for newly added labels that are not covered, sample data of the newly added labels are collected and input into the first AI model to generate the probability output of the newly added label sample data in the basic label space, and to construct an initial feature vector representing the newly added label through cluster analysis to obtain the initial feature vector set of the newly added label. Based on the initial feature vector set of the newly added labels, the first data with the highest probability of model output being in the preset middle interval is selected from the full set of unlabeled data. The normalized cosine similarity between the probability vector of the first data and the initial feature vector of the newly added labels is calculated. When the normalized cosine similarity exceeds the set similarity threshold, the first data is marked as the candidate data of the corresponding newly added labels, thus obtaining the extended labeled dataset. The feature vector of each newly added label is recursively updated based on the extended labeled dataset to obtain the corrected set of feature vectors for the newly added labels. Based on the newly added label feature vector set, the consistency between the model output labels and the manual annotation results is compared, and the data is divided into correct groups or incorrect groups to obtain the data annotation evaluation subset; Based on the data annotation evaluation subset, candidate annotators are evaluated in stages to obtain a list of annotators who meet the annotation criteria.
[0008] Optionally, dividing the data into correct or incorrect groups includes: For the first confidence score data of the basic label, if the manager's labeled label is equal to the AI model's output label, the first confidence score data is divided into the correct group; if the manager's labeled label is not equal to the AI model's output label but the manager's labeled label is equal to the set probability label, the first confidence score data is divided into the undetermined group; if the manager's labeled label is not equal to the AI model's output label and the manager's labeled label is not equal to the set probability label, the first confidence score data is divided into the incorrect group. For the second confidence data of the basic label and the data in the extended labeled dataset, if the manager's labeled label is equal to the AI model's output label, the second confidence data is classified into the correct group; if the manager's labeled label is not equal to the AI model's output label, the second confidence data is classified into the incorrect group, wherein the confidence of the first confidence data is higher than that of the second confidence data.
[0009] Optionally, the phased evaluation includes: error group evaluation, wherein the phased evaluation of candidate annotators based on the data annotation evaluation subset to obtain a list of annotators who meet the annotation criteria includes: In the error group evaluation, a primary label and alternative labels are selected for each data point in the data labeling evaluation subset; If the main label matches the manager's label, then the candidate labeler will receive bonus points. If the main label is inconsistent with the manager's label and the alternative label is inconsistent with the manager's label, then the candidate labeler will have points deducted. Select individuals whose scores are greater than the score threshold from the candidate annotators to generate an annotator list.
[0010] Optionally, the step of obtaining the first labeled dataset by performing a weighted crowdfunding operation using the first AI model based on the labeling task, the label set, and the second AI model includes: Based on the similarity of semantic or probability distribution among the tags in the tag set, a multi-level mapping relationship composite tag system is constructed to obtain a hierarchical composite tag system. The second AI model is used to perform inference on the full set of unlabeled data to generate prediction results and their confidence probabilities under their respective supporting label spaces, thus obtaining the original set of prediction results. The original prediction results are mapped to the converged label space under the composite label system. If the prediction results of all models are consistent or are mapped to the same composite label, the label is determined to be a crowdfunding label; otherwise, the weighted fusion process is entered, and the output is a preliminary consistent crowdfunding label data and an inconsistent subset of data to be processed. Initial weights are assigned to each of the second AI models to obtain dynamic weighted parameters for each of the second AI models at different label levels; For data with inconsistent voting results, a first quality-labeled dataset containing candidate labels is determined based on the dynamic weighting parameters and the probability values output by the model. Cluster the first quality-labeled dataset to obtain the first labeled dataset.
[0011] Optionally, the step of selecting labeled datasets with data quality below a quality threshold from the first labeled dataset for correction, based on the data quality of the first labeled dataset, to obtain a corrected second labeled dataset includes: The first labeled dataset is evaluated in batches. An AI model that supports all labels is used to infer the data for each batch to obtain the model-discriminated labels and probability distributions for each batch of data, and to divide the data into a subset of data with consistent labels and a subset of data with inconsistent labels. Cluster the subset of labeled consistent data, extract a first confidence level sample and a second confidence level sample; and select a reverse bias sample from the subset of labeled inconsistent data, wherein the confidence level of the first confidence level sample is higher than the confidence level of the second confidence level sample. The first confidence level sampling sample, the second confidence level sampling sample, and the reverse bias sample are corrected to obtain the corrected second labeled dataset.
[0012] This disclosure also provides a data annotation apparatus, including: The acquisition module is used to acquire the annotation task, the label set corresponding to the annotation task, the full amount of data to be annotated, and multiple original AI models; The determination module is used to perform an annotation personnel capability verification operation based on the annotation task, the tag set, the full amount of data to be annotated, and the first AI model, and determine the list of annotation personnel who meet the annotation conditions; wherein, the first AI model is an AI model among multiple original AI models that can cover the basic tag subset in the tag set; The annotation module is used to perform a weighted crowdsourcing operation using the first AI model based on the annotation task, the label set, and the second AI model to obtain a first labeled dataset; wherein the second AI model is an AI model that cannot cover the label set among the multiple original AI models; The correction module is used to select, based on the data quality of the first labeled dataset, labeled datasets whose data quality is less than a quality threshold from the first labeled dataset and correct them to obtain a corrected second labeled dataset.
[0013] This disclosure also provides an electronic device, including: processor; Memory used to store processor-executable instructions; The processor is configured to execute instructions to implement the steps in the above method.
[0014] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor of the steps in the above-described method.
[0015] This disclosure also provides a computer program product, including a computer program / instructions, which are executed by a processor through the steps of the methods described above.
[0016] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: In some embodiments of this disclosure, a labeling task, a label set corresponding to the labeling task, a full set of data to be labeled, and multiple original AI models are obtained. Based on the labeling task, label set, full set of data to be labeled, and a first AI model, a capability verification operation for labelers is performed to determine a list of labelers who meet the labeling criteria. The capability verification of labelers is conducted by selecting a suitable first AI model based on the labeling task and label set. The first AI model is an AI model among the multiple original AI models that can cover a subset of basic labels in the label set. Based on the labeling task, label set, and second AI model, a weighted crowdsourcing operation is performed using the first AI model to obtain a first labeled dataset. The second AI model is an AI model among the multiple original AI models that cannot cover the label set. Based on the data quality of the first labeled dataset, labeled datasets with data quality below a quality threshold are selected from the first labeled dataset for correction to obtain a corrected second labeled dataset. The second AI model with incomplete coverage is combined with the first AI model for weighted crowdsourcing and iterative quality correction to improve the efficiency and accuracy of data labeling.
[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0019] Figure 1 A flowchart illustrating a data annotation method provided for an exemplary embodiment of this disclosure; Figure 2 A schematic diagram of the structure of a data annotation device provided for an exemplary embodiment of this disclosure; Figure 3 A schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this disclosure. Detailed Implementation
[0020] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0021] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure.
[0022] It should be noted that the user information involved in this disclosure includes, but is not limited to, user device information and user personal information; the collection, storage, use, processing, transmission, provision and disclosure of user information in this disclosure all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0023] To address the aforementioned technical issues, some embodiments of this disclosure involve obtaining a labeling task, a label set corresponding to the labeling task, a full set of data to be labeled, and multiple original AI models. Based on the labeling task, label set, full set of data to be labeled, and a first AI model, a capability verification operation is performed on the labelers to determine a list of labelers who meet the labeling criteria. The labelers' capabilities are verified by selecting a suitable first AI model based on the labeling task and label set. The first AI model is an AI model among the multiple original AI models that can cover a subset of basic labels in the label set. Based on the labeling task, label set, and second AI model, a weighted crowdsourcing operation is performed using the first AI model to obtain a first labeled dataset. The second AI model is an AI model among the multiple original AI models that cannot cover the label set. Based on the data quality of the first labeled dataset, labeled datasets with data quality below a quality threshold are selected from the first labeled dataset for correction to obtain a corrected second labeled dataset. The second AI model with incomplete coverage is combined with the first AI model for weighted crowdsourcing and iterative quality correction, thereby improving the efficiency and accuracy of data labeling.
[0024] The technical solutions provided by the embodiments of this disclosure are described in detail below with reference to the accompanying drawings.
[0025] Figure 1 This is a flowchart illustrating a data annotation method provided for an exemplary embodiment of this disclosure. Figure 1 As shown, the method includes: S101; Obtain the annotation task, the label set corresponding to the annotation task, the full amount of data to be annotated, and multiple original AI models; S102: Based on the annotation task, label set, full unannotated data, and first AI model, perform annotation personnel capability verification operations to determine the list of annotation personnel who meet the annotation conditions; among them, the first AI model is the AI model among multiple original AI models that can cover the basic label subset in the label set; S103: Based on the annotation task, the label set, and the second AI model, a weighted crowdsourcing operation is performed using the first AI model to obtain the first labeled dataset; wherein, the second AI model is the AI model among multiple original AI models that cannot cover the label set; S104: Based on the data quality of the first labeled dataset, select labeled datasets whose data quality is less than the quality threshold from the first labeled dataset for correction, and obtain the corrected second labeled dataset.
[0026] In this embodiment, the entity executing the above method can be a terminal device or a server.
[0027] The terminal device includes, but is not limited to, mobile stations (MS), mobile terminals, mobile phones, handsets, and portable equipment. This terminal device can communicate with one or more core networks via a radio access network (RAN). For example, the terminal device can be a mobile phone (or "cellular" phone), a computer with wireless communication capabilities, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an AR terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical care, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. The operating systems installed on the terminal device include, but are not limited to, iOS, Android, Windows, Linux, and Mac OS. In different networks, terminals may be called by different names, such as: user equipment, mobile station, user unit, station, cellular phone, personal digital assistant, wireless modem, wireless communication device, handheld device, laptop, cordless phone, wireless local loop station, television, etc. For ease of description, this embodiment will simply refer to it as terminal device.
[0028] In this embodiment, the implementation form of the server is not limited. For example, the server can be a conventional server, a cloud server, a cloud host, a virtual center, or other server devices. The server mainly consists of a processor, hard disk, memory, system bus, and other common computer architecture types.
[0029] In some embodiments of this disclosure, an annotation personnel capability verification operation is performed based on the annotation task, the label set, the full set of data to be labeled, and the first AI model to determine a list of annotation personnel who meet the annotation conditions. One possible approach is to determine whether there is an AI model among multiple original AI models that can cover the label set based on the label set; if there is no AI model among multiple original AI models that can cover the label set, then select a first AI model from the multiple original AI models that can cover a subset of the basic labels; use the first AI model to infer the full set of data to be labeled to obtain a set of probability distribution vectors for each data point in the basic label space; based on the set of probability distribution vectors, for newly added labels that are not covered, collect sample data of the newly added labels and input it into the first AI model to generate the probability output of the newly added label sample data in the basic label space, and construct an initial feature vector representing the newly added labels through cluster analysis to obtain an initial feature vector set of the newly added labels; based on the initial feature vector set of the newly added labels... The initial feature vector set is selected from the full set of unlabeled data, choosing the first data point whose model output has the highest probability of falling within a preset middle interval. The normalized cosine similarity between the probability vector of the first data point and the initial feature vector of the new label is calculated. When the normalized cosine similarity exceeds a set similarity threshold, the first data point is marked as candidate data for the corresponding new label, resulting in an expanded label dataset. The feature vector of each new label is recursively updated based on the expanded label dataset, resulting in a corrected set of new label feature vectors. Based on the set of new label feature vectors, the consistency between the model output label and the manual annotation results is compared, and the data is divided into correct or incorrect groups, resulting in a data annotation evaluation subset. Based on the data annotation evaluation subset, candidate annotators are evaluated in stages to obtain a list of annotators who meet the annotation criteria.
[0030] In the above embodiments, the data is divided into a correct group or an incorrect group. One possible approach is as follows: For the first confidence data of the basic labels, if the manager's labeled label equals the AI model's output label, the first confidence data is divided into the correct group; if the manager's labeled label does not equal the AI model's output label but equals the set probability label, the first confidence data is divided into the undetermined group; if the manager's labeled label does not equal the AI model's output label and does not equal the set probability label, the first confidence data is divided into the incorrect group. For the second confidence data of the basic labels and the data in the extended labeled dataset, if the manager's labeled label equals the AI model's output label, the second confidence data is divided into the correct group; if the manager's labeled label does not equal the AI model's output label, the second confidence data is divided into the incorrect group, wherein the confidence of the first confidence data is higher than that of the second confidence data.
[0031] In the above embodiments, candidate annotators are evaluated in stages based on the data annotation evaluation subset to obtain a list of annotators who meet the annotation criteria. One possible approach is to select a primary label and alternative labels for each piece of data in the error group evaluation subset; if the primary label matches the manager's annotation label, the candidate annotator receives bonus points; if the primary label does not match the manager's annotation label and the alternative labels do not match the manager's annotation label, the candidate annotator receives deduction points; and from the candidate annotators, those with scores greater than a score threshold are selected to generate the list of annotators.
[0032] Specifically, the data annotation evaluation set is the set of annotated samples. For subjective annotation tasks, the administrator (the dataset requester) needs to clearly define the annotation task and provide annotation examples. Taking the audio sentiment annotation task as an example, let the full set of data to be labeled be ALL_Unlabeled_Dataset. In this disclosure and data cleaning process, the data to be labeled, ALL_Unlabeled_Dataset, refers to the cleaned, high-quality data.
[0033] Step 1: The administrator identifies all the labels for this subjective labeling task and selects an available AI model accordingly. Taking audio emotion as an example, all labels could be ['angry', 'disgusted', 'fearful', 'happy', 'neutral', 'other', 'sad', 'surprised'], denoted as LabelExampleA for ease of description; or they could be ['angry', 'happy', 'neutral', 'sad', 'surprised'], denoted as LabelExampleB. A comparison shows that LabelExampleA has more content than LabelExampleB. Based on the labels, a corresponding available AI model is selected. Assume that the publicly available AI model ModelB fully matches and supports the labels in LabelExampleB, but does not support the labels ['disgusted', 'fearful'] in LabelExampleA. That is, there are two possible scenarios: first, there is a publicly available model that covers all the labels determined by the administrator; second, there is an available model that can cover all the labels determined by the administrator.
[0034] Step 2: The handling of the first case is as follows: Use the AI model ModelB to infer the data to be labeled. Suppose that after inferring a certain audio (denoted as unlabeled_wav_a), the following data is obtained [0.568937361, 0, 0, 0, 0.264778376]. This list is denoted as emotProbList_a, where the values in the list represent the probability of a certain emotion. This data shows that unlabeled_wav_a has a 56% probability of being angry and a 26% probability of being surprised. Finally, the model will consider the audio to be an angry emotion. Randomly infer the data to be labeled and select emot_X_up_N + emot_X_mid_M data points and their corresponding labels. Explanation of variable structure: `emot_X` represents a specific emotion (e.g., anger); `up` indicates that the probability of `emot_X` exceeds a certain threshold (e.g., 98%); `mid` indicates that the probability of `emot_X` is the highest among all emotions and falls within a certain threshold range (e.g., 45%-70%); N and M represent the quantity (human-preset values, e.g., 10). The filtering process checks the `emotProbList`. For example, if `emot_X_up` is satisfied, it is saved. At this point, `emot_X_up` still needs N-1 more data entries. This continues until all required data is found. The formula for the total number of data entries is: In this formula, emot_1 represents anger, emot_2 represents happiness, and so on, with emot_5 representing surprise. The x in the formula is an incrementing subscript that iterates through all the emotion tags.
[0035] Step 3: The manager labels the data selected in Step 2, analyzes the model's output labels and the manager's output labels, and categorizes the data into different groups as subsets for data labeling and evaluation. The logic for analyzing labels and dividing data is as follows: For emot_X_up data: When the manager's output label equals the model's output label, the data entry is classified into the "correct" group. When the manager's output label is not equal to the model's output label, but is equal to the model's second-highest probability label, the data should be classified into the "pending" group. When the manager outputs a label that is not equal to the model output label, and is not equal to the model's second-highest probability label, the data entry should be classified into the "error" group. For emot_X_mid data: When the manager's output label equals the model's output label, the data entry is classified into the "correct" group. When the administrator outputs a label that is not equal to the model outputs a label, the data entry is classified into the "Error" group.
[0036] Step 4: The second scenario is handled as follows: The AI model ModelB is still needed for inference on the labeled data. ModelB aims to cover as many sentiment labels as possible. Continuing with the example above, ModelB supports five sentiment labels ['angry', 'happy', 'neutral', 'sad', 'surprised']. Now, two more sentiment labels are added: ['disgusted', 'fearful']. A small number of audio files with the newly added sentiment labels are recorded. These audio files are then fed into the ModelB model for judgment, obtaining the probabilities of all sentiment labels corresponding to each audio file. Let the newly added audio file be wav_disg_1, representing audio file 1 with the disgust sentiment. The sentiment label emotProbList provided by ModelB might be [0.488736293, 0, 0.1, 0.283764476, 0]. For all audio files with the same emotion (e.g., disgust, wav_disg_all), the emotion tag `emotProbList` will form a two-dimensional table. Let the quinary linear formula be y = ax1 + bx2 + cx3 + dx4 + ex5 + f. By calculating the Euclidean distance of each row in the two-dimensional table, the formula is obtained based on the idea of unsupervised clustering (i.e., finding the parameters a / b / c / d / e / f). For example, the formula for disgust is: y = 0.58x1 + 0.04x3 + 0.28x4 + 0.09, which means that anger accounts for 58%, happiness for 0%, neutrality for 4%, sadness for 28%, surprise for 0%, and 9% is the adjustment coefficient used in the normalization process, corresponding to the parameter f in the formula. For the supported tags, select emot_X_up_N + emot_X_mid_M data points and their corresponding tags according to the second step. The following focuses on the selection process of adding new emotions.
[0037] For randomly inferred data to be labeled, when the maximum sentiment probability output by a certain data point is between 0.4 and 0.6, the cosine similarity between that data point and the newly added sentiment is calculated and normalized to [0,1]. For example, the vector for the aversion sentiment, as mentioned earlier, is [0.58,0,0.04,0.28,0,0.09]. That is, the cosine similarity between the two vectors is calculated. When the normalized cosine similarity exceeds a threshold (e.g., 0.7), it is filtered out and labeled as a new sentiment tag (e.g., aversion), thus completing the process of using the ModelB model to generate sentiment tags that were not originally supported. The filtered data and tags are denoted as emot_NewD.
[0038] After a certain amount of data to be labeled and a certain amount of new labels are generated through random inference, the inference process will pause, and recursively iterate through the sentiment vector. The aversion sentiment vector [0.58, 0, 0.04, 0.28, 0, 0.09] calculated earlier was determined with a small amount of data. As the data volume increases, it needs to be updated iteratively. At this point, the data volume is the initial wav_disg_all + wav_disg_New, which represents the audio files with aversion labels output during the inference process. Subsequent inferences will use the updated sentiment vector.
[0039] Step 5: Perform similar logic as in step 3 on the emot_NewD data. If the manager's output label equals the model's output label, it is classified into the "correct" group; otherwise, it is classified into the "incorrect" group. The manager then confirms the final label of the data. Finally, the selected data is used as a subset for data labeling evaluation.
[0040] Step 6: Evaluate the annotators using the data annotation evaluation subset. An interactive annotation system exists at the front end, prioritizing the "Correct" subset. When an annotator's score exceeds 80, the "Incorrect" subset is activated, and a new score is calculated. For the "Correct" subset, a label is considered correct and scores points only if the label output by the annotator matches the label in the data annotation evaluation set; otherwise, points are deducted. The initial score is 60 points. For the "Incorrect" subset, annotators can select two labels: a primary label and a backup label. If the primary label matches, it is considered correct and scores points; if the primary labels differ, but the backup label equals the manager's label, no points are deducted or scored; otherwise, points are deducted. Finally, an annotator participating in the evaluation is considered competent for this subjective annotation work only if they obtain two scores (the first score must be above 80 for the "Incorrect" subset to be activated), and the second score exceeds the threshold.
[0041] The labels provided by the labelers in this publication are more likely to align with the labeling expectations of the administrators, reducing the likelihood of different people assigning different labels to the same data. Because when there are discrepancies, the label given by the administrators should prevail, the labelers' subjective choices should be closer to the administrators' expectations.
[0042] In some embodiments of this disclosure, a first labeled dataset is obtained by performing a weighted crowdsourcing operation using a first AI model based on the annotation task, a label set, and a second AI model. One possible approach is to construct a multi-level mapping composite label system based on the semantic or probability distribution similarity between labels in the label set, resulting in a hierarchical composite label system; use the second AI model to infer the entire set of data to be labeled, generating prediction results and their confidence probabilities under their respective supporting label spaces, resulting in an original set of prediction results; map the original prediction results to the converged label space under the composite label system; if the prediction results of all models are consistent or mapped to the same composite label, then the label is determined to be a crowdsourced label; otherwise, a weighted fusion process is entered, outputting initially consistent crowdsourced label data and inconsistent subsets of data to be processed; initial weights are assigned to each second AI model, obtaining dynamic weighting parameters for each second AI model at different label levels; for data with inconsistent crowdsourcing results, a first quality labeled dataset containing candidate labels is determined based on the dynamic weighting parameters and the probability values output by the models; the first quality labeled dataset is clustered to obtain the first labeled dataset.
[0043] Specifically, for a large amount of data to be labeled, it needs to be automatically labeled by an AI model first, and then manually corrected.
[0044] This disclosure uses multiple AI models for crowdsourcing to determine initial data labels and evaluate label quality. Using multiple AI models for crowdsourcing and selecting the one with the most votes as the final label is a common approach among industry professionals. However, existing technologies all assume that these multiple AI models perfectly support the labeling task, which is a flawed assumption. For example, consider three AI models, ModelA, ModelB, and ModelC. ModelA supports 7 sentiment labels, ModelB supports 5 sentiment labels, and ModelC supports 3 sentiment labels. When faced with a labeling task requiring 7 sentiment labels, ModelB and ModelC are almost unusable directly. Existing technologies assume that all three AI models support 7 sentiment labels. Based on the natural progression of technology, model iteration follows a pattern from ModelC to ModelB and then to ModelA. Subsequent iterations may support models with 9 sentiments, naturally requiring labeled datasets for 9 sentiments. Therefore, the problem of multiple AI models not perfectly supporting the labeling task exists, and this proposal aims to address this issue.
[0045] Step 1: Tag Aggregation. All tags are merged and aggregated, specifically 7 tags into 5, then into 3. The numbers 7, 5, and 3 are related to the selected AI model, representing the number of emotion categories supported by the model. Using the processing logic from Step 4 in S1, the proximity distance (or similarity) of the emotion tags is calculated. The result is: disgust is categorized as anger, and fear as surprise. At this point, if Model A outputs the disgust tag and Model B outputs the anger tag, then Model A and Model B are considered to have the same tag for the same audio. To distinguish between them, this proposal refers to the tags before aggregation as single tags and the tags after aggregation as composite tags. For example, happiness is a single tag if no other tags are incorporated during the aggregation from 7 tags to 5; anger is a composite tag if the disgust tag is incorporated.
[0046] Step 2: For data where the AI models collectively assign a single label, such as all AI models identifying it as "happy," directly label it as the collectively assigned label and consider it a high-quality label. For data where the AI models collectively assign a composite label, such as three AI models providing labels like "disgust," "angry," and "angry," the audio cannot be directly labeled as "angry." It is necessary to determine whether to refine the label and proceed to Step 3.
[0047] Step 3: Each AI model's vote has a weight, which is the result of normalizing the publicly disclosed accuracy minus 15%. For example, if three AI models have accuracies of [92%, 87%, 84%], the data before normalization is [77%, 72%, 69%], and the weights after normalization are [35.32%, 33.03%, 31.65%]. For single labels, the weight remains unchanged. For composite labels, the weight of the composite model is reduced by a coefficient (e.g., 0.5; the coefficient calculation formula can be supplemented, but not supplementing it does not affect the overall logic). For example, Model B's weight is 33.03%. For the "happy" label, its weight remains unchanged; for the "angry" label, its weight is reduced to 33.03%. 0.5 = 16.52%. Then, the final tag score is calculated using the following formula: Where Wi represents the weight of model i, and Pi represents the probability of the model's output label. The aversion score can be calculated using the formula: 35.32%. 94% 100 = 33.2, the score for anger is 16.52%. 87% 100+9.5% 91% 100 = 14.37 + 8.64 = 23.01. The highest score is selected as the label for this data, i.e., the aversion label, and it is considered high-quality labeled data.
[0048] Step 4: Focus on handling cases of inconsistent labels, such as when three AI models provide labels of [fear], [anger], and [neutral]. Neutral and fear are single labels, while anger is a composite label, but anger does not include fear. In other words, the three models provide three different answers. In this case, the label cannot be determined directly by calculating the score using the process in Step 3. Calculate the score for each label according to the process in Step 3 and save the data for later use. However, regardless of the final label, this data will be considered low-quality labeled data and requires manual correction.
[0049] This disclosure involves clustering low-quality data from automated annotation, followed by manual correction of a small number of audio samples after clustering, resulting in the low-quality annotated data obtained in step four above. Existing techniques primarily rely on manual sampling for correction, and the quality of the annotated data is judged based on these manual checks, limiting the number of corrections possible.
[0050] Step 5: Perform k-means clustering analysis on the low-quality labeled data, where k is 16. The value of k is the average of various combinations. For example, selecting 3 out of 7 sentiments has 35 possible combinations; selecting 3 out of 5 sentiments has 10 possible combinations; and selecting 3 out of 3 sentiments has 1 possible combination. (35+10+1) / 3=16.
[0051] Step 6: After clustering, select 3 data points from the center region of each cluster as "cluster representatives", thus obtaining 3 There are 16 data points to be corrected. In step S1, a competent human annotator is evaluated to correct 48 data points, using two or more labels as the labels for each data point in a given cluster. For example, the 48 data points are [M_k_n], where k takes values of (1, 16) to represent the cluster category, and n takes values of (1, 3) to represent the data number in cluster k. When two data points in M_k_n have the same label, all data points in M_k are considered to have that label. In an extreme case, manually labeling M_k_n may result in n different labels; in this case, step seven is executed. Otherwise, the program exits, and the manual correction ends.
[0052] Step 7: Integrate and analyze the data from extreme cases, calculate the k-value for the clustering algorithm, and then repeat steps 5 and 6. The k-value must be less than the k-value from the previous clustering iteration. For example, the second k=10, which is less than the initial k=16.
[0053] In some embodiments of this disclosure, based on the data quality of the first labeled dataset, labeled datasets with data quality below a quality threshold are selected from the first labeled dataset for correction, resulting in a corrected second labeled dataset. One possible approach is to perform quality assessment on the first labeled dataset in batches, use an AI model supporting all labels to infer the model's discriminant labels and probability distribution for each batch of data, and divide the data into a consistent label subset and an inconsistent label subset; cluster the consistent label subset, extracting first-confidence and second-confidence sampled samples; and select reverse-bias samples from the inconsistent label subset, wherein the confidence level of the first-confidence sampled sample is higher than that of the second-confidence sampled sample; and correct the first-confidence sampled sample, the second-confidence sampled sample, and the reverse-bias sampled sample to obtain the corrected second labeled dataset.
[0054] Specifically, this disclosure tests the quality of labeled data, and if it does not meet the requirements, the above process can be recursively iterated. The sampled data is not random, and therefore the quality of the labeled data is not based on a single data set. This allows for the recursive iteration of the above process, making the correction process targeted and effective.
[0055] Step 1: Divide the labeled data into batches (to avoid excessive data volume) and process each batch in the same manner.
[0056] The second step is to use ModelA (an AI model that supports all labels, specifically the model that supports 7 sentiment labels mentioned earlier) to evaluate the batch data and output the data's discrimination labels. This divides the data into two categories: those whose model discrimination labels are equal to the data's labeled labels and those whose discriminant labels are not equal to the data's labeled labels.
[0057] Step 3: Perform k-means clustering on the data whose model-determined labels match the labeled data, with k=7, representing the number of labels. The data is already categorized before clustering (labels determine the cluster size). The significance of this process is to determine the centers and margins of each cluster. Although each data point has the same final label, the sentiment probabilities for each data point are different (thus forming cluster centers and margins), and the accuracy of the model's label determination is not 100%. Randomly select multiple data points from the centers and margins of each cluster as sample samples. The center data is denoted as Ensure_sample, and the margin data as Vague_sample. The center data is generally accurate, while the margin data may be incorrect, thus providing a basis for further correction of the labeled data.
[0058] Step 4: For data whose model discrimination label is not equal to the data annotation label, calculate the cosine similarity and normalize it to [0,1]. For data below the threshold (e.g., 0.4), randomly select multiple data as sampling samples, denoted as Reverse_sample.
[0059] Step 5: Evaluate the accuracy of the sampled samples Ensure_sample, Vague_sample, and Reverse_sample respectively, and make the above corrections for samples with accuracy lower than the labeling requirements.
[0060] Figure 2 This is a schematic diagram of the structure of a data annotation device 20 provided for an exemplary embodiment of this disclosure. For example... Figure 2 As shown, the data annotation device 20 includes: an acquisition module 21, a determination module 22, an annotation module 23, and a correction module 24.
[0061] Among them, the acquisition module 21 is used to acquire the annotation task, the label set corresponding to the annotation task, the full amount of data to be annotated, and multiple original AI models; The determination module 22 is used to perform the annotation personnel capability verification operation based on the annotation task, the label set, the full amount of data to be annotated, and the first AI model, and to determine the list of annotation personnel who meet the annotation conditions; wherein, the first AI model is the AI model among multiple original AI models that can cover the basic label subset in the label set; The annotation module 23 is used to perform a weighted crowdsourcing operation using the first AI model based on the annotation task, the label set, and the second AI model to obtain the first annotation dataset; wherein, the second AI model is an AI model that cannot cover the label set among multiple original AI models; The correction module 24 is used to select labeled datasets whose data quality is less than a quality threshold from the first labeled dataset based on the data quality of the first labeled dataset, and correct them to obtain the corrected second labeled dataset.
[0062] Optionally, when determining the list of annotators who meet the annotation criteria by performing annotator capability verification operations based on the annotation task, label set, full unannotated data, and the first AI model, module 22 is used for: Based on the label set, determine whether there is an AI model among multiple original AI models that can cover the label set; If there is no AI model among the multiple original AI models that can cover the set of labels, then the first AI model that can cover the basic subset of labels is selected from the multiple original AI models. The first AI model is used to infer the full set of unlabeled data to obtain the set of probability distribution vectors for each data point in the basic label space. Based on the set of probability distribution vectors, for newly added labels that are not covered, new label sample data is collected and input into the first AI model to generate the probability output of the new label sample data in the basic label space, and to construct the initial feature vectors representing the new labels through cluster analysis, thus obtaining the initial feature vector set of the new labels. Based on the initial feature vector set of the new label, the first data with the highest probability of model output being in the preset middle interval is selected from the full set of unlabeled data. The normalized cosine similarity between the probability vector of the first data and the initial feature vector of the new label is calculated. When the normalized cosine similarity exceeds the set similarity threshold, the first data is marked as the candidate data of the corresponding new label, thus obtaining the extended labeled dataset. The feature vector of each newly added label is recursively updated based on the extended labeled dataset to obtain the corrected set of feature vectors for the newly added labels. Based on the newly added label feature vector set, the consistency between the model output labels and the manual annotation results is compared, and the data is divided into correct groups or incorrect groups to obtain the data annotation evaluation subset; Based on the data annotation evaluation subset, candidate annotators are evaluated in stages to obtain a list of annotators who meet the annotation criteria.
[0063] Optionally, when determining whether data is divided into correct or incorrect groups, the determining module 22 is used to: For the first confidence score data of the basic labels, if the manager's labeled label is equal to the AI model's output label, the first confidence score data is divided into the correct group; if the manager's labeled label is not equal to the AI model's output label but is equal to the set probability label, the first confidence score data is divided into the undetermined group; if the manager's labeled label is not equal to the AI model's output label and is not equal to the set probability label, the first confidence score data is divided into the incorrect group. For the second confidence data of the basic label and the data in the extended labeled dataset, if the manager's labeled label is equal to the AI model's output label, the second confidence data is classified as the correct group; if the manager's labeled label is not equal to the AI model's output label, the second confidence data is classified as the incorrect group. The confidence of the first confidence data is higher than that of the second confidence data.
[0064] Optionally, the phased evaluation includes: error group evaluation, which, based on the data annotation evaluation subset, determines that module 22, when conducting phased evaluations of candidate annotators and obtaining a list of annotators who meet the annotation criteria, is used for: In the error group evaluation, a primary label and alternative labels are selected for each data point in the data labeling evaluation subset; If the main label matches the label labeled by the manager, the candidate labeler will receive bonus points. If the main label is inconsistent with the manager's label and the alternative label is inconsistent with the manager's label, then the candidate labeler will be penalized. Select individuals whose scores are greater than the score threshold from the candidate annotators to generate a list of annotators.
[0065] Optionally, when the annotation module 23 obtains the first annotation dataset by performing a weighted crowdsourcing operation using the first AI model based on the annotation task, the label set, and the second AI model, it is used to: Based on the similarity of semantic or probability distributions between tags in the tag set, a multi-level mapping relationship composite tag system is constructed to obtain a hierarchical composite tag system. The second AI model is used to perform inference on the full set of unlabeled data to generate prediction results and their confidence probabilities under their respective supporting label spaces, thus obtaining the original set of prediction results. The original prediction results are mapped to the aggregated label space under the composite label system. If the prediction results of all models are consistent or are mapped to the same composite label, the label is determined to be the crowd-voted label; otherwise, the weighted fusion process is entered, and the output is the crowd-voted label data that is initially consistent and the subset of inconsistent data to be processed. Assign initial weights to each second AI model to obtain dynamic weighted parameters for each second AI model at different label levels; For data with inconsistent crowdfunding results, a first quality labeled dataset containing candidate labels is determined based on dynamic weighting parameters and the probability values output by the model. Cluster the first quality labeled dataset to obtain the first labeled dataset.
[0066] Optionally, when the correction module 24 selects labeled datasets with data quality below a quality threshold from the first labeled dataset based on the data quality of the first labeled dataset for correction, and obtains the corrected second labeled dataset, it is used to: The first labeled dataset was quality-assessed in batches. An AI model that supports all labels was used to infer the data for each batch to obtain the model-discriminated labels and probability distributions for each batch of data, and to divide the data into a subset of data with consistent labels and a subset of data with inconsistent labels. Cluster the subset of labeled consistent data and extract samples with first and second confidence levels; and select reverse bias samples from the subset of labeled inconsistent data, wherein the confidence level of the first confidence level sample is higher than that of the second confidence level sample. The first confidence level sample, the second confidence level sample, and the reverse bias sample are corrected to obtain the corrected second labeled dataset.
[0067] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0068] Figure 3 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of the present disclosure. For example... Figure 3 As shown, the electronic device includes a memory 31 and a processor 32. Additionally, the electronic device also includes a power supply component 33 and a communication component 34.
[0069] Memory 31 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device.
[0070] The memory 31 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0071] Communication component 34 is used for data transmission with other devices.
[0072] The processor 32 can execute computer instructions stored in the memory 31 to: acquire a labeling task, a label set corresponding to the labeling task, a full set of data to be labeled, and multiple original AI models; perform a labeling personnel capability verification operation based on the labeling task, the label set, the full set of data to be labeled, and the first AI model to determine a list of labeling personnel who meet the labeling conditions; wherein, the first AI model is an AI model among multiple original AI models that can cover a subset of basic labels in the label set; based on the labeling task, the label set, and the second AI model, perform a weighted crowdfunding operation using the first AI model to obtain a first labeled dataset; wherein, the second AI model is an AI model among multiple original AI models that cannot cover the label set; based on the data quality of the first labeled dataset, select labeled datasets with data quality less than a quality threshold from the first labeled dataset for correction to obtain a corrected second labeled dataset.
[0073] Accordingly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program. When the computer-readable storage medium stores a computer program, and the computer program is executed by one or more processors, it causes one or more processors to perform... Figure 1 Each step in the method embodiment.
[0074] Accordingly, embodiments of this disclosure also provide a computer program product, which includes a computer program / instructions that are executed by a processor. Figure 1 Each step in the method embodiment.
[0075] The above Figure 3 The communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0076] The above Figure 3 The power supply component provides power to the various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.
[0077] The aforementioned electronic devices also include a display screen and audio components.
[0078] The display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also the duration and pressure associated with the touch or swipe operation.
[0079] An audio component may be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals may be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0080] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0081] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0082] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0083] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0084] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0085] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0086] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0087] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0088] The above are merely specific embodiments of this disclosure, enabling those skilled in the art to understand or implement this disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data annotation method, characterized in that, include: The process involves acquiring a labeling task, a set of labels corresponding to the labeling task, a full set of data to be labeled, and multiple original AI models. The labeling task is an audio emotion task, and the full set of data to be labeled is audio emotion data. Based on the annotation task, the tag set, the full amount of data to be annotated, and the first AI model, an annotation personnel capability verification operation is performed to determine a list of annotation personnel who meet the annotation conditions; wherein, the first AI model is an AI model among multiple original AI models that can cover a subset of basic tags in the tag set; Based on the annotation task, the label set, and the second AI model, a weighted crowdsourcing operation is performed using the first AI model to obtain a first labeled dataset; wherein, the second AI model is an AI model among the multiple original AI models that cannot cover the label set; Based on the data quality of the first labeled dataset, a labeled dataset whose data quality is less than the quality threshold is selected from the first labeled dataset and corrected to obtain a corrected second labeled dataset. The step of verifying the labeling personnel's capabilities based on the labeling task, the tag set, the full amount of data to be labeled, and the first AI model to determine the list of labeling personnel who meet the labeling criteria includes: Based on the tag set, determine whether there exists an AI model among the multiple original AI models that can cover the tag set; If there is no AI model among the multiple original AI models that can cover the set of labels, then the first AI model that can cover the basic subset of labels is selected from the multiple original AI models. The first AI model is used to infer the full set of unlabeled data to obtain a set of probability distribution vectors for each data point in the basic label space. Based on the set of probability distribution vectors, for newly added labels that are not covered, sample data of the newly added labels are collected and input into the first AI model to generate the probability output of the newly added label sample data in the basic label space, and to construct an initial feature vector representing the newly added label through cluster analysis to obtain the initial feature vector set of the newly added label. Based on the initial feature vector set of the newly added labels, the first data with the highest probability of model output being in the preset middle interval is selected from the full set of unlabeled data. The normalized cosine similarity between the probability vector of the first data and the initial feature vector of the newly added labels is calculated. When the normalized cosine similarity exceeds the set similarity threshold, the first data is marked as the candidate data of the corresponding newly added labels, thus obtaining the extended labeled dataset. The feature vector of each newly added label is recursively updated based on the extended labeled dataset to obtain the corrected set of feature vectors for the newly added labels. Based on the newly added label feature vector set, the consistency between the model output labels and the manual annotation results is compared, and the data is divided into correct groups or incorrect groups to obtain the data annotation evaluation subset; Based on the data annotation evaluation subset, candidate annotators are evaluated in stages to obtain a list of annotators who meet the annotation criteria.
2. The method according to claim 1, characterized in that, The process of dividing data into correct or incorrect groups includes: For the first confidence score data of the basic label, if the manager's labeled label is equal to the AI model's output label, the first confidence score data is divided into the correct group; if the manager's labeled label is not equal to the AI model's output label but the manager's labeled label is equal to the set probability label, the first confidence score data is divided into the undetermined group; if the manager's labeled label is not equal to the AI model's output label and the manager's labeled label is not equal to the set probability label, the first confidence score data is divided into the incorrect group. For the second confidence data of the basic label and the data in the extended labeled dataset, if the manager's labeled label is equal to the AI model's output label, the second confidence data is classified into the correct group; if the manager's labeled label is not equal to the AI model's output label, the second confidence data is classified into the incorrect group, wherein the confidence of the first confidence data is higher than that of the second confidence data.
3. The method according to claim 1, characterized in that, The phased evaluation includes: error group evaluation, wherein the phased evaluation of candidate annotators is conducted based on the data annotation evaluation subset to obtain a list of annotators who meet the annotation criteria, including: In the error group evaluation, a primary label and alternative labels are selected for each data point in the data labeling evaluation subset; If the main label matches the manager's label, then the candidate labeler will receive bonus points. If the main label is inconsistent with the manager's label and the alternative label is inconsistent with the manager's label, then the candidate labeler will have points deducted. Select individuals whose scores are greater than the score threshold from the candidate annotators to generate an annotator list.
4. The method according to claim 1, characterized in that, The step of obtaining the first labeled dataset by performing a weighted crowdfunding operation using the first AI model based on the labeling task, the label set, and the second AI model includes: Based on the similarity of semantic or probability distribution among the tags in the tag set, a multi-level mapping relationship composite tag system is constructed to obtain a hierarchical composite tag system. The second AI model is used to perform inference on the full set of unlabeled data to generate prediction results and their confidence probabilities under their respective supporting label spaces, thus obtaining the original set of prediction results. The original prediction results are mapped to the converged label space under the composite label system. If the prediction results of all models are consistent or are mapped to the same composite label, the label is determined to be a crowdfunding label; otherwise, the weighted fusion process is entered, and the output is a preliminary consistent crowdfunding label data and an inconsistent subset of data to be processed. Initial weights are assigned to each of the second AI models to obtain dynamic weighted parameters for each of the second AI models at different label levels; For data with inconsistent voting results, a first quality-labeled dataset containing candidate labels is determined based on the dynamic weighting parameters and the probability values output by the model. Cluster the first quality-labeled dataset to obtain the first labeled dataset.
5. The method according to claim 1, characterized in that, The step of selecting labeled datasets with data quality below a quality threshold from the first labeled dataset based on the data quality of the first labeled dataset for correction, to obtain a corrected second labeled dataset, includes: The first labeled dataset is evaluated in batches. An AI model that supports all labels is used to infer the data for each batch to obtain the model-discriminated labels and probability distributions for each batch of data, and to divide the data into a subset of data with consistent labels and a subset of data with inconsistent labels. Cluster the subset of labeled consistent data, extract a first confidence level sample and a second confidence level sample; and select a reverse bias sample from the subset of labeled inconsistent data, wherein the confidence level of the first confidence level sample is higher than the confidence level of the second confidence level sample. The first confidence level sampling sample, the second confidence level sampling sample, and the reverse bias sample are corrected to obtain the corrected second labeled dataset.
6. A data annotation device, characterized in that, include: The acquisition module is used to acquire the annotation task, the tag set corresponding to the annotation task, the full amount of data to be labeled, and multiple original AI models, wherein the annotation task is an audio emotion task, and the full amount of data to be labeled is audio emotion data; The determination module is used to perform an annotation personnel capability verification operation based on the annotation task, the tag set, the full amount of data to be annotated, and the first AI model, and determine the list of annotation personnel who meet the annotation conditions; wherein, the first AI model is an AI model among multiple original AI models that can cover the basic tag subset in the tag set; The annotation module is used to perform a weighted crowdsourcing operation using the first AI model based on the annotation task, the label set, and the second AI model to obtain a first labeled dataset; wherein the second AI model is an AI model that cannot cover the label set among the multiple original AI models; The correction module is used to select, based on the data quality of the first labeled dataset, labeled datasets whose data quality is less than a quality threshold from the first labeled dataset for correction, so as to obtain a corrected second labeled dataset; The determining module is used to determine, based on the label set, whether there exists an AI model among the multiple original AI models that can cover the label set; If there is no AI model among the multiple original AI models that can cover the set of labels, then the first AI model that can cover the basic subset of labels is selected from the multiple original AI models. The first AI model is used to infer the full set of unlabeled data to obtain a set of probability distribution vectors for each data point in the basic label space. Based on the set of probability distribution vectors, for newly added labels that are not covered, sample data of the newly added labels are collected and input into the first AI model to generate the probability output of the newly added label sample data in the basic label space, and to construct an initial feature vector representing the newly added label through cluster analysis to obtain the initial feature vector set of the newly added label. Based on the initial feature vector set of the newly added labels, the first data with the highest probability of model output being in the preset middle interval is selected from the full set of unlabeled data. The normalized cosine similarity between the probability vector of the first data and the initial feature vector of the newly added labels is calculated. When the normalized cosine similarity exceeds the set similarity threshold, the first data is marked as the candidate data of the corresponding newly added labels, thus obtaining the extended labeled dataset. The feature vector of each newly added label is recursively updated based on the extended labeled dataset to obtain the corrected set of feature vectors for the newly added labels. Based on the newly added label feature vector set, the consistency between the model output labels and the manual annotation results is compared, and the data is divided into correct groups or incorrect groups to obtain the data annotation evaluation subset; Based on the data annotation evaluation subset, candidate annotators are evaluated in stages to obtain a list of annotators who meet the annotation criteria.
7. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute instructions to implement the steps of the method as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-5.
9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Data annotation method, device, electronic device and storage medium
CN109242013A
Model training method and apparatus, emotion message generation method and apparatus, device and medium
WO2023159759A1