A collaborative data labeling method based on artificial intelligence pre-labeling
By using an AI-based multidimensional evaluation and optimized data processing workflow, the problem of balancing efficiency and quality caused by the single data annotation method in existing technologies is solved. This achieves rationality in data processing and synergistic optimization of the annotation process, thereby improving annotation efficiency and quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PUYANG MIQI COMM EQUIP CO LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-31
AI Technical Summary
Existing data annotation methods suffer from a lack of diversity in data processing and an unreasonable allocation of tasks, making it difficult to balance annotation efficiency and quality.
A collaborative data annotation method based on artificial intelligence pre-annotation is adopted. The data is quantitatively evaluated through multi-dimensional evaluation parameters, judgment rules are constructed and comprehensive evaluation weights are generated, data priorities are divided, and different annotation tasks are assigned according to the priorities. An annotation result verification and task scheduling mechanism are set up to achieve collaborative optimization of data processing and annotation process.
It improved data utilization efficiency, enhanced annotation quality and overall annotation efficiency, strengthened the stability and adaptability of the annotation system, and ensured the rationality and consistency of the annotation process.
Smart Images

Figure CN122490189A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of collaborative data processing technology, specifically relating to a collaborative data annotation method based on artificial intelligence pre-annotation. Background Technology
[0002] With the continuous development of artificial intelligence technology, data-driven model training methods have been widely applied in fields such as computer vision, natural language processing, and multimodal information processing. In these applications, large-scale, high-quality data annotation is a crucial foundation for ensuring model performance and generalization ability. Simultaneously, as application scenarios expand, data sources are becoming increasingly diverse, including both structured data and unstructured data such as images, text, audio, and video. Data types are showing a trend towards multimodal fusion, and the existing data scale is growing exponentially. The amount of data required to process in a single task is constantly expanding, and the complexity of annotation tasks is continuously increasing, placing higher demands on annotation accuracy, consistency, and processing efficiency. Moreover, different application scenarios have varying requirements for the precision and real-time nature of data annotation, requiring the data annotation process to not only be highly efficient but also possess good adaptability and scalability. This places more stringent demands on data processing workflows and annotation methods.
[0003] In existing technologies, common data annotation methods mainly include a combination of manual annotation and model-based pre-annotation. Basically, data is initially annotated by a pre-trained model, and then the annotation results are manually verified and corrected to improve annotation efficiency and reduce labor costs. In order to improve the collaborative efficiency of the annotation process, some technical solutions use task division and allocation mechanisms to assign different data to different annotators or processing flows. However, in practical applications, these existing methods are usually based on relatively simple evaluation methods or simple rules to process data, which makes it difficult to fully reflect the differences between data. This leads to a lack of rationality in data processing and task allocation, thereby affecting the overall annotation efficiency and annotation quality.
[0004] There is an urgent need for a data annotation method that can achieve more reasonable data evaluation and task allocation during data processing, in order to solve the problems of single data processing methods, unreasonable task allocation, and difficulty in balancing annotation efficiency and annotation quality in existing technologies. Summary of the Invention
[0005] In view of this, the present invention proposes a collaborative data annotation method based on artificial intelligence pre-annotation, which is applied to the field of collaborative data processing technology to solve the existing technical problems of single data processing methods, unreasonable task allocation, and difficulty in balancing annotation efficiency and annotation quality.
[0006] To achieve the above-mentioned technical objectives, the specific technical solution adopted by the present invention is as follows: A collaborative data annotation method based on artificial intelligence pre-annotation, characterized by the following steps: S1. Clean the raw data, unify its format and standardize its structure to form a set of data to be labeled; S2. Perform pre-labeling data screening on the dataset to be labeled, and then quantitatively evaluate the data to be labeled based on preset multidimensional evaluation parameters to obtain multidimensional evaluation results. Construct judgment rules based on the combination relationship between multidimensional evaluation parameters. When at least one evaluation parameter does not meet the preset conditions, perform screening or downgrading processing on the corresponding data. Construct a data value assessment model based on the multidimensional evaluation results, and assign corresponding weight coefficients to each evaluation parameter to generate comprehensive evaluation weights. Sort and screen the data to be labeled based on the comprehensive evaluation weights to generate pre-labeled candidate datasets. The construction of the data value assessment model and the generation of comprehensive evaluation weights are controlled by preset hierarchical judgment rules and constraints. S3. Based on the comprehensive evaluation results of the pre-labeled candidate dataset, the data is prioritized and divided into high-priority data, medium-priority data and low-priority data according to the preset weight threshold range. S4. Assign the data to be labeled to different labeling tasks according to the data priority. High priority data is assigned to manual fine labeling tasks, medium priority data is assigned to manual verification and model-assisted labeling tasks, and low priority data is assigned to automatic labeling tasks. Record and provide feedback on the execution results of each type of task. S5. Perform consistency verification on the annotation results. Determine the annotation deviation by calculating the difference between different annotation results. When the annotation deviation exceeds the preset threshold, the corresponding data is re-annotated or removed. S6. Classify the annotation tasks according to data type, annotation method and priority, and divide the annotation tasks into multiple task packages according to the classification results, so that the annotation tasks are executed in parallel according to the preset scheduling strategy, and the scheduling results are updated in real time.
[0007] Furthermore, in step S2, the multidimensional evaluation parameters are grouped according to their functions and their impact on data value assessment. Based on the grouping results, they participate in data screening, priority correction and sorting calculations. The evaluation parameters of different groups participate in the data value assessment process in a predetermined order.
[0008] Furthermore, in step S2, in the judgment rule, when at least one evaluation parameter used for data screening judgment does not meet the preset conditions, or when at least two evaluation parameters used for priority correction simultaneously meet the preset combination conditions, priority enhancement processing is performed on the corresponding data.
[0009] Furthermore, in step S2, during the ranking calculation process, corresponding weight coefficients are assigned to the evaluation parameters used for ranking calculation. The weight coefficients are adjusted based on historical annotation errors, model prediction deviations, or annotation consistency results. The weight adjustment is performed when a preset trigger condition is met and is limited to a preset adjustment range, so that the comprehensive evaluation weight changes under the constraints of the judgment rules.
[0010] Furthermore, in step S1, the data cleaning process includes identifying abnormal data in the original data based on preset anomaly judgment rules, and filtering noisy data in combination with a noise detection model. The data structure standardization process includes uniform field mapping and format conversion processing for data from different sources.
[0011] Furthermore, in step S3, the priority division is determined based on the combined judgment result of the comprehensive evaluation weight and at least one evaluation parameter, so that the division of data with different priorities does not depend on a single weight threshold.
[0012] Furthermore, in step S4, the task allocation process constructs allocation decision rules based on data priority, data complexity, and historical annotation errors, and dynamically adjusts the allocation of data to be labeled according to the allocation decision rules, so that the task allocation results are updated as data characteristics change.
[0013] Furthermore, in step S5, the consistency verification includes cross-comparison processing based on multiple annotation results and comparative analysis of historical annotation data to determine annotation deviations.
[0014] Furthermore, in step S6, the task package includes the data to be labeled, the pre-labeled results, and the corresponding evaluation information.
[0015] Furthermore, in step S6, the task packages are dynamically scheduled based on task priority, computing resource usage, and task execution status, and the scheduling strategy is optimized based on historical task execution results.
[0016] This invention transforms the data processing in the data annotation process from the traditional single evaluation method to a processing mechanism based on multi-dimensional information comprehensive judgment. By introducing multi-dimensional evaluation and constructing judgment rules before the data enters the annotation process, it achieves refined identification of data value and further links the data evaluation results with the subsequent task allocation process, so that the data processing and annotation execution process form an organic synergy, thereby improving the rationality of data processing and the efficiency of the annotation process at the overall process level.
[0017] By adopting the above technical solution, the present invention can also bring the following beneficial effects: 1. This invention proposes a collaborative data annotation method based on artificial intelligence pre-annotation. By introducing multi-dimensional evaluation parameters during data processing and constructing judgment rules based on the combination relationship between the evaluation parameters, the data can be more reasonably screened and evaluated before entering the subsequent annotation process. This can effectively distinguish the differences in processing value between different data, avoid interference from low-value or abnormal data in the annotation process, and has the advantages of improving data utilization efficiency and improving the rationality of data screening.
[0018] 2. This invention proposes a collaborative data annotation method based on artificial intelligence pre-annotation. By combining data evaluation results with data priority division and task allocation processes, different types of data can be matched with different processing precision and processing methods, thereby enabling annotation resources to be allocated more rationally. This reduces the situation of insufficient processing of high-value data or excessive processing of low-value data, and has the advantages of improving overall annotation efficiency and quality and enhancing the utilization efficiency of annotation resources.
[0019] 3. This invention proposes a collaborative data annotation method based on artificial intelligence pre-annotation. By setting up annotation result verification and task scheduling mechanisms, the annotation process can be dynamically adjusted and optimized during execution, thereby improving the consistency and stability of annotation results and enhancing the collaborative and adaptability of the annotation process. This enables the system to operate stably under different data scales and application scenarios, and has the advantages of improving the stability and adaptability of the annotation system. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating a collaborative data annotation method based on artificial intelligence pre-annotation mentioned in this invention; Figure 2 This is a flowchart illustrating the process of performing data value assessment and generating candidate data in step S2 of this embodiment; Detailed Implementation
[0022] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0023] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0024] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this invention, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.
[0025] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. The drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0026] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details. Example
[0027] like Figure 1 and Figure 2 As shown, a collaborative data annotation method based on artificial intelligence pre-annotation includes the following steps: S1. Perform data preprocessing and construct the dataset; S101. Perform data cleaning on the raw data, identify abnormal data in the raw data based on the preset anomaly judgment rules. The anomaly judgment rules are constructed based on the statistical distribution of historical data, and the anomaly judgment threshold is set to the mean ±2σ range. At the same time, combine the noise detection model to filter the noise data. The noise detection model adopts the isolated forest model or the detection model based on the convolutional neural network. S102. Perform unified processing on the data format and standardized processing on the data structure, including unified field mapping and format conversion processing on data from different sources. The field mapping rules are matched based on the preset data template. S103. Perform preliminary classification on the processed data according to the preset classification rules, and form a set of data to be labeled.
[0028] S2. Perform data value assessment and generate candidate data; S201. Perform pre-labeling data filtering on the dataset to be labeled. The pre-labeling data filtering process includes generating initial labeling results based on the pre-trained model and performing filtering based on the initial labeling results. At the same time, feature extraction is performed on the dataset to be labeled to construct basic feature representations. S202. Quantitatively evaluate the labeled data based on preset multidimensional evaluation parameters to obtain multidimensional evaluation results. The multidimensional evaluation parameters include at least data distribution balance, sample confidence, feature complexity, and sample similarity. The sample confidence is calculated by outputting probability values from the pre-trained model, with a value range of 0 to 1. The feature complexity is quantified by the number of feature dimensions and feature distribution entropy. The sample similarity is calculated by cosine similarity or Euclidean distance. S203. Construct judgment rules based on the combination relationship between multidimensional evaluation parameters, and group the multidimensional evaluation parameters so that different grouped parameters participate in data screening judgment, priority correction and sorting calculation respectively. The different grouped parameters participate in the processing in a preset order. The hierarchical judgment rule includes processing in the order of data screening judgment layer, priority correction layer and sorting calculation layer. S204. In the judgment rule, when at least one evaluation parameter used for data screening judgment does not meet the preset conditions, the corresponding data is subjected to screening or downgrading processing, wherein the preset conditions include the sample confidence level being lower than 0.6 or the data distribution balance deviating from the preset distribution range; when at least two evaluation parameters used for priority correction simultaneously meet the preset combination conditions, the corresponding data is subjected to priority upgrading processing. S205. Construct a data value assessment model based on the multidimensional evaluation results, and assign corresponding weight coefficients to the evaluation parameters used for ranking calculation. The initial value of the weight coefficient is set in the range of 0.2 to 0.4, and is dynamically adjusted based on historical annotation errors, model prediction deviations, or annotation consistency results. When the annotation error of consecutive batches exceeds 5%, the weight adjustment mechanism is triggered, and the adjustment range is limited to ±10%. The constraints include weight change range constraints and priority adjustment number limits to ensure the stability of the evaluation results. S206. Based on the comprehensive evaluation weight, sort and filter the data to be labeled to generate a pre-labeled candidate dataset. The construction of the data value assessment model and the generation of the comprehensive evaluation weight are controlled by the preset hierarchical judgment rules and constraints.
[0029] S3. Prioritize Data Data priority is determined based on the comprehensive evaluation results of the pre-labeled candidate dataset. Data is divided into high-priority, medium-priority, and low-priority data according to a preset weight threshold range. The priority division is determined by a combination of comprehensive evaluation weights and at least one evaluation parameter, so that the division of different priority data does not depend on a single weight threshold. The proportion of high-priority data is 10% to 30%, medium-priority data is 40% to 60%, and low-priority data is 10% to 30%.
[0030] S4. Assign annotation tasks; Data to be labeled is allocated to different labeling tasks based on data priority. High-priority data is assigned to manual fine labeling tasks, medium-priority data is assigned to manual verification and model-assisted labeling tasks, and low-priority data is assigned to automatic labeling tasks. The task allocation process constructs allocation decision rules based on data priority, data complexity, and historical labeling errors. Data complexity is calculated through feature dimensions and feature distribution, and historical labeling errors are obtained based on statistics of historical labeling results. The allocation of data to be labeled is dynamically adjusted according to the allocation decision rules, so that the task allocation results are updated as data characteristics change. At the same time, the execution results of various tasks are recorded and feedback is provided.
[0031] S5. Perform consistency verification of annotation results; The consistency of the annotation results is checked by calculating the difference between different annotation results to determine the annotation deviation. The consistency check includes cross-comparison processing based on multiple annotation results and comparative analysis of historical annotation data. The difference value is obtained by calculating the label matching degree or edit distance. When the annotation deviation exceeds the preset threshold (preferably 10%), the corresponding data is re-annotated or removed.
[0032] S6. Perform task scheduling and optimize the execution process; S601. Classify annotation tasks according to data type, annotation method and priority, and divide the annotation tasks into multiple task packages according to the classification results; S602. Construct a task package structure. The task package contains data to be labeled, pre-labeled results, and corresponding evaluation information. The size of the task package is preferably 100 to 500 data points. S603. Based on task priority, computing resource usage, and task execution status, the task package is dynamically scheduled so that the marked tasks are executed in parallel according to the preset scheduling strategy, and the scheduling results are updated in real time. When the system resource usage exceeds 80%, the scheduling frequency of low-priority tasks is reduced, and the scheduling strategy is optimized based on the historical task execution results.
[0033] In summary, this invention has the advantages of enabling data filtering, fine priority division, high matching degree of annotation task allocation, improved annotation efficiency and high annotation consistency.
[0034] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A collaborative data labeling method based on artificial intelligence pre-labeling, characterized in that, Includes the following steps: S1. Clean the raw data, unify its format and standardize its structure to form a set of data to be labeled; S2. Perform pre-labeling data screening on the dataset to be labeled, then perform quantitative evaluation on the dataset to be labeled based on preset multidimensional evaluation parameters to obtain multidimensional evaluation results, construct judgment rules based on the combination relationship between multidimensional evaluation parameters, and perform screening or downgrading processing on the corresponding data when at least one evaluation parameter does not meet the preset conditions. Construct a data value assessment model based on the multidimensional evaluation results, and assign corresponding weight coefficients to each evaluation parameter to generate a comprehensive evaluation weight. Sort and screen the dataset to be labeled based on the comprehensive evaluation weight to generate a pre-labeled candidate dataset. The construction of the data value assessment model and the generation of the comprehensive evaluation weight are controlled by preset hierarchical judgment rules and constraints. S3. Based on the comprehensive evaluation results of the pre-labeled candidate dataset, the data is prioritized and divided into high-priority data, medium-priority data and low-priority data according to the preset weight threshold range. S4. Assign the data to be labeled to different labeling tasks according to the data priority. High priority data is assigned to manual fine labeling tasks, medium priority data is assigned to manual verification and model-assisted labeling tasks, and low priority data is assigned to automatic labeling tasks. Record and provide feedback on the execution results of each type of task. S5. Perform consistency verification on the annotation results. Determine the annotation deviation by calculating the difference between different annotation results. When the annotation deviation exceeds the preset threshold, the corresponding data is re-annotated or removed. S6. Classify the annotation tasks according to data type, annotation method and priority, and divide the annotation tasks into multiple task packages according to the classification results, so that the annotation tasks are executed in parallel according to the preset scheduling strategy, and the scheduling results are updated in real time.
2. The collaborative data labeling method based on artificial intelligence pre-labeling according to claim 1, characterized in that: In step S2, the multidimensional evaluation parameters are grouped according to their functions and their impact on data value assessment. Based on the grouping results, they are used in data screening, priority correction and sorting calculation. The evaluation parameters of different groups are processed in a preset order during the data value assessment process.
3. The collaborative data annotation method based on artificial intelligence pre-annotation as described in claim 2, characterized in that: In step S2, in the determination rule, when at least one evaluation parameter used for data filtering does not meet the preset conditions, or when at least two evaluation parameters used for priority correction simultaneously meet the preset combination conditions, priority enhancement processing is performed on the corresponding data.
4. The collaborative data annotation method based on artificial intelligence pre-annotation as described in claim 2, characterized in that: In step S2, during the ranking calculation process, corresponding weight coefficients are assigned to the evaluation parameters used for ranking calculation. The weight coefficients are adjusted based on historical annotation errors, model prediction deviations, or annotation consistency results. The weight adjustment is performed when a preset trigger condition is met and is limited to a preset adjustment range so that the comprehensive evaluation weight changes under the constraints of the judgment rules.
5. The collaborative data annotation method based on artificial intelligence pre-annotation as described in claim 1, characterized in that: In step S1, the data cleaning process includes identifying abnormal data in the original data based on preset anomaly judgment rules and filtering noisy data in combination with a noise detection model. The data structure standardization process includes uniform field mapping and format conversion processing for data from different sources.
6. The collaborative data annotation method based on artificial intelligence pre-annotation as described in claim 1, characterized in that: In step S3, the priority division is determined based on the combined judgment result of comprehensive evaluation weight and at least one evaluation parameter, so that the division of data with different priorities does not depend on a single weight threshold.
7. The collaborative data annotation method based on artificial intelligence pre-annotation as described in claim 1, characterized in that: In step S4, the task allocation process constructs allocation decision rules based on data priority, data complexity, and historical annotation errors, and dynamically adjusts the allocation of data to be labeled according to the allocation decision rules, so that the task allocation results are updated as data characteristics change.
8. The collaborative data annotation method based on artificial intelligence pre-annotation as described in claim 1, characterized in that: In step S5, the consistency verification includes cross-comparison processing based on multiple annotation results and comparative analysis of historical annotation data to determine annotation deviations.
9. The collaborative data annotation method based on artificial intelligence pre-annotation as described in claim 1, characterized in that: In step S6, the task package includes data to be labeled, pre-labeling results, and corresponding evaluation information.
10. The collaborative data annotation method based on artificial intelligence pre-annotation as described in claim 1, characterized in that: In step S6, the task packages are dynamically scheduled based on task priority, computing resource usage, and task execution status, and the scheduling strategy is optimized based on historical task execution results.