Data labeling method and apparatus, storage medium, and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TRAVELSKY TECHNOLOGY LIMITED
- Filing Date
- 2026-03-25
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies have low accuracy in data sample labeling results, manual labeling is inefficient and costly, semi-automatic labeling relies on the quality of pre-trained models and is prone to introducing biases, and system deployment and maintenance are complex.
A multimodal large model is used to label the data samples, and the final labeling results are determined by comparing the results and the confidence of the machine learning model. Data sets with high consistency, divergent sample sets, and difficult sample sets are selected, and the final confirmation is carried out by hard voting mechanism and manual labeling.
It improves the accuracy and efficiency of data sample labeling, reduces labor costs, decreases system complexity, and ensures the reliability and interpretability of labeling results.
Smart Images

Figure CN122286302A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and more specifically, to a data annotation method, apparatus, storage medium, and program product. Background Technology
[0002] With the continuous advancement of artificial intelligence technology, high-quality datasets (such as text, image, audio, and video datasets) have become a key foundation for its development, and data annotation, as a crucial step in constructing these datasets, plays an indispensable role. Data annotation closely connects data resources, algorithm models, and practical application scenarios, and is the core driving force for creating high-quality artificial intelligence datasets.
[0003] In related technologies, data annotation methods can be divided into two categories, as follows:
[0004] (1) Manual annotation relies entirely on manual operation, but it is relatively inefficient, has high labor costs, and is difficult to scale to massive data scenarios. Especially in AI (Artificial Intelligence) applications that require rapid iteration, manual annotation often becomes a bottleneck and cannot meet the needs of real-time or large-scale data processing. In addition, the consistency and repeatability of manual annotation are also limited by the subjective judgment of the annotator, which may introduce human bias.
[0005] (2) Semi-automatic annotation relies mainly on manual work, supplemented by AI technology. However, its performance is highly dependent on the quality of the pre-trained model. If the pre-trained model performs poorly in a specific domain, the workload of manual correction is still very large, and it may even introduce the inherent systematic bias of the model, affecting the purity of the final dataset. At the same time, the human-computer interaction in the semi-automatic annotation process often requires complex engineering integration, which increases the complexity of system deployment and maintenance.
[0006] There is currently no effective solution to the above problems. Summary of the Invention
[0007] This invention provides a data annotation method, apparatus, storage medium, and program product to at least solve the technical problem of low accuracy in annotation results for data samples in related technologies.
[0008] According to one aspect of the present invention, a data annotation method is provided, comprising: annotating data samples in a dataset to be annotated using n first models to obtain n initial annotated datasets, wherein the datasets to be annotated include s data samples to be annotated, the model type of the first models includes: multimodal large model, and n and s are positive integers; performing pairwise comparisons on the n initial annotated datasets to obtain comparison results; obtaining the confidence scores of the second model in annotating the data samples in the datasets to be annotated to obtain s confidence scores, wherein the second model includes: machine learning model; and determining a target annotated dataset based on the comparison results and the s confidence scores, wherein the target annotated dataset includes: the final annotation result of each data sample in the datasets to be annotated.
[0009] Further, based on the comparison results and s confidence levels, determining the target labeled dataset includes: based on the comparison results, selecting labeled results with consistent labeling results for the same data sample from n initial labeled datasets to obtain a first labeled dataset; based on the comparison results, selecting data samples with inconsistent labeling results for the same data sample from n initial labeled datasets to obtain a divergent sample set; based on the s confidence levels, determining the labeling results of the data samples in the divergent sample set to obtain a second labeled dataset; and determining the target labeled dataset based on the first labeled dataset and the second labeled dataset.
[0010] Further, based on s confidence levels, the annotation results of the data samples in the divergent sample set are determined to obtain a second labeled dataset, including: based on the s confidence levels, k data samples are selected from the dataset to be labeled to obtain a sample set to be determined, wherein the confidence level of the data samples in the sample set to be determined is lower than the confidence level of the other data samples in the dataset to be labeled except for the k data samples, and k is a positive integer less than s; based on the sample set to be determined, the annotation results of the data samples in the divergent sample set are determined to obtain a second labeled dataset.
[0011] Further, based on the sample set to be determined, the annotation results of the data samples in the divergent sample set are determined to obtain a second labeled dataset, including: calculating the intersection of the divergent sample set and the sample set to be determined to obtain a difficult sample set; calculating the difference between the divergent sample set and the difficult sample set to obtain a target sample set; and determining the annotation results of the data samples in the difficult sample set and the target sample set respectively to obtain the second labeled dataset.
[0012] Further, determining the annotation results of the data samples in the difficult sample set and the target sample set respectively to obtain the second annotation dataset includes: determining the annotation results of all data samples in the target sample set based on the annotation results of each data sample in the n initial annotation datasets to obtain a first annotation result set; sending the difficult sample set to the target object and receiving all annotation results returned by the target object to obtain a second annotation result set, wherein the target object is used to annotate the data samples in the difficult sample set; and determining the second annotation dataset based on the first annotation result set and the second annotation result set.
[0013] Further, based on the annotation results of each data sample in the n initial annotation datasets, the annotation results of all data samples in the target sample set are determined to obtain a first annotation result set, including: for each data sample in the target sample set, the annotation result with the highest number of identical annotation results is selected from the n initial annotation datasets to obtain the target annotation result of the data sample; based on the target annotation result of each data sample in the target sample set, the first annotation result is determined.
[0014] Furthermore, the data samples in the dataset to be labeled include at least one of the following types: text, image, audio, and video.
[0015] According to another aspect of the present invention, a data annotation apparatus is also provided, comprising: an annotation unit, configured to annotate data samples in a dataset to be annotated using n first models respectively, to obtain n initial annotated datasets, wherein the dataset to be annotated includes s data samples to be annotated, the model type of the first models includes: multimodal large model, and n and s are positive integers; a comparison unit, configured to perform pairwise comparisons on the n initial annotated datasets to obtain comparison results; an acquisition unit, configured to acquire the confidence scores of the second model annotating the data samples in the dataset to be annotated, to obtain s confidence scores, wherein the second model includes: machine learning model; and a determination unit, configured to determine a target annotated dataset based on the comparison results and the s confidence scores, wherein the target annotated dataset includes: the final annotation result of each data sample in the dataset to be annotated.
[0016] Further, the determining unit includes: a first filtering subunit, used to filter out, based on the comparison results, annotation results that are consistent in their annotation results for the same data sample from n initial annotation datasets to obtain a first annotation dataset; a second filtering subunit, used to filter out, based on the comparison results, data samples with inconsistent annotation results for the same data sample from n initial annotation datasets to be annotated to obtain a divergence sample set; a first determining subunit, used to determine the annotation results of the data samples in the divergence sample set based on s confidence levels to obtain a second annotation dataset; and a second determining subunit, used to determine the target annotation dataset based on the first annotation dataset and the second annotation dataset.
[0017] Further, the first determining subunit includes: a filtering module, used to filter k data samples from the dataset to be labeled based on s confidence levels to obtain a sample set to be determined, wherein the confidence level of the data samples in the sample set to be determined is lower than the confidence level of other data samples in the dataset to be labeled except for the k data samples, and k is a positive integer less than s; and a determining module, used to determine the labeling results of the data samples in the divergent sample set based on the sample set to be determined to obtain a second labeled dataset.
[0018] Further, the determining module includes: a first calculation submodule, used to calculate the intersection of the divergent sample set and the sample set to be determined, to obtain a difficult sample set; a second calculation submodule, used to calculate the difference between the divergent sample set and the difficult sample set, to obtain a target sample set; and a determining submodule, used to determine the annotation results of the data samples in the difficult sample set and the target sample set respectively, to obtain the second labeled dataset.
[0019] Further, the determining submodule includes: a first determining submodule, which determines the labeling results of all data samples in the target sample set based on the labeling results of each data sample in the n initial labeled datasets, to obtain a first labeled result set; a processing submodule, which sends the difficult sample set to the target object and receives all the labeling results returned by the target object, to obtain a second labeled result set, wherein the target object is used to label the data samples in the difficult sample set; and a second determining submodule, which determines the second labeled dataset based on the first labeled result set and the second labeled result set.
[0020] Furthermore, the first determining submodule also includes: a filtering submodule one, which, for each data sample in the target sample set, filters the labeling result with the highest number of identical labeling results from the n initial labeled datasets to obtain the target labeling result of the data sample; and a determining submodule one, which determines the first labeling result based on the target labeling result of each data sample in the target sample set.
[0021] Furthermore, the data samples in the dataset to be labeled include at least one of the following types: text, image, audio, and video.
[0022] According to another aspect of the present invention, an electronic device is also provided, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the data annotation method of any of the above-mentioned methods by executing the executable instructions.
[0023] According to another aspect of the present invention, a computer-readable storage medium is also provided, which stores a computer program, wherein the computer program controls the device where the computer-readable storage medium is located to execute the data annotation method described above when it is running.
[0024] In this invention, the following method is adopted: n first models are used to annotate data samples in the dataset to be labeled, resulting in n initial labeled datasets. Each dataset includes s data samples to be labeled. The model types of the first models include multimodal large models. The n initial labeled datasets are compared pairwise to obtain comparison results. The confidence scores of the second model in labeling the data samples in the dataset to be labeled are obtained, resulting in s confidence scores. The second model includes machine learning models. Based on the comparison results and the s confidence scores, the target labeled dataset is determined. The target labeled dataset includes the final labeled result for each data sample in the dataset to be labeled. This solves the technical problem of low accuracy in labeling results of data samples in related technologies. In this invention, multiple multimodal large models are used to label data samples, and the final labeled result of the data samples is determined based on the differences in the labeling results and the confidence scores predicted by the machine learning models. This avoids the low accuracy of manual or semi-automatic sample labeling in related technologies, thereby achieving the technical effect of improving the accuracy of data sample labeling. Attached Figure Description
[0025] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0026] Figure 1This is a flowchart of an optional data annotation method according to an embodiment of the present invention;
[0027] Figure 2 This is a schematic diagram of an optional data annotation system according to an embodiment of the present invention;
[0028] Figure 3 This is a flowchart of an optional data annotation method according to an embodiment of the present invention;
[0029] Figure 4 This is a schematic diagram of an optional data annotation device according to an embodiment of the present invention;
[0030] Figure 5 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] For ease of description, some terms or nouns involved in the various embodiments of the present invention will be explained below.
[0034] (1) Data set to be labeled: refers to the original data set that has not been labeled or processed. This dataset may contain one or more modalities of data, such as text, images, audio, and video. It is the starting point for all labeling processes in this method.
[0035] (2) Pre-labeled dataset: also known as the first labeled dataset, refers to the intermediate dataset automatically formed during the parallel labeling and divergence discovery phase, which consists of multiple multimodal large models that have completely consistent labeling results for the same data sample. This dataset represents the highly reliable automatic labeling results that have reached a high degree of consensus within the system, and is a fundamental component of the final labeled dataset.
[0036] (3) Divergence dataset: also known as the divergence sample set, refers to the set of data samples that are selected during the parallel annotation and divergence discovery phase by cross-comparing the annotation results of various multimodal large models, and that show inconsistencies in the annotation results of at least one model. This dataset reveals samples where there are ambiguities or difficulties in understanding different large models, and is one of the key bases for locating annotation difficulties.
[0037] (4) Uncertainty dataset: This refers to the set of samples to be determined, which is the set of k samples with the lowest confidence (k is the number of samples in the divergent dataset) selected by machine learning or interpretable deep learning models based on their own prediction confidence (such as probability entropy, margin value, etc.) during the uncertainty quantification stage. This dataset reflects the samples that the machine learning model judges with ambiguity and insufficient confidence, and is a key basis for identifying labeling difficulties from another dimension.
[0038] (5) Divergence Hard Voting Dataset: This is the target sample set, which refers to the subset of data obtained by taking the difference between the divergence dataset and the difficult sample dataset during the difficult sample localization stage. Specifically, this dataset contains samples that, although there is divergence among the multimodal large models, have not been classified as low-confidence by the machine learning model (i.e., do not belong to the uncertain dataset). For the samples in this dataset, the system will automatically determine their final labels using a hard voting mechanism (such as based on the output results of the majority of large models and the machine learning model / deep learning model).
[0039] (6) Difficult Sample Dataset: The difficult sample set refers to the set of data samples that are finally determined by calculating the intersection of the "disagreement dataset" and the "uncertainty dataset". This dataset meets the dual conditions of "inconsistent conclusions of multiple models" and "low confidence of a single model", and is judged by the system as the most challenging labeled samples. It needs to be submitted to the manual labeling module for final fine labeling and review.
[0040] It should be noted that the user information (including but not limited to user device information, user personal information, etc.), the collected information and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws and standards of the relevant regions, have taken necessary measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse.
[0041] Example 1
[0042] According to an embodiment of the present invention, an alternative data annotation method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0043] Figure 1 This is a flowchart of an optional data annotation method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:
[0044] Step S101: Label the data samples in the dataset to be labeled using n first models respectively, to obtain n initial labeled datasets. The dataset to be labeled includes s data samples to be labeled. The model type of the first model includes: multimodal large model, where n and s are positive integers.
[0045] The data samples in the aforementioned dataset to be labeled can be of various types, including but not limited to: text, images, audio, and video. The model type of the first module can be a multimodal large model, which can be a model with completely different architectures, or multiple variants of the same base model fine-tuned with data from different domains. In resource-constrained scenarios, multiple efficient miniaturized models (such as distillation models) or combinations of expert models optimized for specific modalities can be used to replace the large general-purpose model with a large number of parameters, in order to further reduce costs while ensuring certain performance. In one optional example, n first models can work in parallel to improve sample labeling efficiency.
[0046] In this embodiment, n first models can be used to annotate each data sample in the dataset to be labeled. Each first model annotates all data samples in the dataset to be labeled to obtain an initial labeled dataset. The n first models annotate each data sample in the dataset to be labeled to obtain n initial labeled datasets.
[0047] Step S102: Perform pairwise comparisons on the n initially labeled datasets to obtain the comparison results.
[0048] In this embodiment, pairwise comparisons can be performed on n initial labeled datasets to analyze whether the labeling results of the n first models on the same data sample are the same, and the comparison results can be obtained.
[0049] In one alternative example, in addition to a simple direct comparison of labeled results, a difference measure based on the internal state of the first model (such as attention distribution and feature similarity) can be introduced as an auxiliary judgment to discover potential discrepancies earlier or in more detail.
[0050] Step S103: Obtain the confidence scores of the second model in labeling the data samples in the dataset to be labeled, and obtain s confidence scores. The second model includes a machine learning model.
[0051] The model type of the second module mentioned above can be a machine learning model (or an interpretable deep learning model). In this embodiment, the second model can classify the dataset to be labeled for labeling and calculate the confidence level of the classification prediction result for each sample.
[0052] In one alternative example, other uncertainty estimation methods may be employed, such as Monte Carlo methods, prediction variance generated by deep ensemble models, or probability entropy based on model calibration. For interpretable deep learning models, their interpretability outputs (such as feature importance and decision rules) can themselves serve as a measure of uncertainty or decision credibility.
[0053] The second model can also be a small set of models whose output can be the joint prediction and combined confidence level of multiple models.
[0054] Step S104: Based on the comparison results and s confidence levels, determine the target annotation dataset, wherein the target annotation dataset includes the final annotation result of each data sample in the dataset to be annotated.
[0055] In this embodiment, based on the comparison results, n first model annotation results that are consistent in annotating the same data sample can be selected to obtain the first annotation dataset. The first annotation dataset can represent the automatic annotation results that have reached a high degree of consensus and have high credibility, and can constitute the basic component of the final annotation dataset.
[0056] In this embodiment, based on the comparison results, n data samples with inconsistent annotation results for the same data sample by the first model can be selected from the dataset to be labeled, thus obtaining a divergent sample set.
[0057] For the data samples in the divergence sample set, they can be labeled based on s confidence levels to obtain a second labeled dataset. The first labeled dataset and the second labeled dataset can be combined to form the target labeled dataset.
[0058] Through the above steps, in this embodiment, data samples are labeled using multiple multimodal large models. The final labeling result of the data samples is determined based on the differences in the labeling results and the confidence level of the machine learning model predictions. This avoids the low accuracy issues associated with manual or semi-automatic sample labeling in related technologies, thereby improving the technical effect of data sample labeling accuracy. This solves the technical problem of low labeling accuracy in related technologies.
[0059] Optionally, based on the comparison results and s confidence levels, the target labeled dataset is determined, including: based on the comparison results, selecting labeled results that are consistent for the same data sample from n initial labeled datasets to obtain a first labeled dataset; based on the comparison results, selecting data samples with inconsistent labeled results for the same data sample from n initial labeled datasets to obtain a divergent sample set; based on s confidence levels, determining the labeled results of data samples in the divergent sample set to obtain a second labeled dataset; and based on the first labeled dataset and the second labeled dataset, the target labeled dataset is determined.
[0060] In this embodiment, based on the comparison results, n first model annotation results that are consistent in annotating the same data sample can be selected to obtain the first annotation dataset. The first annotation dataset can represent the automatic annotation results that have reached a high degree of consensus and have high credibility, and can constitute the basic component of the final annotation dataset.
[0061] In this embodiment, based on the comparison results, n data samples with inconsistent annotation results for the same data sample by the first model can be selected from the dataset to be labeled, thus obtaining a divergent sample set.
[0062] For a data sample in the divergent sample set, the annotation method of the data sample can be selected according to the confidence level of the data sample association, and the data sample can be annotated according to the selected annotation method to obtain the second annotated dataset. The annotation method may include, but is not limited to: manual annotation, and voting annotation based on the annotation results of n first models on the data sample.
[0063] The target annotation dataset can be composed of the first annotation dataset and the second annotation dataset.
[0064] In this embodiment, based on the annotation results of the same data sample by n first models and the confidence of the classification annotation, the final annotation result of the data sample is determined, avoiding the situation of low efficiency in pure manual annotation and low efficiency and low accuracy in semi-automatic annotation in the related art, thus achieving the technical effect of improving the efficiency and accuracy of sample annotation.
[0065] Optionally, based on s confidences, determine the annotation results of the data samples in the disagreement sample set to obtain a second annotation data set, including: based on s confidences, screen out k data samples from the data set to be annotated to obtain a sample set to be determined, where the confidences of the data samples in the sample set to be determined are lower than the confidences of the other data samples in the data set to be annotated except for the k data samples, and k is a positive integer less than s; based on the sample set to be determined, determine the annotation results of the data samples in the disagreement sample set to obtain a second annotation data set.
[0066] In this embodiment, based on s confidences, select the k samples with the lowest confidences (k < s) from the data set to be annotated to form a sample set to be determined (i.e., an uncertainty data set). The sample set to be determined can refer to the samples whose labels cannot be determined by the large model, indicating the cognitive fuzzy area of the large model.
[0067] Based on the sample set to be determined, the disagreement sample set can be split into a difficult sample set and a target sample set (i.e., a differential hard voting data set). The samples in the difficult sample set can indicate the samples where the annotation results of the n first models are in disagreement and are judged to have low confidence. Such data samples can be annotated by the target object to obtain the final labels of such data samples. For example, through manual annotation. The data samples in the target sample set can indicate that the annotation results of the n first models are in disagreement but are judged to have high confidence. For the data samples in the target sample set, the hard voting mechanism can be used to automatically determine their final labels. In this way, the annotation results of the data samples in the disagreement sample set can be determined to obtain a second annotation data set.
[0068] In an optional example, when screening the sample set to be determined, the quantity k may not be fixed as the number of samples in the disagreement data set, and k can be dynamically adjusted according to task requirements, manual annotation resources or historical annotation quality feedback (for example, set as a percentage of the total number of samples, or determined based on the threshold of the confidence distribution).
[0069] In this embodiment, for the data samples with annotation disagreements in the multi-modal large model, the annotation method is selected according to the level of confidence and the samples are annotated, achieving the technical effect of improving the accuracy of data annotation.
[0070] Optionally, based on the sample set to be determined, the annotation results of the data samples in the divergent sample set are determined to obtain a second annotation dataset, including: calculating the intersection of the divergent sample set and the sample set to be determined to obtain a difficult sample set; calculating the difference between the divergent sample set and the difficult sample set to obtain a target sample set; and determining the annotation results of the data samples in the difficult sample set and the target sample set respectively to obtain a second annotation dataset.
[0071] In this embodiment, the intersection of the "disagreement sample set" and the "dataset to be determined" can be calculated to obtain the difficult sample set. This dataset simultaneously meets the dual conditions of "inconsistent labeling results across multiple models" and "low confidence of a single machine learning model," and is therefore identified as the most challenging labeled samples. It can be submitted to the target object for final labeling and review. The difference between the "disagreement sample set" and the "difficult sample set" can also be calculated to obtain the target sample set. Specifically, this dataset contains samples that, although there is disagreement among the large multimodal models, have not been classified as low-confidence by the machine learning models. For samples in this dataset, a hard voting mechanism (such as based on the output results of the majority of large models and the machine learning / deep learning models) can be used to automatically determine their final labels.
[0072] In this embodiment, the manually labeled "difficult samples" and their correct labels from each round can also be added to the training set to update (fine-tune) the second model (e.g., machine learning / interpretable deep learning), or even some large models, thereby gradually improving the overall performance of data labeling and reducing the need for manual intervention in subsequent rounds.
[0073] By determining the labels based on the annotation results of the target object and the hard voting mechanism, a second labeled dataset can be obtained, which achieves the technical effect of improving the accuracy of data sample annotation.
[0074] Optionally, the annotation results of the data samples in the difficult sample set and the target sample set are determined separately to obtain a second annotation dataset, including: determining the annotation results of all data samples in the target sample set based on the annotation results of each data sample in the n initial annotation datasets to obtain a first annotation result set; sending the difficult sample set to the target object and receiving all annotation results returned by the target object to obtain a second annotation result set, wherein the target object is used to annotate the data samples in the difficult sample set; and determining the second annotation dataset based on the first annotation result set and the second annotation result set.
[0075] In this embodiment, from the annotation results of each data sample in n initial labeled datasets, all annotation results of n first models for each data sample in the target sample set can be selected, and their final labels can be automatically determined by hard voting (e.g., the annotation result with more identical annotation results of n first models for a data sample is taken as the final annotation result of the data sample) or weighted fusion mechanism to form a first annotation result set. The first annotation result set represents a high-quality automatically labeled subset that the model has sufficient confidence in and does not require human intervention.
[0076] For data samples in a difficult sample set, they can be sent to the target audience (e.g., professional annotators or domain experts) for precise annotation based on business knowledge and context, resulting in a second annotation result set. Integrating the first annotation result set (automatically output by the machine) with the second annotation result set (precisely annotated by the target audience) yields a second annotated dataset, achieving an optimal collaboration model of "machine-driven with human support." This significantly reduces the human burden, avoids the cost of "full human involvement," and enhances the overall reliability and interpretability of the dataset through a dual verification mechanism.
[0077] Optionally, based on the annotation results of each data sample in the n initial annotation datasets, the annotation results of all data samples in the target sample set are determined to obtain the first annotation result set, including: for each data sample in the target sample set, the annotation result with the highest number of identical annotation results is selected from the n initial annotation datasets to obtain the target annotation result of the data sample; based on the target annotation result of each data sample in the target sample set, the first annotation result is determined.
[0078] In this embodiment, a "majority consensus / voting" mechanism can be used to identify the optimal label for the target sample set from the initial annotation results of n large models, thus obtaining the first annotation result set. For example, for each data sample in the target sample set, the n annotation results corresponding to that data sample in the n initial annotation datasets can be counted, and the annotation result (i.e., label) that appears most frequently can be used as the "target annotation result" for that data sample. If multiple labels have the same highest frequency, weighted voting can be further combined with confidence scores or model weights to ensure decision robustness. Instead of using the annotation results of multiple first models to determine the final annotation result for each data sample in the target sample set, the first annotation result can be obtained, which can eliminate the annotation bias of individual models and improve the reliability of automatic annotation.
[0079] Optionally, the data samples in the dataset to be labeled may include at least one of the following types: text, image, audio, and video.
[0080] In this embodiment, the data samples in the dataset to be labeled may be of types including but not limited to: text, images, audio, and video, and may be text, images, audio, and video processed by the civil aviation business system.
[0081] In this embodiment, by organically combining multiple multimodal large-scale model collaboration mechanisms with an uncertainty-based sample selection method, synergistic optimization of multimodal data annotation in terms of quality, efficiency, and cost is achieved. For example, a dual quality assurance mechanism is constructed through deep semantic understanding of multimodal large-scale models and interpretability verification of machine learning models. The uncertainty-based sample selection can accurately identify marginal cases and ambiguous samples, effectively avoiding the propagation of systematic annotation errors in the dataset.
[0082] Example 2
[0083] Embodiment 2 of the present invention provides an optional data annotation system, which can be used to execute the data annotation method provided in Embodiment 1 of the present invention. Figure 2 This is a schematic diagram of an optional data annotation system according to an embodiment of the present invention, such as... Figure 2 As shown, the data annotation system includes: a multimodal large model module, a machine learning / interpretable deep learning module, and a manual annotation module, including:
[0084] (1) Multimodal Large Model Module: This module is used for preliminary and efficient intelligent annotation of the data to be labeled. It integrates multiple heterogeneous or different training objectives of multimodal large models (corresponding to the first model in Example 1), and can simultaneously process cross-modal data such as text, images, audio, and video to achieve a deep understanding of complex semantics and contextual information. During the annotation process, multiple multimodal large models work in parallel to generate preliminary annotation results (corresponding to the initial annotation dataset in Example 1), and automatically identify annotation discrepancies through result comparison to form a "discrepancy dataset (corresponding to the discrepancy sample set in Example 1)". The core function of this module is to utilize the powerful perception and reasoning capabilities of large models to cover most samples that can be clearly classified, providing a high-quality pre-annotation foundation for subsequent steps.
[0085] (2) Machine Learning / Interpretable Deep Learning Module: Provides interpretable and efficient auxiliary annotation and uncertainty assessment. This module typically uses machine learning models with relatively simple structures and traceable decision-making processes (such as decision trees and linear models) or interpretable deep learning models (such as networks with obvious attention mechanisms). It not only independently annotates the same batch of data, but also outputs the prediction confidence for each data sample, thereby constructing an "uncertainty dataset". The key role of this module is: on the one hand, to provide verification and supplementation for the output of large models through interpretable prediction results, thereby enhancing the credibility of the overall system; on the other hand, to accurately identify samples that the model itself cannot grasp by using confidence quantification, providing an important basis for screening difficult samples.
[0086] In this embodiment, the machine learning / interpretable deep learning module is not limited to a single machine learning or interpretable deep learning model, but can also be a small collection of models. Its output can be the joint prediction and combined confidence of multiple models, thereby providing a more robust uncertainty assessment.
[0087] (3) Manual annotation module: As the final step in quality control, this module is responsible for the detailed annotation and review of the "difficult samples" selected by the system. This module can be staffed by domain experts or trained professional annotators who combine business knowledge, contextual information, and human judgment to make the final decision on samples that the machine cannot reach a consensus on. In addition, this module can also undertake the sampling review of some high-value samples to ensure the reliability of the overall annotation results. The manual annotation module not only improves the authority and accuracy of the final dataset, but its annotation results can also be used to update machine learning models or multimodal large models, thereby achieving continuous optimization of system performance.
[0088] The data annotation process of a data annotation system can be divided into the following three stages:
[0089] (1) Parallel annotation and divergence detection stage: First, by Several different multimodal large models independently annotate the dataset to be labeled in parallel, generating... A preliminary labeled dataset (corresponding to the initial labeled data in Example 1). Subsequently, the system cross-references this dataset. The labeled results of each dataset were used to identify all samples with inter-model discrepancies and then summarized to form a dataset containing... The large model of a sample is called "divergent dataset (corresponding to the divergent sample set in Example 1)", and the data annotation results without divergence form "pre-annotated dataset (corresponding to the first labeled dataset in Example 1)".
[0090] (2) Uncertainty Quantification and Difficult Sample Localization Stage: Simultaneously, a machine learning (or interpretable deep learning) model (corresponding to the second model in Example 1) labels the same dataset to be labeled and calculates the prediction confidence of each sample, from which the sample with the lowest confidence is selected. The data samples constitute the "uncertainty dataset (corresponding to the set of samples to be determined in Example 1)". Finally, by calculating the intersection of the large model divergence dataset and the uncertainty dataset, the "difficult sample dataset (corresponding to the difficult sample set)" that ultimately requires manual intervention is accurately located.
[0091] (3) Manual labeling and verification stage: Submit this difficult sample dataset to the manual labeling module, where domain experts will perform the final fine labeling.
[0092] Based on the three stages mentioned above, this embodiment further refines the processing steps using text classification data annotation as a specific example. Figure 3 This is a flowchart of an optional data annotation method according to an embodiment of the present invention, such as... Figure 3 As shown, it includes:
[0093] S31: A multimodal large model is used to annotate the text classification dataset to form a... One labeled dataset;
[0094] S32: Large Model Each labeled dataset is paired to generate a "pre-labeled dataset" and a "divergence dataset," and the number of samples in the text classification divergence dataset is obtained. ;
[0095] S33: A text classification machine learning (or interpretable deep learning) model labels the text classification dataset to be labeled, calculates the confidence score of the text classification prediction for each sample, and obtains... The sample with the lowest confidence level generates the "uncertainty dataset" for text classification;
[0096] S34: The intersection of the text classification "divergence dataset" generated by the multimodal large model and the text classification "uncertainty dataset" of the text classification machine learning (or interpretable deep learning) model is calculated to generate the text classification "difficult sample dataset"; the difference between the "divergence dataset" and the "difficult sample dataset" is the "difference hard vote dataset", and the "difference hard vote dataset" determines the final label through the model's hard bidding mechanism.
[0097] S35: The dataset of difficult text classification samples is handed over to manual annotation, and the manual annotation forms the final annotation results.
[0098] In this embodiment, by organically combining multiple multimodal large-scale model collaboration mechanisms with an uncertainty-based sample selection method, synergistic optimization of multimodal data annotation in terms of quality, efficiency, and cost is achieved. For example, a dual quality assurance mechanism is constructed through deep semantic understanding of multimodal large-scale models and interpretability verification of machine learning models. The uncertainty-based sample selection can accurately identify marginal cases and ambiguous samples, effectively avoiding the propagation of systematic annotation errors in the dataset.
[0099] The data annotation system provided in this embodiment avoids the redundant work of building separate annotation processes for different data types through a multimodal unified processing architecture. The collaborative annotation framework focuses manual annotation resources on the most valuable and challenging samples. A complete and traceable chain from raw data to final annotation results is established, providing a reliable basis for decision-making in high-risk applications. The transparent decision-making and deep reasoning of large models complement each other, significantly improving the interpretability of annotation results.
[0100] Example 3
[0101] Embodiment 3 of the present invention provides an optional data annotation device, wherein each implementation unit in the data annotation device corresponds to each implementation step in Embodiment 1.
[0102] Figure 4 This is a schematic diagram of an optional data annotation device according to an embodiment of the present invention, such as... Figure 4 As shown, it includes: annotation unit 41, comparison unit 42, acquisition unit 43 and determination unit 44.
[0103] The annotation unit 41 is used to annotate the data samples in the dataset to be annotated using n first models to obtain n initial annotated datasets. The dataset to be annotated includes s data samples to be annotated. The model type of the first model includes: multimodal large model, where n and s are positive integers.
[0104] Comparison unit 42 is used to perform pairwise comparisons on n initially labeled datasets to obtain comparison results;
[0105] The acquisition unit 43 is used to acquire the confidence scores of the second model in labeling the data samples in the dataset to be labeled, and obtain s confidence scores. The second model includes a machine learning model.
[0106] The determination unit 44 is used to determine the target labeled dataset based on the comparison results and s confidence levels, wherein the target labeled dataset includes the final labeling result of each data sample in the dataset to be labeled.
[0107] In the data annotation apparatus provided in this embodiment, the annotation unit 41 can annotate data samples in the dataset to be annotated using n first models respectively, to obtain n initial annotated datasets, wherein the dataset to be annotated includes s data samples to be annotated, and the model type of the first model includes: multimodal large model, where n and s are positive integers; the comparison unit 42 is used to compare the n initial annotated datasets pairwise to obtain comparison results; the acquisition unit 43 is used to acquire the confidence scores of the second model in annotating the data samples in the dataset to be annotated, to obtain s confidence scores, wherein the second model includes: machine learning model; the determination unit 44 is used to determine the target annotated dataset based on the comparison results and the s confidence scores, wherein the target annotated dataset includes: the final annotation result of each data sample in the dataset to be annotated. This solves the technical problem of low accuracy in the annotation results of data samples in related technologies. In this embodiment, data samples are labeled using multiple multimodal large models, and the final labeling result of the data samples is determined based on the differences in the labeling results and the confidence level of the machine learning model prediction. This avoids the low accuracy of manual or semi-automatic sample labeling in related technologies, thereby achieving the technical effect of improving the accuracy of data sample labeling.
[0108] Optionally, the determining unit includes: a first filtering subunit, used to filter out consistent annotation results for the same data sample from n initial labeled datasets based on comparison results, to obtain a first labeled dataset; a second filtering subunit, used to filter out inconsistent annotation results for the same data sample from n initial labeled datasets based on comparison results, to obtain a divergent sample set; a first determining subunit, used to determine the annotation results of the data samples in the divergent sample set based on s confidence levels, to obtain a second labeled dataset; and a second determining subunit, used to determine the target labeled dataset based on the first labeled dataset and the second labeled dataset.
[0109] Optionally, the first determining subunit includes: a filtering module, used to filter k data samples from the dataset to be labeled based on s confidence levels to obtain a sample set to be determined, wherein the confidence level of the data samples in the sample set to be determined is lower than the confidence level of the other data samples in the dataset to be labeled except for the k data samples, and k is a positive integer less than s; and a determining module, used to determine the labeling results of the data samples in the divergent sample set based on the sample set to be determined to obtain a second labeled dataset.
[0110] Optionally, the determining module includes: a first calculation submodule for calculating the intersection of the divergent sample set and the sample set to be determined, to obtain the difficult sample set; a second calculation submodule for calculating the difference between the divergent sample set and the difficult sample set, to obtain the target sample set; and a determining submodule for determining the annotation results of the data samples in the difficult sample set and the target sample set, respectively, to obtain the second labeled dataset.
[0111] Optionally, the determination submodule includes: a first determination submodule, which determines the annotation results of all data samples in the target sample set based on the annotation results of each data sample in the n initial annotation datasets, to obtain a first annotation result set; a processing submodule, which sends the difficult sample set to the target object and receives all annotation results returned by the target object, to obtain a second annotation result set, wherein the target object is used to annotate the data samples in the difficult sample set; and a second determination submodule, which determines a second annotation dataset based on the first annotation result set and the second annotation result set.
[0112] Optionally, the first determining submodule further includes: a filtering submodule 1, which, for each data sample in the target sample set, filters the labeling result with the highest number of identical labeling results from the n initial labeled datasets to obtain the target labeling result of the data sample; and a determining submodule 1, which determines the first labeling result based on the target labeling result of each data sample in the target sample set.
[0113] Optionally, the data samples in the dataset to be labeled may include at least one of the following types: text, image, audio, and video.
[0114] The data annotation device described above may also include a processor and a memory. The annotation unit 41, comparison unit 42, acquisition unit 43 and determination unit 44 are all stored in the memory as program units, and the processor executes the program units stored in the memory to realize the corresponding functions.
[0115] The aforementioned processor contains a kernel that retrieves the corresponding program units from memory. One or more kernels can be configured, and by adjusting kernel parameters, data samples can be labeled using multiple multimodal large models. Based on the differences in labeling results and the confidence level of the machine learning model's predictions, the final labeling result of the data samples is determined. This avoids the low accuracy issues associated with manual or semi-automatic sample labeling in related technologies, thus achieving the technical effect of improving the accuracy of data sample labeling.
[0116] The aforementioned memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0117] According to another aspect of the present invention, an electronic device is also provided, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the data annotation method of any of the above-mentioned methods by executing the executable instructions.
[0118] According to another aspect of the present invention, a computer-readable storage medium is also provided, which stores a computer program, wherein the computer program controls the device where the computer-readable storage medium is located to execute the data annotation method described above when it is running.
[0119] Figure 5 This is a schematic diagram of an electronic device according to an embodiment of the present invention, such as... Figure 5 As shown, an embodiment of the present invention provides an electronic device 50, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements any of the above-mentioned data annotation methods.
[0120] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0121] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0122] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0123] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0124] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0125] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0126] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A data annotation method, characterized in that, include: By labeling the data samples in the dataset to be labeled using n first models, n initial labeled datasets are obtained. The dataset to be labeled includes s data samples to be labeled. The model type of the first model includes: multimodal large model, where n and s are positive integers. Perform pairwise comparisons on n initially labeled datasets to obtain the comparison results; Obtain the confidence scores of the second model in labeling the data samples in the dataset to be labeled, and obtain s confidence scores, wherein the second model includes: a machine learning model; Based on the comparison results and the s confidence levels, a target labeled dataset is determined, wherein the target labeled dataset includes the final labeled result of each data sample in the dataset to be labeled.
2. The data annotation method according to claim 1, characterized in that, Based on the comparison results and the s confidence scores, the target labeled dataset is determined, including: Based on the comparison results, from the n initial labeled datasets, the labeled results that are consistent in labeling the same data sample are selected to obtain the first labeled dataset; Based on the comparison results, n data samples from the initial labeled dataset that have inconsistent labeling results for the same data sample are selected from the dataset to be labeled, thus obtaining a divergence sample set; Based on the s confidence levels, the annotation results of the data samples in the divergent sample set are determined to obtain the second labeled dataset; The target labeled dataset is determined based on the first labeled dataset and the second labeled dataset.
3. The data annotation method according to claim 2, characterized in that, Based on the s confidence levels, the annotation results of the data samples in the divergent sample set are determined to obtain the second labeled dataset, including: Based on the s confidence scores, k data samples are selected from the dataset to be labeled to obtain a sample set to be determined. The confidence scores of the data samples in the sample set to be determined are lower than the confidence scores of the other data samples in the dataset to be labeled, excluding the k data samples. k is a positive integer less than s. Based on the sample set to be determined, the annotation results of the data samples in the divergent sample set are determined to obtain the second annotated dataset.
4. The data annotation method according to claim 3, characterized in that, Based on the sample set to be determined, the annotation results of the data samples in the divergent sample set are determined to obtain the second labeled dataset, including: Calculate the intersection of the divergent sample set and the sample set to be determined to obtain the difficult sample set; The difference between the divergent sample set and the problematic sample set is calculated to obtain the target sample set; The annotation results of the data samples in the difficult sample set and the target sample set are determined respectively to obtain the second annotated dataset.
5. The data annotation method according to claim 4, characterized in that, The annotation results of the data samples in the difficult sample set and the target sample set are determined respectively to obtain the second annotated dataset, including: Based on the annotation results of each data sample in the n initial annotation datasets, the annotation results of all data samples in the target sample set are determined to obtain the first annotation result set; The problematic sample set is sent to the target object, and all annotation results returned by the target object are received to obtain a second annotation result set, wherein the target object is used to annotate the data samples in the problematic sample set; Based on the first annotation result set and the second annotation result set, the second annotation dataset is determined.
6. The data annotation method according to claim 5, characterized in that, Based on the annotation results for each data sample in the n initial annotation datasets, the annotation results for all data samples in the target sample set are determined to obtain the first annotation result set, which also includes: For each data sample in the target sample set, the labeling result with the highest number of identical labeling results is selected from the n initial labeled datasets to obtain the target labeling result of that data sample; The first annotation result is determined based on the target annotation result of each data sample in the target sample set.
7. The data annotation method according to claim 1, characterized in that, The data samples in the dataset to be labeled include at least one of the following types: text, image, audio, and video.
8. A data annotation device, characterized in that, include: The annotation unit is used to annotate the data samples in the dataset to be annotated using n first models to obtain n initial annotated datasets. The dataset to be annotated includes s data samples to be annotated. The model type of the first model includes: multimodal large model, where n and s are positive integers. The comparison unit is used to perform pairwise comparisons on n initially labeled datasets to obtain the comparison results. The acquisition unit is used to acquire the confidence scores of the second model in labeling the data samples in the dataset to be labeled, and obtain s confidence scores, wherein the second model includes: a machine learning model; A determining unit is configured to determine a target labeled dataset based on the comparison results and s confidence levels, wherein the target labeled dataset includes the final labeling result of each data sample in the dataset to be labeled.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the data annotation method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the data annotation method according to any one of claims 1 to 7.