Labeled Data Quality Assessment Using Granular Sampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in efficiently and cost-effectively quantifying the quality of labeled data and relabeling it, which is crucial for accurate operation of autonomous vehicles, due to high resource consumption and time consumption in manual verification processes.
Innovation Solution
A method is introduced to identify a labeling quality metric at a lower granularity level, sampling a portion of the data set, applying model-generated pre-labels with zero-margin tolerance, and relabeling based on the metric to ensure ground truth data, thereby reducing resource and time consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual verification processes are used to quantify the quality of labeled data, then the accuracy of quality assessment is improved, but the resource consumption and time consumption increase significantly
Solution Approach 1:
The patent uses model-generated pre-labels as copies of ground truth labels to create synthetic labeled datasets. These synthetic labels serve as proxies for manual verification, allowing quality assessment without requiring actual manual annotation of the entire dataset. The pre-labels replicate the function of ground truth labels while avoiding the time-consuming manual process.
Solution Approach 2:
Instead of manually verifying the entire labeled dataset, the patent performs partial verification by using model-generated pre-labels on a representative subset or the full dataset at a faster rate. This partial action approach maintains quality assessment accuracy while significantly reducing the time investment required compared to complete manual verification.
2Measurement precision
If manual verification processes are used to quantify the quality of labeled data, then the accuracy of quality assessment is improved, but the resource consumption increases significantly
Solution Approach 1:
The patent replaces resource-intensive manual verification with automated model-generated pre-labels. These synthetic labels consume significantly fewer computational resources and human resources while maintaining the ability to assess data quality accurately. The copying approach allows parallel processing and automated quality metrics calculation.
Solution Approach 2:
The patent substitutes the mechanical process of manual human verification with an automated computational system that generates pre-labels using machine learning models. This substitution replaces human cognitive resources with algorithmic processing, dramatically reducing resource consumption while maintaining or improving measurement precision through consistent, scalable automated assessment.
3Manufacturing precision
If ground truth data is created through manual labeling, then the accuracy of the data is improved, but the cost and time required increase
Solution Approach 1:
The patent performs preliminary action by generating pre-labels using trained machine learning models before final data annotation or verification. These pre-labels provide an initial accurate labeling that can be used directly or serve as a foundation for minimal manual verification, significantly reducing the total time required while maintaining high accuracy through the preliminary automated labeling pass.
Solution Approach 2:
The patent uses model-generated pre-labels as accurate copies of what manual labeling would produce. The trained models replicate the labeling accuracy of human annotators while operating at much higher speeds, creating ground truth data that is both accurate and efficiently produced without requiring time-consuming manual intervention for every data point.
4Manufacturing precision
If ground truth data is created through manual labeling, then the accuracy of the data is improved, but the cost and time required increase
Solution Approach 1:
The patent replaces the manual mechanical process of human labeling with automated machine learning model inference. This substitution maintains high data accuracy through well-trained models while dramatically increasing productivity by processing vast amounts of data in parallel without human intervention, achieving both accuracy and efficiency simultaneously.
Solution Approach 2:
The patent changes the operational parameters from manual human processing to automated computational processing. By adjusting model confidence thresholds, sampling rates, and verification criteria, the system optimizes the balance between data accuracy and productivity, achieving high labeling efficiency while maintaining ground truth quality through parameter-tuned automated processes.
Data Source
AI summary
Aspects of the subject technology relate to systems, methods, and computer-readable media for identifying a quality of labeled data. Labeled data of a data set that exists at a specific granularity level can be accessed. The labeled data can be sampled on a lower granularity level relative to the specific granularity level of the data set to generate sampled data of the data set. The sampled data can be labeled to generate ground truth labeled data of the data set. The labeled data can be compared to the ground truth labeled data to identify a labeling quality metric of the labeled data. Relabeling of the data set can be performed based on the labeling quality metric.


