Data-Centric AI Annotation With Quality-Guided Data Collection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data labeling processes in machine learning are expensive and time-consuming, often lacking guidance on which data points to label for quality improvement, leading to inefficient data acquisition and annotation without ensuring meaningful content and quality.

Innovation Solution

A system and method that generates pseudo-labeled data through unsupervised clustering, selects samples based on quality metrics, and prompts users for annotations, recommending additional data collection if quality metrics are not met, thereby optimizing data quality and minimizing costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual data labeling is performed to ensure high data quality, then data quality improves, but time consumption and costs increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary automated labeling using pre-trained machine learning models to generate initial labels before human annotation. This preliminary action provides a head start on the labeling process, reducing the time and effort required for subsequent manual refinement while maintaining high data quality standards.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary automated labeling layer between raw unlabeled data and final human-annotated data. This intermediary uses pre-trained models to generate preliminary labels that serve as a bridge, reducing the gap between automated efficiency and manual quality while minimizing direct human intervention time.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If comprehensive data collection is performed to ensure quality metrics are met, then data quality improves, but data acquisition costs and time increase

Engineering Contradiction:
Improvedata qualityVSAvoiddata acquisition volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system applies local quality by identifying and focusing annotation efforts on specific data samples that have the greatest impact on model performance. Rather than uniformly annotating all data, the system selectively targets samples where annotation provides maximum value, reducing overall annotation volume while maintaining data quality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts annotation thresholds and quality metrics based on the specific characteristics of the dataset and model requirements. By changing parameters such as confidence thresholds and sampling rates, the system optimizes the balance between data quantity collected and quality achieved, minimizing unnecessary data acquisition.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If random data sampling is used for annotation, then annotation process is simple, but data quality and model performance improvement is limited

Engineering Contradiction:
Improveannotation process simplicityVSAvoiddata quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system implements feedback mechanisms where annotation results and model performance metrics are continuously monitored and used to adjust subsequent sampling strategies. This feedback loop enables the system to learn from previous annotation rounds and improve data quality over time while maintaining operational simplicity through automated decision-making.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system transitions from static random sampling to dynamic adaptive sampling where the sampling strategy evolves based on model performance and data characteristics. This dynamic approach automatically adjusts which samples to annotate next, improving data quality without requiring complex manual intervention while maintaining ease of operation through automation.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12417230B2Annotating and collecting data-centric AI quality metrics considering user preferences
Publication Date: 2025.09.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12417230B2 patent drawing
  • US12417230B2 patent drawing
  • US12417230B2 patent drawing

AI summary

A method, computer program, and computer system are provided for collecting and annotating data based on user preference. Unlabeled data corresponding to one or more entries within a dataset is received. Pseudo-labeled data is generated based on the unlabeled data. Based on one or more quality metrics, each entry from among the pseudo-labeled data is determining to be included within a final dataset. A user is prompted for annotations corresponding to entries of the pseudo-labeled data included within the final dataset. A determination is made as to whether additional data is needed based on comparing the final dataset to the one or more quality metrics, and the additional information is collected if the final dataset does not meet the quality metrics.