Automated Data Labeling via AI Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high cost and inefficiency of data labeling for artificial intelligence models, particularly in large-scale projects across various domains, due to the need for extensive human intervention and the development of multiple models, lead to increased labor and time costs.

Innovation Solution

An automated training-based data labeling method that transmits source data to a worker terminal, receives labeled data, generates artificial intelligence models, and designates preprocessing engines based on performance criteria, allowing for automated data labeling with reduced human intervention and improved efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If workers manually label large amounts of data, then data labeling can be performed, but labor costs and time consumption increase significantly

Engineering Contradiction:
Improvedata labeling efficiencyVSAvoidtime consumption
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system enables self-service data labeling where the AI model automatically labels data without requiring human workers. The model processes images independently, making decisions about object detection and classification autonomously, thereby eliminating manual labor and time consumption associated with human annotation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual human labeling with an automated AI-based system. The AI model uses neural networks and machine learning algorithms to substitute human cognitive processes, achieving automated detection and labeling of objects in images without human intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If multiple AI models are developed for different projects, then project-specific accuracy is improved, but development time and cost increase

Engineering Contradiction:
Improveproject-specific accuracyVSAvoidmodel development time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system employs a universal AI model architecture that can serve multiple projects across different domains. The model is designed with multi-functionality to handle various tasks such as object detection, recognition, and classification, allowing it to adapt to different project requirements without requiring complete model redevelopment for each project.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent utilizes parameter changes and model fine-tuning to adapt a single base model to different projects. By adjusting parameters, weights, and training data specific to each project, the system achieves project-specific accuracy without developing entirely new models, thereby reducing development time and cost.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If preprocessing engines are used to process large amounts of data, then data processing speed is improved, but quality control becomes more difficult

Engineering Contradiction:
Improvedata processing speedVSAvoiddata quality control
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system incorporates feedback mechanisms where the AI model continuously learns from labeled data and adjusts its performance. The feedback loop allows the model to improve its accuracy over time by analyzing errors and refining its predictions, ensuring high data quality control even when processing large amounts of data at high speed.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies preliminary action through pre-training the AI model on diverse datasets before deploying it for specific projects. This pre-training phase prepares the model with foundational knowledge and patterns, enabling it to quickly and accurately process large amounts of data while maintaining high quality control without requiring extensive real-time adjustments.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240362536A1Automated training based data labeling method
Publication Date: 2024.10.31 DATAMAKER
  • US20240362536A1 patent drawing
  • US20240362536A1 patent drawing
  • US20240362536A1 patent drawing

AI summary

An automated training-based data labeling method according to the present disclosure is configured to generate an artificial intelligence model in which a processor separates some of labeled data into training data as soon as it receives a certain amount of labeled data from a worker terminal, and automatically performs the data labeling on objects in source data through automated training of the training data. According to the present disclosure, since a proportion of worker participation is reduced when labeling the data for the objects in the source data, it is possible to dramatically reduce operation costs required for the data labeling.