AI Training Data Generation for Balanced Category Performance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The lack of massive, high-quality data and labeled information restricts the performance and generalization capability of artificial intelligence models, particularly in fields like computer vision, due to challenges in data collection under specific conditions such as weather, military, or medical scenarios, leading to biased training and limited applicability.

Innovation Solution

A method involving a generative AI model to create datasets categorized by input prompts, deleting partial data based on inference results, augmenting the dataset with algorithms, training a second AI model, calculating inference performance, and adjusting prompt generation based on performance to enhance data quality and balance across categories.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If massive high-quality data is collected for training AI models, then model performance and generalization capability are improved, but data collection becomes difficult and time-consuming due to specific conditions requirements

Engineering Contradiction:
Improvemodel performanceVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses a generative AI model to create synthetic training data that copies the characteristics and patterns of real-world data without requiring actual data collection. This allows the model to learn from generated samples that simulate various scenarios, weather conditions, and environments, thereby improving model performance while eliminating time-consuming data collection processes

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary data generation and quality assessment before actual training begins. By pre-generating diverse training samples and evaluating their quality using multiple AI models, the system prepares high-quality training data in advance, avoiding the need for time-consuming data collection during the training phase

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If data is collected under specific conditions (weather, military, medical), then model applicability to those scenarios is improved, but data quantity becomes insufficient leading to biased training

Engineering Contradiction:
Improvescenario-specific applicabilityVSAvoiddata quantity
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent transitions from the constraint of real-world data collection to the unlimited dimension of synthetic data generation. By using generative AI to create data across multiple dimensions (different weather conditions, military scenarios, medical situations), the system achieves sufficient data quantity for each specific scenario without being limited by physical data collection constraints, thereby preventing training bias

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system generates data with specific local qualities tailored to each scenario category. By creating synthetic data with targeted characteristics for specific conditions (e.g., snowy weather, military environments, medical scenarios), the system ensures high adaptability to each scenario while maintaining sufficient data quantity through controlled generation processes

Inventive Principle:
Principle #3Local quality

3Measurement precision

If manual labeling is performed for supervised learning, then data quality and accuracy are improved, but the process becomes challenging and time-consuming

Engineering Contradiction:
Improvelabel accuracyVSAvoidlabeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual human labeling with automated synthetic label generation through AI models. The generative model creates ground truth labels alongside synthetic data, copying the labeling function from human annotators to machine-based systems. This maintains high label accuracy through controlled generation while eliminating the time-consuming manual labeling process

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-labeling by using the generative AI model to automatically create labeled training pairs without external human intervention. The model generates both the synthetic data and its corresponding ground truth labels, enabling the training process to be self-sufficient and eliminating dependency on manual human labor for data annotation

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20260080667A1Apparatus and method for training artificial intelligence model
Publication Date: 2026.03.19 AGENCY FOR DEFENSE DEV
  • US20260080667A1 patent drawing
  • US20260080667A1 patent drawing
  • US20260080667A1 patent drawing

AI summary

Disclosed are an apparatus and a method for training an artificial intelligence model. According to the present disclosure, the apparatus for training an artificial intelligence model may generate a dataset that corresponds to input prompts and is classified by a plurality of categories, delete partial data from the dataset based on based on inference results obtained by inputting data included in the dataset into a plurality of pre-trained first artificial intelligence models, augment, by applying a preset algorithm, the dataset from which the partial data is deleted, train a second artificial intelligence model based on the augmented dataset, calculate inference performance of the second artificial intelligence model for each of the plurality of categories based on an inference result obtained by inputting a test dataset into the trained second artificial intelligence model, and adjust generation of the input prompts for the plurality of categories based on the inference performance.