AI Training Data Generation for Balanced Category Performance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The lack of massive, high-quality data and labeled information restricts the performance and generalization capability of artificial intelligence models, particularly in fields like computer vision, due to challenges in data collection under specific conditions such as weather, military, or medical scenarios, leading to biased training and limited applicability.
Innovation Solution
A method involving a generative AI model to create datasets categorized by input prompts, deleting partial data based on inference results, augmenting the dataset with algorithms, training a second AI model, calculating inference performance, and adjusting prompt generation based on performance to enhance data quality and balance across categories.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If massive high-quality data is collected for training AI models, then model performance and generalization capability are improved, but data collection becomes difficult and time-consuming due to specific conditions requirements
Solution Approach 1:
The patent uses a generative AI model to create synthetic training data that copies the characteristics and patterns of real-world data without requiring actual data collection. This allows the model to learn from generated samples that simulate various scenarios, weather conditions, and environments, thereby improving model performance while eliminating time-consuming data collection processes
Solution Approach 2:
The system performs preliminary data generation and quality assessment before actual training begins. By pre-generating diverse training samples and evaluating their quality using multiple AI models, the system prepares high-quality training data in advance, avoiding the need for time-consuming data collection during the training phase
2Adaptability or versatility
If data is collected under specific conditions (weather, military, medical), then model applicability to those scenarios is improved, but data quantity becomes insufficient leading to biased training
Solution Approach 1:
The patent transitions from the constraint of real-world data collection to the unlimited dimension of synthetic data generation. By using generative AI to create data across multiple dimensions (different weather conditions, military scenarios, medical situations), the system achieves sufficient data quantity for each specific scenario without being limited by physical data collection constraints, thereby preventing training bias
Solution Approach 2:
The system generates data with specific local qualities tailored to each scenario category. By creating synthetic data with targeted characteristics for specific conditions (e.g., snowy weather, military environments, medical scenarios), the system ensures high adaptability to each scenario while maintaining sufficient data quantity through controlled generation processes
3Measurement precision
If manual labeling is performed for supervised learning, then data quality and accuracy are improved, but the process becomes challenging and time-consuming
Solution Approach 1:
The patent replaces manual human labeling with automated synthetic label generation through AI models. The generative model creates ground truth labels alongside synthetic data, copying the labeling function from human annotators to machine-based systems. This maintains high label accuracy through controlled generation while eliminating the time-consuming manual labeling process
Solution Approach 2:
The system performs self-labeling by using the generative AI model to automatically create labeled training pairs without external human intervention. The model generates both the synthetic data and its corresponding ground truth labels, enabling the training process to be self-sufficient and eliminating dependency on manual human labor for data annotation
Data Source
AI summary
Disclosed are an apparatus and a method for training an artificial intelligence model. According to the present disclosure, the apparatus for training an artificial intelligence model may generate a dataset that corresponds to input prompts and is classified by a plurality of categories, delete partial data from the dataset based on based on inference results obtained by inputting data included in the dataset into a plurality of pre-trained first artificial intelligence models, augment, by applying a preset algorithm, the dataset from which the partial data is deleted, train a second artificial intelligence model based on the augmented dataset, calculate inference performance of the second artificial intelligence model for each of the plurality of categories based on an inference result obtained by inputting a test dataset into the trained second artificial intelligence model, and adjust generation of the input prompts for the plurality of categories based on the inference performance.


