Training Data Supplementation for Consistent AI Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data labeling methods suffer from inconsistencies due to varying user skill levels, leading to inaccurate labeling and increased costs and efforts in training neural network models.
Innovation Solution
A method that involves receiving a first user input for a first data set, supplementing a second user input based on the first input to obtain a supplemented second data set, reflecting an operator's pattern with high accuracy, thereby reducing labeling costs and efforts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If user inputs are collected for data labeling, then training data can be obtained, but labeling quality becomes inconsistent due to varying user skill levels
Solution Approach 1:
The system provides feedback by comparing user inputs against example labels and supplementing incomplete or incorrect inputs. The computing device receives user inputs, compares them with example labels from high-accuracy operators, and generates supplemented inputs that reflect the desired labeling pattern, thereby improving consistency without requiring all users to have expert skill levels
Solution Approach 2:
The example label acts as an intermediary between the user input and the final training data. Instead of directly requiring users to produce accurate labels, the system uses example labels as a mediator that guides and corrects user inputs, transforming variable-quality user contributions into consistent training data through the supplementation process
2Reliability
If additional operations are performed to identify and supplement inaccurate data, then training accuracy improves, but labeling cost and effort increase
Solution Approach 1:
The system performs self-service by automatically comparing user inputs against example labels and generating supplemented inputs without requiring manual review or additional human intervention. The computing device autonomously identifies inconsistencies and produces corrected data, eliminating the need for costly manual verification while maintaining training accuracy
Solution Approach 2:
The system performs preliminary action by using example labels from high-accuracy operators to pre-establish the desired labeling pattern before users provide their inputs. This preliminary framework allows the system to automatically correct and supplement user inputs in real-time, preventing inaccurate data from entering the training set without requiring subsequent expensive verification operations
Data Source
AI summary
Disclosed is a method for managing a local model and a global model on an AI platform, the method performed by one or more processors of a computing device according to an exemplary embodiment of the present disclosure.the method may include: obtaining a plurality of sample data; receiving a first user input for a first data set included in the plurality of sample data; receiving a second user input for a second data set excluding the first data set among the plurality of sample data; and obtaining a supplemented second data set by supplementing the second user input based on the first user input.


