Data Display Size Decision for Labeling Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In machine learning, the process of labeling data for supervised learning is inefficient due to high costs associated with preparing large volumes of labeled data, especially when not all input data need to be labeled.
Innovation Solution
An information processing apparatus and method that determines and adjusts the display size of data based on its likelihood vector and the class it belongs to, allowing for more efficient labeling by visually highlighting data that is likely to be misclassified, thereby reducing the need for extensive labeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large volume of training data is prepared for highly-accurate classification and recognition, then classification accuracy is improved, but labeling cost increases
Solution Approach 1:
The patent applies local quality by making different types of data have different display sizes based on their likelihood of misclassification. Data with higher misclassification risk (lower likelihood values) are displayed larger, while data with lower misclassification risk are displayed smaller. This selective emphasis on specific data points allows annotators to focus attention where it is most needed, achieving high classification accuracy without labeling all data uniformly at high cost.
2Measurement precision
If all input image data are labeled, then classification accuracy is improved, but labeling time and cost increase
Solution Approach 1:
The patent implements partial action by labeling only a subset of data rather than all input data. By calculating likelihood values and identifying only the data most likely to be misclassified, the system performs labeling on a partial set of critical data points. This approach maintains classification accuracy while significantly reducing the time and resources required compared to labeling all data.
3Loss of information
If data are arranged on likelihood vector for visualization, then misclassification risk is identified, but display complexity increases
Solution Approach 1:
The patent applies dimensionality change by transforming the one-dimensional likelihood value into a two-dimensional visual representation through display size. Instead of showing complex numerical likelihood values or using multiple visual encodings, the system maps the likelihood dimension to the size dimension of displayed data. This simple visual transformation allows annotators to instantly grasp misclassification risk without complex interfaces or additional display elements.
Data Source
AI summary
There is provided an information processing apparatus and an information processing method, the information processing apparatus including: a display size decision unit configured to decide a display size of data that is based on learning of a label corresponding to a class, on the basis of a distance between a position of the data arranged on the basis of a likelihood vector obtained by recognition of the data, and a position of a class to which the data belongs; and a communication control unit configured to cause the display size to be transmitted.


