Feature Space Quadrant Segmentation for Annotation Data Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As the number of data pieces with correct answer information increases, it becomes difficult for users to efficiently select accurate reference information for annotating target data, leading to inefficient annotation processes and potentially inaccurate label additions.
Innovation Solution
An information processing apparatus that estimates the likelihood of labels being added as annotations using a pre-trained model, divides the feature space into quadrants based on these likelihood scores, and presents candidate data from these quadrants to assist users in selecting the most appropriate labels, thereby reducing the workload and improving annotation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the number of pieces of data to which correct answer information is added increases, then the accuracy of machine learning model improves, but it becomes substantially difficult for a user to check all data
Solution Approach 1:
The patent segments the large set of candidate data into multiple groups based on classification results. The presentation unit displays a plurality of groups of candidate data, where each group contains candidate data with similar characteristics. This segmentation allows users to systematically review data in manageable portions rather than being overwhelmed by the entire dataset at once.
Solution Approach 2:
The patent introduces a new dimension of organization by grouping candidate data according to classification categories. Instead of presenting data in a single flat list, the system creates a multi-dimensional structure where data is organized by groups and subgroups, adding a categorical dimension that facilitates easier navigation and comparison.
2Productivity
If random extraction of data is used for presenting reference information, then the workload is reduced, but the reference information may not always serve as a reference when determining which correct answer information to add
Solution Approach 1:
The system provides feedback to users by presenting candidate data that is systematically organized according to classification results. The presentation unit displays multiple groups of candidate data with their classification information, allowing users to see patterns and relationships. This feedback mechanism ensures that the reference information is both efficient to review and reliable for making annotation decisions.
Solution Approach 2:
The classification unit performs preliminary action by automatically categorizing candidate data before presentation to the user. This pre-processing step organizes the data into meaningful groups based on relevant features, so that when users review the candidate data, it is already structured in a way that highlights important patterns and relationships, improving both efficiency and reliability.
Data Source
AI summary
A candidate data determination unit acquires a result of estimation of a score representing a likelihood of a label being added as an annotation to target data. A label candidate input unit receives designation of a candidate label from a user. The candidate data determination unit determines candidate data, from a plurality of pieces of labeled data included in a feature space, the candidate data representing a plurality of pieces of labeled data distributed in respective quadrants into which the feature space is divided, wherein the label is added to the labeled data as the annotation, and wherein the feature space is defined using the score, which represents a likelihood of a label being added as an annotation, as an axis. The candidate data determination unit determines, for each of the plurality of quadrants, candidate data based on the labeled data included in each of the plurality of quadrants.


