Feature Space Quadrant Segmentation for Annotation Data Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As the number of data pieces with correct answer information increases, it becomes difficult for users to efficiently select accurate reference information for annotating target data, leading to inefficient annotation processes and potentially inaccurate label additions.

Innovation Solution

An information processing apparatus that estimates the likelihood of labels being added as annotations using a pre-trained model, divides the feature space into quadrants based on these likelihood scores, and presents candidate data from these quadrants to assist users in selecting the most appropriate labels, thereby reducing the workload and improving annotation accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the number of pieces of data to which correct answer information is added increases, then the accuracy of machine learning model improves, but it becomes substantially difficult for a user to check all data

Engineering Contradiction:
Improveaccuracy of machine learning modelVSAvoidease of checking data
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent segments the large set of candidate data into multiple groups based on classification results. The presentation unit displays a plurality of groups of candidate data, where each group contains candidate data with similar characteristics. This segmentation allows users to systematically review data in manageable portions rather than being overwhelmed by the entire dataset at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of organization by grouping candidate data according to classification categories. Instead of presenting data in a single flat list, the system creates a multi-dimensional structure where data is organized by groups and subgroups, adding a categorical dimension that facilitates easier navigation and comparison.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If random extraction of data is used for presenting reference information, then the workload is reduced, but the reference information may not always serve as a reference when determining which correct answer information to add

Engineering Contradiction:
Improveannotation efficiencyVSAvoidquality of reference information
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system provides feedback to users by presenting candidate data that is systematically organized according to classification results. The presentation unit displays multiple groups of candidate data with their classification information, allowing users to see patterns and relationships. This feedback mechanism ensures that the reference information is both efficient to review and reliable for making annotation decisions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The classification unit performs preliminary action by automatically categorizing candidate data before presentation to the user. This pre-processing step organizes the data into meaningful groups based on relevant features, so that when users review the candidate data, it is already structured in a way that highlights important patterns and relationships, improving both efficiency and reliability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12164528B2Information processing apparatus, information processing method, and storage medium for obtaining annotated training data
Publication Date: 2024.12.10 CANON KK
  • US12164528B2 patent drawing
  • US12164528B2 patent drawing
  • US12164528B2 patent drawing

AI summary

A candidate data determination unit acquires a result of estimation of a score representing a likelihood of a label being added as an annotation to target data. A label candidate input unit receives designation of a candidate label from a user. The candidate data determination unit determines candidate data, from a plurality of pieces of labeled data included in a feature space, the candidate data representing a plurality of pieces of labeled data distributed in respective quadrants into which the feature space is divided, wherein the label is added to the labeled data as the annotation, and wherein the feature space is defined using the score, which represents a likelihood of a label being added as an annotation, as an axis. The candidate data determination unit determines, for each of the plurality of quadrants, candidate data based on the labeled data included in each of the plurality of quadrants.