Feature-Space Data Presentation for Efficient Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional annotation methods for supervised machine learning impose a heavy burden on users, especially when domain information is missing or not correlated with the data, making it difficult to efficiently label multiple pieces of data.

Innovation Solution

A data presentation method involving feature extraction, distance calculation, and rearrangement of data in a feature value space, along with dimensionality reduction and clustering, to facilitate efficient labeling by a user.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If data are rearranged according to domain information (date, time, location, format), then the cognitive burden on user is reduced, but the method cannot be applied when domain information is missing or not correlated with images

Engineering Contradiction:
Improvecognitive burden on userVSAvoidapplicability when domain information is missing
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent changes the basis for data arrangement from domain information parameters (date, time, location) to feature value parameters extracted from the data itself. By calculating feature values and using distance metrics in feature value space, the system adapts to cases where domain information is missing or not correlated, while still achieving efficient data presentation that reduces user cognitive burden.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If many pieces of data are labelled one by one manually, then accurate labels can be assigned, but the workload on user becomes heavy

Engineering Contradiction:
Improvelabeling accuracyVSAvoidannotation speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent performs preliminary actions by automatically calculating feature values and computing distance metrics before the user performs annotation. Data are pre-arranged in feature value space based on their similarity, so when the user views the data, they are presented in an optimized sequence. This preliminary processing reduces the user's workload while maintaining labeling accuracy, as the user can leverage the pre-computed similarity relationships.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If domain information is used to correlate with images, then annotation efficiency is improved, but the method fails when domain information and images are not correlated

Engineering Contradiction:
Improveannotation efficiencyVSAvoideffectiveness when no correlation exists
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces feature values as an intermediary between the raw data and the annotation process. Instead of directly using domain information (which may not correlate with images), the system extracts feature values that capture the essential characteristics of the data. This intermediary representation enables reliable data arrangement and comparison even when original domain information is missing or not correlated, while still improving annotation efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250285020A1Data presentation method, annotation method, and computer program
Publication Date: 2025.09.11 SCREEN HOLDINGS CO LTD
  • US20250285020A1 patent drawing
  • US20250285020A1 patent drawing
  • US20250285020A1 patent drawing

AI summary

In this data presentation method, first, feature values of multiple pieces of data 9 are calculated. Subsequently, a distance d between the pieces of data 9 in a feature value space S defined on the basis of the feature values is calculated. Then, the data 9 is selected from a dataset 90 on the basis of the calculated distance d, and is presented to a user. Thus, the data 9 can be presented to the user in the manner corresponding to the distance d between the pieces of data 9 in the feature value space S. Therefore, the user can efficiently perform a process of labelling each of the multiple pieces of data 9.