Feature-Space Data Presentation for Efficient Annotation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional annotation methods for supervised machine learning impose a heavy burden on users, especially when domain information is missing or not correlated with the data, making it difficult to efficiently label multiple pieces of data.
Innovation Solution
A data presentation method involving feature extraction, distance calculation, and rearrangement of data in a feature value space, along with dimensionality reduction and clustering, to facilitate efficient labeling by a user.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If data are rearranged according to domain information (date, time, location, format), then the cognitive burden on user is reduced, but the method cannot be applied when domain information is missing or not correlated with images
Solution Approach 1:
The patent changes the basis for data arrangement from domain information parameters (date, time, location) to feature value parameters extracted from the data itself. By calculating feature values and using distance metrics in feature value space, the system adapts to cases where domain information is missing or not correlated, while still achieving efficient data presentation that reduces user cognitive burden.
2Measurement precision
If many pieces of data are labelled one by one manually, then accurate labels can be assigned, but the workload on user becomes heavy
Solution Approach 1:
The patent performs preliminary actions by automatically calculating feature values and computing distance metrics before the user performs annotation. Data are pre-arranged in feature value space based on their similarity, so when the user views the data, they are presented in an optimized sequence. This preliminary processing reduces the user's workload while maintaining labeling accuracy, as the user can leverage the pre-computed similarity relationships.
3Productivity
If domain information is used to correlate with images, then annotation efficiency is improved, but the method fails when domain information and images are not correlated
Solution Approach 1:
The patent introduces feature values as an intermediary between the raw data and the annotation process. Instead of directly using domain information (which may not correlate with images), the system extracts feature values that capture the essential characteristics of the data. This intermediary representation enables reliable data arrangement and comparison even when original domain information is missing or not correlated, while still improving annotation efficiency.
Data Source
AI summary
In this data presentation method, first, feature values of multiple pieces of data 9 are calculated. Subsequently, a distance d between the pieces of data 9 in a feature value space S defined on the basis of the feature values is calculated. Then, the data 9 is selected from a dataset 90 on the basis of the calculated distance d, and is presented to a user. Thus, the data 9 can be presented to the user in the manner corresponding to the distance d between the pieces of data 9 in the feature value space S. Therefore, the user can efficiently perform a process of labelling each of the multiple pieces of data 9.


