Feature Vector Visualization for High-Quality Labeling Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for constructing learning data for machine learning models, particularly through crowdsourcing, fail to ensure data quality and often result in biased training due to the indiscriminate labeling of large amounts of raw data, leading to performance degradation.

Innovation Solution

A method and system for visualizing data using a feature embedding model to derive feature vectors, reduce dimensions to three or less, and display them on a plane for user selection, enabling the user to choose suitable data for labeling, thereby improving data quality and reducing training time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If crowdsourcing is used to process large amounts of raw data into learning data, then the time required to construct learning data is shortened, but the quality of the learning data deteriorates due to indiscriminate labeling

Engineering Contradiction:
Improvetime required to construct learning dataVSAvoidquality of learning data
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by performing feature extraction and dimensionality reduction on raw data before the labeling process. The service server derives feature vectors from raw data, reduces their dimensions to 3 or less, and visualizes them on a plane, allowing users to select high-quality data samples before crowdsourcing labeling begins. This preliminary filtering ensures that only suitable data is labeled, maintaining quality while improving efficiency.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If all raw data is used to construct learning data, then the quantity of learning data is increased, but the time and cost for constructing learning data increase

Engineering Contradiction:
Improveamount of learning dataVSAvoidtime for constructing learning data
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent applies the extraction principle by selecting only representative and high-quality data samples from the raw dataset. Through feature vector derivation and dimensionality reduction, the system identifies and extracts data points that best represent the underlying patterns. Users can select a small subset of representative samples (e.g., 100 samples from 10,000 raw data points) for labeling, significantly reducing the time and cost while maintaining data quality.

Inventive Principle:
Principle #2Taking out (Extraction)

3Loss of information

If feature vectors with high dimensions are used for visualization, then the information representation is comprehensive, but the ease of visual inspection and user selection deteriorates

Engineering Contradiction:
Improveinformation representation completenessVSAvoidease of visual inspection
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent applies dimensionality change by reducing high-dimensional feature vectors to 3 or fewer dimensions through dimensionality reduction techniques. This transformation projects the data onto a 2D plane that can be easily visualized and inspected by users. The reduction maintains the essential structure and relationships in the data while making it visually accessible, allowing users to effectively identify and select representative samples for labeling.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12412325B2Method, service server, and computer-readable medium for visualizing data to select data to be used for labeling
Publication Date: 2025.09.09 SELECT STAR INC
  • US12412325B2 patent drawing
  • US12412325B2 patent drawing
  • US12412325B2 patent drawing

AI summary

The present invention relates to a method, a service server, and a computer-readable medium for visualizing data to select data to be used for labeling, and more particularly, to a method, a service server, and a computer-readable medium for visualizing data to select data to be used for labeling, capable of enabling a user to select data to be used as learning data by deriving a plurality of feature vector for a plurality of data included in an original dataset, reducing dimensions of each of the feature vectors to three or less dimensions, displaying the feature vectors with the reduced dimensions on a plane of the three or less dimensions, and providing a visualization interface including the plane of the three or less dimensions to a user terminal.