Feature Vector Visualization for High-Quality Labeling Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for constructing learning data for machine learning models, particularly through crowdsourcing, fail to ensure data quality and often result in biased training due to the indiscriminate labeling of large amounts of raw data, leading to performance degradation.
Innovation Solution
A method and system for visualizing data using a feature embedding model to derive feature vectors, reduce dimensions to three or less, and display them on a plane for user selection, enabling the user to choose suitable data for labeling, thereby improving data quality and reducing training time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If crowdsourcing is used to process large amounts of raw data into learning data, then the time required to construct learning data is shortened, but the quality of the learning data deteriorates due to indiscriminate labeling
Solution Approach 1:
The patent applies preliminary action by performing feature extraction and dimensionality reduction on raw data before the labeling process. The service server derives feature vectors from raw data, reduces their dimensions to 3 or less, and visualizes them on a plane, allowing users to select high-quality data samples before crowdsourcing labeling begins. This preliminary filtering ensures that only suitable data is labeled, maintaining quality while improving efficiency.
2Quantity of substance
If all raw data is used to construct learning data, then the quantity of learning data is increased, but the time and cost for constructing learning data increase
Solution Approach 1:
The patent applies the extraction principle by selecting only representative and high-quality data samples from the raw dataset. Through feature vector derivation and dimensionality reduction, the system identifies and extracts data points that best represent the underlying patterns. Users can select a small subset of representative samples (e.g., 100 samples from 10,000 raw data points) for labeling, significantly reducing the time and cost while maintaining data quality.
3Loss of information
If feature vectors with high dimensions are used for visualization, then the information representation is comprehensive, but the ease of visual inspection and user selection deteriorates
Solution Approach 1:
The patent applies dimensionality change by reducing high-dimensional feature vectors to 3 or fewer dimensions through dimensionality reduction techniques. This transformation projects the data onto a 2D plane that can be easily visualized and inspected by users. The reduction maintains the essential structure and relationships in the data while making it visually accessible, allowing users to effectively identify and select representative samples for labeling.
Data Source
AI summary
The present invention relates to a method, a service server, and a computer-readable medium for visualizing data to select data to be used for labeling, and more particularly, to a method, a service server, and a computer-readable medium for visualizing data to select data to be used for labeling, capable of enabling a user to select data to be used as learning data by deriving a plurality of feature vector for a plurality of data included in an original dataset, reducing dimensions of each of the feature vectors to three or less dimensions, displaying the feature vectors with the reduced dimensions on a plane of the three or less dimensions, and providing a visualization interface including the plane of the three or less dimensions to a user terminal.


