Dimensionality Reduction for Semi-Supervised Learning Performance Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In semi-supervised learning, the lack of labeled data leads to performance deterioration compared to supervised learning, making it difficult to predict the extent of this deterioration and necessitating labeling of all data for performance improvement.

Innovation Solution

An information processing device and method that perform dimensionality reduction on learning data to generate a data distribution diagram, predicting learning performance based on this diagram and labeling status, and controlling the display of this information to minimize labeling burden.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If semi-supervised learning is used with unlabeled data, then labeling cost is reduced, but learning performance deteriorates

Engineering Contradiction:
Improvelabeling costVSAvoidlearning performance
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system performs preliminary dimensionality reduction and clustering analysis on the data distribution before actual learning. This preliminary action generates a data distribution diagram that predicts learning performance in advance, allowing users to understand the impact of using unlabeled data before committing to the learning process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

A data distribution diagram serves as an intermediary between the raw data and the learning process. This diagram visualizes cluster overlaps and data distribution characteristics, mediating the decision-making process about whether to use labeled or unlabeled data, thus resolving the contradiction between labeling cost and performance

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If all data is labeled for supervised learning, then learning performance is improved, but labeling burden increases

Engineering Contradiction:
Improvelearning performanceVSAvoidlabeling burden
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary dimensionality reduction and clustering analysis on the data distribution before actual learning. This preliminary action generates a data distribution diagram that predicts learning performance in advance, allowing users to understand the impact of using unlabeled data before committing to the learning process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of requiring complete labeling of all data, the system allows partial labeling and uses the data distribution diagram to predict whether this partial action will achieve satisfactory performance. This eliminates the need for excessive labeling while maintaining acceptable performance levels

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If dimensionality reduction is performed to generate data distribution diagram, then learning performance prediction capability is improved, but processing complexity increases

Engineering Contradiction:
Improvelearning performance prediction capabilityVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts only the essential features needed for performance prediction through dimensionality reduction. By transforming high-dimensional data into a lower-dimensional space that preserves cluster structure and distribution characteristics, it extracts the minimum necessary information for accurate prediction without unnecessary complexity

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11610148B2Information processing device and information processing method
Publication Date: 2023.03.21 SONY GROUP CORP
  • US11610148B2 patent drawing
  • US11610148B2 patent drawing
  • US11610148B2 patent drawing

AI summary

To previously predict learning performance in accordance with the labeling status of learning data. Provided is an information processing device including a data distribution presentation unit that performs dimensionality reduction on input learning data to generate a data distribution diagram related to the learning data, a learning performance prediction unit that predicts learning performance on the basis of the data distribution diagram and a labeling status related to the learning data, and a display control unit that controls a display related to the data distribution diagram and the learning performance. The data distribution diagram includes overlap information about clusters including the learning data and information about the number of pieces of the learning data belonging to each of the clusters.