Semi-Supervised Dictionary Learning via Boundary-Based Loss Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current semi-supervised learning methods do not fully utilize unsupervised data for dictionary learning, leading to limited accuracy improvements.

Innovation Solution

An information processing device and method that includes a dictionary input circuit, boundary determination circuit, label assignment circuit, loss calculation circuit, and dictionary update circuit to iteratively refine the dictionary using supervised and unsupervised data, with the loss calculation circuit weighting the identification boundary to update the dictionary and improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional semi-supervised learning methods are used to learn dictionary from supervised and unsupervised data, then the identification accuracy is improved compared to using only supervised data, but not all unsupervised data samples are utilized for learning, limiting further accuracy improvement

Engineering Contradiction:
Improveidentification accuracyVSAvoidutilization of unsupervised data samples
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies local quality by differentiating the treatment of unsupervised data samples based on their spatial relationship to the identification boundary. Samples near the boundary are excluded from learning while those far from the boundary are actively used for dictionary learning. This selective approach based on local characteristics resolves the contradiction by utilizing more unsupervised data (increasing quantity) while maintaining accuracy through careful selection (preserving quality).

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the unsupervised data set into two distinct groups: samples near the identification boundary and samples far from it. This segmentation allows the system to process different subsets of data differently, using samples far from the boundary for learning while excluding those near the boundary. This resolves the contradiction by enabling broader data utilization without compromising the reliability of the learning process.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If unsupervised data near the identification boundary is used for learning, then more data samples are utilized, but the reliability of label assignment may deteriorate due to uncertain class membership

Engineering Contradiction:
Improvenumber of unsupervised data samples used for learningVSAvoidlabel assignment reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies local quality by differentiating the treatment of unsupervised data samples based on their spatial relationship to the identification boundary. Samples near the boundary are excluded from learning while those far from the boundary are actively used for dictionary learning. This selective approach based on local characteristics resolves the contradiction by utilizing more unsupervised data (increasing quantity) while maintaining accuracy through careful selection (preserving quality).

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent applies preliminary anti-action by proactively excluding unsupervised data samples near the identification boundary from the learning process before they can potentially introduce noise or incorrect labels. This preventive measure ensures that only reliable samples (those far from the boundary with clear class membership) are used for learning, thereby maintaining label assignment reliability while still utilizing a substantial portion of the unsupervised data.

Inventive Principle:
Principle #9Preliminary anti-action

3Adaptability or versatility

If a general-purpose semi-supervised learning method is developed that works with any identification device, then versatility is improved, but the complexity of the learning system increases compared to device-specific methods

Engineering Contradiction:
Improveapplicability to any identification deviceVSAvoidcomplexity of learning system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies universality by designing a semi-supervised learning method that operates independently of the specific identification device being used. The method takes any identification device's output (identification results and confidence levels) as input and applies a unified approach to determine which unsupervised samples to use for learning. This universal framework resolves the contradiction by enabling broad applicability across different devices while maintaining relatively simple system complexity through the use of a standardized, device-agnostic procedure.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11537930B2Information processing device, information processing method, and program
Publication Date: 2022.12.27 NEC CORP
  • US11537930B2 patent drawing
  • US11537930B2 patent drawing
  • US11537930B2 patent drawing

AI summary

An information processing device which performs semi-supervised learning is provided with: a dictionary input circuit for acquiring a dictionary, a parameter group used by an identification device; a boundary determination circuit which obtains an identification boundary for the dictionary on the basis of the dictionary, supervised data, and labelled unsupervised data; a labelling circuit which labels the unsupervised data in accordance with the identification boundary; a loss calculation circuit which calculates the sum total of supervised-data loss calculated in accordance with the labels assigned in advance and the labels based on the identification boundary, and unsupervised-data loss calculated such that further from the identification boundary the smaller the loss; a dictionary update circuit which updates the dictionary such that the sum-total loss is reduced; and a dictionary output circuit which outputs the updated dictionary.