Semi-Supervised Dictionary Learning via Boundary-Based Loss Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current semi-supervised learning methods do not fully utilize unsupervised data for dictionary learning, leading to limited accuracy improvements.
Innovation Solution
An information processing device and method that includes a dictionary input circuit, boundary determination circuit, label assignment circuit, loss calculation circuit, and dictionary update circuit to iteratively refine the dictionary using supervised and unsupervised data, with the loss calculation circuit weighting the identification boundary to update the dictionary and improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional semi-supervised learning methods are used to learn dictionary from supervised and unsupervised data, then the identification accuracy is improved compared to using only supervised data, but not all unsupervised data samples are utilized for learning, limiting further accuracy improvement
Solution Approach 1:
The patent applies local quality by differentiating the treatment of unsupervised data samples based on their spatial relationship to the identification boundary. Samples near the boundary are excluded from learning while those far from the boundary are actively used for dictionary learning. This selective approach based on local characteristics resolves the contradiction by utilizing more unsupervised data (increasing quantity) while maintaining accuracy through careful selection (preserving quality).
Solution Approach 2:
The patent segments the unsupervised data set into two distinct groups: samples near the identification boundary and samples far from it. This segmentation allows the system to process different subsets of data differently, using samples far from the boundary for learning while excluding those near the boundary. This resolves the contradiction by enabling broader data utilization without compromising the reliability of the learning process.
2Quantity of substance
If unsupervised data near the identification boundary is used for learning, then more data samples are utilized, but the reliability of label assignment may deteriorate due to uncertain class membership
Solution Approach 1:
The patent applies local quality by differentiating the treatment of unsupervised data samples based on their spatial relationship to the identification boundary. Samples near the boundary are excluded from learning while those far from the boundary are actively used for dictionary learning. This selective approach based on local characteristics resolves the contradiction by utilizing more unsupervised data (increasing quantity) while maintaining accuracy through careful selection (preserving quality).
Solution Approach 2:
The patent applies preliminary anti-action by proactively excluding unsupervised data samples near the identification boundary from the learning process before they can potentially introduce noise or incorrect labels. This preventive measure ensures that only reliable samples (those far from the boundary with clear class membership) are used for learning, thereby maintaining label assignment reliability while still utilizing a substantial portion of the unsupervised data.
3Adaptability or versatility
If a general-purpose semi-supervised learning method is developed that works with any identification device, then versatility is improved, but the complexity of the learning system increases compared to device-specific methods
Solution Approach 1:
The patent applies universality by designing a semi-supervised learning method that operates independently of the specific identification device being used. The method takes any identification device's output (identification results and confidence levels) as input and applies a unified approach to determine which unsupervised samples to use for learning. This universal framework resolves the contradiction by enabling broad applicability across different devices while maintaining relatively simple system complexity through the use of a standardized, device-agnostic procedure.
Data Source
AI summary
An information processing device which performs semi-supervised learning is provided with: a dictionary input circuit for acquiring a dictionary, a parameter group used by an identification device; a boundary determination circuit which obtains an identification boundary for the dictionary on the basis of the dictionary, supervised data, and labelled unsupervised data; a labelling circuit which labels the unsupervised data in accordance with the identification boundary; a loss calculation circuit which calculates the sum total of supervised-data loss calculated in accordance with the labels assigned in advance and the labels based on the identification boundary, and unsupervised-data loss calculated such that further from the identification boundary the smaller the loss; a dictionary update circuit which updates the dictionary such that the sum-total loss is reduced; and a dictionary output circuit which outputs the updated dictionary.


