Pattern Recognition Apparatus Using Regularized PLDA Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing pattern recognition systems, such as discriminatively trained probabilistic linear discriminant analysis (DT-PLDA), rely heavily on labeled data and are prone to mislabeling errors, making them less robust and unable to effectively utilize unlabeled data, which is common in real applications.

Innovation Solution

A pattern recognition apparatus and method that calculates similarities among training data, updates PLDA parameters and labels using data statistics, and employs a regularization term to achieve global optimization, reducing dependence on unreliable labels and enabling the use of unlabeled data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If discriminatively trained PLDA (DT-PLDA) is used to improve classification performance, then discriminative power is improved, but robustness against domain mismatch deteriorates due to overfitting on labeled data

Engineering Contradiction:
Improveclassification accuracyVSAvoidrobustness against domain mismatch
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The method performs preliminary clustering on unlabeled data to generate initial labels before discriminative training, preparing the data in advance to enable utilization of unlabeled resources while maintaining system robustness

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The method introduces an intermediary regularization term based on data statistics that mediates between the discriminative training objective and the need for robustness, preventing overfitting while allowing utilization of unlabeled data

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If DT-PLDA is trained with large amounts of labeled data to improve performance, then classification performance is improved, but the ability to utilize unlabeled data deteriorates

Engineering Contradiction:
Improveclassification performanceVSAvoidability to use unlabeled data
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary clustering analysis on unlabeled data to generate initial labels, enabling these data to be incorporated into the training process before the main discriminative training phase

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The method creates a unified training framework that can simultaneously process both labeled and unlabeled data, making the system versatile enough to utilize all available data resources regardless of label availability

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If clustering is used to create labels for unlabeled data, then label availability is improved, but label accuracy deteriorates due to clustering errors

Engineering Contradiction:
Improvelabel availabilityVSAvoidlabel accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The method introduces a regularization term based on data statistics that acts as an intermediary constraint, preventing the model from overfitting to potentially erroneous cluster labels while still utilizing them

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system uses data statistics calculated from the training data as feedback to guide the training process, allowing the model to adjust and correct potential labeling errors through the regularization constraint

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11403545B2Pattern recognition apparatus, method, and program
Publication Date: 2022.08.02 NEC CORP
  • US11403545B2 patent drawing
  • US11403545B2 patent drawing
  • US11403545B2 patent drawing

AI summary

A pattern recognition apparatus for discriminative training includes: a similarity calculator that calculates similarities among training data; a statistics calculator that calculates statistics from the similarities in accordance with current labels for the training data; and a discriminative probabilistic linear discriminant analysis (PLDA) trainer that receives the training data, the statistics of the training data, the current labels and PLDA parameters, and updates the PLDA parameters and the labels of the training data.