Speech Model Pre-Training With Rule-Based Feature Labels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition technologies face challenges in achieving high accuracy with limited transcription data, as semi-supervised learning methods using erroneous transcriptions hinder performance, and self-supervised learning methods struggle with data distribution imbalances when applying pre-trained models to new speech data.

Innovation Solution

A pre-trained model is developed using self-supervised labels derived from feature vectors of speech data through a deterministic conversion rule, excluding statistical inference, and further trained with a small amount of transcribed text to enhance speech processing tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If self-supervised learning methods are used to pre-train models, then the model can be trained without transcription data, but the model struggles with data distribution imbalances when applying to new speech data

Engineering Contradiction:
Improveamount of transcription dataVSAvoidmodel performance on new speech data
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-training the model using self-supervised learning on feature vectors extracted from speech data before fine-tuning with limited transcription data. This preliminary pre-training phase prepares the model to handle various speech patterns and distributions, improving its reliability when applied to new speech data while requiring minimal transcription data.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If semi-supervised learning methods using erroneous transcriptions are used, then the model can be trained with more data, but the erroneous transcriptions hinder performance

Engineering Contradiction:
Improveamount of training dataVSAvoidspeech processing accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent extracts and removes erroneous transcription data from the training set, using only high-quality transcribed text data for fine-tuning. By taking out the harmful erroneous transcriptions while retaining useful clean data, the model achieves high speech processing accuracy without being degraded by poor quality training examples.

Inventive Principle:
Principle #2Taking out (Extraction)

3Manufacturing precision

If more transcribed text data is used for training, then speech processing accuracy improves, but the cost and time for data preparation increases

Engineering Contradiction:
Improvespeech processing accuracyVSAvoiddata preparation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary self-supervised pre-training on large amounts of unlabeled speech data, extracting feature vectors and training the model without requiring transcription. This preliminary action reduces the need for extensive transcribed text data later, significantly reducing data preparation time and cost while still achieving high speech processing accuracy through subsequent fine-tuning with minimal clean data.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260065909A1Speech processing apparatus, speech processing method, and storage medium
Publication Date: 2026.03.05 RICOH CO LTD
  • US20260065909A1 patent drawing
  • US20260065909A1 patent drawing
  • US20260065909A1 patent drawing

AI summary

A speech processing apparatus processing circuitry. The processing circuitry executes a task related to speech processing based on a trained model. The trained model is trained using speech data and one or more labels obtained by converting, according to a predetermined rule, one or more feature vectors extracted from the speech data.