Mixed Language Speech Recognition Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech-to-text systems struggle with accurately recognizing mixed language speech, particularly challenging languages like Cantonese and English, and often incur high costs in terms of financial, computational, and time resources.

Innovation Solution

A system and method that generate enhanced mixed language datasets using simulated conversations between speakers, and employ a fidelity-accuracy-latency (FAL) evaluator to train a speech recognition model, improving its performance in recognizing mixed language words.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing speech-to-text systems are used to recognize mixed language speech, then the system can process speech input, but the recognition accuracy deteriorates significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoidperformance consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system performs preliminary actions by generating synthetic mixed-language speech data before actual speech recognition training. The data generation module creates diverse mixed-language datasets with multiple language combinations and speaking styles, which are then used to pre-train the speech recognition model, improving its ability to handle real mixed-language speech accurately

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters by adjusting language mixing ratios, speaking styles, and data augmentation parameters during the synthetic data generation process. The training module dynamically adjusts learning parameters and data sampling ratios to optimize recognition accuracy for different mixed-language scenarios

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If high-quality mixed language training data is obtained through manual recording, then the data quality improves, but the time and cost resources increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses copying by generating synthetic speech data that mimics real human speech patterns. The data generation module creates artificial mixed-language speech samples that replicate the characteristics of actual speech without requiring manual recording, thus maintaining data quality while eliminating time-consuming data collection processes

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-service by automatically generating its own training data through the data generation module. The module synthesizes mixed-language speech data using language models and speech synthesis techniques, eliminating the need for external data collection efforts and reducing both time and cost resources

Inventive Principle:
Principle #25Self-service

3Ease of manufacture

If existing speech recognition models are trained on pure language data, then the training process is simple, but the model fails to recognize mixed language speech accurately

Engineering Contradiction:
Improvetraining simplicityVSAvoidmixed language recognition accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The system applies dynamics by making the training process adaptive and flexible. The training module dynamically adjusts the composition of training data, incorporating varying ratios of mixed-language samples based on the model's performance. The language mixing ratios and data sampling strategies are dynamically optimized during training to improve mixed-language recognition while maintaining training feasibility

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250131915A1Enhanced speech-to-text performance with mixed languages
Publication Date: 2025.04.24 THE HONG KONG UNIV OF SCI & TECH
  • US20250131915A1 patent drawing
  • US20250131915A1 patent drawing
  • US20250131915A1 patent drawing

AI summary

Training a mixed language speech recognition model, and speech recognition performed by the model, can be enhanced. Model manager can control training and/or operation of the model. Mixed language data generation manager can generate an enhanced mixed language dataset that can enhance training of the model. Fine tuner can facilitate configuring hyperparameters of the model. Audio-based information and transcript that are representative of the dataset can be applied to the model to facilitate model training. FAL evaluator can determine fidelity, accuracy, and latency of performance of speech recognition on the audio-based information and transcript by the model. Based on such determination, mixed language data generation process and/or hyperparameters can be updated to enhance further training of the model to enhance fidelity, accuracy, and/or latency regarding performance of speech recognition by the model. Model manager can control one or more iterations of model training.