Speech Recognition Learning System Speaker Bias Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-speech and speech-recognition systems face challenges in generating natural-sounding speech and accurately recognizing spoken commands, particularly when training data is limited or biased towards a single speaker's voice, leading to suboptimal performance in speech synthesis and recognition.
Innovation Solution
A speech interface learning system that generates pronunciation sequence conversion models by comparing and contrasting pronunciation sequences from speech inputs and text inputs, and adapts acoustic models using audio signal vectors to improve the synthesis and recognition processes, allowing for better performance across various speakers and voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If training data is limited or biased towards a single speaker's voice, then the system can be trained faster and with less data, but the speech synthesis quality and recognition accuracy deteriorate
Solution Approach 1:
The patent creates virtual speech data by copying and transforming existing speech samples through pitch shifting, time stretching, and other signal processing techniques. This allows the system to generate diverse training data from limited original samples, improving speech synthesis quality without requiring additional real-world recording sessions.
Solution Approach 2:
The system transforms speech parameters such as pitch, tempo, and spectral characteristics to generate varied training examples from the same source material. By manipulating these parameters, the system creates artificial diversity in the training data, enabling better generalization across different speaking styles and conditions.
2Ease of manufacture
If training data is biased towards a single speaker's voice, then data collection is simpler and faster, but the system's adaptability to different speakers deteriorates
Solution Approach 1:
The patent introduces acoustic models and pronunciation dictionaries as intermediary components that decouple the training data from specific speaker characteristics. These intermediaries process and normalize speech data, enabling the system to learn speaker-independent representations that generalize across different voices while maintaining ease of data collection.
Solution Approach 2:
The system designs training procedures and models that serve multiple functions: they work effectively for single-speaker training while simultaneously enabling multi-speaker adaptation. The acoustic models are trained to be universally applicable across different speakers, allowing the same training pipeline to serve both purposes.
3Measurement precision
If the system focuses on speaker-specific characteristics, then it can achieve higher accuracy for that specific speaker, but performance on other speakers deteriorates
Solution Approach 1:
The patent segments the speech recognition task into distinct components: acoustic modeling, pronunciation modeling, and language modeling. By separating speaker-specific features from speaker-independent features in this segmentation, the system can optimize for accuracy on specific speakers while maintaining generalizability through the modular architecture.
Data Source
AI summary
An example text-to-speech learning system performs a method for generating a pronunciation sequence conversion model. The method includes generating a first pronunciation sequence from a speech input of a training pair and generating a second pronunciation sequence from a text input of the training pair. The method also includes determining a pronunciation sequence difference between the first pronunciation sequence and the second pronunciation sequence; and generating a pronunciation sequence conversion model based on the pronunciation sequence difference. An example speech recognition learning system performs a method for generating a pronunciation sequence conversion model. The method includes extracting an audio signal vector from a speech input and applying an audio signal conversion model to the audio signal vector to generate a converted audio signal vector. The method also includes adapting an acoustic model based on the converted audio signal vector to generate an adapted acoustic model.


