Speech Recognition Learning System Speaker Bias Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech and speech-recognition systems face challenges in generating natural-sounding speech and accurately recognizing spoken commands, particularly when training data is limited or biased towards a single speaker's voice, leading to suboptimal performance in speech synthesis and recognition.

Innovation Solution

A speech interface learning system that generates pronunciation sequence conversion models by comparing and contrasting pronunciation sequences from speech inputs and text inputs, and adapts acoustic models using audio signal vectors to improve the synthesis and recognition processes, allowing for better performance across various speakers and voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If training data is limited or biased towards a single speaker's voice, then the system can be trained faster and with less data, but the speech synthesis quality and recognition accuracy deteriorate

Engineering Contradiction:
Improvetraining efficiencyVSAvoidspeech synthesis quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent creates virtual speech data by copying and transforming existing speech samples through pitch shifting, time stretching, and other signal processing techniques. This allows the system to generate diverse training data from limited original samples, improving speech synthesis quality without requiring additional real-world recording sessions.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms speech parameters such as pitch, tempo, and spectral characteristics to generate varied training examples from the same source material. By manipulating these parameters, the system creates artificial diversity in the training data, enabling better generalization across different speaking styles and conditions.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If training data is biased towards a single speaker's voice, then data collection is simpler and faster, but the system's adaptability to different speakers deteriorates

Engineering Contradiction:
Improvedata collection simplicityVSAvoidspeaker independence
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent introduces acoustic models and pronunciation dictionaries as intermediary components that decouple the training data from specific speaker characteristics. These intermediaries process and normalize speech data, enabling the system to learn speaker-independent representations that generalize across different voices while maintaining ease of data collection.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system designs training procedures and models that serve multiple functions: they work effectively for single-speaker training while simultaneously enabling multi-speaker adaptation. The acoustic models are trained to be universally applicable across different speakers, allowing the same training pipeline to serve both purposes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If the system focuses on speaker-specific characteristics, then it can achieve higher accuracy for that specific speaker, but performance on other speakers deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoidcross-speaker performance
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the speech recognition task into distinct components: acoustic modeling, pronunciation modeling, and language modeling. By separating speaker-specific features from speaker-independent features in this segmentation, the system can optimize for accuracy on specific speakers while maintaining generalizability through the modular architecture.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10089974B2Speech recognition and text-to-speech learning system
Publication Date: 2018.10.02 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10089974B2 patent drawing
  • US10089974B2 patent drawing
  • US10089974B2 patent drawing

AI summary

An example text-to-speech learning system performs a method for generating a pronunciation sequence conversion model. The method includes generating a first pronunciation sequence from a speech input of a training pair and generating a second pronunciation sequence from a text input of the training pair. The method also includes determining a pronunciation sequence difference between the first pronunciation sequence and the second pronunciation sequence; and generating a pronunciation sequence conversion model based on the pronunciation sequence difference. An example speech recognition learning system performs a method for generating a pronunciation sequence conversion model. The method includes extracting an audio signal vector from a speech input and applying an audio signal conversion model to the audio signal vector to generate a converted audio signal vector. The method also includes adapting an acoustic model based on the converted audio signal vector to generate an adapted acoustic model.