Geographic Acoustic Model Adaptation for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated Speech Recognition (ASR) engines face difficulties in accurately recognizing spoken words due to variations in accents and dialects across different geographic regions, leading to misinterpretations such as 'park' being recognized as 'pork' or 'pack'.

Innovation Solution

The development of geographic-location specific acoustic models that are trained and adapted using geotagged audio signals, allowing the ASR engine to differentiate and accurately recognize speech patterns unique to specific areas by incorporating location metadata into the acoustic models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single universal acoustic model is used for speech recognition, then the system complexity is low, but the recognition accuracy deteriorates due to inability to handle regional accents and dialects

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidacoustic model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the universal acoustic model into multiple geographic-location specific acoustic models, each trained on audio data from specific regions. This segmentation allows the system to handle regional accents and dialects separately, improving recognition accuracy for each location while maintaining manageable model complexity through modular organization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by creating acoustic models with location-specific characteristics. Each geographic-location specific acoustic model is trained on audio data from its corresponding region, giving it specialized knowledge of local accents, dialects, and speech patterns. This local optimization improves recognition accuracy without requiring a complete redesign of the entire speech recognition system.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If geographic-location specific acoustic models are created for multiple regions, then the speech recognition accuracy improves, but the data processing complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies preliminary action by pre-training geographic-location specific acoustic models using geotagged audio data collected from multiple mobile devices in various locations. This preprocessing step creates ready-to-use location-specific models that can be quickly selected and applied during speech recognition operations, avoiding the need for complex real-time data processing and model adaptation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses geotagged audio data as an intermediary to bridge the gap between universal acoustic models and location-specific recognition needs. The geographic location metadata serves as a mediator that links audio data to specific regions, enabling the system to automatically select or adapt the appropriate acoustic model based on the user's location without requiring complex manual configuration.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If acoustic models are adapted using geotagged audio signals from multiple devices, then the regional speech patterns are captured more accurately, but the time and computational resources required for model adaptation increase

Engineering Contradiction:
Improveregional accent recognition accuracyVSAvoidmodel adaptation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent merges audio data from multiple mobile devices in the same geographic location to create comprehensive training sets for each region. By combining data from multiple sources, the system captures diverse speech patterns and accents more efficiently, improving model accuracy while sharing the computational burden across multiple data contributions rather than requiring extensive data collection from a single source.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system applies self-service by automatically collecting, processing, and adapting acoustic models using geotagged audio data from mobile devices. The automatic adaptation process eliminates the need for manual model training and adjustment, reducing the time and human resources required while still achieving accurate regional speech pattern recognition through automated machine learning processes.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8219384B2Acoustic model adaptation using geographic information
Publication Date: 2012.07.10 GOOGLE LLC
  • US8219384B2 patent drawing
  • US8219384B2 patent drawing
  • US8219384B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for enhancing speech recognition accuracy. In one aspect, a method includes receiving an audio signal that corresponds to an utterance recorded by a mobile device, determining a geographic location associated with the mobile device, adapting one or more acoustic models for the geographic location, and performing speech recognition on the audio signal using the one or more acoustic models model that are adapted for the geographic location.