Focused Acoustic Models for Mobile Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

State-of-the-art automatic speech recognition systems require vast computational resources due to their high dimensionality, making them difficult to deploy on mobile platforms without compromising recognition accuracy.

Innovation Solution

The technique involves selecting observation-specific training data to create focused, low-dimensionality acoustic models tailored to each input speech signal, reducing computational requirements while maintaining recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If state-of-the-art acoustic models use numerous parameters to describe speech variations, then recognition accuracy is improved, but computational resource consumption increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the large global acoustic model into multiple smaller focused acoustic models, each tailored to specific observations or phoneme sequences. This segmentation allows the system to use only the necessary subset of parameters for each recognition task, maintaining accuracy while reducing computational resource consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and removes unnecessary parameters from the global acoustic model by identifying and eliminating redundant Gaussian components and phoneme sequences that do not contribute to recognition accuracy. This extraction process reduces model dimensionality and computational requirements while preserving essential recognition capabilities.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If acoustic models are optimized for accuracy with high dimensionality, then recognition precision improves, but deployment on mobile platforms becomes difficult

Engineering Contradiction:
Improverecognition accuracyVSAvoiddeployability on mobile platforms
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent divides the comprehensive acoustic model into multiple smaller focused models that can be selectively deployed on mobile devices. Each focused model contains only the parameters necessary for specific recognition scenarios, making them suitable for mobile platforms with limited computational resources while maintaining recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter configuration by selecting and retaining only the most relevant Gaussian components and phoneme sequences for each focused model. This parameter optimization reduces model size and computational complexity, enabling deployment on mobile platforms without significant compromise to recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If focused acoustic models with reduced dimensionality are used, then computational efficiency improves, but recognition accuracy may be compromised

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by creating focused acoustic models with specialized parameter configurations optimized for specific observations or phoneme sequences. Each model contains precisely the right parameters for its intended purpose, ensuring high recognition accuracy while maintaining computational efficiency through reduced dimensionality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs preliminary analysis to identify and select the most relevant parameters and phoneme sequences before creating focused acoustic models. This preliminary action ensures that each focused model contains only the necessary parameters for accurate recognition, maintaining both computational efficiency and recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8935167B2Exemplar-based latent perceptual modeling for automatic speech recognition
Publication Date: 2015.01.13 APPLE INC
  • US8935167B2 patent drawing
  • US8935167B2 patent drawing
  • US8935167B2 patent drawing

AI summary

Methods, systems, and computer-readable media related to selecting observation-specific training data (also referred to as “observation-specific exemplars”) from a general training corpus, and then creating, from the observation-specific training data, a focused, observation-specific acoustic model for recognizing the observation in an output domain are disclosed. In one aspect, a global speech recognition model is established based on an initial set of training data; a plurality of input speech segments to be recognized in an output domain are received; and for each of the plurality of input speech segments: a respective set of focused training data relevant to the input speech segment is identified in the global speech recognition model; a respective focused speech recognition model is generated based on the respective set of focused training data; and the respective focused speech recognition model is provided to a recognition device for recognizing the input speech segment in the output domain.