Focused Acoustic Models for Mobile Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
State-of-the-art automatic speech recognition systems require vast computational resources due to their high dimensionality, making them difficult to deploy on mobile platforms without compromising recognition accuracy.
Innovation Solution
The technique involves selecting observation-specific training data to create focused, low-dimensionality acoustic models tailored to each input speech signal, reducing computational requirements while maintaining recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If state-of-the-art acoustic models use numerous parameters to describe speech variations, then recognition accuracy is improved, but computational resource consumption increases
Solution Approach 1:
The patent segments the large global acoustic model into multiple smaller focused acoustic models, each tailored to specific observations or phoneme sequences. This segmentation allows the system to use only the necessary subset of parameters for each recognition task, maintaining accuracy while reducing computational resource consumption.
Solution Approach 2:
The patent extracts and removes unnecessary parameters from the global acoustic model by identifying and eliminating redundant Gaussian components and phoneme sequences that do not contribute to recognition accuracy. This extraction process reduces model dimensionality and computational requirements while preserving essential recognition capabilities.
2Measurement precision
If acoustic models are optimized for accuracy with high dimensionality, then recognition precision improves, but deployment on mobile platforms becomes difficult
Solution Approach 1:
The patent divides the comprehensive acoustic model into multiple smaller focused models that can be selectively deployed on mobile devices. Each focused model contains only the parameters necessary for specific recognition scenarios, making them suitable for mobile platforms with limited computational resources while maintaining recognition accuracy.
Solution Approach 2:
The patent changes the parameter configuration by selecting and retaining only the most relevant Gaussian components and phoneme sequences for each focused model. This parameter optimization reduces model size and computational complexity, enabling deployment on mobile platforms without significant compromise to recognition accuracy.
3Productivity
If focused acoustic models with reduced dimensionality are used, then computational efficiency improves, but recognition accuracy may be compromised
Solution Approach 1:
The patent applies local quality by creating focused acoustic models with specialized parameter configurations optimized for specific observations or phoneme sequences. Each model contains precisely the right parameters for its intended purpose, ensuring high recognition accuracy while maintaining computational efficiency through reduced dimensionality.
Solution Approach 2:
The patent performs preliminary analysis to identify and select the most relevant parameters and phoneme sequences before creating focused acoustic models. This preliminary action ensures that each focused model contains only the necessary parameters for accurate recognition, maintaining both computational efficiency and recognition accuracy.
Data Source
AI summary
Methods, systems, and computer-readable media related to selecting observation-specific training data (also referred to as “observation-specific exemplars”) from a general training corpus, and then creating, from the observation-specific training data, a focused, observation-specific acoustic model for recognizing the observation in an output domain are disclosed. In one aspect, a global speech recognition model is established based on an initial set of training data; a plurality of input speech segments to be recognized in an output domain are received; and for each of the plurality of input speech segments: a respective set of focused training data relevant to the input speech segment is identified in the global speech recognition model; a respective focused speech recognition model is generated based on the respective set of focused training data; and the respective focused speech recognition model is provided to a recognition device for recognizing the input speech segment in the output domain.


