User-Specific Acoustic Model Adaptation for Low-Resource Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Acoustic models intended for general use are computationally demanding and require significant memory, making them unsuitable for less capable electronic devices, which cannot perform speech recognition.
Innovation Solution
An electronic device receives speech inputs from a user, processes them using a user-independent acoustic model, and adjusts a user-specific acoustic model to enable speech recognition on devices with limited computational resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a user-independent acoustic model is used for general speech recognition, then speech recognition capability is achieved, but computational demand and memory requirements increase significantly
Solution Approach 1:
The acoustic model is segmented into two distinct components: a user-independent acoustic model that handles general speech recognition, and a user-specific acoustic model that adapts to individual users. This segmentation allows the system to use the lighter user-specific model for most operations while only invoking the more computationally intensive user-independent model when needed for initial adaptation or when user-specific data is insufficient.
Solution Approach 2:
The system applies local quality by making the acoustic model adaptive to specific users rather than uniformly general. The user-specific acoustic model is tailored to individual users' speech characteristics, allowing the system to achieve high recognition accuracy for each user with a computationally lightweight model, rather than using a heavy general model for all users.
2Adaptability or versatility
If a user-independent acoustic model is used for general speech recognition, then speech recognition capability is achieved, but memory requirements increase significantly
Solution Approach 1:
The acoustic model is segmented into two distinct components: a user-independent acoustic model that handles general speech recognition, and a user-specific acoustic model that adapts to individual users. This segmentation allows the system to use the lighter user-specific model for most operations while only invoking the more computationally intensive user-independent model when needed for initial adaptation or when user-specific data is insufficient.
Solution Approach 2:
The system creates a simplified copy of the acoustic model tailored to individual users. The user-specific acoustic model is a compressed, user-adapted version that captures essential speech patterns without requiring the full memory resources of the general model. This copying approach enables devices with limited memory to perform speech recognition.
3Device complexity
If a user-specific acoustic model is used, then computational resources are reduced, but the model requires adjustment and training data
Solution Approach 1:
The system performs preliminary action by automatically collecting and processing user speech data in the background to create the user-specific acoustic model. This adaptation process occurs proactively rather than requiring explicit user intervention or training sessions, reducing the perceived time loss while still achieving personalized recognition.
Solution Approach 2:
The system implements self-service by automatically adapting the user-specific acoustic model using speech data that the user naturally produces during normal device interaction. The system continuously learns and adjusts the model without requiring dedicated training time or user effort, effectively eliminating the time loss associated with model adjustment.
Data Source
AI summary
Systems and processes for providing user-specific acoustic models are provided. In accordance with one example, a method includes, at an electronic device having one or more processors, receiving a plurality of speech inputs, each of the speech inputs associated with a same user of the electronic device; providing each of the plurality of speech inputs to a user-independent acoustic model, the user-independent acoustic model providing a plurality of speech results based on the plurality of speech inputs; initiating a user-specific acoustic model on the electronic device; and adjusting the user-specific acoustic model based on the plurality of speech inputs and the plurality of speech results.


