Acoustic Model Generation via Voice Authentication Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Developing speech recognition systems for new linguistic markets is time-consuming and costly due to the need for large datasets of language and accent-specific speech samples, which are often unavailable.
Innovation Solution
A method for automatically configuring a speech recognition system by processing speech samples obtained during voice authentication to generate and update acoustic models, allowing for the creation of language and accent-specific models without the need for extensive data collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition systems are developed for new linguistic markets using traditional methods, then recognition accuracy can be achieved, but the process becomes time-consuming and costly due to the need for large datasets
Solution Approach 1:
The system uses speech samples naturally provided by users during voice authentication processes to automatically build and update acoustic models. This self-service approach eliminates the need for separate, time-consuming data collection campaigns, as the system configures itself using readily available user speech data
Solution Approach 2:
The speech samples collected for voice authentication serve dual purposes: both security verification and speech recognition model training. This multi-functionality allows the same data to address multiple system requirements simultaneously, reducing overall development time and resource needs
2Measurement precision
If speech recognition systems are developed for new linguistic markets using traditional methods, then recognition accuracy can be achieved, but the cost increases due to extensive data collection requirements
Solution Approach 1:
The system automatically configures its own acoustic models using speech data from authentication processes, eliminating the need for expensive external data collection services and manual processing, thereby significantly reducing development costs
Solution Approach 2:
By making the speech samples serve dual purposes (authentication and model training), the system eliminates redundant data collection activities, reducing overall development costs while maintaining recognition accuracy
3Measurement precision
If language and accent-specific acoustic models are created for each linguistic market, then speech recognition accuracy improves, but the complexity of system configuration increases
Solution Approach 1:
The system dynamically adapts its acoustic models based on the speech samples it receives during authentication. Rather than requiring static pre-configuration for each linguistic market, the system evolves its models automatically, reducing configuration complexity while maintaining accuracy
Solution Approach 2:
The system continuously updates acoustic models using feedback from authentication speech samples. This automated feedback loop allows the system to adapt to different linguistic markets without manual reconfiguration, simplifying the deployment process while preserving recognition accuracy
4Reliability
If extensive speech data collection is performed for each linguistic market, then acoustic model quality improves, but the productivity of system deployment decreases
Solution Approach 1:
The system performs automatic model configuration using authentication data, eliminating the need for slow, manual data collection and processing steps. This self-service approach maintains model quality while dramatically accelerating deployment speed
Solution Approach 2:
The system prepares acoustic models in advance using speech samples collected during normal authentication operations. By continuously pre-processing available data, the system ensures high-quality models are ready when needed, improving deployment speed without sacrificing model quality
Data Source
AI summary
A method for configuring a speech recognition system comprises obtaining a speech sample utilized by a voice authentication system in a voice authentication process. The speech sample is processed to generate acoustic models for units of speech associated with the speech sample. The acoustic models are stored for subsequent use by the speech recognition system as part of a speech recognition process.


