Domain-Specific Speech Recognizers Using Crowd-Sourced Acoustic Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition techniques in voice user interfaces (VUIs) often fail to accurately recognize spoken words, leading to user frustration and reduced interaction due to inaccuracy.
Innovation Solution
Domain-specific speech recognizers are generated using language data specific to an application interface, combined with crowd-sourced acoustic data to build a unified representation for improved speech recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speech recognition techniques are used in VUIs, then the system is simple and easy to implement, but the speech recognition accuracy deteriorates
Solution Approach 1:
The patent segments the general speech recognition problem into domain-specific sub-problems by creating separate speech recognizers for different application domains (e.g., medical, legal, technical). Each domain-specific recognizer is trained on specialized language data and acoustic data from that domain, allowing high accuracy within each domain while keeping individual recognizer complexity manageable.
Solution Approach 2:
The patent changes the parameters of speech recognition by adapting language models and acoustic models to domain-specific characteristics. This includes using domain-specific vocabulary, grammar rules, and acoustic patterns rather than general-purpose models, thereby improving accuracy for specialized terminology and speech patterns in each domain.
2Measurement precision
If domain-specific speech recognizers are generated using domain-specific language data and crowd-sourced acoustic data, then speech recognition accuracy is improved, but the complexity of data collection and model generation increases
Solution Approach 1:
The patent implements self-service by automatically collecting domain-specific language data from application logs and user interactions without manual annotation. The system autonomously gathers speech samples, processes them through crowd-sourcing platforms, and generates domain-specific acoustic models, reducing the need for manual data preparation and model training complexity.
Solution Approach 2:
The patent uses crowd-sourcing platforms as intermediaries to bridge the gap between domain experts and speech recognition system development. These platforms enable automated collection of domain-specific acoustic data from multiple speakers while maintaining quality control, thereby simplifying the data collection process and reducing the complexity of obtaining high-quality domain-specific speech samples.
3Adaptability or versatility
If conventional speech recognizers are used, then the system can handle general speech, but it fails to accurately recognize domain-specific words and phrases
Solution Approach 1:
The patent applies local quality by tailoring speech recognition models to specific domain requirements. Each domain-specific recognizer is customized with domain-specific vocabulary, terminology, and speech patterns, ensuring high accuracy for local domain terms while maintaining the ability to handle general speech through the underlying acoustic model.
Solution Approach 2:
The patent performs preliminary action by pre-training domain-specific language models and acoustic models with domain-specific data before deployment. This advance preparation ensures that the recognizers are already adapted to domain-specific terminology and speech patterns, improving accuracy from the outset rather than requiring post-deployment adjustments.
Data Source
AI summary
Domain-specific speech recognizer generation with crowd sourcing is described. The domain-specific speech recognizers are generated for voice user interfaces (VUIs) configured to replace or supplement application interfaces. In accordance with the described techniques, the speech recognizers are generated for a respective such application interface and are domain-specific because they are each generated based on language data that corresponds to the respective application interface. This domain-specific language data is used to build a domain-specific language model. The domain-specific language data is also used to collect acoustic data for building an acoustic model. In particular, the domain-specific language data is used to generate user interfaces that prompt crowd-sourcing participants to say selected words represented by the language data for recording. The recordings of these selected words are then used to build the acoustic model. The domain-specific speech recognizers are generated by combining a respective domain-specific language model and crowd-sourced acoustic model.


