Voice Dialog Interface Construction via Speech Recognition Library
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems for personal assistants are complex to program and require expertise, limiting user configuration and customization, especially in terms of call-signs and temporally dependent sentences, and lack voice feedback and intuitive interfaces.
Innovation Solution
A method and system for constructing a voice dialog interface that provides a library of programming interfaces to specify call-signs and commands in textual form, trains speech recognizers based on these inputs, recognizes speech inputs, and performs actions with verbal responses, simplifying the process and abstracting complexity to beginner-level programming.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If speech recognition systems use default settings and require extensive knowledge to be programmed, then the system can maintain simplicity in structure, but the ease of operation deteriorates because users need expert knowledge in phonetization, language models, and acoustic models
Solution Approach 1:
The patent introduces an intermediary layer (speech recognition library with simplified API) that mediates between the user and the complex speech recognition system. This library handles phonetization, language models, and acoustic models internally, allowing users to program without expert knowledge while maintaining system simplicity
Solution Approach 2:
The patent creates a simplified interface that copies or mimics the functionality of complex speech recognition systems at a higher level of abstraction. Users interact with simplified commands and configurations rather than implementing the full complexity of speech recognition algorithms
2Adaptability or versatility
If speech recognition systems allow configuration of understood sentences and call-signs, then the adaptability improves, but the device complexity worsens due to additional training and configuration requirements
Solution Approach 1:
The patent performs preliminary actions by pre-training acoustic models and language models during library initialization. Users can then configure call-signs and sentences without undergoing the complex training process themselves, as the system has already prepared the necessary models in advance
Solution Approach 2:
The patent segments the speech recognition system into distinct modular components: a pre-trained speech recognition library, a configuration interface for call-signs and sentences, and an action execution module. This segmentation allows users to configure only the necessary parts without managing the entire system's complexity
3Adaptability or versatility
If speech recognition systems support temporally dependent sentences and voice feedback, then the functionality improves, but the difficulty of detecting and measuring worsens due to increased programming complexity
Solution Approach 1:
The patent implements self-service by providing automated voice feedback through the speech synthesizer that responds to recognized speech. The system automatically handles the complexity of temporal dependency management and voice synthesis, allowing users to program high-level responses without managing the underlying temporal and acoustic complexity
Data Source
AI summary
Disclosed is a method of facilitating construction of a voice dialog interface for an electronic system. The method includes providing a library of programming interfaces configured to specify one or more of a call-sign and at least one command. Each of the call-sign and the at least one command may be specified in textual form. Additionally, the method includes training a speech recognizer based on one or more of the call-sign and the at least one command. Further, the method may include recognizing, using the speech recognizer, a speech input including a vocal representation of one or more of the call-sign and the at least one command. Additionally, the method includes performing at least one action associated with the at least one command based on recognizing the speech input. Further, the at least one action may include providing a verbal response using an integrated speech synthesizer.


