Cross-Modal Speech Recognition Vocabulary Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech-to-text conversion technologies face challenges in accuracy due to speaker independence and large vocabulary sizes, especially in handheld devices, and are hindered by variations in speech patterns and equipment limitations.
Innovation Solution
A network device with a processor, speech synthesizer, and speech recognizer that converts text messages to audible messages and vice versa, using predefined keywords to limit the vocabulary and improve recognition accuracy in cross-modal communications between instant messaging and telephony users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speaker-independent ASR systems with large vocabularies are used, then adaptability to different users is improved, but recognition accuracy deteriorates
Solution Approach 1:
The patent segments the vocabulary into two parts: a large general vocabulary for adaptability and a small context-specific vocabulary for accuracy. The system dynamically selects which vocabulary to use based on the communication context, allowing IM users to access both broad adaptability and precise recognition when needed.
Solution Approach 2:
The patent implements dynamic vocabulary selection where the system can switch between a large general vocabulary and a small context-specific vocabulary based on the communication scenario. This dynamic adjustment allows the system to optimize between adaptability and accuracy depending on the situation.
2Device complexity
If predefined vocabularies are used in traditional ASR systems, then device complexity is reduced, but recognition accuracy deteriorates due to large variations in speech patterns
Solution Approach 1:
The patent changes the vocabulary parameter dynamically based on communication context. Instead of using a fixed predefined vocabulary, the system adjusts the vocabulary size and composition according to whether the user is communicating via IM or telephony, thereby optimizing recognition accuracy for each mode while maintaining manageable system complexity.
3Reliability
If single-speaker-dependent systems are used, then recognition accuracy is improved, but device complexity and training requirements increase
Solution Approach 1:
The patent creates a universal system that can serve both speaker-dependent and speaker-independent scenarios. The system can operate with a small context-specific vocabulary for high accuracy when needed, while also supporting larger vocabularies for general adaptability, eliminating the need for separate speaker-dependent training infrastructure.
Data Source
AI summary
A device and method are provided which preferably establish cross-modal communications and allow telephony users and text-based users, such as Instant Messaging (IM) users, to communicate with each other. The device may include a processor that receives a text message preferably comprising a query, a keyword, and one or more responses to the query. The processor preferably generates a vocabulary containing the one or more responses provided in the text message. The method preferably includes receiving a text message comprising a query, a keyword, and one or more responses to the query. The method may further include converting the text message into an audible message, sending the audible message to a telephony user, receiving an audible response from the telephony user, and generating text from the audible response. The method may further include generating a vocabulary comprising the one or more responses provided in the text message.


