Cross-Modal Speech Recognition Vocabulary Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech-to-text conversion technologies face challenges in accuracy due to speaker independence and large vocabulary sizes, especially in handheld devices, and are hindered by variations in speech patterns and equipment limitations.

Innovation Solution

A network device with a processor, speech synthesizer, and speech recognizer that converts text messages to audible messages and vice versa, using predefined keywords to limit the vocabulary and improve recognition accuracy in cross-modal communications between instant messaging and telephony users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speaker-independent ASR systems with large vocabularies are used, then adaptability to different users is improved, but recognition accuracy deteriorates

Engineering Contradiction:
Improveadaptability to different usersVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the vocabulary into two parts: a large general vocabulary for adaptability and a small context-specific vocabulary for accuracy. The system dynamically selects which vocabulary to use based on the communication context, allowing IM users to access both broad adaptability and precise recognition when needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic vocabulary selection where the system can switch between a large general vocabulary and a small context-specific vocabulary based on the communication scenario. This dynamic adjustment allows the system to optimize between adaptability and accuracy depending on the situation.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If predefined vocabularies are used in traditional ASR systems, then device complexity is reduced, but recognition accuracy deteriorates due to large variations in speech patterns

Engineering Contradiction:
Improvedevice complexityVSAvoidrecognition accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent changes the vocabulary parameter dynamically based on communication context. Instead of using a fixed predefined vocabulary, the system adjusts the vocabulary size and composition according to whether the user is communicating via IM or telephony, thereby optimizing recognition accuracy for each mode while maintaining manageable system complexity.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If single-speaker-dependent systems are used, then recognition accuracy is improved, but device complexity and training requirements increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidhardware and software requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a universal system that can serve both speaker-dependent and speaker-independent scenarios. The system can operate with a small context-specific vocabulary for high accuracy when needed, while also supporting larger vocabularies for general adaptability, eliminating the need for separate speaker-dependent training infrastructure.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7929672B2Constrained automatic speech recognition for more reliable speech-to-text conversion
Publication Date: 2011.04.19 CISCO TECHNOLOGY INC
  • US7929672B2 patent drawing
  • US7929672B2 patent drawing
  • US7929672B2 patent drawing

AI summary

A device and method are provided which preferably establish cross-modal communications and allow telephony users and text-based users, such as Instant Messaging (IM) users, to communicate with each other. The device may include a processor that receives a text message preferably comprising a query, a keyword, and one or more responses to the query. The processor preferably generates a vocabulary containing the one or more responses provided in the text message. The method preferably includes receiving a text message comprising a query, a keyword, and one or more responses to the query. The method may further include converting the text message into an audible message, sending the audible message to a telephony user, receiving an audible response from the telephony user, and generating text from the audible response. The method may further include generating a vocabulary comprising the one or more responses provided in the text message.