Voice Mail Embedding in Spoken Utterance via NLP Parsing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems cannot effectively embed voice mail in a spoken utterance as a single request, combining a communication request and a message, which is essential for efficient voice mail delivery in modern communication systems.
Innovation Solution
A method in a computerized system that receives and records speech input, parses the utterance to separate the communication portion from the message portion, and transmits the voice mail to the designated destination, utilizing advanced natural language processing techniques such as incremental parsing and probability weights to accurately identify and process the communication parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a user sends voice mail through multiple separate steps using current speech recognition systems, then the system can process the communication, but the process becomes time-consuming and inefficient
Solution Approach 1:
The patent combines the communication request and message delivery into a single integrated speech utterance. The system parses the utterance to identify both the communicative act (e.g., 'send voice mail to') and the message content, allowing the user to complete the entire voice mail sending process in one step rather than multiple separate steps.
Solution Approach 2:
The system segments the spoken utterance into distinct components: the communication request portion and the message portion. This segmentation allows the system to accurately identify what action to perform and what content to transmit, enabling efficient single-step processing while maintaining clarity in function separation.
2Measurement precision
If the system uses basic speech recognition without advanced parsing, then the processing is simpler, but the system cannot accurately separate communication requests from message content
Solution Approach 1:
The system performs preliminary parsing of the speech utterance to identify the structure and components of the communication request before processing the message content. This preliminary analysis enables accurate separation of the communicative act from the message, improving precision without requiring complex real-time processing during message transmission.
Solution Approach 2:
The patent introduces an intermediate natural language processing layer that acts as a mediator between the speech input and the voice mail transmission system. This intermediary parses and interprets the utterance structure, enabling accurate identification of communication parameters and message content while managing the complexity of the overall system.
3Loss of information
If the system processes only text transcriptions, then the processing is straightforward, but prosody features such as tone and pauses are lost
Solution Approach 1:
The patent embeds the speech audio signal within the voice mail message structure, nesting the original audio content inside the transmitted message. This allows the system to preserve prosody features, tone, and pauses by maintaining the audio signal integrity while integrating it into the voice mail delivery system.
Solution Approach 2:
The system is designed to handle both text transcription and audio signal processing through a unified natural language processing framework. This multi-functional approach allows the system to process speech inputs while preserving audio characteristics, enabling flexible message delivery that can include both transcribed text and original audio with prosody features.
Data Source
AI summary
A method for processing a voice message in a computerized system. The method receives and records a speech utterance including a message portion and a communication portion. The method proceeds to parse the input to identify and separate the message portion and the communication portion. It then identifies communication parameters, including one or more destination mailboxes, from the communication portion, and it transmits the message portion to the destination mailbox as a voice message.


