Prosodic Mimic Speech Synthesis for Natural Confirmation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice-enabled communication systems, such as mobile telephones, struggle to produce natural-sounding synthesized speech for audible confirmation messages, which can be unintelligible due to the lack of incorporation of prosodic parameters, making them less effective for hands-free operations and user interface interactions.
Innovation Solution
A method and system that capture spoken utterances, extract both prosodic and non-prosodic parameters, and apply these parameters to synthesized speech to generate a prosodic mimic phrase, mimicking the original spoken voice, enhancing the naturalness and intelligibility of synthesized speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If synthesized speech is generated without prosodic parameters, then the speech synthesis process is simple and fast, but the synthesized speech becomes unintelligible and unnatural-sounding
Solution Approach 1:
The patent applies parameter changes by modifying the synthesized speech signal using prosodic parameters extracted from the original utterance. The system changes pitch, timing, and stress parameters of the synthesized speech to match the prosodic characteristics of the original spoken word, thereby improving intelligibility and naturalness without requiring complete re-synthesis of the entire speech signal.
Solution Approach 2:
The patent uses copying by extracting prosodic parameters from the original spoken utterance and applying them to a synthesized version of the same utterance. This creates a copy of the original speech that retains the semantic content while incorporating the prosodic qualities of the original, achieving naturalness without requiring direct replication of the entire speech signal.
2Manufacturing precision
If prosodic parameters are extracted and applied to synthesized speech, then the naturalness and intelligibility of synthesized speech improve, but the processing time and computational resources increase
Solution Approach 1:
The patent applies partial action by extracting only the essential prosodic parameters (pitch, timing, stress) from the original utterance rather than analyzing every aspect of the speech signal in detail. This selective extraction maintains the most important naturalness characteristics while reducing computational complexity and processing time compared to complete prosodic analysis.
Solution Approach 2:
The system performs preliminary extraction of prosodic parameters from the original utterance before generating the final synthesized speech. By preparing the prosodic parameter set in advance and using it to guide the synthesis process, the system reduces overall processing time compared to attempting to generate natural speech without pre-extracted prosodic information.
3Ease of operation
If audible confirmation messages use synthesized speech without prosody, then the system operation is simple, but the user interface effectiveness decreases for hands-free operations
Solution Approach 1:
The patent applies copying by capturing the prosodic characteristics from the user's spoken command and applying them to the audible confirmation message. This creates a confirmation message that mirrors the user's own speech patterns, making it more intelligible and effective for hands-free operation where the user needs to confirm actions without visual feedback.
Solution Approach 2:
The system provides feedback by generating audible confirmation messages that reflect back the prosodic qualities of the user's original command. This feedback mechanism enhances hands-free operation effectiveness by providing clear, natural-sounding confirmation that is easily distinguishable from machine-generated speech, allowing users to verify actions through auditory feedback alone.
Data Source
AI summary
A method and apparatus for synthesizing audible phrases (words) that includes capturing a spoken utterance, which may be a word, and extracting prosodic information (parameters) there from, then applying the prosodic parameters to a synthesized (nominal) word to produce a prosodic mimic word corresponding to the spoken utterance and the nominal word.


