Prosodic Scripting for Audio-Video Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-speech and behavioral rendering systems fail to accurately capture the 'one-to-many' mapping of speech variations, resulting in lifeless and unengaged renderings, with limited communicative intent, and lack coordinated prosodic dynamics for both voice and face, leading to unrealistic behavioral representations.
Innovation Solution
A method for creating a prosodic script that identifies and annotates prosodic speech features and gestures in audio and video streams, converting them into temporally aligned symbols, and using these symbols to train systems for accurate rendering of audio and video streams, ensuring synchronized and meaningful prosodic expressions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If current text-to-speech systems attempt to create one rendering that covers all speech variations, then the system can handle diverse speech patterns, but the rendering becomes lifeless and lacks full communicative intent
Solution Approach 1:
The patent segments speech into discrete prosodic units (intonation contours, stress patterns, pause structures) that can be independently selected and combined. This allows the system to create varied speech renderings by mixing and matching specific prosodic features rather than generating entire speech patterns from scratch, thereby maintaining communicative intent while achieving diversity.
Solution Approach 2:
The patent applies different prosodic features to specific local regions of speech text rather than uniformly across the entire utterance. By identifying and applying intonation contours, stress patterns, and pause structures to specific segments, the system creates natural-sounding speech with proper communicative intent while allowing for variation in different parts of the utterance.
2Adaptability or versatility
If post-hoc changes are applied to spoken speech to address the one-to-many mapping problem, then speech variations can be achieved, but the solution becomes cumbersome and unreliable
Solution Approach 1:
The patent performs prosodic feature selection and annotation during the text-to-speech generation process itself, rather than applying changes after speech is generated. By pre-selecting and embedding the appropriate prosodic features into the speech synthesis pipeline, the system achieves speech variations without requiring complex post-processing operations.
3Ease of manufacture
If only text is provided as input, then the system can render speech, but the speech lacks real effect or intent and sounds like an unengaged voice actor
Solution Approach 1:
The patent introduces prosodic feature annotations as an intermediary layer between the text input and the speech output. These annotations serve as mediators that carry emotional and communicative intent information from the text to the speech synthesis engine, transforming plain text into emotionally engaged speech without requiring direct text-to-speech mapping.
4Adaptability or versatility
If facial expressions and head movements are completely divorced from coordinated prosodic intent, then face rendering can be generated, but the result shows tell-tale signs of fakery and disengages the listener
Solution Approach 1:
The patent merges prosodic feature annotations with facial animation control signals, creating a unified system where facial expressions are directly driven by the same prosodic features that control speech. This ensures that facial movements and speech prosody are coordinated and mutually reinforcing, eliminating the appearance of fakery while maintaining independent control over both modalities.
Data Source
AI summary
A method, apparatus, and system for creating a script for rendering audio and/or video streams include identifying at least one prosodic speech feature in a received audio stream and/or a received language model, creating a respective prosodic speech symbol for each of the at least one identified prosodic speech features, converting the received audio stream and/or the received language model into a text stream, temporally inserting the created at least one prosodic speech symbol into the text stream, identifying in a received video stream at least one prosodic gesture of at least a portion of a body of a speaker of the received audio stream, creating at least one respective gesture symbol for each of the at least one identified prosodic gestures, and temporally inserting the created at least one gesture symbol into the text stream along with the at least one prosodic speech symbol to create a prosodic script.


