Prosodic Scripting for Audio-Video Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech and behavioral rendering systems fail to accurately capture the 'one-to-many' mapping of speech variations, resulting in lifeless and unengaged renderings, with limited communicative intent, and lack coordinated prosodic dynamics for both voice and face, leading to unrealistic behavioral representations.

Innovation Solution

A method for creating a prosodic script that identifies and annotates prosodic speech features and gestures in audio and video streams, converting them into temporally aligned symbols, and using these symbols to train systems for accurate rendering of audio and video streams, ensuring synchronized and meaningful prosodic expressions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If current text-to-speech systems attempt to create one rendering that covers all speech variations, then the system can handle diverse speech patterns, but the rendering becomes lifeless and lacks full communicative intent

Engineering Contradiction:
Improvespeech variation coverageVSAvoidcommunicative intent accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments speech into discrete prosodic units (intonation contours, stress patterns, pause structures) that can be independently selected and combined. This allows the system to create varied speech renderings by mixing and matching specific prosodic features rather than generating entire speech patterns from scratch, thereby maintaining communicative intent while achieving diversity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different prosodic features to specific local regions of speech text rather than uniformly across the entire utterance. By identifying and applying intonation contours, stress patterns, and pause structures to specific segments, the system creates natural-sounding speech with proper communicative intent while allowing for variation in different parts of the utterance.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If post-hoc changes are applied to spoken speech to address the one-to-many mapping problem, then speech variations can be achieved, but the solution becomes cumbersome and unreliable

Engineering Contradiction:
Improvespeech variation capabilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs prosodic feature selection and annotation during the text-to-speech generation process itself, rather than applying changes after speech is generated. By pre-selecting and embedding the appropriate prosodic features into the speech synthesis pipeline, the system achieves speech variations without requiring complex post-processing operations.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If only text is provided as input, then the system can render speech, but the speech lacks real effect or intent and sounds like an unengaged voice actor

Engineering Contradiction:
Improveinput simplicityVSAvoidspeech engagement quality
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent introduces prosodic feature annotations as an intermediary layer between the text input and the speech output. These annotations serve as mediators that carry emotional and communicative intent information from the text to the speech synthesis engine, transforming plain text into emotionally engaged speech without requiring direct text-to-speech mapping.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If facial expressions and head movements are completely divorced from coordinated prosodic intent, then face rendering can be generated, but the result shows tell-tale signs of fakery and disengages the listener

Engineering Contradiction:
Improvefacial animation capabilityVSAvoidbehavioral realism
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent merges prosodic feature annotations with facial animation control signals, creating a unified system where facial expressions are directly driven by the same prosodic features that control speech. This ensures that facial movements and speech prosody are coordinated and mutually reinforcing, eliminating the appearance of fakery while maintaining independent control over both modalities.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240257801A1Method and system for creating a prosodic script
Publication Date: 2024.08.01 SRI INTERNATIONAL
  • US20240257801A1 patent drawing
  • US20240257801A1 patent drawing
  • US20240257801A1 patent drawing

AI summary

A method, apparatus, and system for creating a script for rendering audio and/or video streams include identifying at least one prosodic speech feature in a received audio stream and/or a received language model, creating a respective prosodic speech symbol for each of the at least one identified prosodic speech features, converting the received audio stream and/or the received language model into a text stream, temporally inserting the created at least one prosodic speech symbol into the text stream, identifying in a received video stream at least one prosodic gesture of at least a portion of a body of a speaker of the received audio stream, creating at least one respective gesture symbol for each of the at least one identified prosodic gestures, and temporally inserting the created at least one gesture symbol into the text stream along with the at least one prosodic speech symbol to create a prosodic script.