Context-Aware Dubbed Speech Adjustment by Scene Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for adjusting dubbed speech fail to consider the context of a scene, resulting in unnatural audio for media assets.
Innovation Solution
A media guidance application that decomposes media assets into scenes based on metadata, determines the scene type, and adjusts the intonation of dubbed speech to match the expected characteristics for that scene, using speech templates and vocal characteristics analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the dubbed speech is adjusted purely based on the original speech voice profile, then the voice characteristics are preserved, but the intonation becomes inconsistent with the scene context
Solution Approach 1:
The media asset is decomposed into multiple scenes based on metadata, allowing the system to process and adjust dubbed speech segment by segment. This segmentation enables context-aware intonation adjustment for each scene while preserving the overall voice profile consistency.
Solution Approach 2:
The system modifies intonation parameters of the dubbed speech based on the determined scene type. By changing pitch and tonal parameters according to scene context (e.g., action scenes require higher energy intonation), the system achieves both voice profile consistency and scene-appropriate delivery.
2Measurement precision
If the media asset is decomposed into multiple scenes for context analysis, then the intonation accuracy is improved, but the processing complexity increases
Solution Approach 1:
The media asset is pre-decomposed into scenes based on metadata before the intonation adjustment process begins. This preliminary segmentation allows the system to prepare scene contexts in advance, reducing real-time processing complexity while maintaining precise intonation measurement for each scene.
3Adaptability or versatility
If the dubbed speech is adjusted to match scene context, then the naturalness is improved, but the deviation from the original voice profile increases
Solution Approach 1:
The system applies different intonation characteristics to different scenes while maintaining the overall voice profile. Each scene receives localized intonation adjustment appropriate to its context, but the adjustments are made relative to the original voice profile, ensuring local adaptability without complete deviation from the actor's characteristic voice.
Data Source
AI summary
Systems and methods are disclosed herein for detecting dubbed speech in a media asset and receiving metadata corresponding to the media asset. The systems and methods may determine a plurality of scenes in the media asset based on the metadata, retrieve a portion of the dubbed speech corresponding to the first scene, and process the retrieved portion of the dubbed speech corresponding to the first scene to identify a speech characteristic of a character featured in the first scene. Further, the systems and methods may determine whether the speech characteristic of the character featured in the first scene matches the context of the first scene, and if the match fails, perform a function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene matches the context of the first scene.


