Automatic Dubbing via Voice Print Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The traditional dubbing process for media content is costly and time-consuming, making it impractical for rapidly updated content with lower budgets, as it requires hiring actors to record multiple audio versions.
Innovation Solution
An automatic dubbing method that extracts speeches from media content, generates a voice print model for user voices, and processes these to replace original voices, allowing users to customize audio in real-time using a user device, such as a media player, with modules for speech extraction, voice print model generation, and speech processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional dubbing process with hired actors is used, then audio quality and customization are improved, but cost and time consumption increase significantly
Solution Approach 1:
The patent replaces the mechanical system of human actors recording audio with an automated voice synthesis system. The speech extracting module captures original speech, the voice print model obtaining module creates voice models, and the speech processing module generates replacement speeches automatically, eliminating the need for human dubbing actors and significantly reducing time and cost while maintaining audio quality.
Solution Approach 2:
The system enables self-service dubbing where the media content itself provides the basis for its own dubbing. The speech extracting module extracts speeches from the media content, and the speech processing module uses these extracted speeches along with voice print models to generate replacement speeches, allowing the system to serve itself without external human intervention.
2Adaptability or versatility
If traditional dubbing process is used, then customized audio versions are achieved, but resource requirements and cost increase
Solution Approach 1:
The patent creates voice print models as digital copies of voice characteristics. The voice print model obtaining module generates these models that can be reused across multiple media contents, allowing the same voice model to be copied and applied to different content without requiring new recording sessions, thereby reducing resource requirements while maintaining customization capability.
Solution Approach 2:
The voice print models serve multiple functions and can be applied to various media contents universally. Once a voice print model is created, it can be used for dubbing different media contents with similar voice requirements, making the system versatile and reducing the need for separate resources for each customization task.
3Productivity
If automated voice processing is implemented, then cost and time efficiency are improved, but voice naturalness and quality may deteriorate
Solution Approach 1:
The speech processing module uses feedback mechanisms where the extracted speeches serve as reference input for generating replacement speeches. The system analyzes the characteristics of the original extracted speeches and uses this feedback to guide the voice synthesis process, ensuring that the generated replacement speeches maintain naturalness and quality while achieving automated processing efficiency.
Data Source
AI summary
A method and system for automatic dubbing method is disclosed, comprising, responsive to receiving a selection of media content for playback on a user device by a user of the user device, processing extracted speeches of a first voice from the media content to generate replacement speeches using a set of phenomes of a second voice of the user of the user device, and replacing the extracted speeches of the first voice with the generated replacement speeches in the audio portion of the media content for playback on the user device.


