Voice Generation Using Speaker Segmentation for Synced Dubbing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Converting audio/video files from one language to another language while ensuring that time nodes and voiceprint features align accurately is challenging, as existing methods lack precision and consistency.
Innovation Solution
Perform speaker segmentation to determine starting and ending times of speaking fragments, extract voiceprint feature vectors, convert text to target language, and generate target speech based on these features and times.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual translation and dubbing methods are used, then voiceprint features can be maintained, but the process requires a lot of manpower and time
Solution Approach 1:
The patent replaces manual mechanical translation and dubbing operations with an automated speech processing system that uses voiceprint extraction, text translation, and speech synthesis to generate target language audio, thereby maintaining voiceprint similarity while dramatically improving translation efficiency
Solution Approach 2:
The system enables self-service automatic translation and dubbing by automatically extracting voiceprint features from the original speech, translating the transcript text, and synthesizing target speech that matches the original speaker's voice characteristics without requiring manual intervention
2Productivity
If automatic translation is used, then productivity is improved, but time node synchronization between original and target speech cannot be ensured
Solution Approach 1:
The patent performs preliminary action by extracting and storing the starting and ending times of each speech fragment from the original audio before translation, then uses these pre-established time markers to ensure the translated speech maintains accurate time node synchronization with the original
3Manufacturing precision
If speech segmentation is performed to maintain time nodes, then time node synchronization is improved, but the process complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the original speech into discrete fragments with identified starting and ending times, processing each fragment individually through translation and synthesis, which maintains time node synchronization while managing complexity through systematic division of the processing task
Data Source
AI summary
A voice generation method and apparatus, an electronic device, and a computer readable storage medium. Said method comprises: performing speaker segmentation on an original voice to determine starting time and ending time of each speaking voice segment in the original voice, so as to obtain segmented voices; determining a voiceprint feature vector corresponding to each speaking voice segment in the original voice; converting a text corresponding to each speaking voice segment in the original voice into a target language text, to obtain a target language text corresponding to each speaking voice segment in the original voice; and generating a target voice on the basis of the starting time and the ending time of each speaking voice segment in the original voice, the voiceprint feature vectors corresponding to the speaking voice segments and the target language texts corresponding to the speaking voice segments.


