Voice Generation Using Speaker Segmentation for Synced Dubbing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Converting audio/video files from one language to another language while ensuring that time nodes and voiceprint features align accurately is challenging, as existing methods lack precision and consistency.

Innovation Solution

Perform speaker segmentation to determine starting and ending times of speaking fragments, extract voiceprint feature vectors, convert text to target language, and generate target speech based on these features and times.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual translation and dubbing methods are used, then voiceprint features can be maintained, but the process requires a lot of manpower and time

Engineering Contradiction:
Improvevoiceprint feature similarityVSAvoidtranslation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual mechanical translation and dubbing operations with an automated speech processing system that uses voiceprint extraction, text translation, and speech synthesis to generate target language audio, thereby maintaining voiceprint similarity while dramatically improving translation efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service automatic translation and dubbing by automatically extracting voiceprint features from the original speech, translating the transcript text, and synthesizing target speech that matches the original speaker's voice characteristics without requiring manual intervention

Inventive Principle:
Principle #25Self-service

2Productivity

If automatic translation is used, then productivity is improved, but time node synchronization between original and target speech cannot be ensured

Engineering Contradiction:
Improvetranslation efficiencyVSAvoidtime node synchronization
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent performs preliminary action by extracting and storing the starting and ending times of each speech fragment from the original audio before translation, then uses these pre-established time markers to ensure the translated speech maintains accurate time node synchronization with the original

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If speech segmentation is performed to maintain time nodes, then time node synchronization is improved, but the process complexity increases

Engineering Contradiction:
Improvetime node synchronizationVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the original speech into discrete fragments with identified starting and ending times, processing each fragment individually through translation and synthesis, which maintains time node synchronization while managing complexity through systematic division of the processing task

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12620387B2Voice generation method and apparatus, device, and computer readable medium
Publication Date: 2026.05.05 DOUYIN VISION CO LTD
  • US12620387B2 patent drawing
  • US12620387B2 patent drawing
  • US12620387B2 patent drawing

AI summary

A voice generation method and apparatus, an electronic device, and a computer readable storage medium. Said method comprises: performing speaker segmentation on an original voice to determine starting time and ending time of each speaking voice segment in the original voice, so as to obtain segmented voices; determining a voiceprint feature vector corresponding to each speaking voice segment in the original voice; converting a text corresponding to each speaking voice segment in the original voice into a target language text, to obtain a target language text corresponding to each speaking voice segment in the original voice; and generating a target voice on the basis of the starting time and the ending time of each speaking voice segment in the original voice, the voiceprint feature vectors corresponding to the speaking voice segments and the target language texts corresponding to the speaking voice segments.