Movie human voice dubbing method based on multi-modal thinking chain

By constructing a multimodal thinking chain of film vocal dubbing method, the problem of insufficient visual semantic capture in the existing technology is solved, and the lip synchronization accuracy and emotional similarity are improved. It is suitable for multiple speech synthesis tasks and is easy to integrate.

CN120340470APending Publication Date: 2025-07-18GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510475096.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing film vocal dubbing technology relies on single-modal text input and cannot effectively capture the visual semantics of video scenes. Traditional solutions require manual acquisition of visual features in advance, making it difficult to achieve end-to-end generation, and existing models are difficult to meet the industrial requirements of lip synchronization accuracy and emotional similarity.

Method used

Build a movie dubbing dataset with CoT annotation, integrate multilingual voice library and animation dataset, train TTS and V2S modules, optimize multimodal video comprehension and speech generation models, integrate visual comprehension and speech generation through multi-stage training strategies, and output high-quality synthetic speech.

Benefits of technology

It achieves improved lip synchronous accuracy and emotional similarity, improves the model's performance on low-quality data sets, enhances the stability and scope of dubbing synthesis, and is easy to integrate with existing tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340470A_ABST
    Figure CN120340470A_ABST
Patent Text Reader

Abstract

The invention relates to a movie human voice dubbing method based on a multi-modal thinking chain. The method comprises the following steps: constructing a movie dubbing data set with CoT labels; a multi-language voice library, an animation data set and a multi-speaker data set are integrated, and a TTS voice synthesis module and a V2S video dubbing module are trained; data with noise and unclear semantics are removed; a multi-modal video understanding model and a voice generation model are trained, model parameters are optimized, and the generalization ability of the models is improved; and performing a dubbing task by using the trained model, and outputting high-quality synthetic speech. According to the invention, lip shape synchronization precision improvement and emotion similarity improvement can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of movie voice dubbing, and particularly relates to a movie voice dubbing method based on a multi-modal thought chain. Background Art

[0002] Existing movie voice dubbing technology (movie dubber) has the following defects:

[0003] (1) Existing movie voice dubbing technology relies on unimodal text input and cannot effectively capture the visual semantics of video scenes (such as character gender and age, scene mood and atmosphere, etc.);

[0004] (2) Traditional solutions require manual acquisition of visual features in advance, making it difficult to generate end-to-end, which is very inconvenient;

[0005] (3) Current models (such as HPMDubbing, StyleDubber) are based on fusing face and lip frames in the video with a small speech synthesis model, and it is difficult to meet the requirements of the industrial community.

[0006] Therefore, it is necessary to provide a movie voice dubbing method based on a multi-modal thought chain to improve the lip synchronization accuracy and emotional similarity. Summary of the Invention

[0007] The purpose of the present invention is to provide a movie voice dubbing method based on a multi-modal thought chain to improve the lip synchronization accuracy and emotional similarity.

[0008] To solve the problems existing in the prior art, the present invention provides a movie voice dubbing method based on a multi-modal thought chain, including the following steps:

[0009] Construct a movie dubbing dataset with CoT annotation;

[0010] Integrate multi-language speech databases, animation datasets, and multi-speaker datasets, and train the TTS speech synthesis module and the V2S video dubbing module;

[0011] Remove data with noise and unclear semantics;

[0012] Train a multi-modal video understanding model and a speech generation model, optimize the model parameters, and improve the generalization ability of the model;

[0013] Use the trained model to perform a dubbing task and output high-quality synthesized speech.

[0014] Optionally, in the movie voice dubbing method based on a multi-modal thought chain, the CoT annotation rule is a rule system for annotating text data.

[0015] Optionally, in the movie voice dubbing method based on the multi-modal thought chain, TTS stands for Text To Speech, which is used to automatically convert text into speech; V2S is a high-definition digital input and audio-video synchronous switching video processor.

[0016] Optionally, in the movie voice dubbing method based on the multi-modal thought chain, the multi-modal video understanding model and the speech generation model include a multi-modal reasoning module, a conditional flow matching speech generator, and a dynamic conditional control mechanism.

[0017] Compared with the prior art, the present invention has the following advantages:

[0018] (1) By fusing visual understanding and speech generation through a multi-stage training strategy, it can truly achieve an improvement in lip-sync accuracy and emotional similarity like a real human voice dubbing.

[0019] (2) On the V2C-Animation dataset, the speaker similarity of SPK-SIM reaches 89.74% (a 0.44% improvement compared to F5-TTS), and the emotional similarity of EMO-SIM is 78.88% (a 12.1% improvement).

[0020] (3) The lip-sync error of the GRID dataset LSE-D is reduced to 14.63 (optimized by 1.03); through the hybrid preference optimization strategy, the multi-modal reasoning accuracy is improved by 8.7%.

[0021] (4) Improve the model robustness: enhance the performance of the model on low-quality datasets and improve the stability and reliability of voice dubbing synthesis.

[0022] (5) Wide range of applications: applicable to various speech synthesis and voice dubbing synthesis tasks.

[0023] (6) Easy to integrate: The present invention can be seamlessly integrated with existing voice dubbing tools to enhance the overall performance of the existing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flowchart of the movie voice dubbing method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The following will describe the specific embodiments of the present invention in more detail with reference to the schematic diagrams. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the drawings are all in a very simplified form and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the objectives of the embodiments of the present invention.

[0026] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present application.

[0027] The existing movie dubber technology has the following defects: (1) The existing movie dubber technology relies on single-modal text input and cannot effectively capture the visual semantics of video scenes (such as character gender and age, scene mood and atmosphere, etc.); (2) Traditional solutions require manual acquisition of visual features in advance and are difficult to generate end-to-end, which brings many inconveniences; (3) Current models (such as HPMDubbing and StyleDubber) fuse the face and lip frames in the video with a small speech synthesis model, making it difficult to meet the requirements of the industrial community.

[0028] To solve the problems existing in the prior art, the present invention provides a movie dubbing method based on a multi-modal chain of thought, as Figure 1 shown, the method includes the following steps:

[0029] S1: Collect and preprocess training data, and use models such as VAD and ASR to segment and preliminarily label the collected data (the text content of the speech).

[0030] Specifically, construct a movie dubbing dataset with 7.2 hours of CoT annotation; the CoT annotation rule is a rule system for annotating text data, where CoT is the abbreviation of "Chain-of-Thought". The purpose of the CoT annotation rule is to help annotators accurately annotate text data according to the text context and context, avoiding ambiguity and misjudgment.

[0031] The movie dubbing dataset contains <summary> <reasoning> <conclusion>Structured annotations such as these are used to label the video scene types (dialogue / narration / monologue) and the characteristics and speech styles of the characters in the video by a professional team, for the training of the MLLM model.

[0032] S2: Integrate the EMILIA multilingual speech library (101,654 hours), the V2C-Animation animation dataset (10,217 segments), and the GRID multi-speaker dataset (33 people × 1,000 samples) to train the TTS speech synthesis module and the V2S video dubbing module. Among them, the full name of TTS is TextTo Speech, which is used to automatically convert text into speech; V2S is a high-definition digital input and audio-video synchronous switching video processor.

[0033] S3: Remove data containing noise and unclear semantics to ensure data quality;

[0034] S4: Train the multi-modal video understanding model and the speech generation model, optimize the model parameters, and improve the generalization ability of the model;

[0035] S5: Use the trained model for the dubbing task to output high-quality synthesized speech.

[0036] Preferably, the multi-modal video understanding model and the speech generation model include a multi-modal reasoning module, a conditional flow matching speech generator, and a dynamic conditional control mechanism. Specifically, (a) the multi-modal reasoning module: adopts the InternVL2-8B architecture, fuses supervised learning (CE Loss) and reinforcement learning (DPO / BCO Loss) through the mixed preference optimization (MPO) strategy, and achieves an 85.84% scene classification accuracy rate in the four-step reasoning chain; (b) the conditional flow matching speech generator: trains the TTS large model based on the pre-trained DiT architecture, introduces three-way conditional control (visual features, reasoning conclusions, synthesized text), and trains the model using the optimal transport conditional flow matching (OT-CFM) algorithm; (c) the dynamic conditional control mechanism: designs three adjustment coefficients, r_v (visual weight), r_c (MMLM reasoning conclusion weight), and r_t (text weight), to achieve linear adjustment of the synthesized speech attributes.

[0037] In summary, compared with the prior art, the present invention has the following advantages:

[0038] (1) By integrating visual understanding and speech generation through a multi-stage training strategy, it can truly achieve an improvement in lip-sync accuracy and emotional similarity like real human dubbing.

[0039] (2) On the V2C-Animation dataset, the speaker similarity of SPK-SIM reached 89.74% (an increase of 0.44% compared to F5-TTS), and the emotion similarity of EMO-SIM was 78.88% (an increase of 12.1%);

[0040] (3) The lip synchronization error of LSE-D on the GRID dataset was reduced to 14.63 (optimized by 1.03); through the hybrid preference optimization strategy, the multi-modal reasoning accuracy increased by 8.7%.

[0041] (4) Improve the robustness of the model: enhance the performance of the model on low-quality datasets and improve the stability and reliability of dubbing synthesis.

[0042] (5) Wide range of applications: applicable to various speech synthesis and dubbing synthesis tasks.

[0043] (6) Easy to integrate: The present invention can be seamlessly integrated with existing dubbing tools to enhance the overall performance of the existing system.

[0044] The above are only the preferred embodiments of the present invention and do not impose any limitations on the present invention. Any person skilled in the art, without departing from the technical solution of the present invention, makes any form of equivalent replacement or modification to the technical solution and technical content disclosed by the present invention, and such changes are still within the scope of the technical solution of the present invention and still fall within the protection scope of the present invention.< / conclusion> < / reasoning> < / summary>

Claims

1. A method for dubbing movie voices based on a multi-modal thought chain, characterized in that, It includes the following steps: Construct a movie dubbing dataset with CoT annotations; Integrate multilingual speech libraries, animation datasets, and multi-speaker datasets, and train the TTS speech synthesis module and the V2S video dubbing module; Remove data with noise and unclear semantics; Train the multi-modal video understanding model and the speech generation model, optimize the model parameters, and improve the generalization ability of the model; Use the trained model for the dubbing task and output high-quality synthesized speech.

2. The method for dubbing movie voices based on a multi-modal thought chain according to claim 1, wherein, The CoT annotation rule is a rule system for annotating text data.

3. The method for dubbing movie voices based on a multi-modal thought chain according to claim 1, wherein, The full name of TTS is TextTo Speech, which is used to automatically convert text into speech; V2S is a high-definition digital input and audio-video synchronous switching video processor.

4. The method for dubbing movie human voices based on multi-modal thought chains as claimed in claim 1, wherein, The multi-modal video understanding model and the speech generation model include a multi-modal reasoning module, a conditional flow matching speech generator, and a dynamic conditional control mechanism.