Method for constructing unified video-to-human voice sound effect model

By training V2A capabilities on the VGGSound dataset and combining conditional stream matching objective functions and hybrid expert mechanism, the problem of video-to-voice and sound effects tasks is solved, and end-to-end unified generation and efficient sound effects output are achieved, supporting the generation of different sound types and model deployment.

CN120340532APending Publication Date: 2025-07-18GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510474989.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, video to voice (V2S) and video to audio (V2A) tasks are separated and cross-modal alignment is insufficient, and end-to-end unified generation cannot be achieved, resulting in waste of computing resources.

Method used

V2A capabilities are trained on the VGGSound dataset, the TTS capabilities are optimized using the conditional flow matching objective function, a multi-task instruction set is built and the model capability is specified through a hybrid expert mechanism, and the energy profile predicted by V2A is injected as a speech generation condition to generate aligned speech.

Benefits of technology

It realizes end-to-end voice and sound effect generation under video input conditions, supports output of different sound types, and the framework can be expanded and deployed to save computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340532A_ABST
    Figure CN120340532A_ABST
Patent Text Reader

Abstract

The invention relates to a method for constructing a unified video-to-human voice sound effect model, and the method comprises the following steps: training a V2A capability on a VGGSound data set, enabling the V2A capability to be capable of inputting a video and outputting a sound effect; adopting a conditional flow matching objective function, and optimizing the TTS capability based on an EMILIA data set; a multi-task instruction set is constructed, and different model capabilities are specified through a hybrid expert mechanism; and injecting an energy contour predicted by the V2A as a voice generation condition, and generating aligned voice conforming to the video content. According to the invention, the model can support the generation of end-to-end voice and sound effect under the video input condition, and the problems of video-to-voice (V2S) and video-to-sound effect (V2A) task splitting and insufficient cross-modal alignment in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent voice dubbing, and particularly relates to a method for constructing a unified video-to-human voice audio model. Background Art

[0002] In the prior art, there are problems such as the disconnection between the video-to-speech (V2S) and video-to-audio (V2A) tasks and insufficient cross-modal alignment. Currently, there is no model that can support the unified generation of V2A and V2S end-to-end. Multiple models are required to implement this task, so a lot of models need to be loaded, consuming computing resources.

[0003] Therefore, it is necessary to provide a method for constructing a unified video-to-human voice audio model to achieve the generation of end-to-end speech and audio effects by the model under video input conditions. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for constructing a unified video-to-human voice audio model to achieve the generation of end-to-end speech and audio effects by the model under video input conditions.

[0005] To solve the problems existing in the prior art, the present invention provides a method for constructing a unified video-to-human voice audio model, including the following steps:

[0006] Train the V2A ability on the VGGSound dataset so that it can input a video and output an audio effect;

[0007] Adopt a conditional flow matching objective function to optimize the TTS ability based on the EMILIA dataset;

[0008] Construct a multi-task instruction set and implement the specification of different model capabilities through a mixture of experts mechanism;

[0009] Inject the energy profile predicted by V2A as a speech generation condition to generate aligned speech that conforms to the video content.

[0010] Optionally, in the method for constructing a unified video-to-human voice audio model,

[0011] The full name of V2A is Video-to-Audio, which is used to automatically convert video content into audio;

[0012] The full name of TTS is Text To Speech, which is used to automatically convert text into speech.

[0013] Optionally, in the method for constructing a unified video-to-human voice audio model, the VGGSound dataset is a multi-modal single-label audio dataset, and the EMILIA dataset is a multi-language dataset containing more than 100,000 hours of speech data.

[0014] Optionally, in the method for constructing a unified video-to-audio model, the EMILIA dataset includes multiple language types.

[0015] Optionally, in the method for constructing a unified video-to-audio model, the multi-task instruction set includes instructions for generating audio effects that match the video.

[0016] Optionally, in the method for constructing a unified video-to-audio model, the Mixture of Experts mechanism is abbreviated as MoE. The Mixture of Experts mechanism is an ensemble learning method that combines multiple expert models to form a more complex system.

[0017] Optionally, in the method for constructing a unified video-to-audio model, the unified video-to-audio model includes a V2A module, a TTS module, and a Mixture of modality Fusion dynamic routing.

[0018] Compared with the prior art, the present invention has the following advantages:

[0019] (1) The present invention can realize the generation of end-to-end speech and audio effects under video input conditions, and solve the problems of task fragmentation between video-to-speech (V2S) and video-to-audio (V2A) and insufficient cross-modal alignment in the prior art.

[0020] (2) The present invention can obtain an end-to-end model, which can support the output of different voice types according to different text input contents of the model.

[0021] (3) The framework in the present invention supports expansion, and different models can be selected for deployment according to actual situations, supporting cloud or edge-side deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the unified video-to-audio model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0023] The specific embodiments of the present invention will be described in more detail below with reference to the schematic diagrams. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the drawings are in very simplified forms and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the objectives of the embodiments of the present invention.

[0024] In the description of the present application, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present application.

[0025] In the prior art, there are problems such as the fragmentation of video-to-speech (V2S) and video-to-audio (V2A) tasks and insufficient cross-modal alignment. Currently, there is no model that can support the unified generation of V2A and V2S end-to-end. Multiple models are required to implement this task, so a lot of models need to be loaded, consuming computing resources.

[0026] To solve the problems existing in the prior art, the present invention provides a method for constructing a unified video-to-human audio model, including the following steps:

[0027] Phase 1 (V2A pre-training): Train the V2A ability on the VGGSound dataset so that it can input a video and output an audio effect; among them, the VGGSound dataset is a large multi-modal single-label audio dataset. The VGGSound dataset aims to ensure the quality of the data through the audio-visual correspondence relationship and low label noise, while reducing the effort of manual annotation. The VGGSound dataset is mainly used to promote the development of audio recognition technology and performs well especially in audio classification and audio-visual synchronization tasks. The full name of V2A is Video-to-Audio, which is used to automatically convert video content into audio.

[0028] Phase 2 (TTS pre-training): Adopt a conditional flow matching objective function to optimize the TTS ability based on the EMILIA dataset; the EMILIA dataset is a multilingual dataset containing more than 100,000 hours of speech data. The EMILIA dataset contains various language types, such as covering six languages: Chinese, English, German, French, Japanese, and Korean. The EMILIA dataset mainly consists of real natural speech on the Internet and covers various content types such as talk shows, interviews, debates, sports commentaries, and audiobooks, ensuring that the dataset captures a wide range of real human speaking styles. The full name of TTS is TextTo Speech, which is used to automatically convert text into speech.

[0029] Stage 3 (Dynamic Instruction Training): Construct a multi-task instruction set and specify different model capabilities through the Mixture of Experts mechanism. Among them, the multi-task instruction set includes instructions such as generating sound effects that match the video. The English name of the Mixture of Experts mechanism is Mixture of Experts. The Mixture of Experts mechanism is an ensemble learning method that combines multiple expert models to form a more complex system.

[0030] Stage 4 (V2S Fine-tuning): Inject the energy profile predicted by V2A as the speech generation condition to generate aligned speech that conforms to the video content.

[0031] The present invention can obtain an end-to-end model that can support the output of different sound types according to different text input contents of the model. For example, when inputting a video and "generate sound effects that match the video content", the model will output sound effects that are highly aligned with the video; when inputting a video and dubbing text, the model will output human voices that fit the video. At the same time, this framework supports expansion and can select different models for deployment according to actual situations, supporting cloud or edge-side.

[0032] Preferably, the unified video-to-human voice and sound effect model includes a V2A module, a TTS module, and Mixture of modality Fusion (i.e., MoF) dynamic routing.

[0033] Further, the V2A module: Adopt a CLIP multi-modal encoder to extract spatio-temporal features, optimize the DiT flow matching decoder through contrastive learning, and generate 24kHz high-fidelity environmental sound effects that are dynamically synchronized with the video.

[0034] The TTS module: Construct a DiT flow matching architecture, integrate an emotion-controllable prosody predictor, and achieve text-based 48kHz speech generation.

[0035] MoF dynamic routing: Design a gating network (Gating Network), dynamically activate the V2A energy profile or pure TTS text features through FLAN-T5 instruction parsing, and achieve modality fusion. Achieve high-quality and highly aligned video-to-speech generation.

[0036] The principle of the present invention is as Figure 1 shown, and it can support both Video-to-Audio generation from text / video to audio (the left module in the figure), and Text-to-Speech generation from text to speech (the right module in the figure). In addition, through the modality fusion of V2A audio and TTS text, based on the speech synthesis function of pure text input, it can achieve Video-to-Speech speech generation combined with video information.

[0037] In summary, compared with the prior art, the present invention has the following advantages:

[0038] (1) The present invention can realize the generation of end-to-end speech and sound effects by the model under video input conditions, and solve the problems of the disconnection between video-to-speech (V2S) and video-to-audio (V2A) tasks and insufficient cross-modal alignment in the prior art.

[0039] (2) The present invention can obtain an end-to-end model, which can support the output of different sound types according to different text input contents of the model.

[0040] (3) The framework in the present invention can support expansion, and different models can be selected for deployment according to actual situations, supporting cloud or edge-side deployment.

[0041] The above are only the preferred embodiments of the present invention and do not impose any limitation on the present invention. Any person skilled in the art, without departing from the scope of the technical solution of the present invention, makes any form of equivalent substitution or modification and other changes to the technical solution and technical content disclosed by the present invention, all of which are within the content of the technical solution of the present invention and still fall within the protection scope of the present invention.

Claims

1. A method for constructing a unified video-to-audio model, characterized in that, It includes the following steps: Train the V2A ability on the VGGSound dataset so that it can input videos and output sound effects; Adopt the conditional flow matching objective function to optimize the TTS ability based on the EMILIA dataset; Construct a multi-task instruction set and implement the specification of different model capabilities through the mixture of experts mechanism; Inject the energy profile predicted by V2A as a speech generation condition to generate aligned speech that conforms to the video content.

2. The method for constructing a unified video-to-audio model according to claim 1, wherein The full name of V2A is Video-to-Audio, which is used to automatically convert video content into audio; The full name of TTS is Text To Speech, which is used to automatically convert text into speech.

3. The method for constructing a unified video-to-audio model according to claim 1, wherein The VGGSound dataset is a multi-modal single-label audio dataset, and the EMILIA dataset is a multilingual dataset containing more than 100,000 hours of speech data.

4. The method for constructing a unified video-to-audio model according to claim 3, wherein The EMILIA dataset contains multiple language types.

5. The method for constructing a unified video-to-audio model according to claim 1, wherein The multi-task instruction set includes instructions for generating sound effects that match the video.

6. The method for constructing a unified video-to-audio model according to claim 1, wherein The English name of the mixture of experts mechanism is Mixture of Experts. The mixture of experts mechanism is an ensemble learning method that combines multiple expert models to form a more complex system.

7. The method for constructing a unified video-to-audio model according to claim 2, wherein The unified video-to-audio model includes a V2A module, a TTS module, and a Mixture of modality Fusion dynamic routing.