Video sound effect generation method and device, storage medium and electronic equipment
By acquiring multi-layered visual features of the video to be dubbed and using a diffusion model to generate audio latent vectors, the problem of insufficient matching between sound effects and video content in existing technologies is solved, and higher quality sound effect generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING SHENGSHU TECH CO LTD
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, methods that convert video content into text description information and then generate sound effects have poor matching between sound effects and video content, and cannot fully describe the video content.
By acquiring multi-layered visual features of the video to be dubbed, including image-text correlation features, video-speech correlation features, video self-supervised features, and video-text correlation features, an audio latent vector is generated using a pre-trained diffusion model, and dubbing sound effects are generated through a variational autoencoder and a vocoder.
It improves the matching degree and quality of generated dubbing sound effects with video content, and comprehensively extracts multi-level visual features of the video as conditional guidance to improve the accuracy of sound effect generation.
Smart Images

Figure CN121967778A_ABST
Abstract
Description
Methods, devices, storage media, and electronic devices for generating video audio effects Technical Field
[0001] This disclosure relates to the fields of computer technology and artificial intelligence technology, and in particular to a method, apparatus, storage medium and electronic device for generating video audio effects. Background Technology
[0002] Sound effect generation has wide applications in many fields, including but not limited to game development, video production, and virtual reality. With the rapid development of deep neural networks and generative models, significant progress has been made in the technology of generating corresponding sound effects based on text descriptions, while the technology of generating sound effects based on video content is still in the exploratory stage.
[0003] In related technologies, a multimodal large model is used to convert video content into corresponding text description information, and then the text description information is input into a text-based sound effect generation model to generate corresponding sound effects. However, in this video sound effect generation scheme, the accuracy of the text description information obtained from the video content conversion is low, and it cannot describe the comprehensive video content, resulting in a poor match between the sound effects generated by the sound effect generation model and the video content.
[0004] Therefore, how to generate highly compatible sound effects for videos has become a technical problem that urgently needs to be solved. Summary of the Invention
[0005] The embodiments of this disclosure provide a method, apparatus, storage medium, and electronic device for generating video audio effects.
[0006] According to one aspect of the embodiments of this disclosure, a method for generating video sound effects is provided. The method includes: generating multi-layered visual features of a video to be dubbed, the multi-layered visual features including at least two of the following: image-text correlation features, video-speech correlation features, video self-supervised features, and video-text correlation features; generating an audio latent vector corresponding to the video to be dubbed based on the multi-layered visual features and a pre-trained diffusion model; and generating dubbing sound effects for the video to be dubbed based on the audio latent vector.
[0007] According to another aspect of the present disclosure, a video sound effect generation apparatus is provided. The apparatus includes: a visual feature extraction module, configured to generate multi-layered visual features of a video to be dubbed based on the video to be dubbed, wherein the multi-layered visual features include at least two of the following: image-text correlation features, video-speech correlation features, video self-supervised features, and video-text correlation features; a first generation module, configured to generate an audio latent vector corresponding to the video to be dubbed based on the multi-layered visual features and a pre-trained diffusion model; and a second generation module, configured to generate dubbing sound effects for the video to be dubbed based on the audio latent vector.
[0008] According to another aspect of the present disclosure, a computer-readable storage medium is provided, which stores a computer program for performing the above-described method for generating video audio effects.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the above-described method for generating video and audio effects.
[0010] Based on the video audio effect generation method, apparatus, storage medium, and electronic device provided in the above embodiments of this disclosure, when it is necessary to generate dubbing audio effects for a video, multi-layered visual features can be obtained from the video to be dubbed, and corresponding audio latent vectors can be generated based on the multi-layered visual features and a pre-trained diffusion model, thereby generating the dubbing audio effects for the video to be dubbed. The technical solution of this disclosure can extract more comprehensive and complete multi-layered visual features from the video to be dubbed, and then use the extracted multi-layered visual features as the audio latent vector prediction process of the conditionally guided diffusion model, improving the quality of the generated dubbing audio effects and the relevance of the dubbing audio effects to the video content of the video to be dubbed.
[0011] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0012] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0013] Figure 1 is a system framework diagram to which this disclosure applies.
[0014] Figure 2 is a flowchart illustrating a method for generating video audio effects provided in an exemplary embodiment of this disclosure.
[0015] Figure 3 is a schematic diagram of the training and testing process of the diffusion model provided in an exemplary embodiment of this disclosure.
[0016] Figure 4 is a flowchart illustrating step 202 of the embodiment shown in Figure 2, provided by an exemplary embodiment of this disclosure.
[0017] Figure 5 is a schematic diagram of the fusion of multi-layered visual features provided by an exemplary embodiment of the present disclosure.
[0018] Figure 6 is a schematic flowchart of step 203 of the embodiment shown in Figure 2, provided by an exemplary embodiment of this disclosure.
[0019] Figure 7 is a schematic diagram of the training process of the diffusion model provided in an exemplary embodiment of this disclosure.
[0020] Figure 8 is a schematic diagram of the structure of a video audio effect generation apparatus provided in an exemplary embodiment of the present disclosure.
[0021] Figure 9 is a schematic diagram of the structure of a video audio effect generation apparatus provided in another exemplary embodiment of this disclosure.
[0022] Figure 10 is a structural diagram of an electronic device provided in an exemplary embodiment of the present disclosure. Detailed Implementation
[0023] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0024] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0025] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.
[0026] It should also be understood that in the embodiments disclosed herein, "multiple" can refer to two or more, and "at least one" can refer to one, two or more.
[0027] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0028] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.
[0029] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0030] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0031] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0032] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0033] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0034] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.
[0035] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.
[0036] This disclosure outlines
[0037] Current methods for generating video sound effects typically utilize multimodal large models to convert video content into corresponding text descriptions, and then generate sound effects based on these text descriptions. However, the accuracy of the text descriptions obtained from converting video content is relatively low, failing to fully describe the video content, resulting in poor matching between the generated sound effects and the video content.
[0038] This disclosure can obtain more comprehensive and multi-layered visual features of the video to be dubbed through at least one visual feature extractor, and use the multi-layered visual features as a conditional guide to control the diffusion model to generate dubbing sound effects for the video to be dubbed, thereby improving the matching degree between the generated dubbing sound effects and the video content.
[0039] Exemplary System
[0040] Figure 1 illustrates an exemplary system architecture for a method or apparatus for generating video audio effects to which embodiments of the present disclosure may be applied.
[0041] As shown in Figure 1, the system architecture may include a terminal device 101, a communication network 102, a server 103, and a data storage system 104. The network 102 serves as the medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0042] Users can use terminal device 101 to interact with server 103 through communication network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as applications for generating artificial intelligence (AI) content (e.g., AI videos or AI images), multimedia applications, search applications, web browser applications, shopping applications, instant messaging tools, etc.
[0043] Terminal device 101 can be any electronic device capable of triggering the generation of video and audio effects, including but not limited to mobile terminals such as mobile phones, laptops, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and fixed terminals such as digital televisions, desktop computers, and smart home appliances.
[0044] Server 103 can be a server providing various services, such as a backend server processing the video to be dubbed uploaded by terminal device 101. The backend server can generate dubbing sound effects for the received video to be dubbed. Server 103 can be implemented as a standalone server or a cluster of multiple servers.
[0045] The data storage system 104 is used to store the data that the server needs to process. It can be integrated on the server or placed on a cloud server.
[0046] It should be noted that the video and audio effect generation method provided in the embodiments of this disclosure can be executed by the server 103 or by the terminal device 101. Accordingly, the video and audio effect generation device can be set in the server 103 or in the terminal device 101.
[0047] It should be understood that the number of terminal devices 101, network 102, server 103, and data storage system 104 in Figure 1 is merely illustrative. Depending on implementation needs, any number of terminal devices 101, network 102, server 103, and data storage system 104 can be included. For example, if the generation of video and audio effects does not require remote processing, the above system architecture may exclude the network and server, including only terminal devices 101 and data storage system 104.
[0048] Exemplary methods
[0049] Figure 2 is a flowchart illustrating a video audio effect generation method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as the server 103 shown in Figure 1 or the mobile device 101. As shown in Figure 2, it includes the following steps:
[0050] Step 201: Based on the video to be dubbed, generate multi-level visual features of the video to be dubbed. The multi-level visual features include at least the following two types: image-text correlation features, video-speech correlation features, video self-supervised features, and video-text correlation features.
[0051] Optionally, for each visual feature in the multi-layered visual features, feature extraction can be performed separately using the corresponding visual feature extractor. The visual feature extractor includes at least: an image-text correlation feature extractor, a video-speech correlation feature extractor, a video self-supervised feature extractor, and a video-text correlation feature extractor.
[0052] The various visual feature extractors in this disclosure are used to instruct various visual feature extraction models obtained through different task pre-training, for extracting multi-level visual features of different levels and different emphases from the video to be dubbed.
[0053] The "video to be dubbed" in this disclosure is used to indicate videos that require the addition of dubbing sound effects, such as adding background sound or dubbing.
[0054] Specifically, image-text relevance features are used to indicate the textual description features of each frame of an image in a video to be dubbed. These features can be extracted using a Contrastive Language-Image Pre-Training (CLIP) model, which is pre-trained based on paired image and text data. This disclosure can be used to extract image-text relevance features from each frame of an image in a video to be dubbed, obtaining the visual and textual relationships between each frame. The CLIP model is a deep learning model that combines image and textual information. It is trained by simultaneously considering both the image and its associated textual description, thus representing them in a unified embedding space. The CLIP model provides an embedding space that considers both image and textual information, allowing for comparison and analysis of images and text within the same space, facilitating feature association and learning.
[0055] Video-text relevance features are used to indicate the textual description features of a video to be dubbed or its various video segments. These features can be extracted using a Video-Image CLIP (ViCLIP) pre-trained model, which is based on paired video and text data. The ViCLIP model is a deep learning model that combines video (or video segments) and textual information. It is trained by simultaneously considering both the video and its associated textual descriptions, thus representing them within a unified embedding space. The ViCLIP model provides an embedding space that considers both video and textual information, allowing for comparison and analysis of video and text within the same space, facilitating feature association and learning. In this disclosure, the ViCLIP model can be used to extract video-text relevance features from various video segments in a video to be dubbed, obtaining the interrelationships between video segments and text.
[0056] The CLIP and ViCLIP models described above can be used to obtain textual description features of images or video segments in the video to be dubbed. This disclosure can also use other deep learning models to obtain textual description features of the video to be dubbed, such as Convolutional Neural Network (CNN).
[0057] Video self-supervised features are used to indicate the intrinsic representational features of the video to be dubbed, that is, the inherent features of the video itself. Video self-supervised features can be extracted using models such as Self-Supervised Learning (SSL), Masked Autoencoders (MAE), and Bootstrap Your Own Latent (BYOL). Through these self-supervised models, the video self-supervised features of the video to be dubbed can be obtained. Video self-supervised feature extraction models are deep learning models that extract features by leveraging the inherent structure of video data and the correlation between the learning data. They can learn video features related to user interests; through video self-supervised feature extraction models, key features that reflect the essential content of the video can be extracted. Therefore, video and self-supervised features can be compared and analyzed in the same space, facilitating feature association and learning.
[0058] Video-speech relevance features are used to indicate sound effect-related features in a video to be dubbed. This disclosure uses a Contrastive Audio-Video Pre-Training (CAVP) model to extract video-speech relevance features from a video to be dubbed. The CAVP model in this disclosure uses supervised contrastive learning to mine sound effect-related features in the video to be dubbed through audio-visual pairing. The CAVP model is a deep learning model that combines video and audio information. It is trained by simultaneously considering both video and associated audio, thereby representing them in a unified embedding space. The CAVP model provides an embedding space that considers both video and audio, allowing for comparison and analysis of video and audio within the same space, facilitating feature association and learning.
[0059] In addition to using the aforementioned visual feature extractors to obtain multi-layered visual features in the video to be dubbed, any feature extractor capable of extracting video-related features from the video to be dubbed can also be used to extract features; this disclosure does not limit this.
[0060] Optionally, before generating multi-layered visual features of the video to be dubbed, the video can be processed by frame-by-frame extraction or segmentation to obtain multiple video segments. The segmentation duration can be randomly determined or determined according to business requirements. The video segments obtained through this frame-by-frame extraction or segmentation process are easier to extract feature data from.
[0061] Step 202: Based on multi-level visual features and a pre-trained diffusion model, generate the audio latent vector corresponding to the video to be dubbed.
[0062] In this disclosure, audio latent vectors are used to indicate the spectral characteristics of dubbing sound effects in the latent space.
[0063] Among them, the pre-trained diffusion model is used to indicate the model that can determine and remove the noise to be removed at each time step of random noise based on the multi-layer visual features of the video to be dubbed, and thus obtain the audio latent vector of the video to be dubbed.
[0064] In this disclosure, the pre-trained diffusion model can be trained through the diffusion process (forward process) of the diffusion model, that is, by adding noise to the latent vector of the sample audio corresponding to the sample video.
[0065] The reverse process of the pre-trained diffusion model is used to instruct the process of gradually removing noise to obtain the audio latent vector.
[0066] Optionally, based on a cross-attention mechanism, a cross-attention feature representation of the pre-trained diffusion model is obtained according to multi-layered visual features; and, based on the pre-trained diffusion model and the cross-attention feature representation, the audio latent vector corresponding to the video to be dubbed is obtained. The cross-attention mechanism assigns attention weights to each layer of visual features based on the input multi-layered visual features, enabling the pre-trained diffusion model to focus more on key feature information with higher weights (i.e., higher feature correlation) during the generation process. In this way, the pre-trained diffusion model can better preserve and reflect the key features of the video to be dubbed when generating the audio latent vector, thereby improving the quality of the generated audio and its matching degree with the video to be dubbed.
[0067] In this disclosure, during the denoising process using a pre-trained diffusion model, multi-layered visual features can be used as conditional guidance. Cross-attention between the visual features at each layer and the input latent vector of the cross-attention module of the pre-trained diffusion model is calculated, and the cross-attentions are weighted and fused to obtain a weighted fused feature. The audio latent vector is then obtained based on this weighted fused feature. Specifically, during the weighted fusion process, a weight (i.e., a weighted fusion weight) is assigned to the cross-attention calculated based on each layer of visual features to obtain the weighted fused feature. Optionally, this weighted fusion weight is automatically learned during the pre-training of the diffusion model. A specific implementation of using multi-layered visual features as conditional guidance to generate the audio latent vector can be found in the embodiment shown in Figure 3, which will not be detailed here.
[0068] Step 203: Generate dubbing sound effects for the video to be dubbed based on the audio latent vectors.
[0069] In this embodiment, a pre-trained variational autoencoder (VAE) decoder can be used to decode the audio latent vector representation to obtain the corresponding Mel spectrum data. Then, a Mel spectrum vocoder is used to convert the Mel spectrum data into the final dubbing sound effect (the waveform of the dubbing information).
[0070] Among them, the Mel spectrum vocoder is a model used for analyzing and synthesizing Mel spectrum data.
[0071] In some alternative implementations, after obtaining the dubbing sound effects of the video to be dubbed, the dubbing sound effects can be merged with the video to be dubbed to obtain a video with sound effects.
[0072] The method provided in the above embodiments of this disclosure, when it is necessary to generate dubbing sound effects for a video, can obtain multi-layered visual features from the video to be dubbed, and generate corresponding audio latent vectors based on the multi-layered visual features and a pre-trained diffusion model, thereby generating dubbing sound effects for the video to be dubbed. The technical solution of this disclosure can extract more comprehensive and complete multi-layered visual features from the video to be dubbed based on visual feature extractors with different focuses. Furthermore, the extracted multi-layered visual features can be used as conditional guidance to control the audio latent vector prediction process of the diffusion model, improving the quality of the generated dubbing sound effects, as well as their relevance and matching degree with the video content of the video to be dubbed.
[0073] As shown in Figure 4, based on the embodiment shown in Figure 2 above, the operation of obtaining the cross-attention feature representation of the pre-trained diffusion model based on the cross-attention mechanism and multi-layer visual features may include the following steps.
[0074] Step 221: Calculate the cross-attention between each level of visual features in the multi-level visual features and the input latent vector of the cross-attention module of the pre-trained diffusion model, and obtain multiple cross-attention corresponding to the multi-level visual features.
[0075] Step 222: Perform weighted fusion on multiple cross-attention points to obtain weighted fusion features.
[0076] The weighted fusion feature is the cross-attention feature representation of the pre-trained diffusion model. Furthermore, the weighted fusion weights in this process are automatically learned during the pre-training of the diffusion model.
[0077] In steps 221 and 222, the input latent vector H of the cross-attention module of the pre-trained diffusion model can be compared with the visual features F at each level. i Calculate cross-attention CA(H,F) i Then, a weighted fusion is performed to obtain a new feature representation H. new , with H new Replace the features output by the corresponding cross-attention module in the original hidden diffusion model. H new The calculation is shown in equation (1):
[0078]
[0079] In equation (1), CA(H,F) i α is used to represent the cross-attention between H and the visual features of the i-th level. i Weighting coefficients used to represent visual features at the i-th level.
[0080] For example, see Figure 5, which illustrates the process of generating the audio latent vector corresponding to the video to be dubbed by using multi-level features as conditions. First, the cross-attention between the visual features of each level and the input latent vector of the cross-attention module can be calculated. For example, in Figure 5, the cross-attention between the image text relevance features and the input latent vector of the cross-attention module is calculated (cross-attention layer-1), the cross-attention between the video speech relevance features and the input latent vector of the cross-attention module is calculated (cross-attention layer-2), the cross-attention between the video self-supervised features and the input latent vector of the cross-attention module is calculated (cross-attention layer-3), and the cross-attention between the video text relevance features and the input latent vector of the cross-attention module is calculated (cross-attention layer-4). Then, the weighted fusion feature of each cross-attention is calculated. The weighted fusion feature is used as a condition to remove the noise to be removed at each time step in the random noise, so as to obtain the audio latent vector corresponding to the video to be dubbed.
[0081] The method provided in the above embodiments of this disclosure, when generating audio latent vectors for dubbing sound effects, uses a cross-attention mechanism to assign attention weights to each level of visual features based on the multi-level visual features of the input. This allows the pre-trained diffusion model to pay more attention to key feature information with higher weights during the generation of audio latent vectors, thereby better preserving and reflecting the key features of the video to be dubbed when generating audio latent vectors, thus improving the quality of the generated audio and its matching degree with the video to be dubbed.
[0082] Figure 6 is a schematic flowchart illustrating step 203 of the embodiment shown in Figure 2, provided by an exemplary embodiment of this disclosure. Step 203 includes the following steps:
[0083] Step 231: Use the decoder of the pre-trained variational autoencoder to decode the audio latent vector and obtain the corresponding Mel spectrum.
[0084] In this disclosure, the decoder of a pre-trained variational autoencoder (VAE) is used to reconstruct audio latent vectors into audio spectra. VAE is a deep generative model based on autoencoders, which can map features of various modalities to latent vectors in a low-dimensional space through the encoder, and then reconstruct the latent vectors into data of the corresponding modality through the decoder.
[0085] The VAE encoder maps data from the original space to the latent space to obtain an implicit, continuous representation. Since the latent space is a low-dimensional matrix, it can make the subsequent calculation of the diffusion model more efficient.
[0086] Mel spectrum is a characteristic representation of audio signals, based on a nonlinear Mel scale to represent the spectral characteristics of the audio signal. The Mel scale approximates a linear transformation in the low-frequency range and a logarithmic transformation in the high-frequency range, a characteristic that aligns with how the human ear perceives frequency. This disclosure decodes audio latent vectors into Mel spectra, contributing to a frequency domain representation that better reflects the characteristics of human hearing.
[0087] Step 232: Use a vocoder to convert the Mel spectrum to obtain the dubbing sound effect.
[0088] In this disclosure, a vocoder is a speech synthesis technique used to convert a Mel spectrum into an audible sound waveform, that is, to convert the Mel spectrum back into the original audio signal.
[0089] The method provided in the above embodiments of this disclosure, after generating the audio latent vector of the dubbing sound effect, can further convert the audio latent vector into the dubbing sound effect through the decoder and vocoder of the pre-trained VAE. Subsequently, the dubbing sound effect can be merged with the corresponding video to be dubbed to complete the video sound effect dubbing.
[0090] In some alternative implementations, the pre-trained diffusion model described above can be obtained through training. Figure 7 is a schematic diagram of the training process of the diffusion model provided by an exemplary embodiment of this disclosure. As shown in Figure 7, the training process includes the following steps.
[0091] Step 701: Obtain the training sample set, which includes sample videos and corresponding sample audio.
[0092] In this disclosure, the training sample set may include at least one set of training samples, each set of training samples including a sample video (V) and a corresponding sample audio (A).
[0093] Step 702: Using at least one visual feature extractor, multi-level visual features of the sample video are generated based on the sample video; and the sample audio is encoded using an audio encoder to obtain the sample latent vector corresponding to the sample audio.
[0094] In this embodiment, a pre-trained visual feature extractor is used. Multi-level feature extraction is performed on the sample video V. Among these features, N represents the i-th visual feature extractor. v Let F represent the total number of visual feature extractors. The visual features extracted by the i-th visual feature extractor are denoted as F. i This yields the multi-layered visual features extracted by all visual feature extractors.
[0095] In this embodiment, each sample audio can be converted to the latent space by an audio encoder (such as the encoder of VAE) to obtain the sample latent vector z0 of each sample audio.
[0096] Step 703: Based on the multi-layer visual features and latent vectors of the sample video, noise prediction is performed at each time step to train a pre-trained diffusion model.
[0097] In this disclosure, the time step is used to indicate the discrete time unit of the simulation system's evolution and can be used to control the progress of the diffusion process. Typically, each time step corresponds to a specific noise level. As the time step increases, the noise level gradually increases, and the data distribution gradually transforms into a noise distribution.
[0098] The time step can be the time reference used by the diffusion model when adding noise to the latent vector of the sample video during the diffusion process, or it can be the time reference used by the diffusion model when identifying and removing noise from random noise during the reverse process. The noise to be removed corresponding to each time step can refer to the noise that needs to be removed from the random noise. In practical applications, there can be one noise to be removed at each time step.
[0099] In this embodiment, Gaussian noise is gradually added to the sample latent vector z0 through a diffusion process of T time steps to obtain the final noise z. T During the diffusion process, for the t-th time step, the diffusion model is responsible for applying multi-layered visual features and noise z. t The noise added at the current time step is predicted, and the loss function is calculated as shown in formula (2).
[0100]
[0101] In equation (2), ||·|| represents the square of the L2 norm. This indicates that the model has an input latent vector of z. t The time step and constraints are The prediction noise at step t, ε t This represents real noise.
[0102] Using the loss function in equation (2), the loss function value for the diffusion model to be trained can be determined based on the mean square error between the predicted noise and the actual noise at each time step; and the noise removal model to be trained can be trained based on the loss function value.
[0103] For example, referring to Figure 3, which illustrates the model training process, Mel spectrum data is extracted from sample audio, and sample latent vectors are obtained through an encoder. Multi-layered visual features (including image-text correlation features extracted by an image-text correlation feature extractor, video-speech correlation features extracted by a video-speech correlation feature extractor, video self-supervised features extracted by a video self-supervised feature extractor, and video-text correlation features extracted by a video-text correlation feature extractor) are extracted through a multi-layered visual feature extractor. The prediction noise at each time step is predicted through the multi-layered visual features. Based on the mean square error between the prediction noise and the real noise at each time step, the loss function value of the diffusion model to be trained is determined. Then, the noise removal model to be trained is trained according to the loss function value to obtain the pre-trained diffusion model.
[0104] The method provided in the above embodiments of this disclosure trains a diffusion model to be trained based on sample videos and sample audio. It can utilize the multi-layer visual features and latent vectors of the sample videos to train the diffusion model, ensuring that the error between the noise to be removed and the added Gaussian noise determined by the diffusion model meets the preset requirements. This allows the diffusion model to fully learn the matching relationship between the multi-layer visual features and the latent vectors of the samples, improving the accuracy of the audio latent vectors output through the reverse process of the diffusion model, and thus improving the matching degree between the dubbing sound effects and the video to be dubbed.
[0105] Exemplary device
[0106] Figure 8 is a schematic diagram of the structure of a video audio effect generation apparatus provided in an exemplary embodiment of the present disclosure. As shown in Figure 8, the apparatus may include:
[0107] The visual feature extraction module 81 is used to generate multi-level visual features of the video to be dubbed based on the video to be dubbed. The multi-level visual features include at least the following two types: image-text correlation features, video-speech correlation features, video self-supervised features, and video-text correlation features.
[0108] The first generation module 82 is used to generate the audio latent vector corresponding to the video to be dubbed based on multi-level visual features and a pre-trained diffusion model.
[0109] The second generation module 83 is used to generate dubbing sound effects for the video to be dubbed based on the audio latent vector.
[0110] Figure 9 is a schematic diagram of the structure of a video audio effect generation device provided in another exemplary embodiment of the present disclosure. As shown in Figure 9, based on the embodiment shown in Figure 8, in some embodiments, the first generation module 82 is used to use a pre-trained diffusion model to remove the noise to be removed corresponding to each time step in the random noise by using multi-level visual features as conditional guidance, so as to obtain the audio latent vector corresponding to the video to be dubbed.
[0111] In some implementations, the first generation module 82 includes:
[0112] The first generation submodule 821 is used to obtain the cross-attention feature representation of the pre-trained diffusion model based on the cross-attention mechanism and multi-level visual features.
[0113] The second generation submodule 822 is used to obtain the audio latent vector corresponding to the video to be dubbed based on the pre-trained diffusion model and the cross-attention feature representation.
[0114] In some implementations, the first generation submodule 821 includes:
[0115] Attention unit 8211 is used to calculate the cross attention between each level of visual features and the input latent vector in the multi-level visual features, so as to obtain multiple cross attention corresponding to the multi-level visual features.
[0116] The weighted fusion unit 8212 is used to perform weighted fusion on multiple cross-attentions to obtain weighted fusion features, wherein the weighted fusion features are the cross-attention feature table of the pre-trained diffusion model.
[0117] In some implementations, the second generation module 83 includes:
[0118] The decoding submodule 831 is used to decode the audio latent vector using the decoder of the pre-trained variational autoencoder to obtain the corresponding Mel spectrum.
[0119] The conversion submodule 832 is used to convert the Mel spectrum using a vocoder to obtain dubbing sound effects.
[0120] In some implementations, it also includes: a model training module 84;
[0121] Model training module 84 includes:
[0122] The acquisition submodule 841 is used to acquire the training sample set, which includes sample videos and corresponding sample audio.
[0123] The feature extraction submodule 842 is used to generate multi-level visual features of the sample video based on the sample video using at least one visual feature extractor; and to encode the sample audio using an audio encoder to obtain the sample latent vector corresponding to the sample audio.
[0124] Training submodule 843 is used to perform noise prediction at each time step based on the multi-layer visual features and latent vectors of the sample video, and train a pre-trained diffusion model.
[0125] It should be noted that the modules in this device can be disassembled and / or recombined, and these disassemblies and / or recombinations should be considered as equivalent solutions of this device.
[0126] The exemplary embodiments of this device correspond to the exemplary method section described above, and the relevant content can be referenced and cited interchangeably. The beneficial technical effects corresponding to the exemplary embodiments of this device can be found in the corresponding beneficial technical effects of the exemplary method section described above, and will not be repeated here.
[0127] Exemplary electronic devices
[0128] Figure 10 is a structural diagram of an electronic device provided in an embodiment of the present disclosure, including at least one processor 101 and a memory 102.
[0129] The processor 101 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0130] The memory 102 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 101 may execute one or more computer program instructions to implement the vehicle pose detection method and / or other desired functions of the various embodiments of this disclosure described above.
[0131] In one example, the electronic device may also include an input device 103 and an output device 104, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0132] The input device 103 may also include, for example, a keyboard, a mouse, a touch screen, a pickup device (such as a microphone array), etc.
[0133] The output device 104 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0134] Of course, for simplicity, Figure 10 only shows some of the components of the electronic device that are relevant to this disclosure, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.
[0135] Exemplary systems, computer program products, and computer-readable storage media
[0136] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products, including computer program instructions that, when executed by a processor, cause the processor to perform the steps in the methods for generating video audio effects according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0137] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0138] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the methods for generating video audio effects according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.
[0139] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0140] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0141] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0142] The block diagrams of devices, apparatuses, and devices involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, and devices can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0143] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the method is for illustrative purposes only, and the steps of the method of this disclosure are not limited to the order specifically described above, unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the method according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the method according to this disclosure.
[0144] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0145] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0146] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A method for generating video sound effects, characterized in that, include: Based on the video to be dubbed, multi-layered visual features of the video to be dubbed are generated, wherein the multi-layered visual features are at least This includes the following two types: image-text correlation features, video-speech correlation features, video self-supervised features, and video-text correlation features; Based on the multi-layered visual features and the pre-trained diffusion model, an audio latent vector corresponding to the video to be dubbed is generated; based on the audio latent vector, a dubbing sound effect for the video to be dubbed is generated.
2. The method according to claim 1, characterized in that, The step of generating the audio latent vector corresponding to the video to be dubbed based on the multi-layer visual features and the pre-trained diffusion model includes: using the pre-trained diffusion model, taking the multi-layer visual features as a conditional guide, removing the noise to be removed at each time step of the random noise, and obtaining the audio latent vector corresponding to the video to be dubbed.
3. The method according to claim 2, characterized in that, The step of using the multi-layered visual features as conditions to remove the noise to be removed from the random noise at each time step to obtain the audio latent vector corresponding to the video to be dubbed includes: obtaining the cross-attention feature representation of the pre-trained diffusion model based on the multi-layered visual features, and obtaining the audio latent vector corresponding to the video to be dubbed based on the pre-trained diffusion model and the cross-attention feature representation.
4. The method according to claim 3, characterized in that, The method of obtaining the cross-attention feature representation of the pre-trained diffusion model based on the cross-attention mechanism and the multi-layer visual features includes: calculating the cross-attention between each layer of visual features in the multi-layer visual features and the input latent vector of the cross-attention module of the pre-trained diffusion model, to obtain multiple cross-attentions corresponding to the multi-layer visual features; and performing weighted fusion on the multiple cross-attentions to obtain weighted fusion features, wherein the weighted fusion features are the cross-attention feature representation of the pre-trained diffusion model.
5. The method according to any one of claims 1-4, characterized in that, The step of generating the dubbing sound effect for the video to be dubbed based on the audio latent vector includes: decoding the audio latent vector using a decoder of a pre-trained variational autoencoder to obtain the corresponding Mel spectrum; and converting the Mel spectrum using a vocoder to obtain the dubbing sound effect.
6. The method according to any one of claims 1-5, characterized in that, The pre-trained diffusion model is trained through the following steps: obtaining a training sample set, which includes sample videos and corresponding sample audio; using at least one visual feature extractor, generating multi-level visual features of the sample videos based on the sample videos; and encoding the sample audio using an audio encoder to obtain the sample latent vectors corresponding to the sample audio; and performing noise prediction at each time step based on the multi-level visual features of the sample videos and the sample latent vectors to train the pre-trained diffusion model.
7. A device for generating video audio effects, characterized in that, include: A visual feature extraction module is used to generate multi-layered visual features of the video to be dubbed, based on the video to be dubbed. The multi-layered visual features include at least... This includes the following two types: image-text correlation features, video-speech correlation features, video self-supervised features, and video-text correlation features; The first generation module is used to generate the audio latent vector corresponding to the video to be dubbed based on the multi-layer visual features and the pre-trained diffusion model. The second generation module is used to generate the dubbing sound effects of the video to be dubbed based on the audio latent vector.
8. A computer-readable storage medium storing a computer program for performing the method according to any one of claims 1-6.
9. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1-6.
10. A computer program product comprising computer program instructions, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1-6.