Video-to-audio generation method based on text assistance
Through the text-assisted video-to-audio generation method (TA-V2A), the text-audio-video comparison pre-training module (CVALP) and the potential diffusion model (LDM) are used to solve the problem of insufficient alignment of semantic expression and timing in video-to-audio generation, and high-quality audio generation is achieved.
Patent Information
- Application Number
- CN202510298019.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-18
AI Technical Summary
The existing video-to-audio generation methods have shortcomings in semantic expression and timing alignment, which leads to the mismatch of the generated audio with the video content and the timing alignment accuracy is insufficient, making it difficult to generate high-quality audio.
Using a text-assisted video-to-audio generation method (TA-V2A), features are extracted and aligned by the text-audio-video comparison pre-training module (CVALP), combining the potential diffusion model (LDM) and the dual-guiding mechanism to achieve high-quality and synchronous generation of video content to audio.
It improves the semantic correlation and timing accuracy of generated audio, enhances feature expression capabilities, provides flexible personalized generation interfaces, and improves the quality and practicality of audio generation.
Smart Images

Figure CN120340531A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and multimedia processing, and particularly relates to a method for video-to-audio generation based on text assistance. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, audio generation has become an important field of audio technology research. In the fields of multimedia editing, augmented reality, and automatic content generation, the demand for generating matching audio based on video content is increasing. This demand mainly comes from three aspects: First, in the field of Foley Sound production, traditionally, Foley artists need to edit sound effect libraries or use props for on-site recording in the studio. This process is not only time-consuming and laborious but also difficult to fully match the visual content. Second, in the field of virtual reality, audio is the core element for constructing an immersive experience, and a three-dimensional sound field needs to be generated in real time according to the user's perspective. Third, for the environmental understanding of general intelligent agents, it is necessary to associate visual information with acoustic events to improve the cognitive ability of the environment.
[0003] In the field of video understanding, significant progress has been made in the past decade. In 2014, Karpathy et al. first proposed using deep learning methods for video understanding, pioneering this research direction. Subsequent research has mainly been divided into three stages: The first stage (2014 - 2017), represented by works such as Two Stream Network and LRCN, mainly learned temporal information by introducing an optical flow branch. The second stage (2015 - 2020), represented by works such as C3D and I3D, directly modeled the temporal information in videos using three-dimensional convolutional kernels. The third stage (since 2020), represented by works such as TimeSformer and ViViT, these models enhanced the ability to model long-term temporal information through the self-attention mechanism by introducing the Transformer architecture. In 2023, NExT-GPT proposed by Wu et al. first achieved free conversion of the input and output of multi-modal large language models. However, these methods mainly focus on the understanding and classification of video content, and there is still room for improvement in capturing and expressing acoustic events.
[0004] In the field of audio generation, technological development has also gone through multiple stages. AudioLDM proposed by Liu et al. in 2023 introduced the diffusion model into the field of audio generation and achieved high-quality text-to-audio conversion through contrastive language-audio pre-training. AudioGen proposed by Kreuk et al. in the same year combined an audio representation model and an audio language model, demonstrating excellent audio generation capabilities. In terms of direct video-to-audio conversion, Mo et al. proposed VTA-LDM in 2024, using video frames as auxiliary features to enhance the audio generation process. Ruan et al. achieved joint audio-visual generation using MM-Diffusion through a two-stream U-Net structure. In the latest work, Xu et al. conducted in-depth research on the selection of different modules in the V2A model framework. However, this modality conversion currently faces three main challenges: First, it is necessary to accurately extract semantic information from the video that can reflect acoustic events, and current models often lose sequence context when relying on single-frame features. Second, different from static images, videos contain continuous time series information, requiring the generated audio to be accurately aligned with the time series of the video. Finally, it is necessary to find a suitable feature expression carrier for semantic information and temporal information in the latent space and perform appropriate fusion to generate high-quality audio.
[0005] Although some recent works such as FoleyCrafter by Zhang et al. enhanced semantic expression through text guidance and Diff-Foley by Luo et al. improved the feature alignment effect through contrastive learning pre-training, the overall processing of latent space features and time alignment is still in its infancy. In particular, there are deficiencies in the following aspects: First, there is a lack of in-depth semantic understanding of video content, resulting in potential semantic mismatches between the generated audio and the video content. Second, the temporal alignment accuracy is insufficient, making it difficult to accurately capture the key acoustic event moments in the video. Third, the quality and realism of the generated audio still need to be improved, especially when dealing with complex acoustic scenarios.
[0006] Due to the above reasons and the deficiencies in existing methods, there is an urgent need for a video-to-audio generation method that can simultaneously achieve good semantic expression and accurate temporal alignment to improve the quality and practicality of the generated audio. Summary of the Invention
[0007] Aiming at the problems of insufficient semantic expression and temporal alignment in current video-to-audio generation, the purpose of the present invention is to provide a text-assisted video-to-audio generation method (TA-V2A). This method introduces text assistance in network training, feature generation, and the diffusion process, combines contrastive learning pre-training and multi-condition diffusion generation, so as to achieve high-quality and synchronous generation of video content to audio.
[0008] The present invention first uses a text-audio-video contrast pre-training (CVALP) module, which extracts features using a video encoder, an audio encoder, and a text encoder respectively. For an existing video dataset, we can extract the audio track and expand the original labeled video descriptions in the dataset with the help of a large language model to obtain video-audio-text triple training data (x v ,x a ,x l ), where each triple is generated for a video segment, that is, x v represents the video segment, x a represents the mel spectrogram of x v , and x l is the corresponding text description of x v . The features of each modality are extracted through the encoders f V (·), f A (·), and f L (·) respectively. To align the feature dimensions, temporal pooling is performed on the extracted video features and audio features respectively, and then the extracted text features and the temporally pooled video features and audio features are all mapped to the same feature space to unify the dimensions of the text features, video features, and audio features. Then, an audio-centered contrast loss function is designed, which includes an audio-text alignment term and an audio-video alignment term, and the audio-video alignment is further divided into two parts: semantic alignment and temporal alignment. After completing the encoder training, during inference, the present invention introduces a video-large language model (such as Video-LLaMA2) to automatically generate video content descriptions for the input videos, and uses a portable plug-in prompt fine-tuner (PPPR) for normalization processing to obtain text modality inputs. After extracting multi-modal features using the trained video and text encoders, two feature mixing strategies of weighted average and concatenation are adopted, and the final conditional features are obtained through position encoding and projection layer processing. Then, a latent diffusion model (LDM) is used to generate audio data in a low-dimensional latent space, and a dual guidance method is adopted at each step of the diffusion process: on the one hand, classifier-free guidance (CFG) is used to improve the generation quality, and on the other hand, classifier guidance (CG) is used to enhance temporal alignment. Finally, the latent space features obtained by the diffusion model are decoded into mel spectrograms by the diffusion model decoder, and then the mel spectrograms are converted into audio waveforms by a vocoder.
[0009] The technical solution adopted by the present invention is:
[0010] A text-assisted video-to-audio generation method, the steps of which include:
[0011] 1) Construct a training set, and each training sample in the training set is a triple (x v ,x a ,xl ) where x v represents a video segment, and x a represents the Mel spectrogram of x v , and x l is the corresponding text description of x v ;
[0012] 2) Use the text-audio-video contrast pre-training module to extract the triplet (x v , x a , x l ) video features, audio features, and text features; map the extracted text features and the video features and audio features after temporal pooling to the same feature space to obtain text features, video features, and audio features of the same dimension; the text-audio-video contrast pre-training module includes a video encoder f V (·) for extracting video features, an audio encoder f A (·) for extracting audio features, and a text encoder f L (·) for extracting text features;
[0013] 3) Align the text features, video features, and audio features of the same triplet processed in step 2) with the audio features as the center, and calculate the alignment loss of the corresponding triplet; calculate the total loss value according to the alignment losses of each triplet
[0014] 4) Optimize the text-audio-video contrast pre-training module according to the total loss value ;
[0015] 5) For a target video and its text description; use the optimized video encoder f V (·) to extract the video features of the target video, and use the optimized text encoder f L (·) to extract the text features of the target video; then fuse the video features and text features of the target video, and map the fused features to a low-dimensional space through positional encoding and projection processing in sequence to obtain the conditional features of the target video;
[0016] 6) Use the latent diffusion model to generate the audio data corresponding to the target video according to the conditional features of the target video.
[0017] Furthermore, use the video-large language model to generate the content description of the target video, and then use the portable plug-in prompt fine-tuner to standardize the content description of the target video to obtain the text description of the target video.
[0018] Furthermore, the text description of the target video is the information input manually.
[0019] Furthermore, the total loss value wherein, and respectively represent the audio feature, text feature, and video feature of the i-th triple, N, N S and N T respectively represent the total number of audio-text pairs, the total number of cross-video audio-video pairs, and the total number of audio-video pairs at different times within the same video. The parameters λ and μ are weight coefficients; and are feature vectors from different modalities x and y, correspond to or correspond to or M is the number of cross-modal pairs, and τ controls the softmax smoothness.
[0021] Furthermore, a weighted average method or a concatenation method is used to fuse the video feature and text feature of the target video.
[0022] Furthermore, the feature after fusion by the weighted average method is The feature after fusion by the concatenation method is E concat = concat[P(E v ) ; P(E l )]; wherein, the parameters λ and μ are weight coefficients, and P(E v ) and P(E l ) are linear projections.
[0023] Furthermore, using positional encoding and projection τ θ map the fused feature E concat to a low-dimensional space, denoted as where MLP serves as the projection layer and PE represents positional encoding. This hybrid feature representation is used as the input feature vector for LDM training and inference.
[0024] Furthermore, a classifier guidance CG and a classifier-free guidance CFG are jointly used to control the latent diffusion model to generate the audio data corresponding to the target video according to the conditional features of the target video.
[0025] Furthermore, the formulaic representation for generating the audio data corresponding to the target video is: wherein, γ represents the scale of CG, ω represents the scale of CFG, is the latent representation obtained from the input data z t at the t-th step in the inference stage of the latent diffusion model, E v represents the video feature, Emix is the fused feature, P φ (y|z t , t, E v ) is the alignment classifier for aligning audio-visual pairs, ∈ θ (z t , t, E mix ) is the latent representation for additional inference using the hybrid feature as a condition, is the latent representation for additional inference without using the hybrid feature as a condition, is the average diffusion coefficient.
[0026] Furthermore, the multi-condition inference step of the latent diffusion model is expressed as: where E l represents the text prompt feature, E nl represents the negative text prompt feature, ω v is the hyperparameter corresponding to the video feature, ω l is the hyperparameter corresponding to the text prompt feature, ω nl is the hyperparameter corresponding to the negative text prompt feature, ∈ θ (z t , t) is the latent representation for additional inference without using the hybrid feature as a condition, g(E l , ω l ) is the conditional score estimation difference introduced by the text prompt feature, g(E nl , ω nl ) is the conditional score estimation difference introduced by the negative text prompt feature, g(E v , ω v ) is the conditional score estimation difference introduced by the video feature.
[0027] The beneficial effects of the present invention are:
[0028] Through CVALP pre-training, the tri-modal alignment of video-audio-text features is achieved, improving the semantic relevance and temporal accuracy of the generated audio;
[0029] Adopting the feature mixing strategy, while maintaining the video temporal information, the text semantic information is fused, enhancing the feature expression ability;
[0030] Introducing a dual guidance mechanism, while ensuring the audio generation quality, better controllability is achieved;
[0031] Using the large language model to automatically generate text descriptions and combining with PPPR for standardization, providing a flexible personalized generation interface;
[0032] The overall scheme has achieved better results than the existing methods in terms of objective metrics (such as IS, FID, FAD, MKL) and subjective evaluation.
[0033] The present invention has successfully solved the key technical problems in video-to-audio generation, providing a new technical solution for fields such as multimedia content generation and virtual reality audio synthesis. Through experimental verification, the method of the present invention has achieved significant improvements in aspects such as semantic expression and temporal alignment, laying a foundation for further promoting the development of sound generation technology for human perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a block diagram of the TA-V2A system. DETAILED DESCRIPTION OF THE INVENTION
[0035] The following will provide a detailed description of a text-assisted video-to-audio generation system (TA-V2A) provided by the present invention with reference to the accompanying drawings:
[0036] As Figure 1 shown, the TA-V2A system of the present invention mainly includes the following key modules: a contrastive video-audio-language pre-training (CVALP) module, a feature mixing module, a latent diffusion model (LDM), and an inference guidance module. The working process of the system is as follows:
[0037] First, the system receives a video and a text description as inputs. Among them, the text description can be provided manually or automatically generated by a large language model (LLM). The system will perform automatic enhancement processing on the generated text.
[0038] In the CVALP module, the system uses specific encoders (E v , E a , E l ) to extract features from the video, audio, and text respectively. These features are then aligned and fused into audio-aligned feature E mix . The core of the CVALP module is to achieve multi-modal alignment of video, audio, and text features through contrastive learning. This simultaneous alignment improves feature quality, convergence speed, model robustness, and hierarchical learning.
[0039] During the training phase, given a video-audio-text triple (x v , x a , x l ) that matches the data, where represents a video segment with T v frames and a size of H×W, represents a mel spectrogram with T a time steps and M mel bands, and x l is the corresponding text description. The present invention uses video, audio, and text encoders f V (·), f A(·) and f L (·) Extract features from video and audio where T = 32 is the number of time segments, C = 512 is the feature dimension, and they are extracted from the text To align the feature dimensions, the present invention applies temporal pooling to the video and audio features to obtain E v and This ensures that all features are in the same space for easy fusion and comparison. The present invention defines the contrastive loss function for the i-th cross-modal pair as:
[0040]
[0041] where and are feature vectors from different modalities x and y, M is the number of cross-modal pairs, and τ controls the softmax smoothness. The numerator represents the similarity between correct pairs, while the denominator sums the similarities between and all pairs. This loss function is not permutation symmetric because swapping and will change the calculation, always being the anchor point compared with all . This asymmetry is useful when one modality (such as text) provides stronger semantic guidance for the feature alignment of another modality (such as video or audio).
[0042] For the feature triple (E v , E a , E l ), the present invention defines the audio-centered loss function as follows:
[0043]
[0044] where and represent the i-th triple audio, text, and video embeddings respectively. N, N S and N T represent the total numbers of audio-text pairs, cross-video audio-video pairs, and intra-video different-time audio-video pairs respectively. The parameters λ and μ are weights that determine the relative importance of different loss components: 1 / (λ + μ) reflects the importance of text relative to video, and μ / (1 + λ) represents the emphasis on temporal alignment and semantic expression. They will be determined through experiments. Equation (2) captures the similarity metric between audio and text, while (3) and (4) express the similarity between audio and video from semantic and temporal perspectives. The final loss function (5) integrates these modality pairs. This loss function is used to train the encoder parameters of the three modalities.
[0045] To mix the features extracted from different modalities, the present invention adopts a feature mixing strategy that balances video and text information. Since the video modality contains both semantic and temporal features, while text usually only contains semantic features, blindly aligning the two (e.g., using methods such as cross-attention) may result in the loss of the temporal alignment information learned in CAVP. To solve this problem, the present invention uses a weighted average method or a concatenation method for feature fusion:
[0046]
[0047] E concat = concat[P(E v );P(E l )]#(7)
[0048] where P(E v ) and P(E l ) are linear projections that halve the dimension to keep the size of E concat unchanged.
[0049] The present invention uses positional encoding and projection τ θ to map E concat to an appropriate dimension, denoted as where MLP serves as the projection layer and PE represents the positional encoding. This mixed feature representation is used as the input feature vector for LDM training and inference.
[0050] LDM generates high-dimensional audio data in a low-dimensional latent space. Starting from the encoded Mel spectrogram z0 = ε(x a ), the diffusion process is carried out in the latent space, where noise is gradually added to z0 to form a sequence of latent variables z1, z2, …, z T . Each diffusion step is modeled as:
[0051]
[0052] where t represents the time step, β t is the diffusion coefficient, represents the normal distribution. For data generation, the model reverses the diffusion process, starting from a sample z T drawn from a Gaussian distribution. The neural network p θ predicts the reverse step at each time step:
[0053]
[0054] where μ θ represents the mean of the Gaussian distribution predicted by the neural network with parameters θ, and σ tDenotes the standard deviation, usually related to the time step t. After reverse diffusion, the final latent variable z0 is decoded back to the data space to generate a Mel spectrogram and then converted to audio samples using a vocoder. Here, the encoder ε and the decoder both directly adopt the frozen codecs of Stable Diffusion. The conditional loss function is given as:
[0055]
[0056] ∈ θ is the latent representation obtained from the original training data through the diffusion encoder, and the training process minimizes this loss function to train the parameters in the diffusion model.
[0057] During the inference process, the present invention adopts techniques such as classifier guidance (CG) and classifier-free guidance (CFG) to control the generation process. CG relies on an additional classifier P φ and guides the reverse process at each time step through the gradient of the class label log-likelihood . On the other hand, CFG combines conditional and unconditional score estimates to guide the reverse process. The present invention combines the characteristics of CG and CFG and applies dual guidance to enhance alignment:
[0058]
[0059] where γ and ω represent the scales of CG and CFG respectively, is the latent representation obtained from the input data z t at the t-th step of the inference stage of the latent diffusion model, ∈ θ (z t ,t,E mix ) is the latent representation for additional inference using mixed features as conditions, is the latent representation for additional inference without using mixed features as conditions, is the average diffusion coefficient. It should be noted that CFG uses E mix , while CG uses E v , because the alignment classifier P φ (y∣z t ,E v ) is trained for the alignment of audio-visual pairs. From the perspective of energy-based models (EBMs), multiple conditions can also independently affect the inference process without combination. The conditional probability can be estimated by the following formula:
[0060]
[0061] Here, the present invention defines g to represent the difference between unconditional and conditional score estimates:
[0062] g(c, ω) = ω(∈ θ (z t , t|c) - ∈ θ (z t , t))#(13)
[0063] where c is the condition and ω is the hyperparameter. Then the multi-condition inference step can be expressed as:
[0064]
[0065] where is the latent representation obtained from the input data in the inference stage, E v represents the video feature, E l represents the text prompt feature, E nl represents the negative text prompt feature. Correspondingly, ω v is the hyperparameter corresponding to the video feature, ω l is the hyperparameter corresponding to the text prompt feature, ω nl is the hyperparameter corresponding to the negative text prompt feature, ∈ θ (z t , t) is the latent representation for additional inference without using the mixed feature as a condition, g(E l , ω l ) is the difference in conditional score estimates introduced by the text prompt feature, g(E nl , ω nl ) is the difference in conditional score estimates introduced by the negative text prompt feature, g(E v , ω v ) is the difference in conditional score estimates introduced by the video feature. This method allows for a more flexible inference process by independently considering the effects of multiple conditions. Finally, the latent variables generated by LDM are converted into mel spectrograms by the decoder and then synthesized into the final audio waveform by the vocoder.
[0066] Another important innovation of the present invention is the introduction of a Portable Plug-and-Play Prompt Refiner (PPPR). During the inference process, the manually input prompt is standardized by the PPPR to ensure its correspondence with the AI-generated text used in training. If no manually provided prompt is available, the system will automatically generate a description of the video content using the Video-LlaMA2 model.
[0067] To verify the effectiveness of the present invention, a series of objective and subjective evaluation experiments were conducted. In the objective evaluation, the present invention adopted indicators such as the inception score (IS), Fréchet inception distance (FID), Fréchet audio distance (FAD), mean Kullback-Leibler divergence (MKL), and alignment accuracy (Align). As shown in Table 1, the TA-V2A system of the present invention outperforms the baseline method in most indicators. Especially when using the feature concatenation method, significant improvements are achieved in the FID and FAD indicators.
[0068] Table 1 is a performance comparison table
[0069]
[0070] In the subjective evaluation, the present invention invited 20 participants to rate the generated audio. The ratings were based on the mean opinion score (MOS) on a 5-point scale, covering two aspects: semantic consistency and temporal alignment. As shown in Table 2, when reasoning with user-modified prompts, the TA-V2A system obtained the highest MOS score in terms of semantic consistency. This indicates that the text control interface and PPPR text extension method introduced by the present invention can effectively improve the semantic consistency between the generated audio and video, and better conform to human understanding.
[0071] Table 2 is a semantic consistency comparison table
[0072]
[0073] Taking a badminton game video as an example, the spectrogram of the audio generated by TA-V2A is highly similar to the real audio in terms of temporal and frequency features, accurately capturing key sound events such as hitting the ball, and maintaining precise temporal synchronization with the actions in the video frame.
[0074] In summary, the TA-V2A system proposed by the present invention significantly improves the quality and semantic expression ability of video-to-audio generation through innovative text-aided methods and multi-modal feature fusion technologies. The system not only performs excellently in objective indicators but also demonstrates better semantic consistency and temporal alignment ability in subjective evaluations, providing a new effective solution for the video-to-audio generation task.
[0075] Although specific embodiments and drawings of the present invention are disclosed for illustrative purposes, aiming to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiments and drawings.
Claims
1. A method for video-to-audio generation based on text assistance, the steps of which include: 1) Construct a training set, where each training sample in the training set is a triple (x v , x a , x l ); where x v represents a video clip, x a represents the Mel spectrogram of x v , and x l is the corresponding text description of x v . 2) Use the text-audio-video contrast pre-training module to extract triples (x v , x a , x l ) video features, audio features, and text features; map the extracted text features and the video features and audio features after temporal pooling to the same feature space to obtain text features, video features, and audio features of the same dimension; the text-audio-video contrast pre-training module includes a video encoder f V (·) for extracting video features, an audio encoder f A (·) for extracting audio features, and a text encoder f L (·); 3) Align the text features, video features, and audio features of the same triple processed in step 2) with the audio features as the center, and calculate the alignment loss of the corresponding triple; calculate the total loss value based on the alignment losses of each triple 4) According to the total loss value Optimize the text-audio-video contrast pre-training module; 5) For a target video and its text description; use the optimized video encoder f V (·) Extract the video features of the target video, and use the optimized text encoder f L (·) Extract the text features of the target video; then fuse the video features and text features of the target video, and map the fused features to a low-dimensional space through positional encoding and projection processing in sequence to obtain the conditional features of the target video; 6) Use a latent diffusion model to generate audio data corresponding to the target video according to the conditional features of the target video.
2. The method according to claim 1, wherein Use a video-large language model to generate a content description of the target video, and then use a portable plug-in prompt fine-tuner to standardize the content description of the target video to obtain a text description of the target video.
3. The method according to claim 1, wherein The text description of the target video is information input manually.
4. The method according to claim 1 or 2 or 3, characterized in that, Total loss value Among them, and respectively represent the audio feature, text feature and video feature of the i-th triple. N, N S and N T respectively represent the total number of audio-text pairs, the total number of cross-video audio-video pairs and the total number of audio-video pairs at different times within the same video. The parameters λ and μ are weight coefficients; and are feature vectors from different modalities x and y, corresponding to or corresponding to or M is the number of cross-modal pairs, and τ controls the softmax smoothness.
5. The method according to claim 1 or 2 or 3, characterized in that, Adopt a weighted average method or a splicing method to fuse the video features and text features of the target video.
6. The method according to claim 5, wherein The features after fusion by the weighted average method are The features after fusion by the splicing method are E concat = concat[P(E v )); P(E l ));]; where the parameters λ and μ are weight coefficients, and P(E v ) and P(E l ) are linear projections.
7. The method according to claim 1, characterized in that Using positional encoding and projection τ θ The fused feature E concat is mapped to a low-dimensional space, denoted as where MLP serves as the projection layer and PE represents positional encoding.
8. The method according to claim 1, wherein Adopt classifier guidance CG and classifier-free guidance CFG to jointly control the latent diffusion model to generate audio data corresponding to the target video according to the conditional features of the target video.
9. The method according to claim 8, characterized in that, The formulaic representation for generating the audio data corresponding to the target video is as follows: where γ represents the scale of CG, and ω represents the scale of CFG, is the latent representation obtained from the input data z at the t-th step in the inference stage of the latent diffusion model, E t represents the video feature, E v is the fused feature, P mix is the alignment classifier used for aligning the audio-visual pair, P φ (y∣z t ,t,E v ) is the alignment classifier used for aligning the audio-visual pair, ∈ θ (z t ,t,E mix ) is the latent representation for additional inference using the hybrid feature as a condition, is the latent representation for additional inference without using the hybrid feature as a condition, is the average diffusion coefficient.
10. The method according to claim 9, wherein The multi-condition inference step of the potential diffusion model is expressed as: where E l represents the text prompt feature, E nl represents the negative text prompt feature, ω v is the hyperparameter corresponding to the video feature, ω l is the hyperparameter corresponding to the text prompt feature, ω nl is the hyperparameter corresponding to the negative text prompt feature, g(E v , ω v ) is the latent representation that does not use the mixed feature as a condition for additional inference, g(E θ , t) for ∈ t , t) is the latent representation that does not use the mixed feature as a condition for additional inference, g(E l , a l ) is the conditional score estimation difference introduced by the text prompt feature, g(E nl , ω nl ) is the conditional score estimation difference introduced by the negative text prompt feature, g(E v , ω v ) is the conditional score estimation difference introduced by the video feature.