Method for improving video semantic understanding based on space-time convolution

By adopting a spatiotemporal convolution method based on video to sound effect generation technology, the video understanding ability and the stability of sound effect synthesis are enhanced, and the problems of insufficient spatial and temporal feature extraction and insufficient depth of video semantic understanding in the prior art are solved, thereby achieving higher quality sound effect synthesis.

CN120014516APending Publication Date: 2025-05-16GIANT MOBILE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510098087.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing video-to-sound effect generation technology has shortcomings in spatial and temporal feature extraction, depth of video semantic understanding and scene recognition, resulting in the generated sound effects not being exactly matched with the video content, and lacking realism and compatibility.

Method used

Using a spatiotemporal convolution-based method, by adding visual connectors and visual adapters, combining diffusion models and decoders, the spatiotemporal characteristics of videos are extracted and sound effects are generated, thereby enhancing the stability and reliability of video comprehension capabilities and sound effects synthesis.

Benefits of technology

It significantly improves the accuracy of video semantics understanding and the quality of sound effect synthesis, making the generated sound effects more in line with the video, and enhancing the robustness and scope of application of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014516A_ABST
    Figure CN120014516A_ABST
Patent Text Reader

Abstract

The invention relates to a method for improving video semantic understanding based on space-time convolution. The method comprises the following steps: S1, collecting training data from multiple channels; s2, performing primary processing on the training data; s3, marking the sound effect description of the training data by using a Clap model and / or a manual marking mode to obtain the training data in which the video, the sound effect and the sound effect description are matched; s4, training data are sent to the model for training in a video-sound effect description pairing mode; s41, a visual connector is additionally arranged behind the visual encoder, and S42, after passing through the visual connector, the video modality and the sound effect description are aligned through a visual adapter; s43, obtaining a Mel spectrum through a diffusion model and a decoder, and then obtaining an output sound effect audio through a vocoder; and S5, performing sound effect synthesis by using the trained model, and outputting a synthesized sound effect by taking a video frame as input. According to the invention, through fine processing of the input video frame, the overall quality of the synthesized sound effect is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video semantic understanding, and in particular to a method for improving video semantic understanding based on spatiotemporal convolution. Background Art

[0002] Existing video-to-audio (V2A) methods are mainly based on deep learning frameworks, which generate corresponding sound effects by analyzing the visual features of video content (such as objects, actions, scenes, etc.). However, video content contains not only spatial information, but also dynamic features of the temporal dimension. Some sounds in the video are often associated with specific spatiotemporal events. Therefore, how to more accurately extract the spatiotemporal information of the video to improve the accuracy of video semantic understanding has become a key issue in current V2A research.

[0003] The current V2A technology still faces some challenges, which are mainly reflected in the following aspects:

[0004] (1) Insufficient spatiotemporal feature extraction. Existing V2A methods mostly rely on convolutional neural networks (CNNs) to extract visual features of videos. Although CNNs are effective in extracting spatial features, their ability to capture dynamic features in the temporal dimension is limited. Therefore, existing technologies often have difficulty accurately capturing the temporal continuity of actions and events in videos, resulting in the generated sound effects not fully matching the actual content of the video.

[0005] (2) Insufficient depth of video semantic understanding. Video content is complex and diverse, and a single convolution operation cannot effectively understand high-level semantic information in the video, such as the interaction between characters and the cause-and-effect relationship of events. This lack of semantic understanding can easily lead to a lack of realism in the generated sound effects, making it difficult to express the emotions or atmosphere of the video.

[0006] (3) Lack of scene recognition and understanding. Sound effect generation not only depends on the main actions in the video, but is also closely related to the scene environment. For example, there are significant differences in the background sounds of indoor and outdoor scenes, and existing methods often ignore this factor, resulting in the generated sound effects not being suitable and natural in specific scenes.

[0007] Therefore, it is necessary to provide a method based on spatiotemporal convolution to improve video semantic understanding, which can significantly improve the overall quality of synthesized sound effects by fine-tuning the input video frames. Summary of the invention

[0008] The purpose of the present invention is to provide a method for improving video semantic understanding based on spatiotemporal convolution, which significantly improves the overall quality of synthesized sound effects by fine-tuning the input video frames.

[0009] In order to solve the problems existing in the prior art, the present invention provides a method for improving video semantic understanding based on spatiotemporal convolution, comprising the following steps:

[0010] S1: Collect training data from multiple channels, including video-sound effect pairing data;

[0011] S2: Perform preliminary processing on the collected training data;

[0012] S3: Using the Clap model and / or manual labeling method, the sound effect description is annotated on the training data after preliminary processing to obtain training data that matches the video, the sound effect, and the sound effect description;

[0013] S4: Send the training data to the model for training in the form of video-sound effect description pairs;

[0014] S41: In the model, a visual connector is added after the visual encoder;

[0015] S42: after the visual connector, align the video modality with the sound effect description through the visual adapter;

[0016] S43: Obtain the Mel spectrum through the diffusion model and the decoder, and then obtain the output sound effect audio through the vocoder;

[0017] S5: Use the trained model to synthesize sound effects, take video frames as input, and output synthesized sound effects.

[0018] Optionally, in the method for improving video semantic understanding based on spatiotemporal convolution, channels include: real-world video and sound effects, game video and sound effects, and film and television video and sound effects.

[0019] Optionally, in the method for improving video semantic understanding based on spatiotemporal convolution, the preliminary processing is as follows:

[0020] Segment the video-sound pairing data into 10-second segments;

[0021] Use manual scoring, Clap Score or Clip Score to filter the training data.

[0022] Optionally, in the method for improving video semantic understanding based on spatiotemporal convolution, the visual connector includes spatial convolution-spatiotemporal downsampling-spatial convolution-attention mechanism-projection.

[0023] Optionally, in the method for improving video semantic understanding based on spatiotemporal convolution, video frames are extracted in the following manner: a video frame sequence is extracted from the video at 6 frames per second, and each frame is processed uniformly and filled into a size of 336*336.

[0024] Optionally, in the method for improving video semantic understanding based on spatiotemporal convolution, a visual connector based on spatiotemporal convolution is constructed as follows: a visual connector is constructed, and input video frame features are processed to obtain a video semantic vector of a fixed length.

[0025] Optionally, in the method for improving video semantic understanding based on spatiotemporal convolution, the extracted video semantic vector is injected into the original sound effect generation model through a visual adapter, so that the original sound effect generation model has the ability to generate sound effects based on video semantic features.

[0026] Optionally, in the method for improving video semantic understanding based on spatiotemporal convolution, the model optimizes model parameters by continuously reducing the mean square error between the generated Mel spectrum and the target Mel spectrum.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] (1) Improved model video understanding ability: The V2A model’s ability to understand videos has been improved, making the generated sound effects more consistent with the video.

[0029] (2) Improve model robustness: Enhance the model’s performance on datasets with complex video composition, and improve the stability and reliability of video-based sound synthesis.

[0030] (3) Wide range of applications: Applicable to various sound effects synthesis applications, including animation / animation dubbing, game dubbing, film and television dubbing and other fields.

[0031] (4) Easy to integrate: The present invention can be seamlessly integrated with the existing sound effect synthesis system to enhance the sound effect synthesis performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 A flow chart of model training provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The specific implementation of the present invention will be described in more detail below in conjunction with the schematic diagram. The advantages and features of the present invention will become clearer based on the following description. It should be noted that the drawings are all in a very simplified form and are not in exact proportions, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.

[0034] Hereinafter, if the method described herein includes a series of steps, the order in which the steps are presented herein is not necessarily the only order in which the steps may be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.

[0035] Current V2A technology still faces some challenges, mainly in the following aspects: (1) Insufficient spatiotemporal feature extraction. (2) Insufficient depth of video semantic understanding. (3) Lack of scene recognition and understanding.

[0036] In order to solve the problems existing in the prior art, the present invention provides a method for improving video semantic understanding based on spatiotemporal convolution, comprising the following steps:

[0037] S1: Collect training data from multiple channels, including video-sound effect pairing data;

[0038] For example, channels include: real-world video and sound effects, game video and sound effects, and film and television (TV series, movies, cartoons, commercials, etc.) video and sound effects.

[0039] S2: Perform preliminary processing on the collected training data in the following way:

[0040] Segment the video-sound pairing data into 10-second segments;

[0041] Use manual scoring, Clap Score or Clip Score to filter the training data to remove data with problems such as no sound effects, low match between video and sound effects, poor video quality, and poor sound quality to ensure the quality of training data.

[0042] S3: Using the Clap model and / or manual labeling method, the sound effect description is annotated on the training data after preliminary processing to obtain high-quality and rich training data with matching video, sound effect and sound effect description;

[0043] S4: Figure 1As shown, (1) Visual Encoder: Input: Receive video frames as input; Output: Generate a feature representation with the same length as the number of video frames, containing the visual features of each video frame. (2) Text Encoder: Input: Text description of sound effects; Output: Generate a fixed-length semantic feature representation. (3) Visual Connector: Input: Visual features corresponding to video frames; Output: Generate a fixed-length visual feature representation. (4) Visual Adapter: Input: Fixed-length visual feature representation; Output: Implicit representation vector of visual features. (5) Diffusion: Input: Receive video feature vector and feature vector of text description as input; Output: Implicit representation of Mel spectrum. (6) Variational Auto Encoder Decoder (VAE Decoder): Input: Receive implicit representation of Mel spectrum as input; Output: Mel spectrum. (7) Vocoder: Input: Receive Mel spectrum as input; Output: Sound effect audio.

[0044] The training data is fed into the model for training in the form of video-sound effect description pairs;

[0045] Preferably, the video frames are extracted in the following manner: a video frame sequence is extracted from the video at 6 frames per second, and each frame is uniformly processed and filled into a 336*336 size.

[0046] S41: In the model, in order to improve the model's ability to understand complex video content, a visual connector is added after the visual encoder to form the following Figure 1 As shown in the Visual Connector in Figure 2; further, the visual connector includes spatial convolution-spatiotemporal downsampling-spatial convolution-attention mechanism-projection.

[0047] Preferably, the construction of the visual connector based on spatiotemporal convolution is as follows: construct the visual connector, process the input video frame features, and obtain a video semantic vector of a fixed length.

[0048] S42: After the visual connector, the video modality is aligned with the sound effect description (text modality) through the visual adapter / attention mechanism;

[0049] Preferably, the extracted video semantic vector is injected into the original sound effect generation model through a visual adapter, so that the original sound effect generation model has the ability to generate sound effects based on video semantic features.

[0050] S43: Obtain the Mel spectrum through the diffusion model and the decoder, and then obtain the output sound effect audio through the vocoder;

[0051] Preferably, the model optimizes the model parameters by continuously reducing the mean square error between the generated Mel spectrum and the target Mel spectrum.

[0052] S5: Use the trained model to synthesize sound effects, taking video frames as input, and the model outputs sound effects that are highly consistent with the video and highly natural.

[0053] In summary, compared with the prior art, the present invention has the following advantages:

[0054] (1) Improved model video understanding ability: The V2A model’s ability to understand videos has been improved, making the generated sound effects more consistent with the video.

[0055] (2) Improve model robustness: Enhance the model’s performance on datasets with complex video composition, and improve the stability and reliability of video-based sound synthesis.

[0056] (3) Wide range of applications: Applicable to various sound effects synthesis applications, including animation / animation dubbing, game dubbing, film and television dubbing and other fields.

[0057] (4) Easy to integrate: The present invention can be seamlessly integrated with the existing sound effect synthesis system to enhance the sound effect synthesis performance of the system.

[0058] The above is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any technician in the relevant technical field, without departing from the scope of the technical solution of the present invention, makes any form of equivalent replacement or modification to the technical solution and technical content disclosed in the present invention, which does not depart from the content of the technical solution of the present invention and still falls within the protection scope of the present invention.

Claims

1. A method for improving video semantic understanding based on spatiotemporal convolution, characterized in that: The following steps are involved: S1: Collect training data from multiple channels, including video-sound effect pairing data; S2: Perform preliminary processing on the collected training data; S3: Using the Clap model and / or manual labeling method, the sound effect description is annotated on the training data after preliminary processing to obtain training data that matches the video, the sound effect, and the sound effect description; S4: Send the training data to the model for training in the form of video-sound effect description pairs; S41: In the model, a visual connector is added after the visual encoder; S42: after the visual connector, align the video modality with the sound effect description through the visual adapter; S43: Obtain the Mel spectrum through the diffusion model and the decoder, and then obtain the output sound effect audio through the vocoder; S5: Use the trained model to synthesize sound effects, take video frames as input, and output synthesized sound effects.

2. The method for improving video semantic understanding based on spatiotemporal convolution as claimed in claim 1, characterized in that: Channels include: Real-world video and sound effects, game video and sound effects, film and television video and sound effects.

3. The method for improving video semantic understanding based on spatiotemporal convolution as claimed in claim 1, characterized in that: The initial processing method is as follows: Segment the video-sound pairing data into 10-second segments; Use manual scoring, Clap Score or Clip Score to filter the training data.

4. The method for improving video semantic understanding based on spatiotemporal convolution as claimed in claim 1, characterized in that: The visual connector includes spatial convolution-temporal downsampling-spatial convolution-attention mechanism-projection.

5. The method for improving video semantic understanding based on spatiotemporal convolution as claimed in claim 4, characterized in that: The video frames are extracted in the following manner: a video frame sequence is extracted from the video at 6 frames per second, and each frame is processed uniformly and filled into a 336*336 size.

6. The method for improving video semantic understanding based on spatiotemporal convolution as claimed in claim 5, characterized in that: The construction of a visual connector based on spatiotemporal convolution is as follows: construct a visual connector, process the input video frame features, and obtain a video semantic vector of a fixed length.

7. The method for improving video semantic understanding based on spatiotemporal convolution as claimed in claim 6, characterized in that: The extracted video semantic vector is injected into the original sound effect generation model through a visual adapter, so that the original sound effect generation model has the ability to generate sound effects based on video semantic features.

8. The method for improving video semantic understanding based on spatiotemporal convolution as claimed in claim 1, characterized in that: The model optimizes the model parameters by continuously reducing the mean square error between the generated Mel spectrum and the target Mel spectrum.

Citation Information

Cited By

  • Audio-visual content synchronous sound effect synthesis method

    CN120358379A