Audio and video joint characterization method based on multi-mode association learning
Through the multimodal correlation learning method, the modal internal and external correlation of audio and video features are enhanced, and the high correlation characteristics are selectively fused, which solves the problems of insufficient information of a single visual modality and imbalance of modal information, and improves the accuracy of joint representation of audio and video.
Patent Information
- Application Number
- CN202510643739.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, a single visual mode neural network ignores the information of the audio mode, resulting in limited model understanding capabilities, and there is a problem of modal information imbalance when the audio and video modes are fusion.
The multimodal correlation learning method is adopted to extract audio and video features through the pre-trained CNN network, and the modal enhancement-interactive module is used to enhance modal internal and external feature and associate learning, and the dynamic fusion module is combined with the dynamic fusion module to select high correlation features for fusion.
Effectively utilizing complementary information of audio and video modalities improves the model's understanding ability and fusion effect, solves the problem of modal information imbalance, and improves the accuracy of joint representation of audio and video.
Smart Images

Figure CN120375259A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio - video fusion, and specifically relates to an audio - video joint representation method based on multi - modal correlation learning. Background Art
[0002] With the development of deep learning technology, deep neural networks with powerful feature - processing capabilities have been proposed and widely applied to video tasks, such as abnormal behavior detection in videos, video summary generation, video retrieval and recommendation, etc., and have achieved remarkable results. However, traditional neural networks are trained based on a single visual modality, ignoring the audio modality that also contains rich semantics. Therefore, some additional helpful information in the video will be missed, resulting in limited model understanding ability and application scenarios.
[0003] Some multi - modal - based methods have been proposed in this context. They jointly learn visual and audio representations by leveraging complementary information from different modalities to improve performance. However, they also face some challenges, such as the problem of unbalanced modal information in the audio - video modal fusion process. That is, due to different scenarios, the correlation between video and audio information varies. Direct fusion in scenarios with weak correlation may introduce noise, which becomes a bottleneck problem affecting multi - modal representation ability. Summary of the Invention
[0004] In order to solve the problems of insufficient representation information of a single video modality and unbalanced modal information in audio - video joint representation, the purpose of the present invention is to provide an audio - video joint representation method based on multi - modal correlation learning. The specific technical solutions adopted are as follows:
[0005] The present invention provides an audio - video joint representation method based on multi - modal correlation learning, and the method includes the following steps:
[0006] Obtain video data; separate and cut the video data to generate video modal segments and audio modal segments based on 16 frames; use a pre - trained CNN network to extract the deep features of the video modal segments and the deep features of the audio modal segments respectively;
[0007] Input the deep features into a two - stage modality enhancement - interaction module to enhance the unique features of the modalities under global information and perform inter - modal correlation learning;
[0008] Based on the results of the correlation learning, use a dynamic fusion module to select highly correlated audio and video features for fusion to obtain a fusion result.
[0009] Preferably, the two - stage includes a first stage and a second stage;
[0010] In the first stage, the modal features are enhanced. Based on the multi-head self-attention layer, the in-modal learning is respectively carried out on the depth features of the extracted video modal segments and the depth features of the audio modal segments to obtain video enhanced features and audio enhanced features;
[0011] In the second stage, based on the correlation learning, the multi-head cross-attention layer is used to calculate the cross-attention between the audio and video modalities to generate the inter-modal context features.
[0012] Preferably, the calculation formula of the video enhanced features is:
[0013]
[0014] where F v is the video feature, w v , w q and w k are the weights of value, query and key respectively. Value represents the value vector, query represents the query vector, key represents the key vector, D v is the dimension of the video feature, softmax is the activation function, is the context feature of the video feature, F v is the input video feature, is the video enhanced feature, (w k F v ) T represents the video feature after linear transformation.
[0015] Preferably, the calculation formula of the inter-modal context features is:
[0016]
[0017] where is the inter-modal context feature, is the transposed audio enhanced feature, D a is the feature dimension of the audio enhanced feature.
[0018] Preferably, after generating the inter-modal context features, it further includes: introducing a downsampling and upsampling fully connected layer and a GELU activation function to smooth the inter-modal context features.
[0019] Preferably, based on the result of the correlation learning, a dynamic fusion module is adopted to select the highly correlated audio and video features for fusion to obtain the fusion result, including:
[0020] The dynamic fusion module introduces a learnable selection function, adjusts the weights of audio features by considering the importance of audio features relative to the video modality, combines the selected features with the smoothed inter-modal context features, fuses them with the video enhanced features, and uses a fully connected layer to refine the fused representation to generate the fused result of audio and video.
[0021] Preferably, the specific formula for the fused result of audio and video is:
[0022]
[0023] Among them, F fus represents the fused result of audio and video, f sel represents the selection function, represents the smoothed inter-modal context features, FC represents the fully connected layer, represents the audio enhanced features, F sel represents the audio selected features, Sigmoid represents the normalization function, W sel represents the weight matrix.
[0024] The present invention has at least the following beneficial effects:
[0025] 1. Based on the modality enhancement-interaction module, the present invention enables video features and audio features to be learned within the modality, allows interaction between intra-modal features, so that sufficient global context information can be obtained, and solves the problem of insufficient feature information extracted by the pre-trained network; at the same time, it enables them to selectively and interactively focus on the complementary information of each modality, thereby focusing on task-related information and key features and better understanding the semantic association between modalities.
[0026] 2. The dynamic fusion module adopted by the present invention can dynamically adjust the weights of audio features according to the importance of audio features relative to the video modality, adaptively reduce the noise introduced by relevant audio features, and ensure that only highly relevant audio features important for visual learning are used for joint representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a flowchart of a method for audio-visual joint representation based on multi-modal association learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0029] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, provide a detailed description of an audio-visual joint representation method based on multi-modal correlation learning proposed according to the present invention as follows.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0031] The following will specifically describe the specific solution of an audio-visual joint representation method provided by the present invention in conjunction with the accompanying drawings.
[0032] An embodiment of an audio-visual joint representation method based on multi-modal correlation learning:
[0033] This embodiment proposes an audio-visual joint representation method based on multi-modal correlation learning. As Figure 1 shown, an audio-visual joint representation method based on multi-modal correlation learning in this embodiment includes the following steps:
[0034] Step S1, obtain video data; separate and cut the video data to generate video modality segments and audio modality segments based on 16 frames; use a pre-trained CNN network to extract the depth features of the video modality segments and the depth features of the audio modality segments respectively.
[0035] First, collect the video data to be processed; then, use Adobe Premiere Pro software to separate the video modality and audio modality of the video data, and use a sliding window with a window size of 16 frames to cut the video modality and audio modality to generate non-overlapping video and audio sequences based on a group of 16 frames. Then, input them into a pre-trained CNN convolutional neural network to extract features respectively to generate video features and audio features Among them, represents the feature space set, T represents the sequence length, D V represents the video feature dimension, D A represents the audio feature dimension.
[0036] Step S2, input the depth features into a two-stage modality enhancement-interaction module to enhance the unique features of the modality under global information and perform inter-modal correlation learning.
[0037] Next, the depth features of the extracted video modality segments and the depth features of the audio modality segments will be used for modality features and inter-modal correlation learning will be performed.
[0038] Both stages model the intra-modal and inter-modal feature correlations based on the multi-head attention mechanism. In the first stage, due to the insufficient feature extraction ability of the pre-trained neural network, rich global semantic information cannot be obtained, so the unique features of the modality cannot be fully utilized. To solve this problem, an intra-modal feature enhancement module based on the multi-head self-attention mechanism is designed for intra-modal learning. Taking video features as an example, let F v serve as the input key, query, and value simultaneously for intra-modal self-attention calculation, model the global correlation information of intra-modal features, obtain the corresponding attention weights after softmax activation and combination with the corresponding weights. Finally, weighted and summed with the value to generate the enhanced video modality features, which contain rich context information and highlight the unique feature information of the video modality. The specific formula is:
[0039]
[0040] where F v is the video feature, w v , w q and w k are the weights of value, query, and key respectively, value represents the value vector, query represents the query vector, key represents the key vector, D v is the dimension of the video feature, softmax is the activation function, is the context feature of the video feature, F v is the input video feature, is the enhanced video feature, (w k F v ) T represents the video feature after linear transformation.
[0041] Similarly, the same operation is performed on the audio feature F a to obtain the enhanced audio feature
[0042] In the second stage, to make full use of the complementary information between modalities, an inter-modal interaction module based on the multi-head cross-attention mechanism is designed, which allows inter-modal features to interact and learn, enabling them to selectively and interactively focus on the complementary information of each modality, thus focusing on task-related information and key features and better understanding the semantic associations between modalities. The specific process is that the enhanced video feature serves as the query vector query, the enhanced audio feature serves as the key and value, and the cross-attention mechanism is used to model the inter-modal feature correlation information, and the context feature between the audio and video modalities is generated after softmax activation and combination with the value. The specific formula is:
[0043]
[0044] Among them, is the inter-modal context feature, is the transposed audio enhancement feature, D a is the feature dimension of the audio enhancement feature.
[0045] Finally, in order to ensure smooth interaction between modalities while preserving modal characteristics, a full connection layer of dimensionality reduction - dimensionality increase is introduced, with a GELU activation function embedded in the middle to introduce non-linearity. The specific formula is:
[0046]
[0047] Among them, represents the smoothed inter-modal context feature, Down represents the dimensionality reduction full connection layer, and Up represents the dimensionality increase full connection layer.
[0048] The method provided in this embodiment maps the input to a shared bottleneck representation, which reduces the computational complexity while enhancing the expressive power of the features, thus effectively promoting context-aware fusion.
[0049] Step S3, based on the result of the association learning, a dynamic fusion module is adopted to select highly relevant audio and video features for fusion to obtain a fusion result.
[0050] Since not all audio features can provide effective semantic assistance for the joint representation, in order to reduce the impact of irrelevant audio modal features on the joint representation and maintain the balance after fusion, this embodiment designs a dynamic fusion module to only select audio features highly relevant to the corresponding video for fusion.
[0051] Specifically, a learnable selection function f sel is introduced, which performs a dot product on the learned weight matrix and the audio enhancement feature and is activated through the Sigmoid function to obtain the audio selection feature F sel . The specific formula is:
[0052]
[0053] In order to make full use of the correlation information of the inter-modal features, the smoothed inter-modal context feature is further fused with the selection feature F sel and the enhanced video feature, and the fused representation is refined through a full connection layer. The specific formula is:
[0054]
[0055] Among them, Ffus represents the fusion result of audio and video, f sel represents a selection function represents the smoothed cross-modal context features, and FC represents a fully connected layer represents the audio enhancement features, F sel represents the audio selection features, and Sigmoid represents a normalization function, W sel represents a weight matrix
[0056] In this embodiment, through the method of dynamic fusion, the most relevant information in the two modalities can be captured, so as to effectively filter out audio noise, balance the modal information of video and audio, complementarily represent while retaining the modal characteristics, and finally generate a refined joint representation of audio and video
[0057] The method provided in this embodiment designs a feature enhancement module based on the self-attention mechanism, enabling the video features and audio features to be learned within the modality, allowing interactions between the intra-modal features, thus being able to obtain sufficient global context information and solving the problem of insufficient feature information extracted by the pre-trained network; at the same time, enabling them to selectively and interactively focus on the complementary information of each modality, so as to focus on the task-related information and key features and better understand the semantic associations between modalities
[0058] The dynamic fusion module adopted by the method provided in this embodiment can dynamically adjust the weights of the audio features according to the importance of the audio features relative to the video modality, adaptively reduce the noise introduced by the relevant audio features, and ensure that only the highly relevant audio features important for visual learning are used for joint representation
[0059] This embodiment designs a fully connected layer for dimensionality reduction and then dimensionality increase, with a GELU activation function embedded in the middle to introduce non-linearity, enhancing the expressive power of the features while reducing the computational complexity. And this design ensures smooth interaction between different modalities while retaining the unique information of each modality
[0060] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention
Claims
1. An audio-visual joint representation method based on multi-modal correlation learning, characterized in that The method includes the following steps: Obtain video data; separate and cut the video data to generate video modality segments and audio modality segments based on 16 frames; use a pre-trained CNN network to extract the depth features of the video modality segments and the depth features of the audio modality segments respectively; Input the depth features into a two-stage modality enhancement-interaction module to enhance the unique features of the modalities under global information and perform cross-modal correlation learning; Based on the results of the correlation learning, use a dynamic fusion module to select highly correlated audio and video features for fusion to obtain a fusion result.
2. The audio-visual joint representation method based on multi-modal correlation learning according to claim 1, wherein, The two stages include a first stage and a second stage; In the first stage, the modality features are enhanced. Based on the multi-head self-attention layer, intra-modal learning is performed on the depth features of the extracted video modality segments and the depth features of the audio modality segments respectively to obtain video enhanced features and audio enhanced features; In the second stage, based on the correlation learning, the multi-head cross-attention layer is used to perform cross-attention calculation between the audio and video modalities to generate cross-modal context features.
3. The audio-visual joint representation method based on multimodal correlation learning according to claim 2, wherein The calculation formula for the video enhanced features is: Among them, F v is the video feature, w v , w q and w k are the weights of value, query, and key respectively. Value represents the value vector, query represents the query vector, key represents the key vector, D v is the dimension of the video feature, softmax is the activation function, is the context feature of the video feature, F v is the input video feature, is the video enhancement feature, (w k F v ) T represents the video feature after linear transformation.
4. The audio-visual joint representation method based on multi-modal correlation learning according to claim 3, characterized in that, The calculation formula for the cross-modal context features is: Among them, is the inter-modal context feature, is the transposed audio enhancement feature, D a is the feature dimension of the audio enhancement feature.
5. The audio-visual joint representation method based on multi-modal correlation learning according to claim 2, wherein After generating the cross-modal context features, it further includes: introducing a downsampling and upsampling fully connected layer and a GELU activation function to smooth the cross-modal context features.
6. The audio-visual joint representation method based on multi-modal association learning according to claim 5, characterized in that The step of using a dynamic fusion module to select highly correlated audio and video features for fusion to obtain a fusion result based on the results of the correlation learning includes: The dynamic fusion module introduces a learnable selection function, adjusts the weights of the audio features by considering the importance of the audio features relative to the video modality, combines the selected features with the smoothed cross-modal context features, and fuses them with the video enhanced features, and uses a fully connected layer to refine the fusion representation to generate the fusion result of the audio and video.
7. The audio-visual joint representation method based on multi-modal correlation learning according to claim 6, wherein The specific formula for the fusion result of the audio and video is: Among them, F fus represents the fusion result of audio and video, and f sel represents the selection function, represents the smoothed inter-modal context feature, FC represents the fully connected layer, represents the audio enhancement feature, and F sel represents the audio selection feature, Sigmoid represents the normalization function, and W sel represents the weight matrix.