Audio-visual event identification and positioning method, system and device based on bidirectional collaborative attention guidance and medium
Through the bidirectional collaborative guidance attention model, the two-way guidance of auditory and visual features is used, combined with the bilinear pooling method, the problem of noise interference in audio-visual event positioning is solved, and the accuracy of event recognition and positioning is improved.
Patent Information
- Application Number
- CN202510542675.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
AI Technical Summary
The existing audio-visual event positioning methods ignore the guiding role of visual mode on audio mode, resulting in noise and weak signal interference, affecting the multimodal fusion process, and reducing the accuracy of audio-visual event recognition and positioning.
The two-way collaborative guidance attention model is adopted to guide visual features and visual features to guide auditory features through auditory features, combined with the bilinear pooling method and attention mechanism, balance multimodal input, reduce noise interference, and improve feature correlation.
It improves the accuracy of audio-visual event recognition and positioning, enhances the robustness of visual and auditory modes, reduces noise interference, and achieves more efficient feature fusion and event recognition.
Smart Images

Figure CN120451866A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of audio-visual event positioning, and specifically relates to an audio-visual event recognition and positioning method, system, device and medium based on bidirectional collaborative attention guidance. Background Art
[0002] The task of audiovisual event localization aims to use machines to locate events / actions in videos and identify their categories. In recent years, it has become a challenging task in video understanding. Most existing methods rely primarily on unimodal event recognition and localization, which primarily relies on the video modality. Representative methods include recurrent neural networks, disentangled representation learning, generative adversarial networks, and attention mechanisms. These methods process the video modality into RGB frames as input to locate and identify events. However, the video modality is easily disturbed by the visual background, which can lead to significant errors in methods that use visual information to identify and localize events.
[0003] Inspired by the human multimodal perception system, researchers are exploring the complementary role of audio in video, leading to the gradual evolution of audiovisual event localization from unimodal event recognition to multimodal audiovisual event recognition. Compared to unimodal event localization, multimodal audiovisual event localization requires the simultaneous integration of visual and auditory modalities. For example, in the case of a dog barking event, while a puppy may be difficult to observe in a visual scene, the sound of the barking dog can draw our attention to the event and locate the area where it occurred.
[0004] The multimodal audiovisual event localization task aims to integrate cues from both video and audio modalities and guide behavior, enabling intelligent systems to perceive their surroundings through both hearing and vision, much like humans do. This overcomes the limitations of single-modality information constraints and improves recognition and localization accuracy. Attention is a key mechanism for integrating and capturing audiovisual relationships. It mimics the human perceptual system, capturing the dependencies between visual and auditory modalities and achieving a more complete representation of events. Therefore, utilizing attention to learn and capture audiovisual relationships is crucial in audiovisual recognition and localization.
[0005] Multimodal feature fusion is the core research direction of audiovisual event localization tasks. The fusion methods mainly include attention-based methods and bilinear pooling-based methods.
[0006] 1. Existing methods only consider attention from sound to vision, ignoring the modeling of visual-guided attention to sound, thus overlooking some cross-modal relationships. This kind of audiovisual attention is crucial for improving the robustness of both visual and auditory modalities.
[0007] 2. Existing Methods Noise and inherent weak signals in video and audio can interfere with the fusion process, affecting the process of the visual modality guiding the auditory modality and vice versa.
[0008] Patent application CN 118427394 A discloses a method for audiovisual event recognition based on cross-modal perceptual fusion. This invention incorporates an audio-guided spatial attention module, which leverages the audio signal's guiding power to direct visual attention, precisely highlighting information features and the most effective visual information areas for video event recognition while minimizing background interference. This approach emphasizes the guiding role of the audio signal on the video signal but neglects the guiding role of the video signal on the audio signal, thus limiting audiovisual learning capabilities.
[0009] Patent application publication number CN 117037046 A discloses an audiovisual event detection method, apparatus, storage medium, and electronic device. This invention discloses an audiovisual event detection method that extracts target video and audio from target audio and video data, segments the target video and audio, and fused features of the video and audio features of the same event segment to represent the audiovisual event semantics of the audio and video pair, ultimately determining the audiovisual event detection result. This invention effectively fuses audio and video features, but fails to account for the impact of noise and inherent weak signals in the video and audio during the fusion process, resulting in reduced audiovisual event detection accuracy. Summary of the Invention
[0010] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a method, system, device and medium for audio-visual event recognition and positioning based on bidirectional collaborative attention guidance. The method uses auditory guidance to guide visual attention while introducing sound guidance to guide visual attention. The bidirectional guidance method allows the visual modality and the auditory modality to serve as each other's guidance signals, balances multimodal input, and makes full use of the complementary semantic information of the auditory and visual modalities; by combining the bilinear pooling method and improved attention to explore the correlation between auditory features and visual features, enhance the relationship between modalities in the channel dimension and spatial dimension, reduce the noise of the visual modality and the auditory modality, realize effective audio-visual relationship learning, and thus improve the accuracy of audio-visual event recognition and positioning.
[0011] In order to achieve the above object, the technical solution adopted by the present invention is:
[0012] A method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance includes the following steps:
[0013] Step 1: Build an audiovisual event recognition and localization model based on bidirectional collaborative attention guidance;
[0014] Step 2: Design a loss function, and continuously train and optimize the audiovisual event recognition and localization model based on bidirectional collaborative attention guidance constructed in step 1 through the loss function. When the loss function is minimized, the optimal audiovisual event recognition and localization model based on bidirectional collaborative attention guidance is obtained.
[0015] Step 3: Input the target video into the optimal audiovisual event recognition and positioning model based on bidirectional collaborative attention guidance obtained in step 2 to obtain the optimal target event recognition accuracy and positioning information of the target event.
[0016] The implementation method of step 1 includes:
[0017] Step 1.1: Obtain the target video, preprocess the target video to obtain the video signal and audio signal, and use the convolutional neural network to extract the visual features of the video signal and the auditory features of the audio signal;
[0018] Step 1.2: Construct a bidirectional collaborative guided attention module, input the visual features of the video signal and the auditory features of the audio signal in step 1.1 into the bidirectional collaborative guided attention module, and obtain visual features guided by auditory features and calibrated with visual features, and auditory features guided by visual features and calibrated with auditory features;
[0019] Step 1.3: The visual features obtained in step 1.2, which have been guided by auditory features and calibrated by visual features, and the auditory features obtained in step 1.2, which have been guided by visual features and calibrated by auditory features, are respectively input into two bidirectional long short-term memory (Bi-LSTM) neural networks for spatiotemporal feature extraction and modeling, thereby obtaining visual features and auditory features with spatiotemporal characteristics.
[0020] Step 1.4: Fuse the spatiotemporal visual and auditory features obtained in step 1.3, and input the fused features into the classification and recognition module to identify and locate the event categories in the target video; thus, obtaining an audiovisual event recognition and location model based on bidirectional collaborative attention guidance.
[0021] The implementation method of step 1.1 includes:
[0022] Step 1.11: Get the target video The target video is decomposed into T continuous segments of 1 second in length. Each segment consists of a synchronized video signal segment and an audio signal segment. t V , S t A Represented as video signal segments and audio signal segments respectively;
[0023] Step 1.12: Sample each video signal segment at a fixed number of frames per second to obtain a fixed number of images for each video signal segment. Use a convolutional neural network (VGG-19) to extract feature maps of the fixed number of images for each video signal segment. Globally merge the feature maps of the fixed number of images per second to obtain a global feature map, which is the segment-level feature of each video signal segment.
[0024] Step 1.13: Convert each audio signal segment into a log-mel spectrogram and use a convolutional neural network (CNN) to extract segment-level features of each audio signal segment.
[0025] Step 1.14: Arrange the segment-level features of each video signal segment extracted in step 1.12 and the segment-level features of each audio signal segment extracted in step 1.13 in temporal order to form the visual features of the video signal and the auditory features of the audio signal.
[0026] The bidirectional collaborative attention guidance module described in step 1.2 includes auditory features guiding the attention of visual features and visual features guiding the attention of auditory features;
[0027] The auditory feature-guided visual feature attention is composed of an auditory feature-guided visual feature branch and a visual feature calibration branch; including a bilinear pooling block based on multimodal matrix decomposition, channel attention, spatial attention, and other parts including ReLU activation function, fully connected layer (FC), matrix multiplication, element-wise multiplication and addition operations;
[0028] The bilinear pooling block based on multimodal matrix decomposition includes a dot product, a fully connected layer (FC), regularization (Dropout), feature pooling (Sum pooling), and normalization (Normalization); the bilinear pooling block based on multimodal matrix decomposition simultaneously processes the auditory features of the audio signal and the visual features of the video signal, the auditory features of the audio signal and the visual features of the audio signal respectively pass through the fully connected layer (FC) and perform a dot product operation, and the result after the dot product operation is successively subjected to regularization (Dropout), feature pooling (Sum pooling), and normalization operations to obtain linear fusion features;
[0029] The workflow of the auditory feature-guided visual feature branch includes: simultaneously inputting the visual features of the video signal and the auditory features of the audio signal in step 1.1 into two bilinear pooling blocks based on multimodal matrix decomposition, respectively passing the auditory features of the audio signal and the visual features of the audio signal through a fully connected layer (FC) and performing a dot product operation, and performing regularization (Dropout), feature pooling (Sum pooling), and normalization on the results after the dot product operation, and outputting two linear fusion features; the process can be described as follows:
[0030]
[0031]
[0032] in, represents the output linear fusion feature, k represents the pooling weight, D(·) represents regularization (Dropout), SP1(·) represents feature pooling, F1 and F2 represent fully connected layers;
[0033] Linear fusion features Input to ReLU activation function and fully connected layer (FC), and then input to spatial attention and channel attention respectively. The output of channel attention is combined with linear fusion feature. Perform element-wise multiplication, perform matrix multiplication on the result of element-wise multiplication and the result of spatial attention, and output the result The process is described as:
[0034]
[0035] Among them, W t S , W t C are the weights of spatial attention and channel attention respectively;
[0036] The visual feature calibration branch is composed of channel attention and spatial attention. The workflow of the visual feature calibration branch includes: inputting the visual features of the video signal in step 1.1 into the spatial attention to obtain features The features Input into channel attention to obtain features The result of channel attention and spatial attention results Perform matrix multiplication and generate visual calibration features through ReLU activation function and fully connected layer (FC); the process is described as:
[0037]
[0038] in, is the visual calibration feature, Sigmoid(·) represents the activation function;
[0039] The function of guiding visual features by adaptively adjusting auditory features through visual calibration features; the visual calibration features are controlled by setting the hyperparameter β The weight of , will be controlled by the hyper parameter β to add the visual calibration features and the auditory features to guide the visual features. The process is described as:
[0040]
[0041] Among them, β is a hyperparameter, v t Represents visual features guided by auditory features and calibrated with visual features;
[0042] The visual feature-guided auditory feature attention is composed of a visual feature-guided auditory feature branch and an auditory feature calibration branch; it includes a bilinear pooling block based on multimodal matrix decomposition, channel attention, spatial attention, and temporal attention. Other parts also include a ReLU activation function, a fully connected layer (FC), a global pooling operation (Global Pooling), and an addition operation;
[0043] The workflow of the visual feature-guided auditory feature branch includes: the visual features of the video signal in step 1.1 are input into the spatial attention and channel attention, and the process is expressed as:
[0044]
[0045] Among them, W1 and W2 represent the spatial attention and channel attention weights respectively, and F sq Indicates information compression operation;
[0046] In step 1.1, the auditory features of the audio signal are input into the temporal attention to extract and mine the auditory feature representation. The process is described as follows:
[0047]
[0048] Q = a t W Q ,K=a t W K ,V=a t W V ,
[0049] Among them, W Q , W K , W V represents the weight matrix, Q, K, V represent the query, key, and value generated according to the input transformation;
[0050] Input the visual features of the video signal in step 1.1 into the output results of spatial attention and channel attention, perform a global pooling operation, and output the result after global pooling The result after using global pooling As a query vector, we extract the auditory features related to the video, and use the visual features to guide the auditory features. The process is described as follows:
[0051]
[0052] in, is the auditory feature guided by the visual feature, represents the weight matrix, q, k, v represent the query, key and value generated according to the input transformation;
[0053] Auditory features guided by visual features The final auditory features guided by visual features are generated through the ReLU activation function and the fully connected layer (FC);
[0054] The auditory feature calibration branch: the auditory features of the audio signal in step 1.1 and the visual features guided by the auditory features and calibrated by the visual features are input into the bilinear pooling block based on multimodal matrix decomposition. The auditory features of the audio signal and the visual features of the video signal are respectively passed through the fully connected layer (FC) and dot product operation. The results after the dot product operation are successively regularized (Dropout), feature pooling (Sum pooling), and normalization (Normalization) to output the auditory calibration features.
[0055] Adaptively adjust visual features to guide auditory features through auditory calibration features and set hyperparameters Controlling auditory features guided by visual features The weights will be passed through the hyperparameters The auditory features guided by the visual features are added to the auditory calibration features. The process is described as follows:
[0056]
[0057] in, is a hyperparameter, a t Represents auditory features guided by visual features and calibrated with auditory features.
[0058] The specific method of step 1.4 includes:
[0059] Step 1.41: The visual features and auditory features with spatiotemporal characteristics output by the bidirectional long short-term memory (Bi-LSTM) neural network in step 1.3 are added together to obtain fused features.
[0060] Step 1.42: Input the fused features obtained in step 1.41 into the classification and recognition module. The classification and recognition module consists of two fully connected layers. It performs audio-visual event recognition based on the fused features to obtain the identification and location information of the target event.
[0061] In step 2, the loss function is calculated as follows: according to different audiovisual event localization tasks, the loss function is divided into a loss function under full supervision and a loss function under weak supervision;
[0062] In the fully supervised setting, the loss function is composed of the event-related score s r , the known segment-level event labels Y in the dataset tc , background label y t ={y t ||yt ∈{0,1},t=1,…,T} and the segment-level event labels O predicted by the model in the fully supervised setting tc Calculation; the calculation process is:
[0063] s r =δ(AV·F1),
[0064]
[0065] Among them, F1 represents the fully connected layer, AV represents the visual features and auditory features with spatiotemporal characteristics output by the bidirectional long short-term memory (Bi-LSTM) neural network, and the fusion features are obtained by feature addition. δ represents the activation function, and L BCE (s r ,y t ) represents the event-related score s r and background label y t ={y t ||y t Cross entropy loss ∈{0,1},t=1,...,T};
[0066] The loss function in the weakly supervised setting is the segment-level event labels predicted by the model in the weakly supervised setting. and the known video-level labels Y in the dataset c Calculation, the calculation process is:
[0067]
[0068] Among them, s represents the softmax function and λ represents the hyperparameter.
[0069] The present invention also provides an audiovisual event recognition and positioning system based on bidirectional collaborative attention guidance, comprising:
[0070] A model building module for building an audiovisual event recognition and localization model based on bidirectional collaborative attention guidance;
[0071] The model training module is used to implement the designed loss function, and continuously train and optimize the audiovisual event recognition and localization model based on bidirectional collaborative attention guidance through the loss function. When the loss function is minimized, the optimal audiovisual event recognition and localization model based on bidirectional collaborative attention guidance is obtained;
[0072] The recognition and positioning result output module is used to input the target video into the optimal audio-visual event recognition and positioning model based on bidirectional collaborative attention guidance to obtain the optimal target event recognition accuracy and target event positioning information.
[0073] The present invention also provides an audiovisual event recognition and positioning device based on bidirectional collaborative attention guidance, comprising:
[0074] Memory: a computer-readable device storing a computer program for the above-mentioned method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance;
[0075] Processor: used to implement the method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance when executing the computer program.
[0076] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance.
[0077] Compared with the prior art, the present invention has the following beneficial effects:
[0078] 1. The present invention proposes a method for audiovisual event recognition and positioning based on bidirectional collaborative attention guidance. By constructing a bidirectional collaborative attention guidance module, the complementary information of the visual modality and auditory modality in the video is fully utilized, and the accuracy of audiovisual event recognition and positioning is improved through audiovisual learning.
[0079] 2. Based on auditory features guiding visual features, this invention adds a branch for visual features to guide auditory features, achieving bidirectional collaborative attention modeling of visual and auditory features. Furthermore, by adding visual and auditory feature calibration branches, the model focuses on sound-related visual areas and visually related sound segments.
[0080] 3. The present invention combines the bilinear pooling method based on multimodal matrix decomposition and spatial-channel attention to explore the correlation features of auditory features and visual features, realizes efficient feature fusion through the bilinear pooling method based on multimodal matrix decomposition, and efficiently models the semantic features of auditory and visual features through spatial-channel attention, thereby reducing the noise generated in the process of fusion and guidance of visual features and auditory features.
[0081] In summary, the present invention first extracts visual and auditory features from video and audio signals through the VGG-19 network. These features are then fed into a bidirectional collaborative attention guidance module to obtain visual features calibrated with sound features and auditory features, and then these two features are fed into a bidirectional long short-term memory (Bi-LSTM) neural network to model the temporal characteristics of these features. The visual and auditory features with spatiotemporal characteristics are then summed and fused, and finally fed into a recognition and classification module to obtain the target time category and location information. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 It is the overall network structure diagram of the present invention.
[0083] Figure 2 It is the bidirectional collaborative attention guidance module of the present invention. DETAILED DESCRIPTION
[0084] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings.
[0085] The present invention first extracts visual and auditory features from video and audio signals using a VGG-19 network. These features are then fed into a bidirectional collaborative attention-guided module, which fuses the visual and auditory features using bilinear pooling. Attention is then used to query auditory information from the visual features, reducing visual noise. Visual global spatial features are then used to query auditory information from visual events, reducing audio noise. The visual and auditory modalities are allowed to serve as guidance signals for each other, and multimodal input is balanced through visual and auditory feature calibration branches. The visual features, guided by sound features and calibrated by visual features, and the auditory features, guided by visual features and calibrated by auditory features, are then fed into a bidirectional long short-term memory (Bi-LSTM) neural network to extract temporal characteristics. The outputs of the Bi-LSTM neural network are then combined and fused, and finally fed into a recognition and classification module to obtain the target's temporal category and location information. Simulation results in a test environment demonstrate that the present invention overcomes the shortcomings of existing technologies and achieves excellent recognition and localization results.
[0086] like Figure 1 As shown, a method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance includes the following steps:
[0087] Step 1: Build an audiovisual event localization and recognition model based on bidirectional collaborative attention guidance;
[0088] Step 1.1: Obtain the target video, pre-process the target video, decompose the target video into multiple continuous segments, obtain multiple synchronized video signal segments and audio signal segments from the decomposed video continuous segments, and extract the visual features of the video signal and the auditory features of the audio signal through a convolutional neural network;
[0089] The specific implementation steps of step 1.1 include:
[0090] Step 1.11: Get the target video The target video is decomposed into T continuous segments of 1 second in length. Each segment consists of a synchronized video signal segment and an audio signal segment. t V , S t A Represented as video signal segments and audio signal segments respectively;
[0091] Step 1.12: Obtain 16 frames of images for each video signal segment at a sampling frequency of 16 frames per second. Use a convolutional neural network (VGG-19) to extract feature maps of the 16 frames of images for each video signal segment. Combine the 16 frame feature maps per second through global merging to obtain segment-level features for each video signal segment.
[0092] Step 1.13: Sample each audio signal segment at 16000 Hz and convert it into a log-mel spectrogram. Use a convolutional neural network (CNN) to extract segment-level features for each audio signal segment.
[0093] Step 1.14: Arrange the segment-level features of each video signal segment extracted in step 1.12 and the segment-level features of each audio signal segment extracted in step 1.13 in temporal order to form visual features of the video signal and auditory features of the audio signal.
[0094] Step 1.2: Construct a bidirectional collaborative guided attention module, input the visual features of the video signal and the auditory features of the audio signal in step 1.1 into the bidirectional collaborative guided attention module, and obtain visual features guided by auditory features and calibrated by visual features, and auditory features guided by visual features and calibrated by auditory features.
[0095] The bidirectional collaborative attention guidance module described in step 1.2 includes auditory features guiding the attention of visual features and visual features guiding the attention of auditory features;
[0096] The auditory feature-guided visual feature attention is composed of an auditory feature-guided visual feature branch and a visual feature calibration branch; including a bilinear pooling block based on multimodal matrix decomposition, channel attention, spatial attention, and other parts including ReLU activation function, fully connected layer (FC), matrix multiplication, element-wise multiplication and addition operations;
[0097] The bilinear pooling block based on multimodal matrix decomposition includes a dot product, a fully connected layer (FC), regularization (Dropout), feature pooling (Sum pooling), and normalization (Normalization); the bilinear pooling block based on multimodal matrix decomposition simultaneously processes the auditory features of the audio signal and the visual features of the video signal, the auditory features of the audio signal and the visual features of the audio signal respectively pass through the fully connected layer (FC) and perform a dot product operation, and the result after the dot product operation is successively subjected to regularization (Dropout), feature pooling (Sum pooling), and normalization operations to obtain linear fusion features;
[0098] The workflow of the auditory feature-guided visual feature branch includes: simultaneously inputting the visual features of the video signal and the auditory features of the audio signal in step 1.1 into two bilinear pooling blocks based on multimodal matrix decomposition, respectively passing the auditory features of the audio signal and the visual features of the audio signal through a fully connected layer (FC) and performing a dot product operation, and performing regularization (Dropout), feature pooling (Sum pooling), and normalization on the results after the dot product operation, and outputting two linear fusion features; the process can be described as follows:
[0099]
[0100] in represents the linear fusion feature of the output, D(·) represents regularization (Dropout), SP1(·) represents feature pooling, F1 and F2 represent fully connected layers;
[0101] Linear fusion features Input to ReLU activation function and fully connected layer (FC), and then input to spatial attention and channel attention respectively. The output of channel attention is combined with linear fusion feature. Perform element-wise multiplication, perform matrix multiplication on the result of element-wise multiplication and the result of spatial attention, and output the result The process is described as:
[0102]
[0103] Among them, W t S, W t C are the weights of spatial attention and channel attention respectively;
[0104] The visual feature calibration branch is mainly composed of channel attention and spatial attention. The workflow of the visual feature calibration branch includes: inputting the visual features of the video signal in step 1.1 into the spatial attention to obtain features The features Input into channel attention to obtain features The result of channel attention and spatial attention results Perform matrix multiplication and generate visual calibration features through ReLU activation function and fully connected layer (FC); the process is described as:
[0105]
[0106] in, is the visual calibration feature, Sigmoid(·) represents the activation function;
[0107] The function of guiding visual features by adaptively adjusting auditory features through visual calibration features is controlled by setting the hyperparameter β. The weight of , will be controlled by the hyper parameter β to add the visual calibration features and the auditory features to guide the visual features. The process is described as:
[0108]
[0109] in, is the visual calibration feature, β is a hyperparameter, v t Represents visual features guided by auditory features and calibrated with visual features;
[0110] The visual feature-guided auditory feature attention is composed of a visual feature-guided auditory feature branch and an auditory feature calibration branch; it includes a bilinear pooling block based on multimodal matrix decomposition, channel attention, spatial attention, and temporal attention. Other parts also include a ReLU activation function, a fully connected layer (FC), a global pooling operation (Global Pooling), and an addition operation;
[0111] The workflow of the visual feature-guided auditory feature branch includes: inputting the visual features of the video signal in step 1.1 into the spatial attention, and then inputting the spatial attention result into the channel attention. The global spatial information and channel information of the visual features are compressed to improve the feature representation. This process is expressed as:
[0112]
[0113] Among them, W1 and W2 represent the spatial attention and channel attention weights respectively, and F sq Indicates information compression operation;
[0114] Auditory features are input into temporal attention to extract and mine auditory feature representations. The process is described as follows:
[0115]
[0116] Q = a t W Q ,K=a t W K ,V=a t W V ,
[0117] Among them, W Q , W K , W V represents the weight matrix, Q, K, V represent the query, key, and value generated according to the input transformation;
[0118] In order to achieve the guidance of auditory features by visual features, the present invention uses the output results of the spatial attention and channel attention input of the visual features of the video signal in step 1.1 to perform a global pooling operation and output the result after global pooling. The result after using global pooling As a query vector, we extract the auditory features related to the video, and use the visual features to guide the auditory features. The process is described as follows:
[0119]
[0120] in, is the auditory feature guided by the visual feature, represents the weight matrix, q, k, v represent the query, key and value generated according to the input transformation;
[0121] Auditory features guided by visual features The final auditory features guided by visual features are generated through the ReLU activation function and the fully connected layer (FC);
[0122] The auditory feature calibration branch: the auditory features of the audio signal in step 1.1 and the visual features guided by the auditory features and calibrated by the visual features are input into the bilinear pooling block based on multimodal matrix decomposition. The auditory features of the audio signal and the visual features of the video signal are respectively passed through the fully connected layer (FC) and dot product operation. The results after the dot product operation are successively regularized (Dropout), feature pooling (Sum pooling), and normalization (Normalization) to output the auditory calibration features.
[0123] Adaptively adjust visual features to guide auditory features through auditory calibration features and set hyperparameters Controlling auditory features guided by visual features The weights will be passed through the hyperparameters The auditory features guided by the visual features are added to the visual calibration features. The process is described as follows:
[0124]
[0125] in, is a hyperparameter, a t Represents auditory features guided by visual features and calibrated with auditory features.
[0126] Step 1.3: The visual features obtained in step 1.2, which have been guided by auditory features and calibrated by visual features, and the auditory features obtained in step 1.2, which have been guided by visual features and calibrated by auditory features, are respectively input into two bidirectional long short-term memory (Bi-LSTM) neural networks for spatiotemporal feature extraction and modeling, thereby obtaining visual features and auditory features with spatiotemporal characteristics.
[0127] Step 1.4: Fuse the spatiotemporal visual and auditory features obtained in step 1.3, and input the fused features into the classification and recognition module to identify and locate the event categories in the target video; thus, a model for audiovisual event recognition and location based on bidirectional collaborative attention guidance is obtained. Specifically:
[0128] Step 1.41: The visual features and auditory features with spatiotemporal characteristics output by the bidirectional long short-term memory (Bi-LSTM) neural network in step 1.3 are added together to obtain fused features.
[0129] Step 1.42: Input the fused features obtained in step 1.41 into the classification and recognition module. The classification and recognition module consists of two fully connected layers. It performs audio-visual event recognition based on the fused features to obtain the identification and location information of the target event.
[0130] Step 2: Design a loss function, and continuously train and optimize the audiovisual event recognition and localization model based on bidirectional collaborative attention guidance constructed in step 1 through the loss function. When the loss function is minimized, the optimal audiovisual event recognition and localization model based on bidirectional collaborative attention guidance is obtained.
[0131] The loss function is calculated as follows: according to different audiovisual event localization tasks, the loss function is divided into a loss function under full supervision and a loss function under weak supervision;
[0132] In the fully supervised setting, the loss function is composed of the event-related score s r, the known segment-level event labels Y in the dataset tc , background label y t ={y t ||y t ∈{0,1},t=1,…,T} and the segment-level event labels O predicted by the model in the fully supervised setting tc Calculation; the calculation process is:
[0133] s r =δ(AV·F1),
[0134]
[0135] Among them, F1 represents the fully connected layer, AV represents the visual features and auditory features with spatiotemporal characteristics output by the bidirectional long short-term memory (Bi-LSTM) neural network, and the fusion features are obtained by feature addition. δ represents the activation function, and L BCE (s r ,y t ) represents the event-related score s r and background label y t ={y t ||y t Cross entropy loss ∈{0,1},t=1,...,T};
[0136] The loss function in the weakly supervised setting is the segment-level event labels predicted by the model in the weakly supervised setting. and the known video-level labels Y in the dataset c Calculation, the calculation process is:
[0137]
[0138] Where s represents the softmax function and λ represents the hyperparameter.
[0139] In step 3, the target video is fed into the optimal audiovisual event recognition and localization model based on bidirectional collaborative attention guidance obtained in step 2 to obtain the optimal target event recognition accuracy and location information. This recognition accuracy and location information are used to evaluate the effectiveness of the audiovisual event recognition and localization model based on bidirectional collaborative attention guidance.
[0140] The effects of the present invention are further described below in conjunction with simulation experiments.
[0141] 1. Simulation Experiment Conditions
[0142] The hardware platform of the simulation experiment of the present invention is: the processor is Intel i7-8750H, and the GPU is NUVIDA GeForce RTX 4060.
[0143] The software platform of the simulation experiment platform of the present invention is: Ubuntu20.04 operating system and PyCharm 2022, PyTorch 1.12.0, and CUDA11.2.
[0144] 2. Simulation steps
[0145] This simulation step uses the public dataset The Audio-Visual Event dataset to evaluate the method of the present invention. The data in this dataset is input into this embodiment. The VGG-19 network and the CNN network extract visual features and auditory features respectively, and input them into the bidirectional collaborative guidance attention module. The visual features and auditory features after bidirectional collaborative guidance are input into the bidirectional long short-term memory (Bi-LSTM) neural network module, and finally input into the classification and recognition module. The model is continuously trained and optimized through the loss function to achieve final accurate recognition and positioning.
[0146] 3. Simulation content and results analysis
[0147] The simulation experiments in this paper use the publicly available Audio-Visual Event dataset, which contains 4,143 video clips in 28 categories, covering a variety of audiovisual activities, including human speech, car driving, airplane roars, and animal sounds. Each video clip lasts 10 seconds and is annotated at the second level. For weakly supervised tasks, 178 unannotated noise samples are also introduced. The event classification of each video clip is predicted, and the overall recognition and localization accuracy of the two AVE tasks is used as the performance evaluation metric. The results are shown in Table 1.
[0148] The simulation effect of the present invention is further described below in conjunction with Table 1.
[0149] Table 1
[0150]
[0151] Table 1 compares the recognition and localization accuracy of the proposed method with existing methods, using the publicly available AVE dataset. As shown in Table 1, the proposed method improves recognition and localization accuracy by 2.9% and 1.6% under full and weak supervision, respectively. This improvement also demonstrates the effectiveness and technical advantages of the proposed method compared to the CMAN and AVSDN methods.
[0152] The present invention also provides an audiovisual event recognition and positioning system based on bidirectional collaborative attention guidance, comprising:
[0153] A model building module, used to implement the construction of an audiovisual event recognition and positioning model based on bidirectional collaborative attention guidance in step 1;
[0154] A model training module is used to implement the loss function designed in step 2, and continuously train and optimize the audiovisual event recognition and localization model based on bidirectional collaborative attention guidance constructed in step 1 through the loss function. When the loss function is minimized, the optimal audiovisual event recognition and localization model based on bidirectional collaborative attention guidance is obtained;
[0155] The recognition and positioning result output module is used to implement step 3 by inputting the target video into the optimal audio-visual event recognition and positioning model based on bidirectional collaborative attention guidance obtained in step 2, so as to obtain the optimal target event recognition accuracy and positioning information of the target event.
[0156] The present invention also provides an audiovisual event recognition and positioning device based on bidirectional collaborative attention guidance, comprising:
[0157] Memory: a computer-readable device storing a computer program for the above-mentioned method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance;
[0158] Processor: used to implement the method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance when executing the computer program.
[0159] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance.
Claims
1. A method for audiovisual event recognition and location based on bidirectional collaborative attention guidance, characterized in that: The following steps are involved: Step 1: Build an audiovisual event recognition and localization model based on bidirectional collaborative attention guidance; Step 2: Design a loss function, and continuously train and optimize the audiovisual event recognition and localization model based on bidirectional collaborative attention guidance constructed in step 1 through the loss function. When the loss function is minimized, the optimal audiovisual event recognition and localization model based on bidirectional collaborative attention guidance is obtained. Step 3: Input the target video into the optimal audiovisual event recognition and positioning model based on bidirectional collaborative attention guidance obtained in step 2 to obtain the optimal target event recognition accuracy and positioning information of the target event.
2. The method for audiovisual event recognition and positioning based on bidirectional collaborative attention guidance according to claim 1, characterized in that: The implementation method of step 1 includes: Step 1.1: Obtain the target video, preprocess the target video to obtain the video signal and audio signal, and use the convolutional neural network to extract the visual features of the video signal and the auditory features of the audio signal; Step 1.2: Construct a bidirectional collaborative guided attention module, input the visual features of the video signal and the auditory features of the audio signal in step 1.1 into the bidirectional collaborative guided attention module, and obtain visual features guided by auditory features and calibrated with visual features, and auditory features guided by visual features and calibrated with auditory features; Step 1.3: The visual features obtained in step 1.2, which have been guided by auditory features and calibrated by visual features, and the auditory features obtained in step 1.2, which have been guided by visual features and calibrated by auditory features, are respectively input into two bidirectional long short-term memory (Bi-LSTM) neural networks for spatiotemporal feature extraction and modeling, thereby obtaining visual features and auditory features with spatiotemporal characteristics. Step 1.4: Fuse the spatiotemporal visual and auditory features obtained in step 1.3, and input the fused features into the classification and recognition module to identify and locate the event categories in the target video; thus, obtaining an audiovisual event recognition and location model based on bidirectional collaborative attention guidance.
3. The method for audiovisual event recognition and positioning based on bidirectional collaborative attention guidance according to claim 2, characterized in that: The implementation method of step 1.1 includes: Step 1.11: Get the target video The target video is decomposed into T continuous segments of 1 second in length. Each segment consists of a synchronized video signal segment and an audio signal segment. t V , S t A Represented as video signal segments and audio signal segments respectively; Step 1.12: Sample each video signal segment at a fixed number of frames per second to obtain a fixed number of images for each video signal segment. Use a convolutional neural network (VGG-19) to extract feature maps of the fixed number of images for each video signal segment. Globally merge the feature maps of the fixed number of images per second to obtain a global feature map, which is the segment-level feature of each video signal segment. Step 1.13: Convert each audio signal segment into a log-mel spectrogram and use a convolutional neural network (CNN) to extract segment-level features for each audio signal segment. Step 1.14: Arrange the segment-level features of each video signal segment extracted in step 1.12 and the segment-level features of each audio signal segment extracted in step 1.13 in temporal order to form the visual features of the video signal and the auditory features of the audio signal.
4. The method for audiovisual event recognition and positioning based on bidirectional collaborative attention guidance according to claim 2, characterized in that: The bidirectional collaborative attention guidance module described in step 1.2 includes auditory features guiding the attention of visual features and visual features guiding the attention of auditory features; The auditory feature-guided visual feature attention is composed of an auditory feature-guided visual feature branch and a visual feature calibration branch; including a bilinear pooling block based on multimodal matrix decomposition, channel attention, spatial attention, and other parts including ReLU activation function, fully connected layer (FC), matrix multiplication, element-wise multiplication and addition operations; The bilinear pooling block based on multimodal matrix decomposition includes a dot product, a fully connected layer (FC), regularization (Dropout), feature pooling (Sum pooling), and normalization (Normalization); the bilinear pooling block based on multimodal matrix decomposition simultaneously processes the auditory features of the audio signal and the visual features of the video signal, the auditory features of the audio signal and the visual features of the audio signal respectively pass through the fully connected layer (FC) and perform a dot product operation, and the result after the dot product operation is successively subjected to regularization (Dropout), feature pooling (Sum pooling), and normalization operations to obtain linear fusion features; The workflow of the auditory feature-guided visual feature branch includes: simultaneously inputting the visual features of the video signal and the auditory features of the audio signal in step 1.1 into two bilinear pooling blocks based on multimodal matrix decomposition, respectively passing the auditory features of the audio signal and the visual features of the audio signal through a fully connected layer (FC) and performing a dot product operation, and performing regularization (Dropout), feature pooling (Sum pooling), and normalization on the results after the dot product operation, and outputting two linear fusion features; the process can be described as follows: in, represents the output linear fusion feature, k represents the pooling weight, D(·) represents regularization (Dropout), SP1(·) represents feature pooling, F1 and F2 represent fully connected layers; Linear fusion features Input to ReLU activation function and fully connected layer (FC), and then input to spatial attention and channel attention respectively. The output of channel attention is combined with linear fusion feature. Perform element-wise multiplication, perform matrix multiplication on the result of element-wise multiplication and the result of spatial attention, and output the result The process is described as: Among them, W t S , W t C are the weights of spatial attention and channel attention respectively; The visual feature calibration branch is composed of channel attention and spatial attention. The workflow of the visual feature calibration branch includes: inputting the visual features of the video signal in step 1.1 into the spatial attention to obtain features The features Input into channel attention to obtain features The result of channel attention and spatial attention results Perform matrix multiplication and generate visual calibration features through ReLU activation function and fully connected layer (FC); the process is described as: in, is the visual calibration feature, Sigmoid(·) represents the activation function; The function of guiding visual features by adaptively adjusting auditory features through visual calibration features; controlling visual calibration features by setting hyperparameter β The weight of , will be controlled by the hyper parameter β to add the visual calibration features and the auditory features to guide the visual features. The process is described as: Among them, β is a hyperparameter, v t Represents visual features guided by auditory features and calibrated with visual features; The visual feature-guided auditory feature attention is composed of a visual feature-guided auditory feature branch and an auditory feature calibration branch; it includes a bilinear pooling block based on multimodal matrix decomposition, channel attention, spatial attention, and temporal attention. Other parts also include a ReLU activation function, a fully connected layer (FC), a global pooling operation (Global Pooling), and an addition operation; The workflow of the visual feature-guided auditory feature branch includes: the visual features of the video signal in step 1.1 are input into the spatial attention and channel attention, and the process is expressed as: Among them, W1 and W2 represent the spatial attention and channel attention weights respectively, and F sq Indicates information compression operation; The auditory features of the audio signal in step 1.1 are input into the temporal attention to extract and mine the auditory feature representation. The process is expressed as: Q=a t W Q ,K=a t W K ,V=a t W V , Among them, W Q , W K , W V represents the weight matrix, Q, K, V represent the query, key, and value generated according to the input transformation; Input the visual features of the video signal in step 1.1 into the output results of spatial attention and channel attention, perform a global pooling operation, and output the result after global pooling The result after using global pooling As a query vector, we extract the auditory features related to the video, and use the visual features to guide the auditory features. The process is described as follows: in, is the auditory feature guided by the visual feature, W v Q , W a K , represents the weight matrix, q, k, v represent the query, key and value generated according to the input transformation; Auditory features guided by visual features The final auditory features guided by visual features are generated through the ReLU activation function and the fully connected layer (FC); The auditory feature calibration branch: the auditory features of the audio signal in step 1.1 and the visual features guided by the auditory features and calibrated by the visual features are input into the bilinear pooling block based on multimodal matrix decomposition. The auditory features of the audio signal and the visual features of the video signal are respectively passed through the fully connected layer (FC) and dot product operation. The results after the dot product operation are successively regularized (Dropout), feature pooling (Sum pooling), and normalization (Normalization) to output the auditory calibration features. Adaptively adjust visual features to guide auditory features through auditory calibration features and set hyperparameters Controlling auditory features guided by visual features The weights will be passed through the hyperparameters The auditory features guided by the visual features are added to the auditory calibration features. The process is described as follows: in, is a hyperparameter, a t Represents auditory features guided by visual features and calibrated with auditory features.
5. The method for audiovisual event recognition and positioning based on bidirectional collaborative attention guidance according to claim 2, characterized in that: The specific method of step 1.4 includes: Step 1.41: Obtain fusion features by adding the visual features and auditory features with spatiotemporal characteristics output by the bidirectional long short-term memory (Bi-LSTM) neural network in step 1.3; Step 1.42: Input the fused features obtained in step 1.41 into the classification and recognition module. The classification and recognition module consists of two fully connected layers. It performs audio-visual event recognition based on the fused features to obtain the identification and location information of the target event.
6. The method for audiovisual event recognition and location based on bidirectional collaborative attention guidance according to claim 1, characterized in that: In step 2, the loss function is calculated as follows: according to different audiovisual event localization tasks, the loss function is divided into a loss function under full supervision and a loss function under weak supervision; In the fully supervised setting, the loss function is composed of the event-related score s r , the known segment-level event labels Y in the dataset tc , background label y t ={y t ||y t ∈{0,1},t=1,…,T} and the segment-level event labels O predicted by the model in the fully supervised setting tc Calculation; the calculation process is: s r =δ(AV·F1), Among them, F1 represents the fully connected layer, AV represents the visual features and auditory features with spatiotemporal characteristics output by the bidirectional long short-term memory (Bi-LSTM) neural network, and the fusion features are obtained by feature addition, δ represents the activation function, and L BCE (s r ,y t ) represents the event-related score s r and background label y t ={y t ||y t Cross entropy loss ∈{0,1},t=1,…,T}; The loss function in the weakly supervised setting is the segment-level event labels predicted by the model in the weakly supervised setting. and the known video-level labels Y in the dataset c Calculation, the calculation process is: Among them, s represents the softmax function and λ represents the hyperparameter.
7. An audiovisual event recognition and positioning system based on bidirectional collaborative attention guidance based on the method according to any one of claims 1 to 6, characterized in that: include: A model building module for building an audiovisual event recognition and localization model based on bidirectional collaborative attention guidance; The model training module is used to implement the designed loss function, and continuously train and optimize the audiovisual event recognition and localization model based on bidirectional collaborative attention guidance through the loss function. When the loss function is minimized, the optimal audiovisual event recognition and localization model based on bidirectional collaborative attention guidance is obtained; The recognition and positioning result output module is used to input the target video into the optimal audio-visual event recognition and positioning model based on bidirectional collaborative attention guidance to obtain the optimal target event recognition accuracy and target event positioning information.
8. An audiovisual event recognition and positioning device based on bidirectional collaborative attention guidance, characterized in that: include: Memory: a computer-readable device storing a computer program for the method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance according to any one of claims 1 to 6; Processor: used to implement the method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance as described in any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, can implement the method for identifying and locating audiovisual events based on bidirectional collaborative attention guidance as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Audio-visual event detection method and device, storage medium and electronic equipment
CN117037046A
Audio-visual event identification method based on cross-modal relation perception fusion
CN118427394A