Audio-visual event positioning method and computer equipment

Through the audio-visual event positioning model of the progressive fusion strategy, the problem of low fusion efficiency of audio-visual modal features is solved, the accuracy of audio-visual event positioning is improved, and the complementarity between audio-visual modals is enhanced.

CN120032297APending Publication Date: 2025-05-23HENAN UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510119534.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the existing audio-visual event positioning methods, the audio-visual modal feature fusion efficiency is low, the positioning accuracy is not high, and the fine-grained audio-visual cross-modal interaction information is ignored, resulting in poor audio-visual event positioning performance.

Method used

The audio-visual event positioning model adopts a progressive fusion strategy, including a single-modal feature extraction module, a multi-modal collaborative state space module, a feature fusion module and a multi-modal enhancement state space module, learn the shared global context information and modal-specific features between audio-visual modes through collaborative state space modeling, and gradually enhance the accuracy and consistency of feature representation.

Benefits of technology

It significantly improves the overall performance of audio-visual event positioning tasks, enhances the complementarity between audio-visual modes, and improves the accuracy and robustness of audio-visual event positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032297A_ABST
    Figure CN120032297A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of audio-visual event positioning, and particularly relates to an audio-visual event positioning method and computer equipment. Inputting visual and audio data of a video into the trained audio-visual event positioning model to obtain an audio-visual event positioning result; wherein the audio-visual event positioning model comprises a single-mode feature extraction module, a multi-mode collaborative state space module, a feature fusion module, a multi-mode enhanced state space module and an event prediction module. Wherein the multi-mode collaborative state space module can learn global context information shared between audio and visual modes and specific feature information of each mode, and the multi-mode enhanced state space module can learn global context information of a feature fusion result. According to the method, efficient fusion of visual and audio modes can be realized, mining of fine-grained information is optimized, and the overall performance of an audio-visual event positioning task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of audio-visual event positioning, and in particular relates to an audio-visual event positioning method and computer equipment. Background Art

[0002] With the rapid development of information technology and communication technology, video has been deeply integrated into all aspects of human production and life. For example, surveillance cameras record a large amount of video data every day, and network platforms receive a large amount of uploaded videos every day. Against the background of the rapid growth of video scale and the increasing diversity of forms, how to use computers to realize intelligent automatic video analysis and processing has become a problem that needs to be solved urgently in academia and industry. This demand has promoted the research of many video understanding tasks, among which video event detection is a typical and important direction. By accurately detecting key events in videos, not only can the efficiency of emergency response be improved, but also the workload of manual review can be effectively reduced, thereby providing more reliable and efficient technical support for related fields.

[0003] However, most current methods only use single-modal information in video samples for event detection. When the video is blocked or dimly lit, it is impossible to accurately determine whether there is a target event in the video. To solve this problem, the Audio-Visual Event Localization (AVEL) task was born. By fusing the visual information and audio information in the video, complementary features are extracted from cross-modal data to more accurately determine whether there is a target event in the video.

[0004] As a hot research topic in the field of computer intelligent perception, the audiovisual event localization task has potential application value in many fields. For example, real-time detection of abnormal events in surveillance videos and timely warning, classification and rapid retrieval of multimedia data, and intelligent applications in medical and transportation scenarios. To accurately locate and predict events in video sequences, it is necessary to rely on a variety of artificial intelligence capabilities, such as semantic understanding, time series modeling, target detection, multimodal fusion, etc. Effective modeling of video sequences and fusion of different information in video samples can enhance the expressive ability of features and improve the performance of audiovisual event localization. The audiovisual event localization (AVEL) task aims to identify audiovisual events of interest in unconstrained long video sequences by integrating information from visual and audio modalities and predict their categories. The core goal of this task is to find semantically matching segments between audiovisual modalities from video sequences containing sound and visual events, and classify and predict the categories of audiovisual events based on this. Therefore, the effective fusion of visual and audio modalities has become the core link in solving the AVEL task.

[0005] Figure 1The typical architecture of most deep learning-based audiovisual event localization models is shown, which mainly includes the following three modules: (1) Visual and audio feature extraction module: In terms of visual feature extraction, deep convolutional neural networks (CNN) are usually used to process visual data; in terms of audio feature extraction, the audio input is first converted into Mel-spectrogram features, and then high-level semantic features are extracted through deep convolutional networks. (2) Multimodal information fusion module: Commonly used methods are divided into feature-level fusion and attention fusion. Feature-level fusion integrates features of different modalities through mathematical calculation methods. Commonly used methods include concatenation and weighted summation; attention fusion assigns weights according to the importance of the input data, captures the strongly correlated areas between modalities, and uses these significant features for task reasoning. (3) Event prediction module: Multi-layer perceptron (MLP) is usually used.

[0006] The Chinese invention patent application with application publication number CN115393968A and application publication date 2022.11.25 uses an architecture similar to the above architecture. It discloses an audiovisual event localization method that integrates self-supervised multimodal features. It uses an audiovisual event localization model to process the image and sound signals in the target video to obtain the event type at each moment in the target video. Among them, the audiovisual event localization model includes a visual-auditory feature extraction module, an audiovisual fusion module and a classification module. The visual-auditory feature extraction module extracts visual and auditory features from the image and sound signals, and then the audiovisual fusion module calculates the similarity between asynchronous visual features and auditory features based on the cosine distance, and corrects the similarity of the features according to the law of temporal correlation attenuation and then fuses the features to reduce the reliance on additional manual supervision of the background, and finally uses the classification model to classify the fused features. This scheme has some limitations, including that the method tends to update the features of the visual and visual modalities only by considering the situation where the correlation between visual and auditory signals is weaker when the time interval is longer, ignoring the fine-grained audiovisual cross-modal interaction information, and insufficient mining of details, resulting in poor performance of the audiovisual event localization model and affecting the accuracy of audiovisual event localization. Summary of the invention

[0007] The purpose of the present invention is to provide an audiovisual event localization method and a computer device to solve the problems of low efficiency of audiovisual modality feature fusion and low positioning accuracy in audiovisual event localization tasks in the prior art.

[0008] In order to solve the above technical problems, the present invention provides an audio-visual event localization method, which comprises: inputting visual and audio data of a video into a trained audio-visual event localization model to obtain an audio-visual event localization result; wherein the audio-visual event localization model comprises a single-modal feature extraction module, a multi-modal collaborative state space module, a feature fusion module, a multi-modal enhanced state space module and an event prediction module;

[0009] The unimodal feature extraction module is used to extract visual and audio global features from the visual and audio data respectively;

[0010] The multimodal collaborative state space module is used to perform collaborative state space modeling based on the extracted visual and audio global features and input the modeled visual and audio features into the feature fusion module for feature fusion; the collaborative state space modeling is used to learn the global context information shared between the audiovisual modalities and the feature information specific to each modality;

[0011] The multimodal enhanced state space module is used to perform enhanced state space modeling based on the feature fusion results and input the modeled features into the event prediction module for audiovisual event positioning; the enhanced state space modeling is used to learn the global context information of the feature fusion results.

[0012] Furthermore, the multimodal collaborative state space module includes a visual local feature extraction unit, an audio local feature extraction unit and a collaborative Bi-Mamba, and the process of collaborative state space modeling includes: 1) according to the visual and audio global features, using the visual local feature extraction unit and the audio local feature extraction unit to extract visual and audio local features respectively; 2) inputting the extracted visual and audio local features into the collaborative Bi-Mamba, and using the output of the collaborative Bi-Mamba as the input of the visual local feature extraction unit and the audio local feature extraction unit to extract visual and audio local features again, and repeating step 2) until the number of repetitions reaches the required number of collaborative state space modeling; the collaborative Bi-Mamba includes two Bi-Mambas, one Bi-Mamba is used to process visual local features, and the other Bi-Mamba is used to process audio local features; the two collaborative Bi-Mambas share forward and backward state transfer matrices.

[0013] Furthermore, both the visual local feature extraction unit and the audio local feature extraction unit are residual blocks.

[0014] Furthermore, the multimodal enhanced state space module includes a fused local feature extraction unit and Bi-Mamba, and the process of performing enhanced state space modeling includes: 1) using the fused local feature extraction unit to extract multimodal local temporal features based on the fused features input by the feature fusion module; 2) inputting the extracted multimodal local temporal features into Bi-Mamba, using the output of Bi-Mamba as the input of the fused local feature extraction unit to extract multimodal local temporal features again, and repeating step 2) until the number of repetitions reaches the required number of times for enhanced state space modeling.

[0015] Furthermore, the local feature extraction units are fused into a residual block.

[0016] Furthermore, the unimodal feature extraction module includes a pre-trained convolutional neural network and 1D convolution, and its process of extracting global visual features includes: first inputting the visual data into the pre-trained convolutional neural network for feature extraction, and then subjecting the output of the convolutional neural network to dimensional space conversion through 1D convolution to obtain global visual features.

[0017] Furthermore, the process of extracting global audio features by the unimodal feature extraction module includes: processing the audio data into a multi-frame audio sequence, performing time-frequency conversion on each frame of the audio sequence, encoding the frequency domain signal using a Mel filter to obtain a Mel spectrum feature, and then performing a logarithmic transformation to obtain a time series audio feature; inputting the time series audio feature into a pre-trained VGGish network for feature extraction, and then performing a 1D convolution on the output of the VGGish network for dimensional space conversion to obtain the global audio feature.

[0018] Furthermore, the event prediction module includes two fully connected neural networks and a softmax function connected in sequence.

[0019] Furthermore, the loss function used to train the audio-visual event localization model is the cross entropy loss function.

[0020] In order to solve the above technical problem, the present invention further provides a computer device, including a processor, wherein the processor is used to execute a computer program to implement the steps of the audiovisual event locating method introduced above.

[0021] The present invention is an improved invention creation, and its beneficial effects are as follows: the present invention improves the audio-visual event localization model used for audio-visual event localization. After using a single-modal feature extraction module to extract visual and audio global features, the model adopts a progressive fusion strategy to fuse the extracted visual and audio global features. Specifically, firstly, a multimodal collaborative state space module is used to learn the global context information shared between audio-visual modalities and the feature information specific to each modality, so as to accurately capture the key fine-grained information in the target event and the unique features of each modality, and enhance the accuracy and reliability of feature representation. Then, the learning results are feature fused through a feature fusion module, and finally, the global context information of the feature fusion results is learned through a multimodal enhanced state space module to enhance multimodal consistency, and gradually enhance the complementarity between audio and video modalities to achieve efficient fusion of visual and audio modalities. The progressive fusion strategy is intended to optimize the mining of fine-grained information and improve the overall performance of the audio-visual event localization task. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a basic framework diagram of the audio-visual event localization model commonly used in the prior art;

[0023] Figure 2 is the state space model diagram;

[0024] Figure 3 is a processing flow chart of the audio-visual event localization model of the present invention;

[0025] Figure 4 : is a network structure diagram of the residual block used in the present invention;

[0026] Figure 5 This is a network architecture diagram of the collaborative bidirectional Bi-Mamba of the present invention;

[0027] Figure 6 This is a network architecture diagram of VGG-19 used in the present invention;

[0028] Figure 7 is a network architecture diagram of the audio-visual event localization model of the present invention;

[0029] Figure 8 It is an example diagram of audio-visual event positioning of the present invention. DETAILED DESCRIPTION

[0030] The main idea of ​​the present invention is to design an audio-visual event localization model based on a progressive fusion strategy, which specifically includes a unimodal feature extraction module, a multimodal collaborative state space module, a feature fusion module, a multimodal enhanced state space module and an event prediction module, wherein the multimodal collaborative state space module, the feature fusion module and the multimodal enhanced state space module are modules for implementing a progressive fusion strategy. The multimodal collaborative state space module is used for collaborative state space modeling, and collaborative state space modeling is used to learn the global context information shared between audio-visual modalities and the feature information specific to each modality. The feature fusion module is used to perform feature fusion on the output of the multimodal collaborative state space module, and the multimodal enhanced state space module is used to learn the global context information of the feature fusion result to enhance multimodal consistency. The present invention gradually enhances the complementarity between audio and video modalities, while retaining specific features within the modality, significantly improving the audio and video fusion effect. Based on this, an audio-visual event localization method and a computer device of the present invention can be realized. In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0031] An audio-visual event location method embodiment:

[0032] The method described in this embodiment requires the use of the Mamba model. The Mamba model is an emerging basic model that has shown important value in multiple research and application fields such as computer vision, natural language processing, and medical treatment because of its powerful modeling capabilities for complex long sequence data and nearly linear scalability. The core advantage of the Mamba model is that it introduces a simple but effective selection mechanism that can reparameterize the state space model (SSM) according to the input, so that the model can effectively filter out irrelevant information while retaining necessary and relevant data. In addition, the Mamba model also includes a hardware-aware algorithm that uses a scan operation instead of a convolution to loop the calculation model, thereby significantly improving the training speed and computational efficiency of the model.

[0033] Figure 2 The structure of SSM is shown. This model is derived from classical control theory, has linear scalability, and is suitable for long sequence modeling. SSM consists of a hidden state Input features Mapping to output Where N and L represent the number of hidden states and sequence length respectively. The process of the continuous SSM system is as follows: h′(t)=Ah(t)+Bx(t), y(t)=Ch(t), where is the state matrix, and are the input and output projection matrices respectively. The Mamba model further introduces a time scale parameter Δ and discretizes the continuous parameters A and B into The discrete state space equation is: After discretization, the state space equation can be transformed into the following recursive form: y(t)=Ch t The equation can also be equivalently expressed in convolution form: y = x*K, Where * represents the convolution operation, is the global convolution kernel.

[0034] The entire process of the audio-visual event positioning method is introduced in detail below.

[0035] Step 1: Get a video containing audio and visual modality data {S v ,S a}.

[0036] In this embodiment, the video length is 10s, and the input data is preprocessed, specifically, the entire video is divided into 10 non-overlapping audio-visual segments with equal time intervals, and the sampling interval is set to 1s, that is, and Represents the visual track and audio track in the video respectively.

[0037] Step 2: Input the visual and audio modality data obtained in step 1 into the trained audio-visual event localization model to obtain the audio-visual event localization result, which includes the position of the audio-visual event in the video and the type of the audio-visual event.

[0038] The specific architecture of the audio-visual event localization model is as follows: Figure 7 The overall processing flow is as shown in Figure 3 As shown, it specifically includes a single-modal feature extraction module and a multi-modal collaborative state space module (i.e. Figure 7 Multimodal collaborative state space model in ), feature fusion module ( Figure 7 Not marked in Figure 7 The link between the multimodal collaborative state space model and the multimodal enhanced state space model in Figure 7The multimodal enhanced state space model in the multimodal collaborative state space module) and the event prediction module. The unimodal feature extraction module is used to extract visual global features and audio global features from the visual data and audio data respectively. The multimodal collaborative state space module is used to perform collaborative state space modeling based on the extracted visual global features and audio global features and input the modeled visual features and audio features into the feature fusion module; the collaborative state space modeling is used to learn the global context information shared between the audio-visual modalities and the feature information specific to each modality. The feature fusion module is used to fuse the visual features and audio features output by the multimodal collaborative state space module and input the fusion results into the multimodal enhanced state space module. The modality enhanced state space module is used to perform enhanced state space modeling based on the feature fusion results and input the modeled features into the prediction module; the enhanced state space modeling is used to learn the global context information of the feature fusion results to enhance multimodal consistency. The event prediction module is used to locate audio-visual events based on the output of the modality enhanced state space module. The processing process of each module in the audio-visual event localization model is described in detail below.

[0039] 1. Single modal feature extraction module.

[0040] The unimodal feature extraction module involves visual global feature extraction and audio global feature extraction.

[0041] 1) Visual semantic feature extraction (i.e. visual global feature extraction).

[0042] First, the input visual data S v Use as Figure 6 The pre-trained convolutional neural network shown extracts the visual feature vector F v The convolutional neural network can be specifically selected from the VGG-19 network. For each 1-second visual segment, the sampling rate is 16 samples / s. The pool5 feature map of the VGG-19 network is extracted from the sampled 16 frames of RGB video images. Then, the global average pooling operation is applied to these 16 frames of images to generate a feature map with an image size of 7×7 and a channel number of 512. The visual feature dimension obtained is Then, the extracted visual feature vector is transformed by 1D convolution so that F v and F a (The next paragraph will introduce F a ) is transformed to the same dimensional space, and after 1D convolution transformation, the final visual feature representation is obtained as

[0043] 2) Audio semantic feature extraction (i.e. audio global feature extraction).

[0044] First, the audio clips are pre-emphasized, framed, and windowed, and then each frame is subjected to a Fast Fourier Transform (FFT) calculation to generate a frequency domain representation. Next, an 80-bin Mel filter bank is used to encode the frequency domain signal, and the Mel filter bank features are extracted from each short time frame, followed by a logarithmic transformation to form a temporal audio representation. Subsequently, an audio feature vector is extracted through a pre-trained VGGish network, whose dimension is Get F a ; Finally, the audio feature vector F a The final audio features are obtained after 1D convolution transformation

[0045] 2. Multimodal collaborative state space module.

[0046] The multimodal collaborative state space module includes a visual local feature extraction unit, an audio local feature extraction unit, and a collaborative bidirectional Bi-Mamba. The visual local feature extraction unit and the audio local feature extraction unit are both residual blocks. The network architecture of the residual block is as follows: Figure 4 As shown; the architecture of the collaborative bidirectional Bi-Mamba is as follows Figure 5 As shown, it includes two bidirectional Bi-Mambas, one of which is used to process local visual features, and the other is used to process local audio features, and the two coordinated bidirectional Bi-Mambas share the forward and backward state transfer matrices. The specific processing process is as follows:

[0047] 2.1) Use residual blocks with small convolution kernels to capture C v and C a The local temporal information of the residual operation is obtained by v and audio feature R a , the calculation process is as follows:

[0048] R v =H v (C v )+C v

[0049] R a =H a (C a )+C a

[0050] G v (C v )=DP(σ(BN(Conv 1 (C v ))))

[0051] H a (C a )=DP(σ(BN(Conv 2 (C a ))))

[0052] In the formula, Conv i (·) represents a convolutional neural network, DP(·) and BN(·) represent the Dropout operation and batch normalization operation respectively, and σ(·) represents the ReLU (Rectified Linear Units) activation function.

[0053] 2.2) R v and R a Sent to the collaborative bidirectional Bi-Mamba, modeled by bidirectional SSM R v and R a The global context information is used to obtain the visual features after selective attention. and audio features The forward formula of collaborative SSM is as follows:

[0054]

[0055] In the formula, Respectively represent the forward input, output and hidden features of the visual modality, Represent the forward input, output and hidden features of the audio modality, A f represents the forward state transfer matrix shared by audio and visual modalities, and Represent the forward input and output matrices of the visual modality, respectively. and Represent the forward input and output matrices of the audio modality respectively.

[0056] The backward formula of collaborative SSM is as follows:

[0057]

[0058] In the formula, Represent the backward input, output and hidden features of the visual modality respectively, Represent the backward input, output and hidden features of the audio modality, A b represents the backward state transfer matrix shared by audio and visual modalities, and Respectively represent the backward input and output matrices of the visual modality, and Represent the backward input and output matrices of the audio modality respectively.

[0059] 2.3) Order m=m+1. When m is less than 3 (3 represents the required number of collaborative state space modeling, which can be adjusted according to actual conditions), return to step 2.1) until m=3 is output, indicating that three collaborative state space modelings have been completed.

[0060] 3. Feature fusion module.

[0061] The feature fusion module is used to fuse the output of the multimodal collaborative state space module. and Get the fused feature vector F av , the fusion method is:

[0062]

[0063] In the formula, Concat[] represents the concatenation operation.

[0064] 4. Multimodal enhanced state space module.

[0065] The multimodal enhanced state space module includes a fused local feature extraction unit and a bidirectional Bi-Mamba, and the fused local feature extraction unit is a residual block. The specific processing process is as follows:

[0066] 4.1) Use residual blocks with small convolution kernels to capture F av The multi-modal local time series information is obtained by the residual operation of the fusion feature R av , the calculation process is as follows:

[0067] R av =H av (F av )+F av

[0068] H av =DP(σ(BN(Conv 3 F av ))))

[0069] 4.2) The extracted R av Sent to the bidirectional Bi-Mamba, modeled by SSM R av The global context information is obtained by selectively paying attention to the fused feature vector

[0070] 4.3) Order n=n+1, when n is less than 3 (3 represents the number of times the enhanced state space modeling is required, which can be adjusted according to the actual situation), return to step 4.1) until n=3 is output, indicating that the enhanced state space modeling has been completed three times.

[0071] 5. Event prediction module.

[0072] Using the last fused feature vector output by the loop Perform event category prediction and obtain event category prediction vector The event with the largest score in each time segment is selected as the final event of the time period. The prediction process is as follows:

[0073]

[0074] In the formula, FC i represents a fully connected neural network, and Softmax represents the softmax function.

[0075] 6. Calculate the losses.

[0076] During the training process, the prediction vector and label Y t,c The cross entropy loss function is used above, and the specific loss function is as follows:

[0077]

[0078] In the formula, Denotes the cross entropy loss. The batch stochastic gradient descent optimization method is used to reduce the loss and optimize the network parameters.

[0079] Figure 8 The attention heat map and prediction results generated by the model during event localization are shown. The model of the present invention can accurately locate the "Baby cry" event between the 4th and 9th frames of the video and successfully predict its category.

[0080] A computer device embodiment:

[0081] A computer device embodiment of the present invention includes a memory, a processor, an internal bus, a computer program stored in the memory, a communication interface, an input device and a display screen. The processor and the memory communicate and exchange data with each other through the internal bus.

[0082] Among them, the processor is used to provide computing and control capabilities, and executes the computer program to implement the steps of the method described in the embodiment of the audio-visual event positioning method of the present invention, and can be a processing device such as a microprocessor MCU, a programmable logic device FPGA, etc. The memory can be a non-volatile storage medium or an internal memory. The non-volatile storage medium stores an operating system and a computer program, and the internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless method can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0083] In summary, the present invention has the following characteristics: the present invention proposes an audio-visual event localization method based on a progressive fusion strategy for audio-visual event localization tasks, aiming to efficiently process visual and audio data in long video sequences and accurately capture the key information of target events. Specifically, first, the present invention realizes the modeling of long audio-visual sequences with linear time complexity through a state space model (SSM), realizes the dynamic time series analysis of the visual and audio modes of the input video, learns the time series changes of each modal feature and the key features related to the target event, accurately captures long-distance dependencies, and significantly reduces the calculation and storage costs; moreover, through the dynamic parameterization mechanism of the bidirectional Bi-Mamba model, the state space matrix shared across modalities is used to realize the efficient fusion of visual and audio modes, and effectively filter out redundant information unrelated to the event, and enhance the accuracy and reliability of feature representation. Secondly, the present invention combines convolutional neural networks (CNN, including VGG-19 and VGGish in this embodiment) and bidirectional Bi-Mamba models to realize hierarchical context modeling from local to global scales, which can effectively extract and integrate key content information in long audio and video sequences. Furthermore, the two-stage progressive fusion strategy can gradually enhance the complementarity between audio and video modalities, while retaining specific features within the modality, significantly improving the audio and video fusion effect. In addition, the present invention supports end-to-end training, adaptively learning audio-visual fusion features that are highly associated with target events, significantly improving the accuracy and robustness of event positioning, and showing high computational efficiency when dealing with large-scale video data.

[0084] Specific implementation methods are given above, but the present invention is not limited to the described implementation methods. The basic idea of ​​the present invention lies in the above basic scheme. For ordinary technicians in this field, it does not take creative work to design various deformed models, formulas, and parameters according to the teachings of the present invention. Changes, modifications, substitutions, and variations of the implementation methods without departing from the principles and spirit of the present invention still fall within the scope of protection of the present invention.

Claims

1. A method for locating an audiovisual event, characterized in that: The method comprises: inputting visual and audio data of a video into a trained audio-visual event localization model to obtain an audio-visual event localization result; wherein the audio-visual event localization model comprises a single-modal feature extraction module, a multi-modal collaborative state space module, a feature fusion module, a multi-modal enhanced state space module and an event prediction module; The unimodal feature extraction module is used to extract visual and audio global features from the visual and audio data respectively; The multimodal collaborative state space module is used to perform collaborative state space modeling based on the extracted visual and audio global features and input the modeled visual and audio features into the feature fusion module for feature fusion; the collaborative state space modeling is used to learn the global context information shared between the audiovisual modalities and the feature information specific to each modality; The multimodal enhanced state space module is used to perform enhanced state space modeling based on the feature fusion results and input the modeled features into the event prediction module for audiovisual event positioning; the enhanced state space modeling is used to learn the global context information of the feature fusion results.

2. The audiovisual event location method according to claim 1, characterized in that: The multimodal collaborative state space module includes a visual local feature extraction unit, an audio local feature extraction unit and a collaborative Bi-Mamba, and the process of collaborative state space modeling includes: 1) according to the visual and audio global features, the visual local feature extraction unit and the audio local feature extraction unit are used to extract visual and audio local features respectively; 2) the extracted visual and audio local features are input to the collaborative Bi-Mamba, and the output of the collaborative Bi-Mamba is used as the input of the visual local feature extraction unit and the audio local feature extraction unit to extract visual and audio local features again, and step 2) is repeated until the number of repetitions reaches the number required for collaborative state space modeling; The collaborative Bi-Mamba includes two Bi-Mambas, one Bi-Mamba is used to process the local visual features, and the other Bi-Mamba is used to process the local audio features; the two collaborative Bi-Mambas share the forward and backward state transfer matrices.

3. The audiovisual event location method according to claim 2, characterized in that: Both the visual local feature extraction unit and the audio local feature extraction unit are residual blocks.

4. The audiovisual event location method according to claim 1, characterized in that: The multimodal enhanced state space module includes a fused local feature extraction unit and Bi-Mamba, and the process of performing enhanced state space modeling includes: 1) using the fused local feature extraction unit to extract multimodal local temporal features according to the fused features input by the feature fusion module; 2) inputting the extracted multimodal local temporal features into Bi-Mamba, using the output of Bi-Mamba as the input of the fused local feature extraction unit to extract multimodal local temporal features again, and repeating step 2) until the number of repetitions reaches the required number of times for enhanced state space modeling.

5. The audiovisual event location method according to claim 4, characterized in that: The local feature extraction units are fused into residual blocks.

6. The audiovisual event location method according to claim 1, characterized in that: The unimodal feature extraction module includes a pre-trained convolutional neural network and 1D convolution. The process of extracting global visual features includes: first inputting the visual data into the pre-trained convolutional neural network for feature extraction, and then subjecting the output of the convolutional neural network to dimensional space conversion through 1D convolution to obtain global visual features.

7. The audiovisual event location method according to claim 1, characterized in that: The process of extracting global audio features by the unimodal feature extraction module includes: processing audio data into multi-frame audio sequences, performing time-frequency conversion on each frame of audio sequence, encoding the frequency domain signal using a Mel filter to obtain Mel spectrum features, and then performing a logarithmic transformation to obtain time series audio features; inputting the time series audio features into a pre-trained VGGish network for feature extraction, and then performing a 1D convolution on the output of the VGGish network for dimensional space conversion to obtain global audio features.

8. The method for locating audiovisual events according to claim 1, characterized in that: The event prediction module consists of two fully connected neural networks and a softmax function connected in sequence.

9. The method for locating an audiovisual event according to any one of claims 1 to 8, characterized in that: The loss function used to train the audio-visual event localization model is the cross entropy loss function.

10. A computer device comprising a processor, characterized in that: The processor is used to execute a computer program to implement the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Audiovisual event positioning method fusing self-supervised multi-modal features

    CN115393968A