Dense audio-visual event positioning method and device, equipment and storage medium

By employing multimodal early fusion and multi-stage semantic guidance, the cross-modal semantic gap problem was solved, enabling precise localization of audiovisual events and accurate prediction of temporal boundaries, thus enhancing the model's reasoning ability in complex scenarios.

CN120953655APending Publication Date: 2025-11-14BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510932338.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies neglect cross-modal semantic bridging in intermediate layers, leading to a modal semantic gap problem. This makes it difficult for models to effectively distinguish between event-related content and unrelated background content, affecting the model's efficiency in focusing on task-related content.

Method used

We employ a multimodal early fusion and multi-stage semantic guidance approach. We use a single-modal attention mechanism, an audio-visual guidance mechanism, and a cross-modal pyramid mechanism to perform temporal modeling on the initial feature data. We use a multi-stage semantic guidance method to constrain the classification loss function of the multimodal fusion data. We further refine the data through multimodal temporal aggregation and multi-event dependency extraction methods. Finally, we use a decoder for prediction.

Benefits of technology

It significantly improves the model's reasoning ability in long videos with multiple concurrent events, reduces missed detections, improves time localization accuracy, and enhances the model's understanding of event-related content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953655A_ABST
    Figure CN120953655A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a dense audio-visual event positioning method, device and equipment and a computer storage medium. The method comprises the following steps: preprocessing video data to be positioned, and performing feature extraction on the preprocessed video data; time modeling is carried out based on a single-mode attention mechanism, an audio visual guidance mechanism and a cross-mode pyramid mechanism, classification loss function constraint is carried out by using a multi-stage semantic guidance method, refining processing is carried out based on a multi-mode time aggregation method and a multi-event dependence extraction method, and comprehensive feature data is obtained; and processing the comprehensive feature data based on a decoder to obtain a predicted event category and a time boundary. A middle layer cross-modal semantic gap is gradually bridged through multi-modal early fusion and multi-stage semantic guidance; a multi-event dependency relationship in a complex scene is adaptively captured by using a hybrid dependency expert module, and accurate event category prediction and time boundary positioning are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and computer storage medium for locating dense audiovisual events. Background Technology

[0002] With the rapid growth of multimedia data, audiovisual event localization plays a crucial role in video analytics. Joint audio and video analysis provides a powerful tool for various tasks, but effectively bridging different modalities to improve model performance remains a challenge. Early research typically extracted audio and visual features from sub-networks, which may lead to insufficient feature fusion and inadequate extraction of task-specific information. Subsequent research attempted to guide attention to task-related regions through audiovisual similarity or utilize contrastive learning for self-supervised synchronous tasks. However, most of these methods only impose semantic constraints on the final output, neglecting cross-modal semantic bridging in intermediate layers. This results in a modal semantic gap, making it difficult for the model to effectively distinguish between event-related content and unrelated background content, thus affecting the model's efficiency in focusing on task-related content.

[0003] In summary, designing an efficient and accurate method for locating dense audiovisual events is a problem that urgently needs to be solved. Summary of the Invention

[0004] Therefore, the technical problem to be solved by the present invention is to overcome the problem of the modal semantic gap caused by the neglect of cross-modal semantic bridging of the intermediate layer in the prior art, which makes it difficult for the model to effectively distinguish between event-related content and unrelated background content, and affects the model's efficiency in focusing on task-related content.

[0005] To address the aforementioned technical problems, this invention provides a method for locating dense audiovisual events, comprising:

[0006] The video data to be located is preprocessed, and features are extracted from the preprocessed video data to obtain initial feature data.

[0007] The initial feature data is temporally modeled based on a single-modal attention mechanism, an audio-visual guidance mechanism, and a cross-modal pyramid mechanism to obtain multimodal fusion data;

[0008] A multi-stage semantic guidance method is used to constrain the classification loss function of the multimodal fusion data to obtain constrained data;

[0009] The constraint data is refined based on the multimodal time aggregation method and the multi-event dependency extraction method to obtain comprehensive feature data;

[0010] The decoder processes the comprehensive feature data to obtain the predicted event category and time boundary.

[0011] Preferably, the step of performing temporal modeling on the initial feature data based on a single-modal attention mechanism, an audio-visual guidance mechanism, and a cross-modal pyramid mechanism to obtain multimodal fusion data includes:

[0012] A unimodal attention mechanism is used to perform temporal modeling of audio and visual features, and audio and visual features are projected into a higher-dimensional space. Information is then aggregated sequentially through a multi-head self-attention network and a feedforward network to obtain preliminary aggregated data.

[0013] The initial aggregated data is supplemented using an audio-visual guidance mechanism. After the audio and visual data are mapped to the same space using an aligner, temporal downsampling processing is performed using a cross-modal pyramid mechanism to obtain multimodal fusion data.

[0014] Preferably, the formula for calculating the initial aggregated data obtained by sequentially aggregating information through multi-head self-attention and a feedforward network is as follows:

[0015]

[0016]

[0017] in, These are the outputs for audio-only and video-only features, respectively. These are learnable parameters.

[0018] Preferably, the formula for calculating the classification loss function constraint is:

[0019]

[0020] in, For multi-stage classification loss function, α i These are the weights of the loss, i∈1,2,3, with α1, α2, and α3 increasing sequentially.

[0021] Preferably, the refinement of the constraint data based on the multimodal time aggregation method and the multi-event dependency extraction method to obtain comprehensive feature data includes:

[0022] Audio and visual features are aggregated, and cross-modal features are obtained by concatenating the audio and visual features. The cross-modal features are then input into the projection layer to obtain preliminary fused features.

[0023] The dependencies of the preliminary fusion features are obtained based on the hybrid expert layer, and the output of the hybrid expert layer is added to the output of the audio features and visual features aggregation to obtain comprehensive feature data.

[0024] Preferably, the decoder includes a classification head and a regression head, each consisting of three layers of one-dimensional convolution and one layer of activation function. The classification head uses a Sigmoid activation function to output the event class probability, and the regression head uses a ReLU activation function to output the temporal boundary distance.

[0025] Preferably, it further includes: training in an end-to-end manner, using three loss functions to constrain event classification, temporal boundary estimation, and multi-stage semantic guidance respectively, wherein the total loss function is:

[0026]

[0027] in, and Used to constrain the final audiovisual event detection and regression, while Cross-modal semantic bridging for the middle layer.

[0028] The present invention also provides a dense audiovisual event localization device, comprising:

[0029] The feature extraction module preprocesses the video data to be located and extracts features from the preprocessed video data to obtain initial feature data.

[0030] The fusion module performs temporal modeling on the initial feature data based on a single-modal attention mechanism, an audio-visual guidance mechanism, and a cross-modal pyramid mechanism to obtain multimodal fused data.

[0031] The constraint module uses a multi-stage semantic guidance method to constrain the classification loss function of the multimodal fusion data to obtain constraint data;

[0032] The optimization processing module refines the constraint data based on the multimodal time aggregation method and the multi-event dependency extraction method to obtain comprehensive feature data;

[0033] The decoding module processes the comprehensive feature data based on the decoder to obtain the predicted event category and time boundary.

[0034] The present invention also provides a dense audiovisual event localization device, comprising:

[0035] Memory, used to store computer programs;

[0036] A processor is used to implement the steps of the above-described method for locating dense audiovisual events when executing the computer program.

[0037] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for locating dense audiovisual events.

[0038] The technical solution of the present invention has the following advantages compared with the prior art:

[0039] This invention discloses a dense audiovisual event localization method that gradually bridges the cross-modal semantic gap in the intermediate layer through early multimodal fusion and multi-stage semantic guidance. It utilizes a hybrid dependency expert module to adaptively capture multi-event dependencies in complex scenes, ultimately achieving accurate event category prediction and temporal boundary localization. By adaptively revealing multi-event dependencies through a hybrid expert layer, the model's reasoning ability in complex scenes is enhanced. This significantly improves the model's reasoning ability in long videos with concurrent multi-event scenarios, reducing missed detections and improving temporal localization accuracy. Attached Figure Description

[0040] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein:

[0041] Figure 1 This is a flowchart of a first specific embodiment of a dense audiovisual event localization method provided by the present invention;

[0042] Figure 2 This is an overall architecture diagram of a dense audiovisual event localization method provided by the present invention;

[0043] Figure 3 It is a semantic constraint distinction structure diagram;

[0044] Figure 4 This is a flowchart of the model training process;

[0045] Figure 5 This is a flowchart of the model inference process;

[0046] Figure 6 This is a structural block diagram of a dense audiovisual event positioning device provided in an embodiment of the present invention. Detailed Implementation

[0047] The core of this invention is to provide a method, apparatus, device, and computer storage medium for locating dense audiovisual events. By using multimodal early fusion and multi-stage semantic guidance, it significantly compensates for the cross-modal semantic gap in the intermediate layer, improving the model's ability to understand event-related content. By adaptively revealing multi-event dependencies through a hybrid expert layer, it enhances the model's reasoning ability in complex scenarios.

[0048] To enable those skilled in the art to better understand the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] Please refer to Figure 1 , Figure 1 The flowchart illustrates the implementation of a dense audiovisual event localization method provided by this invention; the specific operation steps are as follows:

[0050] Step S101: Preprocess the video data to be located, and extract features from the preprocessed video data to obtain initial feature data;

[0051] Step S102: Based on the single-modal attention mechanism, audio-visual guidance mechanism and cross-modal pyramid mechanism, perform temporal modeling on the initial feature data to obtain multimodal fusion data;

[0052] Specifically, a single-modal attention mechanism is used to perform temporal modeling of audio and visual features, and the audio and visual features are projected into a higher-dimensional space respectively. Information is then aggregated sequentially through a multi-head self-attention network and a feedforward network to obtain preliminary aggregated data.

[0053] The initial aggregated data is supplemented by an audio-visual guidance mechanism. After the audio and visual data are mapped to the same space using an aligner, the time downsampling process is performed using a cross-modal pyramid mechanism to obtain multimodal fusion data.

[0054] The calculation formula is as follows:

[0055]

[0056] in, These are the outputs for audio-only and video-only features, respectively. These are learnable parameters.

[0057] Step S103: Use a multi-stage semantic guidance method to constrain the classification loss function of the multimodal fusion data to obtain constrained data;

[0058] The formula for calculating the classification loss function constraint is as follows:

[0059]

[0060] in, For multi-stage classification loss function, α i These are the weights of the loss, i∈1,2,3, with α1, α2, and α3 increasing sequentially.

[0061] Step S104: Refine the constraint data based on the multimodal temporal aggregation method and the multi-event dependency extraction method to obtain comprehensive feature data;

[0062] Audio and visual features are aggregated, and cross-modal features are obtained by concatenating the audio and visual features. The cross-modal features are then input into the projection layer to obtain preliminary fused features.

[0063] The dependencies of the preliminary fusion features are obtained based on the hybrid expert layer, and the output of the hybrid expert layer is added to the output of the audio features and visual features aggregation to obtain comprehensive feature data.

[0064] Step S105: Process the comprehensive feature data based on the decoder to obtain the predicted event category and time boundary.

[0065] The classification head and regression head each consist of three layers of one-dimensional convolution and one layer of activation function. The classification head uses the Sigmoid activation function to output the event class probability, and the regression head uses the ReLU activation function to output the time boundary distance.

[0066] This method employs an end-to-end training approach, utilizing three loss functions to constrain event classification, temporal boundary estimation, and multi-stage semantic guidance, respectively. The total loss function is:

[0067]

[0068] in, and Used to constrain the final audiovisual event detection and regression, while Cross-modal semantic bridging for the middle layer.

[0069] This embodiment provides a method for intensive audiovisual event localization. Through early multimodal fusion and multi-stage semantic guidance, it gradually bridges the cross-modal semantic gap in the intermediate layer. It utilizes a hybrid dependency expert module to adaptively capture multi-event dependencies in complex scenes, ultimately achieving accurate event category prediction and temporal boundary localization. By adaptively revealing multi-event dependencies through a hybrid expert layer, the model's reasoning ability in complex scenes is enhanced. This significantly improves the model's reasoning ability in long videos with concurrent multi-event scenarios, reducing missed detections and improving temporal localization accuracy.

[0070] Based on the above embodiments, this embodiment describes a method for locating dense audiovisual events, as follows: Figure 2 As shown, the details are as follows:

[0071] Raw audio and video inputs are fed into a frozen encoder for feature extraction. Subsequently, ESI performs cross-modal early fusion and multi-stage semantic guidance, focusing on event-related content in intermediate layers. Next, a MoDE consisting of multiple expert hybrid layers is used to extract multi-event dependencies from the integrated features. Finally, the audiovisual features are decoded to obtain event categories and temporal boundaries.

[0072] Overall architecture:

[0073] ESG-Net consists of four key components: 1) a pre-trained audio encoder and visual encoder for extracting features from the raw video; 2) an early semantic interaction module for cross-modal semantic bridging; 3) a hybrid dependency expert module for capturing multi-event dependencies; and 4) a decoder for locating audiovisual events. These five key components are executed sequentially.

[0074] Audio encoder and visual encoder: We employ a frozen encoder to extract audio features F A ={f A1 ,f A2 ,…,f AT} and visual features F V ={f V1 ,f V2 ,…,f VT Here, the multimodal features F A , The features are extracted synchronously, where T represents the length of the video and D is the dimension of these features.

[0075] Early Semantic Interaction Module: To achieve cross-modal semantic bridging in the intermediate layer, ESI includes multimodal early fusion and multi-stage semantic guidance. Multimodal early fusion includes unimodal attention, audio (visual) guided mixing, and a cross-modal pyramid module. The outputs of these three parts are... Where T m It is the maximum length of all videos, T l It is the total length of the multi-level pyramid features. Then, they are used for multi-stage semantic guidance through several classification loss functions.

[0076] Specifically, the lack of cross-modal semantic bridging in the intermediate layer creates a modal semantic gap, hindering the understanding of audiovisual events in previous methods. To address this issue, we propose multimodal early fusion and multi-stage semantic guidance in ESI to achieve a hierarchical understanding of audiovisual events, such as... Figure 3 As shown, most previous methods only apply semantic constraints to the final output. We consider implementing event-related semantic bridging between different modalities in the intermediate layer through early cross-modal fusion and semantic guidance.

[0077] Early multimodal fusion:

[0078] Before performing cross-modal semantic bridging, we need to fuse multimodal information early so that the model can detect events that are both visible and audible in intermediate layers. Our early multimodal fusion can be divided into unimodal attention, audio (visual) guided mixing, and cross-modal pyramid, which are used for unimodal temporal modeling, multimodal information fusion, and multi-temporal resolution modeling, respectively.

[0079] a. Single-modal attention:

[0080] To capture long-term temporal relationships between unimodal segments and filter out some noise, we perform unimodal temporal modeling, as is the case with most methods. Specifically, we first optimize the original features through a projection layer, where F... A ′=f proj (F A ) and F V ′=f proj (F V Then the mapped audio and visual features F A ′, Used for time aggregation via single-mode transformer blocks. Each single-mode transformer block includes a multi-head self-attention (MSA) and a feedforward network (FFN), which can be represented as:

[0081]

[0082] in These are the outputs for audio-only and video-only features, respectively. These are learnable parameters.

[0083] b. Audio (visual) guided hybrid:

[0084] Our audio (visual) guided mixing is designed to enhance audio (visual) features To obtain supplementary information from the audiovisual representation, we first need to obtain early multimodal representations containing synchronized (frame-aligned) multimodal information. However, due to the original audio and visual features F... A ,F V Significant differences exist between them, and simply concatenating them can introduce a large amount of noise and lead to insufficient fusion. To address this issue, we first designed an aligner, which includes two projection layers and a feature optimization attention block, mapping them to the same representation space. The aligner is as follows:

[0085] F g =f proj (F A )⊙f proj (F V )

[0086] F g =MSA(F) g W q g ,F g W k g ,F g W v g )

[0087] in W represents an early multimodal representation. q g W k g ,

[0088] Then, similar to unimodal attention, early multimodal representation F g ′ is also used for time modeling:

[0089]

[0090] in Now, early multimodal representations It contains frame-aligned audio and visual information. We introduce an audio (visual) guided blending to extract audio (visual) representations. Related multimodal supplementary information. The audio (visual) guided mix includes a multi-head cross-attention (MCA) and an FFN, as shown below:

[0091]

[0092]

[0093] in Provide query vectors, and and Provides key and value vectors.

[0094] c. Cross-modal pyramid.

[0095] We further perform cross-modal fusion to enable the model to better understand audiovisual events at different temporal resolutions. Specifically, we use step size... For feature F AV ,F VA Perform time downsampling, where l c This is the index of the current pyramid block. After downsampling of audio and visual features, there is a cross-attention process. The output of the cross-modal pyramid module is... Where Tl It is the total length of the multi-level pyramid features.

[0096] 2) Multi-stage semantic guidance.

[0097] In this section, we aim to enhance audiovisual event localization by enabling the model to understand event-related content in intermediate layers. However, even with early multimodal fusion, the model still struggles to clearly understand the rules governing audiovisual event localization. Due to the significant differences between audio and video features, the lack of intermediate layers with cross-modal semantic bridging can lead to a modal semantic gap. We directly promote the understanding of audiovisual events through explicit, event-related semantic guidance and multi-stage constraints. Specifically, we enhance the understanding of audiovisual events through the output of early multimodal fusion. Decoding is performed using an MLP. Each MLP consists of two linear layers and a PReLU activation function. Then, we apply a focus loss function to the output of each MLP for semantic consistency constraints. Note that here we only detect the event category for each segment, without regressing the temporal boundaries. The total multi-stage classification loss is... as follows:

[0098]

[0099] Where α i These are the weights of the loss, i∈1,2,3. α1, α2, and α3 increase sequentially to ensure the effectiveness of multi-stage optimization from unimodal to multimodal information.

[0100] Hybrid Dependency Expert Module (MoDE): [This will...] and Connection as Subsequently, MoDE performs cross-modal temporal aggregation and event dependency modeling on top of it. Specifically, MoDE consists of two branches: one for deep temporal modeling, and the other for revealing multi-event relationships through the cascading of MoE layers. The outputs of the two branches... and Then add them together to obtain Used for final decoding.

[0101] Specifically,

[0102] (3) Dependency expert hybrid

[0103] The dependency expert hybrid module is designed for final cross-modal fusion and extraction of multi-event correlations. MoDE mainly consists of a multimodal temporal aggregation module and a multi-event dependency extraction module.

[0104] 1) Multimodal temporal aggregation

[0105] Before locating audiovisual events, we need to aggregate the results of ESI. And time modeling is applied to obtain a comprehensive cross-modal representation. By doing so, the model can capture long-term cross-modal relationships and filter out irrelevant noise. Specifically, firstly... and These are connected together to form Z. At this point, Z already contains comprehensive cross-modal information and has undergone multi-stage semantic guidance. However, and Containing some redundant information, we first input Z into a projection layer to obtain preliminary fusion. Here, C is the dimension of Z′, which is also the number of event categories. Then, to further refine the cross-modal information, we adopt temporal modeling, the specific method of which is the same as the temporal modeling described above, consisting of N1+N2 multi-head self-attention units. It is the output after N1 time attention blocks, Z t It is the final output after passing through N1+N2 time attention blocks.

[0106] 2) Multi-event dependency extraction

[0107] In audiovisual scenarios, dependencies exist between concurrent events (e.g., thunder is often accompanied by rain and wind). However, these dependencies are complex and variable across different scenarios, and events may overlap to varying degrees. We aim for the model to automatically reveal event dependencies based on different inputs. We use a series of MoE layers, each containing n experts, selecting only one expert at a time. By combining different experts across m layers, we can adaptively reveal complex multi-event dependencies (each MoE layer has n experts, resulting in n...). m (Number of possible combinations). Specifically, to ensure that only one expert is selected at a time for each layer, we designed a MoE gate with Gumbel-Softmax. The output of the i-th layer is:

[0108]

[0109] Where E i =E i,1 E i,2 ,…,E i,n This is the expert set for the i-th MoE layer. The input to the first layer is the intermediate output mentioned in multimodal time aggregation. E i Each expert in the process consists of a convolutional layer and an adjacency matrix. It consists of a LeakyReLU activation function:

[0110]

[0111] Where E i,j He is the j-th expert at the i-th level. This is the output of the (i-1)th MoE layer. Finally, the output Z of the last MoE layer... e With Z t Combined to obtain Used for decoding.

[0112] Decoder: The decoder consists of a classification head and a regression head. Each head comprises three layers of one-dimensional convolutions and one activation function (Sigmoid for the classification head, ReLU for the regression head). It integrates features from all pyramid levels. It is input into the decoder to predict the probability p(c) of the event category. n ) and from the current time t to the event time boundary distance Please note that the regression head here is class-aware, and its output is only valid when an audiovisual event is detected.

[0113] This embodiment also proposes the training and inference methods used in this method, as follows:

[0114] Training, such as Figure 4 As shown, firstly, the model parameters are initialized, including loading the pre-trained parameters of the feature extractor and randomly initializing other parameters. Raw feature extraction is performed using the frozen audio / video feature extractor. At this point, early multimodal fusion is performed, including unimodal temporal modeling, audio (video)-guided hybrid modeling, and multi-temporal resolution pyramid modeling. While the multimodal data has undergone preliminary fusion, semantic supervision is still lacking. Therefore, event detection supervision signals with different weights are added to the outputs of the three stages of early multimodal fusion. Subsequently, the obtained features are subjected to multimodal temporal aggregation, enhancing the model's event detection capabilities for long-duration videos. Next, multi-event dependency modeling is performed to help the model infer concurrent events. Finally, the result is decoded by the decoder, and the loss function is calculated by label alignment to update the parameters of the entire model.

[0115] We employ an end-to-end training approach, where three loss functions are used to constrain the following aspects: focus loss for event classification. Generalized IoU loss for time boundary estimation And for multi-stage semantically guided classification The total loss function is expressed as:

[0116]

[0117] and Used to constrain the final audiovisual event detection and regression, while Cross-modal semantic bridging is used for the intermediate layer. Note that without sufficient features for cross-modal modeling, it is impossible to regress the accurate temporal boundaries of events. This excludes the generalized IoU loss used for boundary regression.

[0118] Reasoning, such as Figure 5 As shown, a frozen audio / video feature extractor is used for initial feature extraction. Early multimodal fusion is then performed, including unimodal temporal modeling, audio (video)-guided hybrid modeling, and multi-temporal resolution pyramid modeling. Subsequently, the obtained features are subjected to multimodal temporal aggregation, enhancing the model's event detection capabilities for long-duration videos. Next, multi-event dependency modeling is performed to aid the model in inferring concurrent events. Finally, the decoder outputs the final inference result.

[0119] During inference, given all segments of the video, the model outputs candidate objects that include event category and time boundaries, each accompanied by a confidence score. All candidates are then processed using multi-class Soft-NMS to filter out highly overlapping time boundaries within the same event category.

[0120] This embodiment provides a method for locating dense audiovisual events. It employs an event-aware semantic guidance network (ESG-Net) that progressively focuses on audio and video content across multiple intermediate layers and captures relationships between multiple events to achieve accurate localization of dense events. Results show that ESG-Net outperforms the baseline model by 2.1% and the state-of-the-art method by 0.7% on the I3D+VGGish backbone network, and outperforms the baseline model by 3.7% and the state-of-the-art method by 0.6% on the ONE-PEACE backbone network. In summary, the event-aware semantic guidance network can detect events more accurately, significantly reducing missed detections, and provides more accurate temporal localization of audiovisual events.

[0121] Please refer to Figure 6 , Figure 6 This invention provides a structural block diagram of a dense audiovisual event localization device; the specific device may include:

[0122] The feature extraction module 100 preprocesses the video data to be located and extracts features from the preprocessed video data to obtain initial feature data.

[0123] The fusion module 200 performs temporal modeling on the initial feature data based on a single-modal attention mechanism, an audio-visual guidance mechanism, and a cross-modal pyramid mechanism to obtain multimodal fused data.

[0124] The constraint module 300 uses a multi-stage semantic guidance method to constrain the classification loss function of the multimodal fusion data to obtain constraint data;

[0125] The optimization processing module 400 refines the constraint data based on the multimodal time aggregation method and the multi-event dependency extraction method to obtain comprehensive feature data;

[0126] The decoding module 500 processes the comprehensive feature data based on the decoder to obtain the predicted event category and time boundary.

[0127] This embodiment provides a dense audiovisual event localization device for implementing the aforementioned dense audiovisual event localization method. Therefore, the specific implementation of the dense audiovisual event localization device can be found in the previous embodiment section of the dense audiovisual event localization method. For example, the feature extraction module 100, fusion module 200, constraint module 300, optimization processing module 400, and decoding module 500 are respectively used to implement steps S101, S102, S103, S104, and S105 in the aforementioned dense audiovisual event localization method. Therefore, its specific implementation can be referred to the description of the corresponding embodiments, which will not be repeated here.

[0128] A specific embodiment of the present invention also provides a dense audiovisual event localization device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the above-described dense audiovisual event localization method.

[0129] A specific embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for locating dense audiovisual events.

[0130] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0131] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0132] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0134] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for locating dense audiovisual events, characterized in that, include: The video data to be located is preprocessed, and features are extracted from the preprocessed video data to obtain initial feature data. The initial feature data is temporally modeled based on a single-modal attention mechanism, an audio-visual guidance mechanism, and a cross-modal pyramid mechanism to obtain multimodal fusion data; A multi-stage semantic guidance method is used to constrain the classification loss function of the multimodal fusion data to obtain constrained data; The constraint data is refined based on the multimodal time aggregation method and the multi-event dependency extraction method to obtain comprehensive feature data; The decoder processes the comprehensive feature data to obtain the predicted event category and time boundary.

2. The method for locating dense audiovisual events according to claim 1, characterized in that, The process of performing temporal modeling on the initial feature data based on a single-modal attention mechanism, an audio-visual guidance mechanism, and a cross-modal pyramid mechanism to obtain multimodal fusion data includes: A unimodal attention mechanism is used to perform temporal modeling of audio and visual features, and the audio and visual features are projected into a higher-dimensional space respectively. Information is then aggregated through multi-head self-attention and feedforward networks to obtain preliminary aggregated data. The initial aggregated data is supplemented using an audio-visual guidance mechanism. After mapping the audio and visual data to the same space using an aligner, temporal downsampling processing is performed using a cross-modal pyramid mechanism. Obtain multimodal fusion data.

3. The method for locating dense audiovisual events according to claim 2, characterized in that, The information is aggregated sequentially through a multi-head self-attention network and a feedforward network, and the formula for calculating the preliminary aggregated data is as follows: in, These are the outputs for audio-only and video-only features, respectively. These are learnable parameters.

4. The method for locating dense audiovisual events according to claim 1, characterized in that, The formula for calculating the classification loss function constraint is as follows: in, For multi-stage classification loss function, α i These are the weights of the loss, i∈1,2,3, with α1, α2, and α3 increasing sequentially.

5. The method for locating dense audiovisual events according to claim 1, characterized in that, The constraint data is refined using the multimodal time aggregation method and the multi-event dependency extraction method to obtain comprehensive feature data, including: Audio and visual features are aggregated, and cross-modal features are obtained by concatenating the audio and visual features. The cross-modal features are then input into the projection layer to obtain preliminary fused features. The dependencies of the preliminary fusion features are obtained based on the hybrid expert layer, and the output of the hybrid expert layer is added to the output of the audio features and visual features aggregation to obtain comprehensive feature data.

6. The method for locating dense audiovisual events according to claim 1, characterized in that, The decoder includes a classification head and a regression head, each consisting of three layers of one-dimensional convolution and one layer of activation function. The classification head uses a Sigmoid activation function to output the event class probability, and the regression head uses a ReLU activation function to output the temporal boundary distance.

7. The method for locating dense audiovisual events according to claim 1, characterized in that, Also includes: Training is performed end-to-end, using three loss functions to constrain event classification, temporal boundary estimation, and multi-stage semantic guidance, respectively. The total loss function is: in, and Used to constrain the final audiovisual event detection and regression, while Cross-modal semantic bridging for the middle layer.

8. A dense audiovisual event positioning device, characterized in that, include: The feature extraction module preprocesses the video data to be located and extracts features from the preprocessed video data to obtain initial feature data. The fusion module performs temporal modeling on the initial feature data based on a single-modal attention mechanism, an audio-visual guidance mechanism, and a cross-modal pyramid mechanism to obtain multimodal fused data. The constraint module uses a multi-stage semantic guidance method to constrain the classification loss function of the multimodal fusion data to obtain constraint data; The optimization processing module refines the constraint data based on the multimodal time aggregation method and the multi-event dependency extraction method to obtain comprehensive feature data; The decoding module processes the comprehensive feature data based on the decoder to obtain the predicted event category and time boundary.

9. A dense audiovisual event positioning device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the dense audiovisual event localization method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the dense audiovisual event localization method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Audio-visual event positioning system and method based on cross-modal consistency and time sequence multi-granularity cooperation

    CN119152337A

  • Recurrent multimodal attention system based on expert gated networks

    US20190354797A1