Bimodal relational network for audiovisual event localization

By combining visual and audio features through a bimodal relationship network, and utilizing cross-modal relationships for event localization, the difficulty of localization caused by visual background interference in existing technologies is solved, and more accurate event recognition is achieved.

CN116171473BActive Publication Date: 2026-06-02INTERNATIONAL BUSINESS MACHINE CORPORATION

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2021-07-05
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing methods struggle to effectively utilize the cross-modal relationships between visual and audio information, making it difficult to accurately locate events in videos against complex visual backgrounds.

Method used

A bimodal relational network is employed, which combines visual and audio features with an audio-guided visual attention module and intramodal/extramodal relational blocks to perform event localization using cross-modal relations.

Benefits of technology

It improves the accuracy and robustness of event localization, enabling accurate identification of audible and visible events in videos against complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116171473B_ABST
    Figure CN116171473B_ABST
Patent Text Reader

Abstract

A dual modality relationship network for audiovisual event localization can be provided. A video feed for audiovisual event localization can be received. Based on a combination of extracted audio features and video features of the video feed, information features and regions in the video feed can be determined by running a first neural network. Based on the information features and regions in the video feed determined by the first neural network, relationship-aware video features can be determined by running a second neural network. Based on the information features and regions in the video feed, relationship-aware audio features can be determined by running a third neural network. A dual modality representation can be obtained based on the relationship-aware video features and the relationship-aware audio features by running a fourth neural network. The dual modality representation can be input to a classifier to identify an audiovisual event in the video feed.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This application generally relates to computers and computer applications, and more specifically to artificial intelligence, machine learning, neural networks, and audio-visual learning and audio-visual event localization.

[0002] Event localization is a challenging task in video understanding, requiring machines to locate events or actions and identify categories in unconstrained video. Some existing methods use only red-green-blue (RGB) frames or optical flow as input to localize and identify events. However, strong visual background interference and large variations in visual content can make it difficult to localize events using only visual information.

[0003] Audiovisual event localization (AVE) has garnered increasing attention. AVE localization requires machines to determine the presence of audible and visible events within video clips and to what category those events belong. AVE localization can be challenging due to: 1) the complex visual backgrounds in unconstrained videos making AVE localization difficult, and 2) the need for machines to consider information from both modalities (i.e., audio and video) simultaneously and leverage their relationships. Establishing connections between complex visual scenes and intricate sound is crucial. Some methods in this task process the two modalities independently and fuse them only before the final classifier. Existing methods primarily focus on capturing temporal relationships within clips within a single modality as potential cues for event localization. Summary of the Invention

[0004] This disclosure is provided to aid in understanding computer systems, computer applications, machine learning, neural networks, audiovisual learning, and audiovisual event localization, and is not intended to limit the scope of this disclosure or the invention. It should be understood that various aspects and features of this disclosure may be used advantageously alone in some cases, or in combination with other aspects and features of this disclosure in other instances. Therefore, variations and modifications can be made to computer systems, computer applications, machine learning, neural networks, and / or their methods of operation to achieve different effects.

[0005] A system and method may be provided that can implement a bimodal relational network for audiovisual event localization. In one aspect, the system may include a hardware processor and a memory coupled to the hardware processor. The hardware processor may be configured to receive a video feed for audiovisual event localization. The hardware processor may also be configured to determine information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The hardware processor may also be configured to determine relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to determine relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may further be configured to obtain a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The hardware processor may also be configured to input the bimodal representation into a classifier to identify audiovisual events in the video feed.

[0006] In another aspect, the system may include a hardware processor and a memory coupled to the hardware processor. The hardware processor may be configured to receive a video feed for locating audiovisual events. The hardware processor may also be configured to determine information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The hardware processor may also be configured to determine relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to determine relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to obtain a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The hardware processor may also be configured to input the bimodal representation into a classifier to identify audiovisual events in the video feed. The hardware processor may also be configured to run a first convolutional neural network with at least the video portion of the video feed to extract video features.

[0007] In another aspect, the system may include a hardware processor and a memory coupled to the hardware processor. The hardware processor may be configured to receive a video feed for locating audiovisual events. The hardware processor may also be configured to determine information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The hardware processor may also be configured to determine relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to determine relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to obtain a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The hardware processor may also be configured to input the bimodal representation into a classifier to identify audiovisual events in the video feed. The hardware processor may also be configured to run a second convolutional neural network with at least the audio portion of the video feed to extract audio features.

[0008] In another aspect, the system may include a hardware processor and a memory coupled to the hardware processor. The hardware processor may be configured to receive a video feed for locating audiovisual events. The hardware processor may also be configured to determine information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The hardware processor may also be configured to determine relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to determine relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may further be configured to obtain a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The hardware processor may also be configured to input the bimodal representation into a classifier to identify audiovisual events in the video feed. This bimodal representation may be used as the final layer of the classifier in identifying audiovisual events.

[0009] In another embodiment, the system may include a hardware processor and a memory coupled to the hardware processor. The hardware processor may be configured to receive a video feed for locating audiovisual events. The hardware processor may also be configured to determine information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The hardware processor may also be configured to determine relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to determine relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may further be configured to obtain a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The hardware processor may also be configured to input the bimodal representation into a classifier to identify audiovisual events in the video feed. The classifier's identification of audiovisual events in the video feed includes identifying the location of the audiovisual event in the video feed and the category of the audiovisual event.

[0010] In another embodiment, the system may include a hardware processor and a memory coupled to the hardware processor. The hardware processor may be configured to receive a video feed for locating audiovisual events. The hardware processor may also be configured to determine information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The hardware processor may also be configured to determine relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to determine relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to obtain a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The hardware processor may also be configured to input the bimodal representation into a classifier to identify audiovisual events in the video feed. The second neural network may acquire both temporal information from the video features and cross-modal information between the video and audio features when determining the relation-aware video features.

[0011] In another aspect, the system may include a hardware processor and a memory coupled to the hardware processor. The hardware processor may be configured to receive a video feed for locating audiovisual events. The hardware processor may also be configured to determine information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The hardware processor may also be configured to determine relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to determine relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The hardware processor may also be configured to obtain a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The hardware processor may also be configured to input the bimodal representation into a classifier to identify audiovisual events in the video feed. The third neural network may acquire both temporal information from the audio features and cross-modal information between the video features and the audio features when determining the relation-aware audio features.

[0012] In one aspect, a method may include receiving a video feed for locating audiovisual events. The method may further include determining information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The method may further include determining relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include determining relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include obtaining a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The method may further include inputting the bimodal representation into a classifier to identify audiovisual events in the video feed.

[0013] In another aspect, the method may include receiving a video feed for locating audiovisual events. The method may further include determining information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The method may further include determining relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include determining relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include obtaining a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The method may further include inputting the bimodal representation into a classifier to identify audiovisual events in the video feed. The method may further include running a first convolutional neural network with at least a video portion of the video feed to extract video features.

[0014] In another aspect, the method may include receiving a video feed for locating audiovisual events. The method may further include determining information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The method may further include determining relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include determining relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include obtaining a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The method may further include inputting the bimodal representation into a classifier to identify audiovisual events in the video feed. The method may further include running a second convolutional neural network with at least the audio portion of the video feed to extract audio features.

[0015] In another aspect, the method may include receiving a video feed for locating audiovisual events. The method may further include determining information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The method may further include determining relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include determining relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include obtaining a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The method may further include inputting the bimodal representation into a classifier to identify audiovisual events in the video feed. The bimodal representation may be used as the final layer of the classifier in identifying audiovisual events.

[0016] In another aspect, the method may include receiving a video feed for locating audiovisual events. The method may further include determining information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The method may further include determining relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include determining relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include obtaining a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The method may further include inputting the bimodal representation into a classifier to identify audiovisual events in the video feed. The classifier identifying audiovisual events in the video feed may include identifying the location of the audiovisual event in the video feed and the category of the audiovisual event.

[0017] In another aspect, the method may include receiving a video feed for locating audiovisual events. The method may further include determining information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The method may further include determining relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include determining relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include obtaining a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The method may further include inputting the bimodal representation into a classifier to identify audiovisual events in the video feed. The second neural network may acquire both temporal information from the video features and cross-modal information between the video and audio features when determining the relation-aware video features.

[0018] In another aspect, the method may include receiving a video feed for locating audiovisual events. The method may further include determining information features and regions in the video feed by running a first neural network based on a combination of extracted audio and video features from the video feed. The method may further include determining relation-aware video features by running a second neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include determining relation-aware audio features by running a third neural network based on the information features and regions in the video feed determined by the first neural network. The method may further include obtaining a bimodal representation based on the relation-aware video features and relation-aware audio features by running a fourth neural network. The method may further include inputting the bimodal representation into a classifier to identify audiovisual events in the video feed. The third neural network, in determining the relation-aware audio features, acquires both temporal information from the audio features and cross-modal information between the video features and the audio features.

[0019] A computer-readable storage medium may also be provided that stores a program of instructions that can be executed by a machine to perform one or more methods described herein.

[0020] Other features, structures, and operations of various embodiments are described in detail below with reference to the accompanying drawings. In the drawings, the same reference numerals denote the same or functionally similar elements. Attached Figure Description

[0021] Figure 1 This is an illustrative example of an audiovisual event localization task.

[0022] Figure 2 This is a diagram illustrating the bimodal relationship network in the embodiment.

[0023] Figure 3 This is another diagram illustrating the bimodal relationship network in the embodiment.

[0024] Figure 4 An audio-guided spatial channel attention (AGSCA) module is shown in one embodiment.

[0025] Figure 5 The Cross-Modal Relationship Attention (CMRA) mechanism in the embodiment is illustrated.

[0026] Figure 6 Example positioning results output by the method and / or system in the embodiments are shown.

[0027] Figure 7 This is a flowchart illustrating a method for locating audiovisual events in an embodiment.

[0028] Figure 8 This is a diagram illustrating components of a system in one embodiment, which can implement a bimodal relational network for audiovisual event localization.

[0029] Figure 9 A schematic diagram of an example computer or processing system that can implement a bimodal relational network system is shown in one embodiment. Detailed Implementation

[0030] Systems, methods, and techniques can be provided that can identify the presence of both audible and visual events in a given untrimmed video sequence with visual and audio (audio) channels, and determine the category of the event. For example, a machine can be trained to locate audiovisual events. When identifying audiovisual events in a video sequence, these systems, methods, and techniques consider cross-modal or intermodal relationship information between the visual scene and the audio signal.

[0031] In an embodiment, the bimodal relational network is an end-to-end network for performing audiovisual event localization tasks and may include an audio-guided visual attention module, intramodal relational blocks, and intermodal relational blocks. In an embodiment, the audio-guided visual attention module is used to highlight informational regions for reducing visual background interference. In an embodiment, the intramodal and intermodal relational blocks may individually utilize intramodal and intermodal relational information to facilitate presentation learning (e.g., for audiovisual representation learning), which facilitates the recognition of both audible and visible events. On one hand, the bimodal relational network can reduce visual background interference by highlighting certain regions and improve the quality of representations of both modalities by considering intramodal and intermodal relations as potentially useful information. On the other hand, this bimodal relational network enables the capture of valuable intermodal relations between visual scenes and sound, which is largely unavailable in existing methods. For example, the method in the embodiment may feed extracted visual and audio features into the audio-guided visual attention module to emphasize informational regions for background interference reduction. This method can prepare intra-modal and inter-modal relational blocks to independently utilize correspondence information used for audio / visual representation learning. It can also merge relation-aware visual and audio features to obtain a comprehensive bimodal representation for a classifier.

[0032] Machines can be made capable of performing event localization tasks. Machines performing event localization automatically locate events and identify their categories in unconstrained video. Most existing methods utilize only the visual information of the video, ignoring its audio information. However, reasoning using both visual and audio content can aid event localization, for example, because audio signals often carry useful cues for reasoning. Furthermore, audio information can guide the machine or machine model to pay more attention to or focus on informational areas of the visual scene, which can help reduce interference from the background. In embodiments, relation-aware networks utilize both audio and visual information for accurate event localization, for example, providing technological improvements in the machine when identifying audio-visual events in a video stream. In embodiments, to reduce interference introduced by the background, systems, methods, and techniques can implement an audio-guided spatial channel attention module to guide the model to focus on event-related visual regions. Systems, methods, and techniques can also utilize relation-aware modules to establish connections between visual and audio modalities. For example, systems, methods, and techniques learn representations of video and / or audio segments by aggregating information from other modalities based on cross-modal relationships. Based on relation-aware representations, systems, methods, and techniques can perform event localization by predicting event relevance scores and classification scores. In various embodiments, the neural network can be trained to perform event localization in a video stream. Various implementations of the neural network operations can be used, such as different activation functions and optimizations, such as gradient optimization.

[0033] Systems, methods, and techniques consider cross-modal or intermodal relationship information between visual scenes and audio signals, for example, for AVE localization. Cross-modal relationships are the audiovisual correlations between audio and video segments. Figure 1 This is an illustrative example of an audiovisual event localization task. In this embodiment, machine 102 takes a video sequence 104 having a visual channel 106 and an acoustic channel 108 as input. Machine 102 includes, for example, a hardware processor. The hardware processor may include, for example, components such as programmable logic devices, microcontrollers, memory devices, and / or other hardware components that can be configured to perform the corresponding tasks described in this disclosure. Machine 102 is asked to determine whether an event that is both audible and visible exists in the segment and to what category the event belongs. In one aspect, the challenge is that the machine needs to consider information from both modalities simultaneously and utilize their relationship. For example, as Figure 1 As shown, the video sequence may include the sound of a train horn while simultaneously visualizing a moving train, for example, shown as a frame or clip at 110b. This audiovisual correlation suggests both audible and visible events. Therefore, cross-modal or intermodal relationships also contribute to the detection of audiovisual events.

[0034] Self-attention mechanisms can be used to capture intra-modal relationships between words in Natural Language Processing (NLP). It first transforms input features into query, key, and value (i.e., memory) features. Then, it computes the attention output using a weighted sum of all values ​​in the memory, where the weights (i.e., relationships) are learned from the keys and queries in the memory. However, in NLP applications, directly applying self-attention to event localization cannot leverage cross-modal relationships between visual and acoustic content because the query and memory are derived from the same modality. Conversely, if the memory captures features from both modalities, a query (from one of the two modalities) can enable the exploration of cross-modal relationships without losing relevant intra-modal information.

[0035] In embodiments, the systems, methods, and techniques provide a relation-aware module to establish connections between visual and audio information by leveraging intermodal relationships. In these embodiments, this module encapsulates an attention mechanism known as cross-modal relational attention. Unlike self-attention, in cross-modal relational attention, the query is derived from one modality, while the key and value are derived from both modalities. In this way, a single segment from a modality can aggregate useful information from all relevant segments from both modalities based on learned intramodal and intermodal relationships. Simultaneously viewing a visual scene and listening to sound (i.e., utilizing information from both modalities simultaneously) is more efficient and effective than perceiving them separately for locating audible and visible events. In one aspect, the systems, methods, and techniques can utilize both useful relationships to facilitate representation learning and further improve the performance of AVE localization.

[0036] In embodiments, because strong visual background interference can hinder accurate event localization, systems, methods, and techniques can highlight informational visual regions and features to reduce interference. For example, systems, methods, and techniques may include an audio-guided spatial channel attention module that utilizes audio information to establish visual attention at both the spatial and channel levels. Systems, methods, and techniques integrate these components and provide a cross-modal relationship-aware network that significantly outperforms state-of-the-art techniques in supervised and weakly supervised AVE localization tasks on the AVE dataset.

[0037] In embodiments, the system, method, and techniques may include an audio-guided spatial channel attention module (AGSCA) to utilize the guiding capabilities of audio signals for visual attention, which can accurately highlight informational features and sound regions; and a relation-aware module to utilize intra-modal and inter-modal relations for event localization. In embodiments, a cross-modal relation-aware network (also known as a bimodal relation network) may be established for supervised and weakly supervised AVE localization tasks.

[0038] Audiovisual learning can be useful in many fields, such as action recognition, sound source localization, and audiovisual event localization. For example, this work uses audio to build a preview mechanism to reduce temporal redundancy; sparse temporal sampling strategies can fuse multiple modalities to enhance action recognition; audio can be used as a supervisory signal for learning visual models in an unsupervised manner; the Speech2Face framework can be presented, which uses speech-face correlations to generate facial images behind speech; and to leverage readily available large-scale unlabeled videos, this work utilizes audiovisual correspondences to learn audiovisual representations in a self-supervised manner.

[0039] Another work on audiovisual event localization uses two Long Short-Term Memory (LSTMs) to model the temporal dependencies of audio and video segment sequences separately, then simply fuses the audio and visual features via fusion and average pooling for event category prediction. Yet another work first processes the audio and visual modalities separately, then fuses the features of both modalities via an LSTM that operates in a sequence-to-sequence manner. Another work proposes a dual-attention matching module that uses global and local information obtained through intramodal relation modeling to measure cross-modal similarity via inner product operations. Cross-modal similarity is directly used as the final event relevance prediction. These methods primarily focus on leveraging intramodal relations as potential cues, ignoring equally valuable cross-modal relation information for event localization. In contrast to these methods, the systems, methods, and techniques described in this embodiment provide or implement cross-modal relation-aware networks, enabling, for example, bridging connections between visual and audio modalities by simultaneously utilizing both intramodal and intermodal relevant information.

[0040] Attention mechanisms mimic human visual perception. They attempt to automatically focus on certain highly activated parts of the input. Attention mechanisms have many variations, including self-attention. Unlike self-attention, which focuses on capturing intramodal relationships, the systems, methods, and techniques described in this embodiment can provide cross-modal relationship attention, enabling audiovisual representation learning to utilize both intramodal and intermodal relationships simultaneously.

[0041] In this disclosure, the following symbols are used. It is a video sequence with T non-overlapping segments. Here, Vt and At represent the visual content and corresponding audio content of the t-th segment, respectively.

[0042] For example, Figure 1 The video clips 110a, 110b, 110c, 110d, 110e, and 110f are shown. (For example...) Figure 1 As illustrated in the example, given a video sequence S 104, the AVE localization request machine predicts the event label (including background) for each segment St, relying on Vt and At. An audiovisual event is defined as an event that is both audible and visible (i.e., hearing a sound emitted by an object and simultaneously seeing the object). If segment St is not both audible and visible, it should be predicted as background. The challenge in this task is that the machine needs to analyze the two modalities and capture their relationship. In embodiments, systems, methods, and techniques can use cross-modal relevant information to improve performance. In embodiments, this task can be performed in different settings. For example, in one embodiment, the task can be performed in a supervised setting. In another embodiment, the task can be performed in a weakly supervised setting. In a supervised setting, systems, methods, and techniques can access segment-level labels during the training phase. Segment-level labels indicate the category (including background) of the corresponding segment. In one embodiment, non-background category labels are given only when the sound and the corresponding sounding object are presented. In a weakly supervised setting, in one embodiment, systems, methods, and techniques can access only video-level labels during training, and the systems, methods, and techniques are designed to predict the category of each segment during testing. Video-level tags indicate whether a video contains audiovisual events and what category those events belong to.

[0043] In the embodiments, the systems, methods, and techniques address the problem that most existing event localization methods ignore information from audio signals in video; however, this can help mitigate interference from complex backgrounds and provide more cues for reasoning. For example, one approach utilizes both visual and audio information for event localization and evaluates it on an audiovisual event localization task, which requires the machine to localize events that are both audible and visible in unretouched video. This task is challenging because unretouched video often contains complex backgrounds, and establishing connections between complex visual scenes and intricate sounds is important. To address these challenges, in the embodiments, the systems, methods, and techniques provide an audio-guided attention module to highlight certain spatial regions and features to reduce background interference. In the embodiments, the systems, methods, and techniques also incorporate a relation-aware module to utilize inter-modal and intra-modal relationships for audiovisual event localization.

[0044] Figure 2 This is a diagram illustrating the bimodal relationship network in the embodiments. The components shown include computer-implemented components, such as those implemented and / or operating on or coupled to one or more hardware processors. One or more hardware processors or processors may include, for example, components such as programmable logic devices, microcontrollers, memory devices, and / or other hardware components that can be configured to perform the corresponding tasks described in this disclosure. Coupled memory devices may be configured to selectively store instructions executable by one or more hardware processors. Processors may be central processing units (CPUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), other suitable processing components or devices, or one or more combinations thereof. Processors may be coupled to memory devices. Memory devices may include random access memory (RAM), read-only memory (ROM), or other memory devices, and may store data and / or processor instructions for implementing various functions associated with the methods and / or systems described herein. Processors may execute computer instructions stored in memory or received from another computer device or medium. Modules used herein may be implemented as software, hardware components, programmable hardware, firmware, or any combination thereof executable on one or more hardware processors.

[0045] A bimodal relational network is also known as a cross-modal relational perception network. In an embodiment, the bimodal relational network 200 is an end-to-end network for performing audiovisual event localization tasks and may include an audio-guided visual attention module 212, intramodal relational blocks 214, 216, and intermodal relational blocks 218, 220. The audio-guided visual attention module 212 may include a neural network (e.g., referred to as a first neural network for explanation or illustration). In an embodiment, the audio-guided visual attention module 212 is used to highlight informational regions for reducing visual background interference.

[0046] In embodiments, intramodal and intermodal relation blocks 214, 216, 218, and 220 can individually utilize intramodal and intermodal relation information to facilitate presentation learning, for example, for audiovisual representation learning, which facilitates the recognition of events that are both audible and visible. Intramodal and intermodal relation blocks 214 and 218 may include neural networks (e.g., referred to as a second neural network for interpretation). Intramodal and intermodal relation blocks 216 and 220 may include neural networks (e.g., referred to as a third neural network for interpretation). In one aspect, the bimodal relation network 200 can reduce visual background interference by highlighting certain areas and improve the quality of representations of both modalities by utilizing intramodal and intermodal relations as potentially useful information. In another aspect, the bimodal relation network enables the capture of valuable intermodal relations between visual scene 202 and sound 204.

[0047] For example, the method in the embodiment can feed the extracted visual and audio features into an audio-guided visual attention module 212 to emphasize informational regions for background interference reduction. For example, video features fed into the audio-guided visual attention module 212 can be extracted by inputting the input video 202 into a convolutional neural network 206, which is trained to extract video features. The input audio 204 can be processed using a log-Mel spectrogram representation 208, which can be input into a convolutional neural network 210 and trained to extract audio features for feeding into the audio-guided visual attention module 212. The input video 202 and input audio 204 are components of a video feed, stream, or sequence. The method can prepare intra-modal and inter-modal relation blocks 214, 216, 218, and 220 to utilize correspondence information learned for the audio / visual representations, respectively. For example, intramodal relation block 214 and intermodal relation block 218 generate relation-aware features 222; intramodal relation block 216 and intermodal relation block 220 generate relation-aware features 224. The audio-video interaction module 226 can combine the relation-aware visual and audio features 222 and 224 to obtain a comprehensive bimodal representation for a classifier. The audio-video interaction module 226 may include a neural network (e.g., referred to as a fourth neural network for interpretation). The comprehensive bimodal representation output by the audio-video interaction module 226 can be fed into a classifier (e.g., a neural network) for event classification 230 and / or event-related prediction 228.

[0048] As an example, the input AVE dataset (e.g., video and audio inputs 202, 204) can contain videos covering a wide range of domain events (e.g., human activity, animal activity, musical performances, and vehicle sounds). Events can involve multiple categories (e.g., church bells, crying, dog barking, fried food, violin playing, and / or others). As an example, a video can contain a single event and can be divided into multiple time interval segments (e.g., ten one-second segments) for processing by a bimodal relational network. In one embodiment, the video and audio scenes (e.g., video and audio inputs 202, 204) in the video sequence are aligned. In another embodiment, the video and audio scenes (e.g., video and audio inputs 202, 204) in the video sequence do not need to be aligned.

[0049] As an example, CNN 206 can be a convolutional neural network, such as, but not limited to, VGG-19, a residual neural network (e.g., ResNet-151), and can be pre-trained on ImageNet, for example, as a visual feature extractor. For instance, 16 frames can be selected as input within each segment. As an example, the output of a pool5 layer with dimensions 7×7×512 in VGG-19 can be considered as visual features. For ResNet-151, the output of a conv5 layer with dimensions 7×7×2048 can be considered as visual features. Frame-level features within each segment can be averaged over time to obtain segment-level features.

[0050] By way of example, input audio 204 (which may be the original audio) can be converted into a log-Mel spectrogram 208. This method and / or system can, for example, use a VGG-like network pre-trained on AudioSet to extract acoustic features of size 128 for each segment.

[0051] Figure 3 This is another diagram illustrating the bimodal relationship network in the embodiments. The components shown include computer-implemented components, such as those implemented and / or operating on or coupled to one or more hardware processors. One or more hardware processors or processors may include, for example, components such as programmable logic devices, microcontrollers, memory devices, and / or other hardware components that can be configured to perform the corresponding tasks described in this disclosure. Coupled memory devices may be configured to selectively store instructions executable by one or more hardware processors. Processors may be central processing units (CPUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), other suitable processing components or devices, or one or more combinations thereof. Processors may be coupled to memory devices. Memory devices may include random access memory (RAM), read-only memory (ROM), or other memory devices, and may store data and / or processor instructions for implementing various functions associated with the methods and / or systems described herein. Processors may execute computer instructions stored in memory or received from another computer device or medium. Modules used herein may be implemented as software, hardware components, programmable hardware, firmware, or any combination thereof executable on one or more hardware processors.

[0052] The bimodal relational network is also known as the cross-modal relational perception network (CMRAN). Input video 302 is fed into or input to a convolutional neural network (CNN) 306, for example, trained to extract video features. Input audio 304 can be processed using a log-Mel spectrogram representation 308, which can be input to a convolutional neural network (CNN) 310 and trained to extract audio features for feeding into an audio-guided spatial channel attention module (AGSCA) (e.g., in...). Figure 2 The audio features in CNN 312 (also known as the audio-guided visual attention module) are used. Using video features extracted from CNN 306 and audio features extracted from CNN 310, the audio-guided spatial channel attention module (AGSCA) (e.g., in...) Figure 2 The audio-guided visual attention module 312 (also known as the audio-guided visual attention module) is used to guide visual attention at the spatial and channel levels (e.g., video channels) using audio information (e.g., output by the CNN 310), thereby generating enhanced visual features 314. The CNN 310 extracts audio features 316. Two relation-aware modules 322 and 324 capture intra-modal and inter-modal relations for the two modalities (video and audio), respectively, thereby generating relation-aware visual features 322 and relation-aware audio features 324. The cross-modal relation-aware visual features 322 and the cross-modal relation-aware audio features 324 are combined via the audio-video interaction module 326 to generate a joint bimodal representation, which can be input into a classifier for event-related prediction 328 and / or event classification 330.

[0053] Given a video sequence S, a method and / or system, for example, forwards each audiovisual pair {V} via a pre-trained CNN backbone 306, 308. t A t}302, 304 are used to extract fragment-level features The method and / or system forwards audio and visual features via AGSCA module 312 to obtain enhanced visual features 314. Utilizing audio features 316 and enhanced visual features 314, the method and / or system prepares two relation-aware modules, a video relation-aware module 318 and an audio relation-aware module 320, which respectively wrap cross-modal or bimodal relation attention around audio and visual features. The method and / or system feeds visual and audio features 314, 316 into relation-aware modules 318, 329 to utilize two relations for both modalities. Relationship-aware visual and audio features 322, 324 are fed into an audio-video interaction module 326, thereby generating a comprehensive joint bimodal representation for one or more event classifiers 330 or predictions 328.

[0054] Audio-guided spatial channel attention

[0055] Audio signals can guide visual modeling. Channel attention enables the discarding of irrelevant features and improves the quality of visual representation. The Audio-Guided Spatial Channel Attention Module (AGSCA) 312 in embodiments seeks to optimize audio guidance capabilities for visual modeling. In one aspect, in embodiments, AGSCA 312 utilizes audio signals to guide visual attention in both the spatial and channel dimensions, rather than engaging audio features only in the spatial dimension of visual attention. This emphasizes informative features and spatial regions to improve localization accuracy. Known methods or techniques can be used to perform channel and spatial attention sequentially.

[0056] Figure 4 It is shown in one embodiment, for example in Figure 3 The audio-guided spatial channel attention (AGSCA) module is shown at position 312. In this embodiment, AGSCA utilizes audio guidance capabilities to guide visual attention at both the channel level (left portion) and the spatial level (right portion). Given audio features... 402 and visual characteristics 404, where H and W are the height and width of the feature map, respectively. AGSCA generates a channel-wise attention map. 406 adaptively emphasizes information features. Then, AGSCA generates a spatial attention map for channel attention features 410. 408 highlights the sound-producing area, thereby generating visual features that attract attention in the channel space. 412. The attention process can be summarized as follows:

[0057]

[0058] in Represents matrix multiplication, and This indicates element-wise multiplication.

[0059] Channel-by-channel attention 406 generates attention map Furthermore, spatial attention 408 generates an attention map.

[0060] Channel-by-channel attention

[0061] In one embodiment, a method and / or system models the dependencies between channels of features using the guidance of an audio signal. In another embodiment, the method and / or system uses a fully connected layer with non-linearity to transform audio and visual features into a common space, thereby generating an audio guidance map. and having d vThe transformed visual features are of size ×(H*W). In an embodiment, the method and / or system spatially compresses the transformed visual features using global average pooling. Then, the method and / or system combines the visual features with... via element-wise multiplication. Integration to utilize The method and / or system generate a channel attention map by modeling the relationships between channels through visual features fused from two fully connected layers with nonlinear forwarding. In the embodiments, details are shown below:

[0062]

[0063] in, and It is a fully connected layer with a Rectified Linear Unit (ReLU) as the activation function. These are learnable parameters, where d = 256 is the hidden dimension, δa indicates global average pooling, and σ represents the sigmoid function.

[0064] Spatial attention

[0065] This method and / or system also utilizes the guiding ability of audio signals to guide visuospatial attention 408. Spatial attention 408 follows a similar pattern to channel-by-channel attention 406. On one hand, input visual features... 410 is channel attention.

[0066] In an embodiment, the method and / or system formulates the spatial attention process as follows:

[0067]

[0068]

[0069] in, It is a fully connected layer with ReLU as the activation function. The learnable parameter is d = 256, where δ represents the hyperbolic tangent function. Spatial attention maps are used. The method and / or system are based on In v t The above performs a weighted summation to highlight information areas and reduce spatial dimensions, thereby generating a channel spatial attention visual feature vector. 412 is used as the output.

[0070] Cross-modal relationship attention

[0071] In this embodiment, cross-modal relation attention is a component of the relation-aware module (e.g., in...). Figure 3(Shown at 318 and 320). Given visual and acoustic features, the method and / or system can leverage cross-modal relationships to establish bridges between two modalities without ignoring intra-modal relationship information. For this task, in embodiments, the method and / or system implements or provides a cross-modal relationship attention (CMRA) mechanism. Figure 5 The Cross-Modal Relation Attention (CMRA) mechanism in the embodiment is illustrated. Bars in different shades represent fragment-level features from different modalities. CMRA simultaneously utilizes intra-modal and inter-modal relations of audio or video fragment features and enables adaptive learning of the balance between these two relations. Query 502 is derived from features of one modality (e.g., audio or video) and is denoted as q1. For example, input features may include audio and video features as shown at 512. Key-value pairs 504 and 506 are derived from features of two modalities (e.g., audio and video), and the method and / or system encapsulates them into a key matrix K. 1,2 Sum matrix V 1,2 In an embodiment, the method and / or system treats the dot product operation as a pairwise relation function. The method and / or system then computes q1 with all keys K. 1,2 The dot product of each element is divided by the square root of its shared feature dimension dm, and the softmax function is applied to obtain the value V. 1,2 Attention weights. These are determined by q1 and K. 1,2 The relationship of learning (i.e., attention weights) is a weighted average of all values ​​V (508). 1,2 The sum of these values ​​is used to calculate the output of interest, 510.

[0072] In this embodiment, CMRA is defined as:

[0073]

[0074] Indices 1 and 2 represent different modalities. Since q1 comes from audio or visual features, and K... 1,2 and V 1,2 Drawing from both audio and visual features, CMRA enables adaptive learning of both intra-modal and inter-modal relationships, as well as a balance between them. Individual segments from modalities within a video sequence enable the extraction of useful information from all relevant segments of both modalities based on the learned relationships, which facilitates audiovisual representation learning and further improves the performance of AVE localization.

[0075] The following provides examples of specific instances of CMRA in AVE localization. Without loss of generality, for illustrative purposes, the following description uses visual features as queries. Given audio features... and visual features This method and / or system projects v onto the query features using a linear transformation, denoted as: Then, the method and / or system temporarily cascade v with a to obtain the original memory. Then, the method and / or system will m a,v Linear transformation into key features Sum value characteristics Cross-modal attention output v q Calculated as

[0076]

[0077] Among them W Q W K W V It has d m ×d m The size is a learnable parameter. While visual features v are used as the query for illustrative purposes in this example, it should be noted that audio features can also be used as queries that utilize audio features for relationships. In contrast, self-attention can be considered a special case of CMRA when the memory contains only modal features identical to the query. In an embodiment, CMRA can be implemented in the relationship-aware module described below.

[0078] Relationship-aware module

[0079] In an embodiment, the relationship-aware module (e.g., in...) Figure 3 The blocks shown at 318 and 320 involve cross-modal relational modules and internal temporal relational blocks, denoted as M, respectively. cmra and B self . Figure 2 Examples of cross-modal relational blocks at 218 and 220, and internal temporal relational blocks (also known as intra-modal relational blocks) at 214 and 216 are also shown. In the embodiment, module M cmra Includes a cross-modal relational attention mechanism (CMRA) to leverage relations. B self Used as M cmra The assistant. In this embodiment, the video / audio relationship-aware module in the example architecture is a relationship-aware module that uses visual or audio features as queries during CMRA operations.

[0080] For illustrative purposes, visual features from the AGSCA module Used as a query (e.g., in) Figure 3 (The video relationship-aware module shown at point 318). Given the visual feature v to be queried and the audio feature to be included in memory. The method and / or system transforms them into a common space via linear layers. For example, the transformed visual and audio features are represented as having T×d. m F of the same size v and F aThen, B self As input F a Pre-exploration of internal temporal relationships yields a representation as The self-attentional audio characteristics. M cmra As input F v as well as With the help of CMRA, we explore intramodal and intermodal relationships of visual features and generate relation-aware visual features v o (For example, in) Figure 3 (As shown at position 322) as output. The entire process can be summarized as follows:

[0081]

[0082]

[0083]

[0084] in and These are learnable parameters.

[0085] Cross-modal relationship module.

[0086] In this embodiment, the cross-modal relationship module M uses CMRA operation. cmra This is used to utilize inter-modal and intra-modal relationships. In an embodiment, the method and / or system performs CMRA in a multi-head setup as follows:

[0087] H = Concat(h1,...,h) n W h ,

[0088]

[0089] Where || represents a time cascading operation, W i Q W i K W i V W h is the parameter to be learned, and n represents the number of parallel CMRA modules. To avoid transmission losses from CMRA, the method and / or system can use F v Added as a residual join to H along with the layer normalized to

[0090] Hr = LayerNorm(H+F) v (8)

[0091] To further integrate information from several parallel CMRA operations, this method and / or system utilizes ReLU to forward H through two linear layers. r In the embodiment, the output v o Detailed calculations can be given as follows

[0092] v o =LayerNorm(O f +H r ),

[0093]

[0094] Where δ represents the ReLU function, and W3 and W4 are the learnable parameters of the two linear layers.

[0095] Internal time relation block

[0096] In an embodiment, the method and / or system uses self-attention in M cmra Replace CMRA in the middle to obtain the internal time relation block B. self Block B self Pre-focus on exploring the internal temporal relationships of a portion of the memory characteristics to assist M cmra .

[0097] Audio-video interaction module

[0098] The relation-aware module outputs cross-modal relation-aware visual and acoustic representations, which are represented as follows: and exist Figure 2 It is shown at positions 222 and 224, and also at... Figure 3 Points 322 and 324 are shown in the diagram. In an embodiment, the audio-video interaction module obtains a comprehensive representation of two modalities from one or more classifiers. In an embodiment, the audio-video interaction module attempts to capture the resonance between the visual and acoustic channels by combining v0 and a0.

[0099] In an embodiment, the method and / or system uses element-wise multiplication to fuse v o and a o To obtain a joint representation of these two modes, denoted as f av Then, the method and / or system utilize f av Come to participate in visual representation (vo) and acoustic representation (a) o , where v o and a o Visual and acoustic information are provided separately for better visual understanding and acoustic perception. This operation can be viewed as a variant of CMRA, where the query is a fusion of memory features. The method and / or system then adds residual connections and layer normalization to the attention output, similar to a relation-aware module.

[0100] In the embodiment, the fully bimodal representation O av The calculation is as follows:

[0101] O av =LayerNorm(O+f av ),

[0102]

[0103]

[0104] in This represents element-wise multiplication, and These are the parameters to be learned.

[0105] Positioning of oversight and weak oversight audiovisual events

[0106] Localization of supervision

[0107] In an embodiment, the audio-video interaction module (e.g., in...) Figure 2 It is shown at position 226 in the text, and also at... Figure 3 (As shown at position 336) to obtain T×d m Size characteristics O av In an embodiment, the method and / or system decomposes the localization into predicting two scores. One is a confidence score that determines whether the audiovisual event exists in the t-th video segment. The other is the event category score, where C represents the number of foreground categories. Confidence score. Calculated as

[0108]

[0109] Among them W s These are learnable parameters, and σ represents the sigmoid function. For category scores... The methods and / or systems in the embodiments for fusing feature O av Max pooling is performed to generate feature vectors.

[0110] Event category classifier (e.g., in Figure 3 (As shown at position 330) as input o av To predict event category scores

[0111]

[0112] Among them, W c This is the parameter matrix to be learned.

[0113] During the inference phase, the final prediction is made by and Confirmed. If Then the t-th segment is predicted to be event-related, where the event category is based on... if Then the t-th segment is predicted as background.

[0114] During training, the system and / or method may have fragment-level labels, including event-related labels and event-category labels. The overall objective function is the sum of the cross-entropy loss for event classification and the binary cross-entropy loss for event-related prediction.

[0115] Weak supervision positioning

[0116] In a weakly supervised approach, the method and / or system can also make predictions as described above. and In one respect, since the method and / or system can only access video-level tags, the method and / or system can... Repeated T times, and for Repeat C times, and then fuse them via element-wise multiplication to produce a joint fraction. In an embodiment, the method and / or system may formulate the problem as a multi-instance learning (MIL) problem and aggregate fragment-level predictions. Video-level predictions are obtained via MIL pooling during training. During inference, in this embodiment, the prediction process can be the same as that for a supervised task.

[0117] For example, training settings may include setting the hidden dimension dm in the relation-aware module to 256. For CMRA and self-attention in the relation-aware module, the system and / or method may set the number of parallel heads to 4. The batch size is 32. As an example, the method and / or system may apply Adam as an optimizer to iteratively update the weights of the neural network based on the training data. As an example, the method and / or system may set the initial learning to 5 × 10⁻⁶. -4 It is gradually decayed by multiplying by 0.5 at epochs 10, 20, and 30. Another optimizer can be used.

[0118] Figure 6 Example localization results output by the methods and / or systems in the embodiments are shown. The methods and / or systems correctly predict the event category of each segment (e.g., as background (BG) or cat screaming), and thus accurately localize the cat screaming event.

[0119] Figure 7This is a flowchart illustrating a method for audiovisual event localization according to an embodiment. In the embodiment, the bimodal relational network described herein can perform audiovisual event localization. The method can be run or performed by one or more processors (such as hardware processors), or run or performed on one or more processors. At 702, the method includes receiving a video feed for audiovisual event localization. At 704, the method includes determining informational features and regions in the video feed by running a first neural network based on a combination of extracted audio features and video features from the video feed. For example, an audio-guided visual attention module that may include the first neural network may be run.

[0120] At 706, the method includes, based on information features and regions in the video feed determined by a first neural network, determining relation-aware video features by running a second neural network. At 708, based on information features and regions in the video feed determined by the first neural network, determining relation-aware audio features by running a third neural network. For example, intra-modal and inter-modal modules can be implemented and / or run (e.g., referenced above). Figure 2 (As described in 214, 216, 218, and 220). In an embodiment, the second neural network acquires both temporal information in the video features and cross-modal information between the video features and the audio features when determining relation-aware video features. In an embodiment, the third neural network acquires both temporal information in the audio features and cross-modal information between the video features and the audio features when determining relation-aware audio features.

[0121] At 710, the method includes: obtaining a bimodal representation based on relation-aware video features and relation-aware audio features by running a fourth neural network. For example, an audio-video interaction module (e.g., as described above with reference to 226) can be implemented and / or run.

[0122] At 712, the method includes inputting a bimodal representation into a classifier to identify audiovisual events in a video feed. In an embodiment, the bimodal representation is used as the final layer of the classifier in identifying audiovisual events. The classifier for identifying audiovisual events in a video feed may include identifying the location where the audiovisual event occurs in the video feed and the category of the audiovisual event.

[0123] In one embodiment, a convolutional neural network (e.g., referred to as a first convolutional neural network for interpretation) may be run at least with the video portion of the video feed to extract video features. In another embodiment, a convolutional neural network (e.g., referred to as a second convolutional neural network for interpretation) may be run at least with the audio portion of the video feed to extract audio features.

[0124] Figure 8This diagram illustrates the components of a system in one embodiment that can implement a bimodal relational network for audiovisual event localization. One or more hardware processors 802, such as a central processing unit (CPU), graphics processing unit (GPU) and / or field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), and / or another processor, can be coupled to a storage device 804 to implement the bimodal relational network and perform audiovisual event localization. The storage device 804 may include random access memory (RAM), read-only memory (ROM), or another storage device and may store data and / or processor instructions for implementing various functions associated with the methods and / or systems described herein. One or more processors 802 can execute computer instructions stored in the storage device 804 or received from another computer device or medium. The storage device 804 may, for example, store instructions and / or data for the functions of one or more hardware processors 802 and may include an operating system and other programs containing instructions and / or data. One or more hardware processors 802 can receive input including a video feed, from which video and audio features can be extracted, for example. For example, at least one hardware processor 802 may use the methods and techniques described herein to perform audiovisual event localization. In one aspect, data such as input data and / or intermediate data can be stored in storage device 806 or received from a remote device via network interface 808, and can be temporarily loaded into memory device 804 for implementing a bimodal relational network and performing audiovisual event localization. The learned model (such as a neural network model) in the bimodal relational network can be stored on memory device 804, for example, for execution by one or more hardware processors 802. One or more hardware processors 802 can be coupled to interface devices such as network interface 808 for communicating with a remote system, for example, via a network, and input / output interface 810 for communicating with input and / or output devices such as a keyboard, mouse, display, and / or others.

[0125] Figure 9 A schematic diagram of an example computer or processing system that can implement a bimodal relational network system in one embodiment is shown. The computer system is merely one example of a suitable processing system and is not intended to impose any limitation on the scope of use or functionality of the embodiments of the methods described herein. The processing system shown can operate with many other general-purpose or special-purpose computing system environments or configurations. Figure 9Examples of well-known computing systems, environments, and / or configurations of the processing system shown may include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.

[0126] A computer system can be described within the general context of computer system executable instructions (such as program modules) executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc., that perform specific tasks or implement specific abstract data types. Computer systems can be implemented in distributed cloud computing environments, where tasks are performed by remote processing devices linked via communication networks. In distributed cloud computing environments, program modules can reside in local and remote computer system storage media, including memory storage devices.

[0127] The components of the computer system may include, but are not limited to, one or more processors or processing units 12, system memory 16, and a bus 14 that couples the various system components, including system memory 16, to the processor 12. The processor 12 may include one or more modules 30 that perform the methods described herein. Modules 30 may be programmed into an integrated circuit of the processor 12, or loaded from memory 16, storage device 18, or network 24, or a combination thereof.

[0128] Bus 14 can represent one or more of several types of bus architectures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, and processor or local buses using any of the various bus architectures. By way of example and not limitation, such architectures include Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MCA) buses, Enhanced ISA (EISA) buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses.

[0129] Computer systems may include a variety of computer system-readable media. Such media can be any available media that can be accessed by a computer system, and can include volatile and non-volatile media, removable and non-removable media.

[0130] System memory 16 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory or others. The computer system may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 18 may be provided for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard disk drive"). Although not shown, disk drives for reading from or writing to removable non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable non-volatile optical disks (such as CD-ROMs, DVD-ROMs, or other optical media) may be provided. In such a case, each may be connected to bus 14 via one or more data media interfaces.

[0131] The computer system may also communicate with one or more external devices 26 (such as a keyboard, pointing device, display 28, etc.); and / or any device that enables the computer system to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication may occur via input / output (I / O) interface 20.

[0132] Furthermore, the computer system can communicate with one or more networks 24 (such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet)) via network adapter 22. As shown, network adapter 22 communicates with other components of the computer system via bus 14. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with the computer system. Examples include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archiving storage systems.

[0133] This invention can be a system, method, and / or computer program product with any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the invention.

[0134] Computer-readable storage media can be tangible means for retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital universal disk (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or protrusions in slots having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.

[0135] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.

[0136] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Smalltalk, C++, etc.) and procedural programming languages ​​(such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuitry in order to perform aspects of this invention.

[0137] The present invention will now be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0138] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.

[0139] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce computer-implemented processing, such that the instructions executed on the computer, other programmable apparatus, or other device perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the figures. For example, two blocks shown consecutively may actually be completed as a single step, executed simultaneously, substantially simultaneously, or with partial or complete temporal overlap, or the blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0141] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, unless the context explicitly indicates otherwise, the singular forms “a,” “an,” and “the” are intended to also include the plural forms. As used herein, the term “or” is an inclusive operator and may mean “and / or” unless the context explicitly or explicitly indicates otherwise. It should also be understood that, when used herein, the terms “comprise,” “comprises,” “comprising,” “includes,” “including,” and / or “having” may specify the presence of the stated feature, integral, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or combinations thereof. As used herein, the phrase “in an embodiment” does not necessarily refer to the same embodiment, although it may refer to the same embodiment. As used herein, the phrase “in one embodiment” does not necessarily refer to the same embodiment, although it may refer to the same embodiment. As used herein, the phrase “in another embodiment” does not necessarily refer to different embodiments, although it may refer to different embodiments. Furthermore, the embodiments and / or the components of the embodiments can be freely combined with each other, unless they are mutually exclusive.

[0142] All the means or steps plus functional elements (if any) in the following claims are intended to include any structure, material, action, and equivalent for performing the function in combination with other claimed elements as specifically claimed. The description of the invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the invention. Embodiments were chosen and described in order to best explain the principles and practical application of the invention, and to enable others skilled in the art to understand various embodiments of the invention with various modifications suitable for the intended particular use.

Claims

1. A system for implementing a bimodal relational network for audiovisual event localization, comprising: Hardware processor; A memory coupled to the hardware processor; The hardware processor is configured to: Receive video feeds for locating audiovisual events; Based on the combination of extracted audio and video features from the video feed, the information features and regions in the video feed are determined by running a first neural network; Based on the information features and regions in the video feed determined by the first neural network, relation-aware video features are determined by running a second neural network, which is configured to implement a cross-modal relational attention mechanism and learn the relation-aware video features using at least one query derived from video features and key-value pairs derived from both video features and audio features associated with the video feed. Based on the information features and regions in the video feed determined by the first neural network, relation-aware audio features are determined by running a third neural network, which is configured to implement a cross-modal relational attention mechanism and learn the relation-aware audio features using at least one query derived from the audio features and key-value pairs derived from both the video features and the audio features associated with the video feed. By running a fourth neural network, a bimodal representation is obtained based on the relation-aware video features and the relation-aware audio features; The bimodal representation is input into a classifier to identify audiovisual events in the video feed; In cross-modal relational attention mechanisms, at least one query is used in the attention mechanism. It is derived from a modality, while the key used in the attention mechanism is... Sum It is derived from two modalities in cross-modal relational attention. calculate With all keys The dot product is calculated by dividing each dot product by the square root of the shared feature dimension dm, and then applying the softmax function to obtain the value. Attention weights, where, by representation from and All values ​​of attention weighted in the learning relationship The sum is used to calculate the output of interest. In this process, individual fragments from one modality simultaneously aggregate useful information from all relevant fragments in both modalities.

2. The system according to claim 1, wherein, The hardware processor is also configured to run a first convolutional neural network with at least a portion of the video feed to extract the video features.

3. The system according to claim 1, wherein, The hardware processor is also configured to run a second convolutional neural network with at least the audio portion of the video feed to extract the audio features.

4. The system according to claim 1, wherein, The bimodal representation is used as the last layer of the classifier in recognizing the audiovisual events.

5. The system according to claim 1, wherein, The classifier identifies audiovisual events in the video feed by identifying the location where the audiovisual event occurs in the video feed and the category of the audiovisual event.

6. The system according to claim 1, wherein, When determining the relationship-aware video features, the second neural network acquires both the temporal information in the video features and the cross-modal information between the video features and the audio features.

7. The system according to claim 1, wherein, The third neural network acquires both the temporal information in the audio features and the cross-modal information between the video features and the audio features when determining the relation-aware audio features.

8. A computer-based method for implementing a bimodal relational network for audiovisual event localization, comprising: Receive video feeds for locating audiovisual events; Based on the combination of extracted audio and video features from the video feed, the information features and regions in the video feed are determined by running a first neural network; Based on the information features and regions in the video feed determined by the first neural network, relation-aware video features are determined by running a second neural network, which is configured to implement a cross-modal relational attention mechanism and learn the relation-aware video features using at least one query derived from video features and key-value pairs derived from both video features and audio features associated with the video feed. Based on the information features and regions in the video feed determined by the first neural network, relation-aware audio features are determined by running a third neural network, which is configured to implement a cross-modal relational attention mechanism and learn the relation-aware audio features using at least one query derived from the audio features and key-value pairs derived from both the video features and the audio features associated with the video feed. By running a fourth neural network, a bimodal representation is obtained based on the relation-aware video features and the relation-aware audio features; The bimodal representation is input into a classifier to identify audiovisual events in the video feed; In cross-modal relational attention mechanisms, at least one query is used in the attention mechanism. It is derived from a modality, while the key used in the attention mechanism is... Sum It is derived from two modalities in cross-modal relational attention. calculate With all keys The dot product is calculated by dividing each dot product by the square root of the shared feature dimension dm, and then applying the softmax function to obtain the value. Attention weights, where, by representation from and All values ​​of attention weighted in the learning relationship The sum is used to calculate the output of interest. In this process, individual fragments from one modality simultaneously aggregate useful information from all relevant fragments in both modalities.

9. The method of claim 8, further comprising running a first convolutional neural network at least with a portion of the video feed to extract the video features.

10. The method of claim 8, further comprising running a second convolutional neural network at least with the audio portion of the video feed to extract the audio features.

11. The method according to claim 8, wherein, The bimodal representation is used as the last layer of the classifier in recognizing the audiovisual events.

12. The method according to claim 8, wherein, The classifier identifies audiovisual events in the video feed by identifying the location where the audiovisual event occurs in the video feed and the category of the audiovisual event.

13. The method according to claim 8, wherein, When determining the relationship-aware video features, the second neural network acquires both the temporal information in the video features and the cross-modal information between the video features and the audio features.

14. The method according to claim 8, wherein, The third neural network acquires both the temporal information in the audio features and the cross-modal information between the video features and the audio features when determining the relation-aware audio features.

15. A computer program product comprising a computer-readable storage medium having program instructions embodied therein, the program instructions being readable / executable by a device to cause the device to: Receive video feeds for locating audiovisual events; Based on the combination of extracted audio and video features from the video feed, the information features and regions in the video feed are determined by running a first neural network; Based on the information features and regions in the video feed determined by the first neural network, relation-aware video features are determined by running a second neural network, which is configured to implement a cross-modal relational attention mechanism and learn the relation-aware video features using at least one query derived from video features and key-value pairs derived from both video features and audio features associated with the video feed. Based on the information features and regions in the video feed determined by the first neural network, relation-aware audio features are determined by running a third neural network, which is configured to implement a cross-modal relational attention mechanism and learn the relation-aware audio features using at least one query derived from the audio features and key-value pairs derived from both the video features and the audio features associated with the video feed. By running a fourth neural network, a bimodal representation is obtained based on the relation-aware video features and the relation-aware audio features; as well as The bimodal representation is input into a classifier to identify audiovisual events in the video feed; In cross-modal relational attention mechanisms, at least one query is used in the attention mechanism. It is derived from a modality, while the key used in the attention mechanism is... Sum It is derived from two modalities in cross-modal relational attention. calculate With all keys The dot product is calculated by dividing each dot product by the square root of the shared feature dimension dm, and then applying the softmax function to obtain the value. Attention weights, where, by representation from and All values ​​of attention weighted in the learning relationship The sum is used to calculate the output of interest. In this process, individual fragments from one modality simultaneously aggregate useful information from all relevant fragments in both modalities.

16. The computer program product according to claim 15, wherein, The device is also configured to run a first convolutional neural network on at least a portion of the video fed by the video to extract the video features.

17. The computer program product according to claim 15, wherein, The device is also configured to run a second convolutional neural network on at least the audio portion fed by the video to extract the audio features.

18. The computer program product according to claim 15, wherein, The bimodal representation is used as the last layer of the classifier in recognizing the audiovisual events.

19. The computer program product according to claim 15, wherein, The classifier identifies audiovisual events in the video feed by identifying the location where the audiovisual event occurs in the video feed and the category of the audiovisual event.

20. The computer program product according to claim 15, wherein, The second neural network acquires both the time information in the video features and the cross-modal information between the video features and the audio features when determining the relationship-aware video features, and the third neural network acquires both the time information in the audio features and the cross-modal information between the video features and the audio features when determining the relationship-aware audio features.