Audio-visual event identification system, method, and program

JP7679142B2Active Publication Date: 2025-05-19INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023507362
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-10
Filing Date
2021-07-05
Publication Date
2025-05-19
Estimated Expiration
2041-07-05

Smart Images

  • Figure 0007679142000075
    Figure 0007679142000075
  • Figure 0007679142000076
    Figure 0007679142000076
  • Figure 0007679142000077
    Figure 0007679142000077
Patent Text Reader

Abstract

A system and method for identifying audio-visual events includes receiving a video feed for audio-visual event localization, determining informative features and regions within the video feed by operating a first neural network based on a combination of extracted audio and video features of the video feed, determining relationship-aware video features by operating a second neural network based on the informative features and regions within the video feed determined by the first neural network, determining relationship-aware audio features by operating a third neural network based on the informative features and regions within the video feed, and operating a fourth neural network to obtain dual-modality representations based on the relationship-aware video and audio features, and inputting the dual-modality representations into a classifier to identify audio-visual events within the video feed.
Need to check novelty before this filing date? Find Prior Art

Claims

1. a hardware processor; a memory coupled to the hardware processor; Equipped with The hardware processor includes: receiving a video feed for audio-visual event location; determining useful features and regions within the video feed by operating a first neural network based on a combination of extracted audio and video features of the video feed; determining relationship-aware video features by operating a second neural network based on the informative features and regions in the video feed determined by the first neural network, the second neural network implementing a cross-modality relationship attention mechanism and configured to learn the relationship-aware video features using at least one query derived from video features and key-value pairs derived from both video and audio features associated with the video feed; determining relationship-aware audio features by operating a third neural network based on the informative features and regions in the video feed determined by the first neural network, the third neural network implementing a cross-modality relationship attention mechanism and configured to learn the relationship-aware audio features using at least one query derived from audio features and the key-value pairs derived from both video and audio features associated with the video feed; obtaining a dual-modality representation based on the relationship-aware video features and the relationship-aware audio features by operating a fourth neural network; inputting the dual-modality representation into a classifier to identify audio-visual events within the video feed; wherein in the cross-modality relational attention mechanism, the at least one query q 1 used in the attention mechanism is derived from one modality, and the key-value pair K 1,2 and V 1,2 are derived from two modalities; The dot products of q1 with all keys K1,2 are computed, each of the computed dot products is divided by the square root of the shared feature dimension dm, and a softmax function is applied to obtain attention weights for the values ​​V1,2, and the attentioned output is computed by summing over all values ​​V1,2 weighted by the attention weights that represent the relationship learned from q1 and K1,2; wherein each individual segment from one modality simultaneously aggregates useful information from all related segments from two modalities; A system configured to run

2. 2. The system of claim 1, wherein the hardware processor is further configured to operate a first convolutional neural network with at least a video portion of the video feed to extract the video features.

3. 2. The system of claim 1, wherein the hardware processor is further configured to operate a second convolutional neural network with at least an audio portion of the video feed to extract the audio features.

4. The system of claim 1 , wherein the dual-modality representation is used as a final layer of the classifier in identifying the audio-visual event.

5. 2. The system of claim 1, wherein the classifier identifying the audio-visual event in the video feed includes identifying a location in the video feed where the audio-visual event occurs and a category of the audio-visual event.

6. 2. The system of claim 1, wherein the second neural network captures both temporal information in the video features and cross-modality information between the video features and the audio features in determining the relationship-aware video features.

7. 2. The system of claim 1, wherein the third neural network obtains both temporal information in the audio features and cross-modality information between the video features and the audio features in determining the relationship-aware audio features.

8. A method for computer-based information processing, comprising the steps of: receiving a video feed for audio-visual event location; determining useful features and regions within the video feed by operating a first neural network based on a combination of extracted audio and video features of the video feed; determining relationship-aware video features by operating a second neural network based on the informative features and regions in the video feed determined by the first neural network, the second neural network implementing a cross-modality relationship attention mechanism and configured to learn the relationship-aware video features using at least one query derived from video features and key-value pairs derived from both video and audio features associated with the video feed; determining relationship-aware audio features by operating a third neural network based on the informative features and regions in the video feed determined by the first neural network, the third neural network implementing a cross-modality relationship attention mechanism and configured to learn the relationship-aware audio features using at least one query derived from audio features and the key-value pairs derived from both video and audio features associated with the video feed; obtaining a dual-modality representation based on the relationship-aware video features and the relationship-aware audio features by operating a fourth neural network; inputting the dual-modality representation into a classifier to identify audio-visual events within the video feed; wherein in the cross-modality relational attention mechanism, the at least one query q 1 used in the attention mechanism is derived from one modality, and the key-value pair K 1,2 and V 1,2 are derived from two modalities; The dot products of q1 with all keys K1,2 are computed, each of the computed dot products is divided by the square root of the shared feature dimension dm, and a softmax function is applied to obtain attention weights for the values ​​V1,2, and the attentioned output is computed by summing over all values ​​V1,2 weighted by the attention weights that represent the relationship learned from q1 and K1,2; wherein each individual segment from one modality simultaneously aggregates useful information from all related segments from two modalities; A method comprising:

9. 10. The method of claim 8, further comprising operating a first convolutional neural network with at least a video portion of the video feed to extract the video features.

10. 10. The method of claim 8, further comprising operating a second convolutional neural network with at least an audio portion of the video feed to extract the audio features.

11. The method of claim 8 , wherein the dual-modality representation is used as a final layer of the classifier in identifying the audio-visual event.

12. 9. The method of claim 8, wherein the classifier identifying the audio-visual event in the video feed includes identifying a location in the video feed where the audio-visual event occurs and a category of the audio-visual event.

13. 10. The method of claim 8, wherein the second neural network captures both temporal information in the video features and cross-modality information between the video features and the audio features in determining the relationship-aware video features.

14. 10. The method of claim 8, wherein the third neural network captures both temporal information in the audio features and cross-modality information between the video features and the audio features in determining the relationship-aware audio features.

15. A computer program causing a computer to execute the method according to any one of claims 8 to 14.

16. A storage medium having the computer program according to claim 15 stored therein, the computer readable storage medium.

Citation Information

Patent Citations

  • Multimodal Data Fusion Using Recurrent Neural Networks

    JP2023501469A

  • Automatic Video Event Detection and Indexing

    US20080193016A1

  • System and Method for Deriving Timeline Metadata for Video Content

    US20150254341A1