Audio-visual event identification system, method, and program
Patent Information
- Application Number
- JP2023507362
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-08-10
- Filing Date
- 2021-07-05
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2041-07-05
Smart Images

Figure 0007679142000075 
Figure 0007679142000076 
Figure 0007679142000077
Abstract
Claims
1. a hardware processor; a memory coupled to the hardware processor; Equipped with The hardware processor includes: receiving a video feed for audio-visual event location; determining useful features and regions within the video feed by operating a first neural network based on a combination of extracted audio and video features of the video feed; determining relationship-aware video features by operating a second neural network based on the informative features and regions in the video feed determined by the first neural network, the second neural network implementing a cross-modality relationship attention mechanism and configured to learn the relationship-aware video features using at least one query derived from video features and key-value pairs derived from both video and audio features associated with the video feed; determining relationship-aware audio features by operating a third neural network based on the informative features and regions in the video feed determined by the first neural network, the third neural network implementing a cross-modality relationship attention mechanism and configured to learn the relationship-aware audio features using at least one query derived from audio features and the key-value pairs derived from both video and audio features associated with the video feed; obtaining a dual-modality representation based on the relationship-aware video features and the relationship-aware audio features by operating a fourth neural network; inputting the dual-modality representation into a classifier to identify audio-visual events within the video feed; wherein in the cross-modality relational attention mechanism, the at least one query q 1 used in the attention mechanism is derived from one modality, and the key-value pair K 1,2 and V 1,2 are derived from two modalities; The dot products of q1 with all keys K1,2 are computed, each of the computed dot products is divided by the square root of the shared feature dimension dm, and a softmax function is applied to obtain attention weights for the values V1,2, and the attentioned output is computed by summing over all values V1,2 weighted by the attention weights that represent the relationship learned from q1 and K1,2; wherein each individual segment from one modality simultaneously aggregates useful information from all related segments from two modalities; A system configured to run
2. 2. The system of claim 1, wherein the hardware processor is further configured to operate a first convolutional neural network with at least a video portion of the video feed to extract the video features.
3. 2. The system of claim 1, wherein the hardware processor is further configured to operate a second convolutional neural network with at least an audio portion of the video feed to extract the audio features.
4. The system of claim 1 , wherein the dual-modality representation is used as a final layer of the classifier in identifying the audio-visual event.
5. 2. The system of claim 1, wherein the classifier identifying the audio-visual event in the video feed includes identifying a location in the video feed where the audio-visual event occurs and a category of the audio-visual event.
6. 2. The system of claim 1, wherein the second neural network captures both temporal information in the video features and cross-modality information between the video features and the audio features in determining the relationship-aware video features.
7. 2. The system of claim 1, wherein the third neural network obtains both temporal information in the audio features and cross-modality information between the video features and the audio features in determining the relationship-aware audio features.
8. A method for computer-based information processing, comprising the steps of: receiving a video feed for audio-visual event location; determining useful features and regions within the video feed by operating a first neural network based on a combination of extracted audio and video features of the video feed; determining relationship-aware video features by operating a second neural network based on the informative features and regions in the video feed determined by the first neural network, the second neural network implementing a cross-modality relationship attention mechanism and configured to learn the relationship-aware video features using at least one query derived from video features and key-value pairs derived from both video and audio features associated with the video feed; determining relationship-aware audio features by operating a third neural network based on the informative features and regions in the video feed determined by the first neural network, the third neural network implementing a cross-modality relationship attention mechanism and configured to learn the relationship-aware audio features using at least one query derived from audio features and the key-value pairs derived from both video and audio features associated with the video feed; obtaining a dual-modality representation based on the relationship-aware video features and the relationship-aware audio features by operating a fourth neural network; inputting the dual-modality representation into a classifier to identify audio-visual events within the video feed; wherein in the cross-modality relational attention mechanism, the at least one query q 1 used in the attention mechanism is derived from one modality, and the key-value pair K 1,2 and V 1,2 are derived from two modalities; The dot products of q1 with all keys K1,2 are computed, each of the computed dot products is divided by the square root of the shared feature dimension dm, and a softmax function is applied to obtain attention weights for the values V1,2, and the attentioned output is computed by summing over all values V1,2 weighted by the attention weights that represent the relationship learned from q1 and K1,2; wherein each individual segment from one modality simultaneously aggregates useful information from all related segments from two modalities; A method comprising:
9. 10. The method of claim 8, further comprising operating a first convolutional neural network with at least a video portion of the video feed to extract the video features.
10. 10. The method of claim 8, further comprising operating a second convolutional neural network with at least an audio portion of the video feed to extract the audio features.
11. The method of claim 8 , wherein the dual-modality representation is used as a final layer of the classifier in identifying the audio-visual event.
12. 9. The method of claim 8, wherein the classifier identifying the audio-visual event in the video feed includes identifying a location in the video feed where the audio-visual event occurs and a category of the audio-visual event.
13. 10. The method of claim 8, wherein the second neural network captures both temporal information in the video features and cross-modality information between the video features and the audio features in determining the relationship-aware video features.
14. 10. The method of claim 8, wherein the third neural network captures both temporal information in the audio features and cross-modality information between the video features and the audio features in determining the relationship-aware audio features.
15. A computer program causing a computer to execute the method according to any one of claims 8 to 14.
16. A storage medium having the computer program according to claim 15 stored therein, the computer readable storage medium.
Citation Information
Patent Citations
Multimodal Data Fusion Using Recurrent Neural Networks
JP2023501469A
Automatic Video Event Detection and Indexing
US20080193016A1
System and Method for Deriving Timeline Metadata for Video Content
US20150254341A1