Dark Video Activity Recognition via Audio-Visual Modulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video processing techniques struggle to accurately identify activities in videos captured under poor lighting conditions, particularly in dark environments, as they rely heavily on image enhancement or infrared sensors, which are sensitive to environmental factors and limited in low-light scenarios.

Innovation Solution

A method that combines video and audio features using a darkness-aware evaluation model to modulate and predict activities, incorporating cross-modal attention mechanisms to adjust channel attentions and classification boundaries based on darkness-aware features, effectively leveraging both visual and auditory cues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Illumination intensity

If image enhancement techniques are used to process dark video, then visual information can be recovered to some extent, but the technique fails when the environment is too dark

Engineering Contradiction:
Improvevideo brightnessVSAvoidactivity recognition accuracy
Core Design Contradiction:
Illumination intensityVSReliability

Solution Approach 1:

The patent introduces audio signals as an intermediary modality to bridge the gap when visual information becomes insufficient in dark environments. The audio-visual model uses audio features as a complementary source of information that remains reliable regardless of lighting conditions, mediating the activity recognition process when vision alone fails

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a composite audio-visual feature representation that combines visual features and audio features into a unified model. This composite approach leverages the strengths of both modalities, allowing the system to maintain reliability across varying lighting conditions by integrating multiple information sources

Inventive Principle:
Principle #40Composite materials

2Illumination intensity

If infrared sensors are used to capture dark environments, then activity information can be obtained, but the sensors become sensitive to environmental factors such as temperature, smoke, dust and haze

Engineering Contradiction:
Improvedark environment detectionVSAvoidenvironmental sensitivity
Core Design Contradiction:
Illumination intensityVSObject-affected harmful factors

Solution Approach 1:

The patent uses audio signals as an alternative intermediary that is not affected by environmental factors like temperature, smoke, dust, and haze. Instead of relying on infrared sensors that are sensitive to these conditions, the model incorporates audio as a robust complementary modality that maintains reliability in challenging environments

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the operational parameters by shifting from purely visual or infrared sensing to a multi-modal approach that includes audio frequency analysis. This parameter change allows the system to operate effectively in dark environments without being constrained by the environmental sensitivities of infrared sensors

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If only visual content is processed for activity recognition, then the system is simpler to implement, but accuracy deteriorates in dark environments

Engineering Contradiction:
Improveprocessing systemVSAvoidactivity detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent creates a composite audio-visual processing system that integrates both modalities into a unified model. This composite approach maintains reasonable system complexity while significantly improving activity detection accuracy in dark environments by leveraging the complementary information from audio signals

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The patent designs a multi-functional model that can process both visual and audio inputs through a shared architecture. This universal approach allows the system to adapt to different lighting conditions by switching between or combining modalities, improving accuracy without proportionally increasing complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If audio signals are used to aid crowd counting in challenging vision conditions, then prediction can be maintained when image quality is poor, but audio signals alone cannot make predictions when image quality is too poor to rely on

Engineering Contradiction:
Improveprediction reliabilityVSAvoidmodality dependency
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a composite audio-visual model that dynamically adapts to input quality. When visual input is sufficient, the model relies on visual features; when visual input deteriorates, the model increasingly weights audio features. This composite approach maintains prediction reliability across the full range of image quality conditions

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The patent implements a dynamic fusion mechanism where the relative contribution of audio and visual modalities changes based on input quality. The model adaptively adjusts the weighting and integration of audio and visual features in real-time, allowing it to maintain reliability across varying conditions without being locked into a fixed modality dependency

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11960576B2Activity recognition in dark video based on both audio and video content
Publication Date: 2024.04.16 INCEPTION AI IP LTD
  • US11960576B2 patent drawing
  • US11960576B2 patent drawing
  • US11960576B2 patent drawing

AI summary

Videos captured in low light conditions can be processed in order to identify an activity being performed in the video. The processing may use both the video and audio streams for identifying the activity in the low light video. The video portion is processed to generate a darkness-aware feature which may be used to modulate the features generated from the audio and video features. The audio features may be used to generate a video attention feature and the video features may be used to generate an audio attention feature. The audio and video attention features may also be used in modulating the audio video features. The modulated audio and video features may be used to predict an activity occurring in the video.