Video-Audio Activity Detection for Poor-Visibility Areas

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video surveillance systems struggle to identify activity in areas with poor visibility due to conditions like low lighting, obstructions, or adverse weather, making it difficult to visually detect human activity that may indicate potential problems.

Innovation Solution

A system utilizing a video camera and an audio sensor to capture and process video to identify visible events of interest, determine sound profiles, and compare them with audio streams to filter out irrelevant sounds, allowing for the detection of human activity through audio analysis even when video capture is impaired.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If video surveillance is used to monitor activity in an area, then visual identification of events is improved, but under conditions of low lighting, heavy rain, or obstructions, the ability to identify activity deteriorates

Engineering Contradiction:
Improvevisual identification accuracyVSAvoidactivity detection reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system merges video surveillance with audio surveillance by combining visual event detection with audio event detection and sound profile analysis. When video identification fails due to adverse conditions, the system switches to or supplements with audio-based activity identification, creating a multi-modal surveillance system that maintains reliability across varying environmental conditions.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces an intermediary decision-making layer that determines whether to rely on video or audio based on current conditions. The processor analyzes video legibility and selectively activates audio analysis when visual identification is compromised, acting as an intermediary that bridges the gap between visual and auditory detection modalities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If audio analysis is used to identify activity when video is unavailable, then activity detection capability is improved, but system complexity increases

Engineering Contradiction:
Improvesurveillance capability under adverse conditionsVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-establishing sound profiles for various events of interest and pre-determining the conditions under which audio analysis should be activated. During operation, when video legibility is compromised, the system can immediately switch to audio analysis using pre-configured sound profiles, avoiding the need for complex real-time decision-making and reducing operational complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts and separates audio analysis as an independent module that can operate autonomously when video surveillance is ineffective. By extracting the audio processing chain (audio sensor, sound profile matching, event identification) as a distinct functional unit, the system manages complexity through modular design, allowing each modality to be developed, optimized, and maintained independently.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12437539B2System and method for identifying activity in an area using a video camera and an audio sensor
Publication Date: 2025.10.07 HONEYWELL INTERNATIONAL INC
  • US12437539B2 patent drawing
  • US12437539B2 patent drawing
  • US12437539B2 patent drawing

AI summary

Identifying activity in an area even during periods of poor visibility using a video camera and an audio sensor are disclosed. The video camera is used to identify visible events of interest and the audio sensor is used to capture audio occurring temporally with the identified visible events of interest. A sound profile is determined for each of the identified visible events of interest based on sounds captured by the audio sensor during the corresponding identified visible event of interest. Then, during a time of poor visibility, a subsequent sound event is identified in a subsequent audio stream captured by the audio sensor. One or more sound characteristics of the subsequent sound event are compared with the sound profiles associated with each of the identified visible events of interest, and if there is a match, one or more matching sound profiles are filtered out from the subsequent audio stream.