Dual-Camera Activity Recognition for Real-Time AV Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current autonomous vehicle (AV) activity recognition systems face challenges such as intraclass variation, interclass similarity, complex backgrounds, multi-subject interactions, and low-quality video feeds, particularly in low light conditions and high motion blur, leading to computational complexity and resource-intensive operations, making real-time processing difficult.

Innovation Solution

The system employs a combination of neuromorphic event-based cameras and frame-based RGB video cameras, using a two-stream neural network model for adaptive sampling and data fusion to recognize activities, where the neuromorphic camera provides high-speed temporal information and the RGB camera provides spatio-temporal data, enabling efficient activity recognition and navigation control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional frame-based RGB video cameras are used for activity recognition, then scene-level contextual information is captured, but computational complexity increases and real-time processing becomes difficult

Engineering Contradiction:
Improvescene-level contextual informationVSAvoidcomputational complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments the visual information processing by using two separate camera streams: neuromorphic cameras capture temporal changes (motion events) while frame-based cameras capture spatial context. This segmentation allows each stream to be processed independently and efficiently, reducing overall computational complexity while preserving both motion and contextual information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts only the essential features from each camera type: temporal motion information is extracted from neuromorphic event streams, while spatial contextual information is extracted from frame-based RGB images. This selective extraction reduces data processing requirements compared to analyzing complete video frames for all information.

Inventive Principle:
Principle #2Taking out (Extraction)

2Productivity

If conventional frame-based cameras with fixed shutter speed are used, then data is captured at regular intervals, but motion blur increases and salient data is lost between intervals

Engineering Contradiction:
Improvedata capture efficiencyVSAvoidmotion capture accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system replaces fixed shutter speed with dynamic event-triggered sampling. Neuromorphic cameras continuously monitor for changes and capture events only when motion occurs, adapting the sampling rate to the actual dynamics of the scene. This dynamic approach eliminates motion blur by using extremely short exposure times while capturing all salient motion events.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from neuromorphic camera events to control the sampling of frame-based cameras. When motion events are detected, the frame-based camera is triggered to capture images at appropriate moments, ensuring salient data is captured without redundant fixed-interval sampling.

Inventive Principle:
Principle #23Feedback

3Speed

If neuromorphic event-based cameras are used alone, then high-speed temporal information is captured, but scene-level contextual information is insufficient

Engineering Contradiction:
Improvetemporal information speedVSAvoidscene-level contextual information
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The system merges data from two complementary camera types: neuromorphic event-based cameras provide high-speed temporal information about motion changes, while frame-based RGB cameras provide spatial and contextual information. The fusion of these two streams in the neural network achieves both high temporal resolution and comprehensive scene understanding.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system adds a temporal dimension to the visual processing by incorporating event-based temporal information alongside spatial information from frame-based cameras. This multi-dimensional approach (spatial + temporal) enables the system to process both context and motion simultaneously with enhanced accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Ease of manufacture

If fixed interval shutter speed sampling is used, then data acquisition is simple, but redundant information increases and processing efficiency decreases

Engineering Contradiction:
Improvedata acquisition simplicityVSAvoidprocessing efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The neuromorphic camera system operates autonomously by self-determining when to capture data based on detected changes in the scene. Instead of relying on external fixed-interval triggering, the camera adapts its sampling rate to scene dynamics, capturing data only when necessary and reducing redundant information generation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11987264B2Method and system for recognizing activities in surrounding environment for controlling navigation of autonomous vehicle
Publication Date: 2024.05.21 WIPRO LTD
  • US11987264B2 patent drawing
  • US11987264B2 patent drawing
  • US11987264B2 patent drawing

AI summary

A method and activity recognition system for recognising activities in surrounding environment for controlling navigation of an autonomous vehicle is disclosed. The activity recognition system receives first data feed from neuromorphic event-based camera and second data feed from frame-based RGB video camera. The first data feed comprises high-speed temporal information encoding motion associated with change in surrounding environment at each spatial location, and second data feed comprises spatio-temporal data providing scene-level contextual information associated with surrounding environment. An adaptive sampling of second data feed is performed with respect to foreground activity rate based on amount of foreground motion encoded in first data feed. Further, the activity recognition system recognizes activities associated with at least one object in surrounding environment by identifying correlation between both data feed by using two-stream neural network model. Thereafter, based on the determined activities, the activity recognition system controls the navigation of the autonomous vehicle.