Dual-Camera Activity Recognition for Real-Time AV Navigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current autonomous vehicle (AV) activity recognition systems face challenges such as intraclass variation, interclass similarity, complex backgrounds, multi-subject interactions, and low-quality video feeds, particularly in low light conditions and high motion blur, leading to computational complexity and resource-intensive operations, making real-time processing difficult.
Innovation Solution
The system employs a combination of neuromorphic event-based cameras and frame-based RGB video cameras, using a two-stream neural network model for adaptive sampling and data fusion to recognize activities, where the neuromorphic camera provides high-speed temporal information and the RGB camera provides spatio-temporal data, enabling efficient activity recognition and navigation control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional frame-based RGB video cameras are used for activity recognition, then scene-level contextual information is captured, but computational complexity increases and real-time processing becomes difficult
Solution Approach 1:
The system segments the visual information processing by using two separate camera streams: neuromorphic cameras capture temporal changes (motion events) while frame-based cameras capture spatial context. This segmentation allows each stream to be processed independently and efficiently, reducing overall computational complexity while preserving both motion and contextual information.
Solution Approach 2:
The system extracts only the essential features from each camera type: temporal motion information is extracted from neuromorphic event streams, while spatial contextual information is extracted from frame-based RGB images. This selective extraction reduces data processing requirements compared to analyzing complete video frames for all information.
2Productivity
If conventional frame-based cameras with fixed shutter speed are used, then data is captured at regular intervals, but motion blur increases and salient data is lost between intervals
Solution Approach 1:
The system replaces fixed shutter speed with dynamic event-triggered sampling. Neuromorphic cameras continuously monitor for changes and capture events only when motion occurs, adapting the sampling rate to the actual dynamics of the scene. This dynamic approach eliminates motion blur by using extremely short exposure times while capturing all salient motion events.
Solution Approach 2:
The system uses feedback from neuromorphic camera events to control the sampling of frame-based cameras. When motion events are detected, the frame-based camera is triggered to capture images at appropriate moments, ensuring salient data is captured without redundant fixed-interval sampling.
3Speed
If neuromorphic event-based cameras are used alone, then high-speed temporal information is captured, but scene-level contextual information is insufficient
Solution Approach 1:
The system merges data from two complementary camera types: neuromorphic event-based cameras provide high-speed temporal information about motion changes, while frame-based RGB cameras provide spatial and contextual information. The fusion of these two streams in the neural network achieves both high temporal resolution and comprehensive scene understanding.
Solution Approach 2:
The system adds a temporal dimension to the visual processing by incorporating event-based temporal information alongside spatial information from frame-based cameras. This multi-dimensional approach (spatial + temporal) enables the system to process both context and motion simultaneously with enhanced accuracy.
4Ease of manufacture
If fixed interval shutter speed sampling is used, then data acquisition is simple, but redundant information increases and processing efficiency decreases
Solution Approach 1:
The neuromorphic camera system operates autonomously by self-determining when to capture data based on detected changes in the scene. Instead of relying on external fixed-interval triggering, the camera adapts its sampling rate to scene dynamics, capturing data only when necessary and reducing redundant information generation.
Data Source
AI summary
A method and activity recognition system for recognising activities in surrounding environment for controlling navigation of an autonomous vehicle is disclosed. The activity recognition system receives first data feed from neuromorphic event-based camera and second data feed from frame-based RGB video camera. The first data feed comprises high-speed temporal information encoding motion associated with change in surrounding environment at each spatial location, and second data feed comprises spatio-temporal data providing scene-level contextual information associated with surrounding environment. An adaptive sampling of second data feed is performed with respect to foreground activity rate based on amount of foreground motion encoded in first data feed. Further, the activity recognition system recognizes activities associated with at least one object in surrounding environment by identifying correlation between both data feed by using two-stream neural network model. Thereafter, based on the determined activities, the activity recognition system controls the navigation of the autonomous vehicle.


