Cross-Camera Object Tracking Using Biometric Identifiers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video surveillance systems face challenges in efficiently tracking objects across multiple cameras due to cumbersome setup processes and high processing and memory requirements, particularly when dealing with a large number of cameras.
Innovation Solution
A method and apparatus for tracking objects of interest (OI) in a video surveillance system that involves generating a unique biometric identifier for an object based on its image, monitoring multiple video streams for matching identifiers, and automatically switching the primary viewing stream to the stream containing the object, using machine learning algorithms for object detection and tracking.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional tracking systems are used to track objects across multiple cameras, then object tracking capability is achieved, but setup complexity and processing requirements increase significantly
Solution Approach 1:
The patent extracts and stores only the bounding box coordinates and visual features of detected objects in a database, rather than storing complete video streams or complex tracking data. This selective extraction reduces the amount of data that needs to be processed and coordinated across multiple cameras, thereby reducing setup complexity while maintaining object tracking capability.
Solution Approach 2:
The system segments the tracking task by processing each camera's video stream independently to detect and extract object features, then consolidating results at a central processing unit. This segmentation allows each camera to operate autonomously with simpler processing requirements, while the overall system achieves coordinated multi-camera tracking without requiring complex pre-configuration of camera relationships.
2Reliability
If conventional tracking systems process multiple video streams from many cameras, then comprehensive object detection is achieved, but processing power and memory requirements increase
Solution Approach 1:
The patent extracts only the essential object information (bounding boxes, visual features, and location data) from each video stream rather than processing and storing entire video streams. This selective extraction dramatically reduces the data volume that requires processing and memory storage, enabling comprehensive object detection across multiple cameras with reduced computational resource requirements.
Solution Approach 2:
The system performs partial processing by analyzing only the portions of video streams that contain objects of interest, rather than processing every pixel and frame completely. By focusing computational resources on object detection and feature extraction rather than complete video analysis, the system achieves reliable object detection while reducing overall processing power consumption.
3Area of stationary object
If cameras are positioned to cover all areas with minimal overlap, then surveillance coverage is optimized, but tracking coordination becomes more difficult
Solution Approach 1:
The patent segments the tracking function from the camera positioning requirements by allowing cameras to be positioned for optimal surveillance coverage independently of tracking coordination needs. Each camera processes its own video stream to extract object data, eliminating the need for complex inter-camera coordination while maintaining comprehensive coverage. The segmentation of detection and tracking functions resolves the conflict between coverage optimization and coordination complexity.
Data Source
AI summary
Example implementations include a method, apparatus and computer-readable medium in a video surveillance system for tracking an of-interest (OI) object, comprising receiving, from a user interface, a request to track an object in a first video stream from a plurality of video streams captured from a first camera of a plurality of cameras installed in an environment. The implementations further include extracting at least one image of the object from the first video stream. Additionally, the implementations further include generating a unique biometric identifier of the object based on the at least one image. Additionally, the implementations further include detecting, using the unique biometric identifier, the object in a second video stream captured from a second camera of the plurality of cameras, and outputting, on the user interface, the second video stream in response to detecting the object in the second video stream.


