Multi-Camera Crowd Counting via Ground-Plane Density Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing crowd counting methods are inadequate for large and crowded areas, as they often struggle with wide scenes, occlusions, and scale variations, limiting their accuracy and applicability.
Innovation Solution
A multi-view counting method using multiple cameras with overlapping fields of view, which extracts features from each camera image using deep neural networks, aligns and normalizes them, and fuses the information to generate a scene-level ground-plane density map, addressing issues of occlusions and scale variations through late-fusion, naïve early-fusion, and multi-view multi-scale early-fusion models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional single-camera or simple gate methods are used, then the system is simple to operate, but the measurement precision deteriorates in large and crowded areas
Solution Approach 1:
The patent divides the large target area into multiple sub-regions, each captured by a separate camera. Each camera generates an independent density map for its field of view, and these maps are then fused to create a comprehensive scene-level density map. This segmentation approach maintains measurement precision across large areas while managing system complexity through modular processing.
Solution Approach 2:
The patent transforms 2D image data from multiple cameras into a 3D ground-plane representation by projecting detected objects and their density maps onto a unified ground plane coordinate system. This dimensional transformation enables accurate fusion of multi-view data while accounting for perspective variations and occlusions, significantly improving counting precision in crowded areas.
2Measurement precision
If multiple cameras with overlapping fields of view are deployed, then the measurement precision improves for large areas, but the device complexity increases
Solution Approach 1:
The patent introduces a ground-plane projection as an intermediary representation that mediates between multiple camera views. Each camera's density map is projected onto the ground plane, allowing consistent fusion of overlapping fields of view. This intermediary approach simplifies the fusion process by providing a unified coordinate system, reducing the complexity of handling multi-camera data while maintaining high precision.
Solution Approach 2:
The patent merges density maps from multiple cameras by projecting them onto a common ground plane and combining the information. Objects detected by multiple cameras are consolidated into a single count on the ground plane, eliminating duplicates caused by overlapping fields of view. This merging process improves scene-level counting accuracy while using efficient computational techniques to manage system complexity.
3Measurement precision
If density maps from multiple views are fused, then the measurement precision improves, but the loss of information increases due to occlusions and scale variations
Solution Approach 1:
The patent addresses occlusion and scale variation problems by transforming 2D image data into a 3D ground-plane representation. Objects at different scales in 2D images are projected to consistent representations on the ground plane, preserving their true spatial relationships. This dimensional change recovers information that would otherwise be lost due to perspective distortion and occlusion, improving density estimation accuracy.
Solution Approach 2:
The patent applies local quality adjustments by processing each camera's density map independently before fusion, preserving local features and characteristics specific to each view. The ground-plane projection maintains local density information while integrating global context from multiple views. This approach prevents information loss by ensuring that local details are not averaged out during the fusion process.
Data Source
AI summary
A system and a method for counting objects includes the steps of: obtaining a plurality of images representing the objects to be counted in a target area; generating a map for each of the plurality of images representing an identification of each of the corresponding objects; and fusing the plurality of maps being generated to obtain a scene-level density map representing a count of the objects in the target area.


