Vehicle Occupant Counting With Edge Vision and Privacy Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer vision systems for home environments are either too simple and unpredictable or too complex and uneconomical, failing to provide accurate, scalable, and privacy-guaranteed real-time data analytics for detected people or objects.
Innovation Solution
A computer-vision system that generates a digital representation of people or objects from pixel streams, determines attributes, and controls networked devices, using an ASIC-based engine for real-time metadata processing without continuous video output.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sophisticated video analysis is performed on servers, then analysis accuracy is improved, but scalability and economic viability deteriorate due to linear scaling of storage and computational costs
Solution Approach 1:
The system segments the video analysis process into two parts: simple preprocessing is performed locally at the camera using dedicated vision processors, while sophisticated analysis is distributed across a server network. This segmentation allows the system to maintain high accuracy through server-based processing while achieving scalability by distributing computational load across multiple nodes rather than concentrating it on a single server.
Solution Approach 2:
The patent introduces dedicated vision processors as intermediary devices between the camera sensor and the server network. These intermediaries perform initial processing and filtering of video data, reducing the computational burden on servers while maintaining analysis accuracy. The vision processors act as mediators that prepare data for server analysis, enabling the system to scale efficiently.
2Loss of energy
If simple video analysis is performed inside cameras, then processing cost is reduced, but analysis accuracy and predictability deteriorate
Solution Approach 1:
The analysis task is segmented between local vision processors in cameras and remote servers. The local processors handle basic preprocessing and feature extraction at low cost, while servers perform sophisticated analysis to ensure high accuracy. This segmentation allows the system to achieve both cost efficiency and high analysis quality.
Solution Approach 2:
The system maintains continuous analysis by performing preliminary processing continuously at the camera level while simultaneously transmitting data to servers for ongoing sophisticated analysis. This continuous multi-level processing ensures that analysis accuracy is maintained without interruption while keeping processing costs distributed and manageable.
3Measurement precision
If full-frame video is transmitted to remote servers for analytics, then analysis capability is improved, but real-time performance and scalability deteriorate due to high computational requirements
Solution Approach 1:
The system extracts and transmits only the essential features and processed data from the video stream to remote servers, rather than transmitting complete full-frame video. This extraction of critical information maintains analysis capability while dramatically reducing transmission bandwidth and server processing requirements, enabling real-time performance.
Solution Approach 2:
Dedicated vision processors in cameras perform preliminary processing of video data before transmission, extracting key features and preparing data for server analysis. This preliminary action reduces the computational burden on remote servers, enabling them to perform sophisticated analysis in real-time without being overwhelmed by raw video data processing requirements.
4Loss of information
If cloud cameras transmit full-frame video, then video quality is preserved, but storage and computational costs increase linearly with the number of sensors
Solution Approach 1:
The system extracts only the essential visual information and processed features from full-frame video for transmission and storage, rather than preserving and storing complete video frames. This extraction maintains the necessary information for analysis while dramatically reducing storage requirements, allowing the system to scale without linear increases in storage cost.
Solution Approach 2:
The video data is segmented into different processing stages: full-resolution processing occurs locally at the camera for quality preservation, while only processed feature data is transmitted to servers for storage and further analysis. This segmentation allows the system to maintain video quality where needed while reducing overall storage requirements for scaling.
Data Source
AI summary
A computer vision-based monitoring system for counting the number of occupants in a road vehicle, including a server located external to the road vehicle and forming part of a distributed computing infrastructure; a first camera positioned in, or configured to be attached to, the road vehicle and configured to capture an image of the environment external to the road vehicle; a second camera positioned in, or configured to be attached to, the road vehicle and configured to capture an image of one or more occupants in the road vehicle; a computer vision sub-system connected to the cameras, configured to be locatable in the road vehicle, and includes an edge layer of the distributed computing infrastructure; a vehicle occupant counting sub-system trained using machine learning, wherein the sub-system is configured to use its training to count the number of occupants in the road vehicle.


