Edge Vision Face Selection for Scalable Cloud Authentication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer vision systems for home environments are either too simple and unpredictable or too complex and uneconomical, failing to provide accurate, scalable, and real-time data analytics on detected people or objects due to high computational and bandwidth requirements, privacy concerns, and limited scalability.
Innovation Solution
A computer-vision system that generates a digital representation of a person or object from a pixel stream, determines attributes, and outputs data to a cloud-based analytics system for identification and authentication, using ASIC or SoC in cameras to process raw image data in real-time, selecting the 'best face' for recognition, and sending only small image crops to the cloud, reducing bandwidth and compute requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sophisticated video analysis is performed on servers, then analysis accuracy is improved, but scalability and economy deteriorate due to linear scaling of storage and computational costs
Solution Approach 1:
The system segments video processing into two parts: simple real-time processing at the camera edge (detecting motion, generating alerts) and sophisticated analysis only when needed on the server. This segmentation allows the system to maintain high accuracy for complex analysis while avoiding the need to process all video data centrally, thus improving scalability and reducing computational costs.
Solution Approach 2:
The system performs preliminary filtering and processing at the edge devices before data reaches the server. Motion detection and basic video analysis are conducted locally, so only relevant data requiring sophisticated analysis is transmitted to the server. This preliminary action reduces the computational burden on servers and improves overall system scalability.
2Productivity
If simple video analysis is performed in cameras, then processing cost is reduced, but analysis reliability and predictability deteriorate
Solution Approach 1:
The system segments analysis tasks by complexity: simple, reliable processing (motion detection, basic object recognition) is performed at the camera edge, while complex analysis (detailed video interpretation, pattern recognition) is performed on the server when needed. This segmentation enables the camera to operate independently with reliable simple functions while maintaining the option for sophisticated analysis.
Solution Approach 2:
The system introduces an intermediary layer that coordinates between edge cameras and central servers. This intermediary manages data flow, determining when simple edge processing suffices and when sophisticated server analysis is required, thereby maintaining reliability while improving processing efficiency.
3Loss of information
If full-frame video is transmitted to cloud servers, then complete data is available for analysis, but bandwidth requirements and computational costs increase linearly
Solution Approach 1:
The system extracts only the essential information needed for analysis from full video frames at the edge devices. Motion detection, object presence, and key event data are extracted and transmitted to the server, while the bulk video data remains local. This extraction maintains data completeness for analysis while dramatically reducing bandwidth requirements and computational energy consumption.
4Measurement precision
If state-of-the-art vision algorithms are deployed on CPUs and GPUs, then analysis capability is improved, but processing cost and power consumption increase significantly
Solution Approach 1:
The system segments computational tasks by complexity and energy requirements. Low-power processors at edge devices handle simple vision tasks (motion detection, basic recognition), while high-performance CPUs and GPUs on servers handle sophisticated analysis only when needed. This segmentation maintains high analysis capability for complex tasks while minimizing overall power consumption by avoiding continuous use of high-power processors.
Data Source
AI summary
A computer-vision system or engine that (a) generates from a pixel stream a digital representation of a person and (b) determines attributes or characteristics of the person from that digital representation and (c) based on those attributes or characteristics, outputs data to a cloud-based analytics system that enables that analytics system to identify and also to authenticate the person. The attributes or characteristics of the person include their pose, and the system or engine analyses that pose to extract a facial image from a video stream that is the best facial image for use by the cloud-based analytics system to identify and authenticate the person. The computer-vision system or engine outputs the facial image to the cloud-based analytics system, but does not output the full-frame real-time video to the cloud-based analytics system.


