Edge Computer Vision Processing for Scalable Face Authentication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer vision systems for home environments are either too simple and unpredictable or un-scalable and uneconomical, failing to provide accurate, real-time data analytics on detected people or objects due to high computational and storage costs that grow linearly with the number of users, cameras, and resolution.
Innovation Solution
A computer-vision system that generates a digital representation of a person from a pixel stream, determines attributes, and outputs data to a cloud-based analytics system for identification and authentication, using ASIC or SoC in cameras to process raw image sensor data in real-time, selecting the 'best face' for recognition, and sending only small image crops to the cloud, thus reducing bandwidth and compute requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sophisticated video analysis is performed on servers, then analysis accuracy is improved, but scalability and economy deteriorate due to linearly growing storage and computational costs
Solution Approach 1:
The system segments the video analysis process into two parts: simple real-time processing at camera edges using dedicated vision processors, and sophisticated analysis on servers. This segmentation allows accurate analysis to be performed only when needed while maintaining scalability through distributed edge processing.
Solution Approach 2:
The system performs preliminary video analysis and filtering at the camera edge before transmitting data to servers. This preliminary action reduces the volume of data requiring sophisticated server-side processing, improving scalability while maintaining analysis accuracy for critical events.
2Loss of energy
If simple video analysis is performed in cameras, then processing cost is reduced, but analysis accuracy and predictability deteriorate
Solution Approach 1:
The system introduces dedicated vision processors as intermediaries between standard camera sensors and server systems. These vision processors provide predictable real-time processing at the camera edge, filtering and preprocessing video data before transmission, thus improving both cost efficiency and analysis reliability.
Solution Approach 2:
The vision processors embedded in cameras perform self-service by autonomously conducting initial video analysis, event detection, and filtering without requiring constant server intervention. This self-service capability reduces processing costs while maintaining predictable accuracy for basic functions.
3Loss of information
If full-frame video is transmitted to cloud servers, then complete information is available for analysis, but bandwidth and storage requirements increase linearly
Solution Approach 1:
The system extracts only relevant information from full-frame video at the camera edge, such as detected objects, events, or anomalies, and transmits only this extracted data to cloud servers. This extraction approach maintains information completeness for analysis while dramatically reducing data transmission volume and storage requirements.
Solution Approach 2:
The system transmits partial video information (only relevant portions or processed data) rather than complete full-frame video. This partial action approach provides sufficient information for most analysis purposes while reducing bandwidth and storage requirements.
4Measurement precision
If high-resolution video processing is performed, then detection accuracy is improved, but computational cost and power consumption increase
Solution Approach 1:
The system applies different processing qualities to different regions or portions of video data. High-resolution processing is applied only to regions containing detected objects or events of interest, while other regions receive minimal or no processing. This local quality approach maintains detection accuracy for critical areas while reducing overall power consumption.
Data Source
AI summary
A computer-vision system or engine that (a) generates from a pixel stream a digital representation of a person and (b) determines attributes or characteristics of the person from that digital representation and (c) based on those attributes or characteristics, outputs data to a cloud-based analytics system that enables that analytics system to identify and also to authenticate the person. The attributes or characteristics of the person include their pose, and the system or engine analyses that pose to extract a facial image from a video stream that is the best facial image for use by the cloud-based analytics system to identify and authenticate the person. The computer-vision system or engine outputs the facial image to the cloud-based analytics system, but does not output the full-frame real-time video to the cloud-based analytics system.


