Edge Vision Face Selection for Scalable Cloud Authentication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computer vision systems for home environments are either too simple and unpredictable or too complex and uneconomical, failing to provide accurate, scalable, and real-time data analytics on detected people or objects due to high computational and bandwidth requirements, privacy concerns, and limited scalability.

Innovation Solution

A computer-vision system that generates a digital representation of a person or object from a pixel stream, determines attributes, and outputs data to a cloud-based analytics system for identification and authentication, using ASIC or SoC in cameras to process raw image data in real-time, selecting the 'best face' for recognition, and sending only small image crops to the cloud, reducing bandwidth and compute requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If sophisticated video analysis is performed on servers, then analysis accuracy is improved, but scalability and economy deteriorate due to linear scaling of storage and computational costs

Engineering Contradiction:
Improvevideo analysis accuracyVSAvoidcomputational infrastructure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments video processing into two parts: simple real-time processing at the camera edge (detecting motion, generating alerts) and sophisticated analysis only when needed on the server. This segmentation allows the system to maintain high accuracy for complex analysis while avoiding the need to process all video data centrally, thus improving scalability and reducing computational costs.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary filtering and processing at the edge devices before data reaches the server. Motion detection and basic video analysis are conducted locally, so only relevant data requiring sophisticated analysis is transmitted to the server. This preliminary action reduces the computational burden on servers and improves overall system scalability.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If simple video analysis is performed in cameras, then processing cost is reduced, but analysis reliability and predictability deteriorate

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidanalysis reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system segments analysis tasks by complexity: simple, reliable processing (motion detection, basic object recognition) is performed at the camera edge, while complex analysis (detailed video interpretation, pattern recognition) is performed on the server when needed. This segmentation enables the camera to operate independently with reliable simple functions while maintaining the option for sophisticated analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary layer that coordinates between edge cameras and central servers. This intermediary manages data flow, determining when simple edge processing suffices and when sophisticated server analysis is required, thereby maintaining reliability while improving processing efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If full-frame video is transmitted to cloud servers, then complete data is available for analysis, but bandwidth requirements and computational costs increase linearly

Engineering Contradiction:
Improvedata completenessVSAvoidbandwidth and computational energy
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The system extracts only the essential information needed for analysis from full video frames at the edge devices. Motion detection, object presence, and key event data are extracted and transmitted to the server, while the bulk video data remains local. This extraction maintains data completeness for analysis while dramatically reducing bandwidth requirements and computational energy consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

4Measurement precision

If state-of-the-art vision algorithms are deployed on CPUs and GPUs, then analysis capability is improved, but processing cost and power consumption increase significantly

Engineering Contradiction:
Improvevision analysis capabilityVSAvoidprocessing power consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system segments computational tasks by complexity and energy requirements. Low-power processors at edge devices handle simple vision tasks (motion detection, basic recognition), while high-performance CPUs and GPUs on servers handle sophisticated analysis only when needed. This segmentation maintains high analysis capability for complex tasks while minimizing overall power consumption by avoiding continuous use of high-power processors.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260038274A1Computer vision systems
Publication Date: 2026.02.05 UNIFAI HLDG LTD
  • US20260038274A1 patent drawing
  • US20260038274A1 patent drawing
  • US20260038274A1 patent drawing

AI summary

A computer-vision system or engine that (a) generates from a pixel stream a digital representation of a person and (b) determines attributes or characteristics of the person from that digital representation and (c) based on those attributes or characteristics, outputs data to a cloud-based analytics system that enables that analytics system to identify and also to authenticate the person. The attributes or characteristics of the person include their pose, and the system or engine analyses that pose to extract a facial image from a video stream that is the best facial image for use by the cloud-based analytics system to identify and authenticate the person. The computer-vision system or engine outputs the facial image to the cloud-based analytics system, but does not output the full-frame real-time video to the cloud-based analytics system.