A scalable real-time video surveillance system with advanced face detection and recognition
AZ20260036BActive Publication Date: 2026-07-31İKTEX MƏHDUD MƏSULİYYƏTLİ CƏMİYYƏTİ
0 Cites 0 Cited by
Patent Information
- Application Number
- AZ20250090
- Authority / Receiving Office
- AZ · AZ
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2026-07-31
- Estimated Expiration
- 2045-05-15
Abstract
A fundamentally designed distributed artificial intelligence (AI) control system is presented to provide real-time face detection and high-accuracy recognition on thousands of RTSP video streams. The system uses the power of OpenCV technology to acquire video and perform comprehensive pre-image processing, and implements a highly optimized YOLO-based detection model. This model is trained using PyTorch, converted to ONNX format, and accelerated with TensorRT, resulting in ultra-low latency (FP16: ~11.9 ms / detection, FP32: ~25.5 ms / detection). The seamless integration with ByteTrack technology ensures consistent, unique face identification and frame-to-frame tracking accuracy. The main innovation is in the face recognition subsystem, which consists of a small-scale and high-performance ResNet-based descriptor extraction method. Developed through advanced architectural optimizations and training methods, this new model achieves a record-breaking Rank-1 accuracy of 98% under challenging observation conditions, outperforming traditional systems in both accuracy and efficiency. Trained on 100 million distinct images representing 3 million distinct individuals, the compact model produces highly selective 512-dimensional face embeddings at a rate of over 400 embeddings per second with the NVIDIA RTX A5000 GPU. These embeddings are efficiently stored in a PostgreSQL database, enhanced with pg_vector and the Hierarchical Navigable Small World (HNSW) index, which allows for L2 distance nearest neighbor search across billions of embeddings with ultra-low latency (~25 ms). The similarity scores are converted into user-friendly percentage values using a custom formula. A comprehensive web interface and API, powered by OpenCV-based preprocessing, simplifies camera configuration, enables real-time monitoring with high visual clarity, and accurately performs large-scale image and vector search. The overall architecture of the system achieves high scalability, real-time responsiveness, and industry-leading facial recognition accuracy with a highly compact and efficient model, using Kafka-dedicated processes, robust load balancing with HAProxy, and dedicated GPU-optimized servers.
Need to check novelty before this filing date? Find Prior Art