Video Facial Recognition via Continuous Face Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in quickly and accurately recognizing humans in videos, as they struggle to efficiently form continuous face sequences and perform reliable facial recognition across multiple video frames with varying displacement and position changes.

Innovation Solution

The method involves forming face sequences by identifying face images in continuous video frames that meet a predetermined displacement requirement, using a face library for facial recognition, and incorporating face features and key points to enhance recognition accuracy through clustering processing and weighted averaging of face features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If facial recognition is performed on each video frame independently, then recognition can be performed quickly, but recognition accuracy deteriorates due to displacement and position changes of faces across frames

Engineering Contradiction:
Improverecognition speedVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary face detection and tracking to establish face sequences across multiple video frames before performing recognition. By pre-organizing face images into continuous sequences based on temporal and spatial relationships, the system prepares structured data that improves recognition accuracy without sacrificing speed, as the sequencing work is done in advance of the actual recognition process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges multiple face images from continuous video frames into face sequences, combining temporal information across frames. This merging process creates a more robust representation of each person by aggregating multiple observations, which improves recognition reliability while maintaining efficiency through batch processing of the combined sequences

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If face sequences are formed using strict displacement requirements, then face sequence accuracy is improved, but the number of detected faces decreases due to position variations

Engineering Contradiction:
Improveface sequence accuracyVSAvoidnumber of detected faces
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent implements dynamic displacement thresholds that adapt based on video characteristics and face movement patterns. Rather than using fixed strict requirements, the system adjusts the displacement criteria dynamically to accommodate varying position changes across different video sequences, thereby maintaining both accuracy and detection completeness

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the displacement parameter thresholds based on the specific video content and face movement characteristics. By adjusting these parameters dynamically, the system optimizes the balance between maintaining accurate face sequences and ensuring that faces with varying position changes are not missed, thus preserving both precision and quantity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11068697B2Methods and apparatus for video-based facial recognition, electronic devices, and storage media
Publication Date: 2021.07.20 BEIJING SENSETIME TECH DEV CO LTD
  • US11068697B2 patent drawing
  • US11068697B2 patent drawing
  • US11068697B2 patent drawing

AI summary

Methods and apparatuses for video-based facial recognition, devices, media, and programs can include: forming a face sequence for face images, in a video, appearing in multiple continuous video frames and having positions in the multiple video frames meeting a predetermined displacement requirement, wherein the face sequence is a set of face images of a same person in the multiple video frames; and performing facial recognition for the face sequence by using a preset face library at least according to face features in the face sequence.