Face Processing Frame Selection for Presentation Attack Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing face processing technologies face inefficiencies in frame selection for video streams, leading to redundant processing and vulnerability to presentation attacks, as they often require complete video transmission and fail to effectively distinguish between human and non-human inputs.
Innovation Solution
A method for selecting frames from a video stream based on liveness events and image quality indicators, where frames with movements or low quality are excluded, and frames are spaced apart to prevent redundancy and enhance presentation attack detection, using a challenge-response strategy or passive liveness detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If complete video stream is transmitted remotely for face processing, then face processing can be performed with high processing power and reference images, but transmission cost and time increase significantly
Solution Approach 1:
The patent extracts only the essential frame information needed for face processing rather than transmitting the complete video stream. By selecting and transmitting only relevant frames containing face images, the system maintains processing accuracy while significantly reducing transmission time and cost.
Solution Approach 2:
The patent applies different processing quality levels to different parts of the video data. Instead of uniform high-quality transmission of all frames, it selectively transmits only frames with adequate face quality at appropriate resolution, optimizing the balance between processing reliability and transmission efficiency.
2Reliability
If all frames are processed for face recognition, then no frames are missed, but processing power consumption increases
Solution Approach 1:
The patent extracts and processes only the subset of frames that contain face images rather than processing all frames in the video stream. This selective extraction approach maintains face recognition accuracy while significantly reducing processing power consumption.
Solution Approach 2:
The patent applies partial action by processing only the necessary portion of the video data - specifically frames where faces are detected - rather than processing every frame. This partial processing approach is sufficient for achieving accurate face recognition without the excessive power consumption of full-stream processing.
3Loss of information
If consecutive frames are selected from video stream, then informative content is captured, but redundant frames increase processing load
Solution Approach 1:
The patent applies preliminary filtering actions before frame selection by detecting face presence in frames first. This preliminary detection step allows the system to pre-identify which frames contain faces and select only those frames, avoiding the selection of redundant frames without faces and reducing the overall number of frames requiring processing.
4Loss of information
If frames with movements are included in processing, then dynamic face expressions are captured, but presentation attack detection becomes vulnerable
Solution Approach 1:
The patent performs preliminary detection of movement in frames before selecting them for processing. By identifying and excluding frames with significant movement in advance, the system prevents presentation attacks that rely on dynamic manipulation while still capturing adequate static and subtle expression information for legitimate face recognition.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method used in a personal equipment for selecting frames (150i) used in face processing, comprising: a) Capturing (102) with a camera (10) video data featuring a face of an individual; b) Determining (106) with a processing unit (11) at least one image 5quality indicator (qij) for at least some frames (150i) in said video data; c) Using said quality indicator (qij) for selecting (110) a subset of said frames (Sel1, Sel2); d) Detecting (108) a sequence of frames (fperiod) corresponding to a movement of a body portion in said video data and/or corresponding to a response window (rw) during which a response to a challenge should be given; e) Adding (110) to said subset at least one second frame (Sel3, Sel4) within a predefined interval (Int1, Int2) before or after said sequence; f) Storing (114) said subset of frames (Sel1, Sel2, Sel3, Sel4) in a memory (12).