Audio-Video Liveness Detection Against Photo and Video Spoofing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing liveness detection schemes in face recognition systems are ineffective against photo and video attacks, leading to security vulnerabilities and reduced accuracy.
Innovation Solution
A liveness detection method that combines audio and video data processing, including time-frequency analysis and motion trajectory extraction, to calculate global attention information from both modalities and fuse them for accurate liveness determination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional single-modality liveness detection is used, then the detection process is simple, but the detection accuracy is poor and vulnerable to photo and video attacks
Solution Approach 1:
The patent combines audio and video modalities into a unified liveness detection system. The audio processing branch extracts features from reflected audio signals, while the video processing branch extracts features from video data. These two branches are merged through feature fusion to produce a comprehensive liveness detection result, making the system resistant to photo and video attacks that cannot simultaneously spoof both modalities.
2Reliability
If multi-modal audio-video fusion is implemented, then the security against spoofing attacks is improved, but the computational complexity increases
Solution Approach 1:
The patent divides the liveness detection system into separate audio processing and video processing branches. Each branch independently processes its respective modality through dedicated neural network components, extracting features separately before fusion. This segmentation allows for optimized computation in each branch while maintaining the security benefits of multi-modal fusion.
Solution Approach 2:
The patent introduces a global attention mechanism that operates across the temporal dimension of the extracted features. By calculating global attention information from audio features and video features separately, then fusing these attention-weighted representations, the system captures long-range dependencies and global contextual information without requiring excessive computational resources.
3Measurement precision
If global attention mechanism is applied to both audio and video features, then the detection precision is enhanced, but the processing time increases
Solution Approach 1:
The patent performs feature extraction and global attention calculation in a preliminary manner during the detection process. By pre-computing the global attention information from extracted features before final fusion, the system prepares optimized representations that reduce computational burden in the decision-making stage, thereby balancing precision with processing efficiency.
Data Source
AI summary
A liveness detection method includes: obtaining a reflected audio signal and video data of a object in response to receiving a liveness detection request; performing signal processing and time-frequency analysis on the reflected audio signal to obtain time-frequency information of a processed audio signal, and extracting motion trajectory information of the object from the video data; respectively extract features from the time-frequency information and the motion trajectory information to obtain an audio feature and a motion feature of the object; calculating first global attention information of the object according to the audio feature, and calculating second global attention information of the object according to the motion feature; and fusing the first global attention information with the second global attention information to obtain fused global information, and determining a liveness detection result of the object based on the fused global information.


