Liveness Detection via Audio-Visual Inconsistency Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing biometric verification systems in eKYC procedures are vulnerable to presentation attacks, such as 2D and 3D attacks, and current liveness detection methods like thermal imaging and rPPG are costly, complex, and have low accuracy.
Innovation Solution
A liveness detection verification method and system that uses an audio-visual similarity check based on a random phrase challenge, employing a speech recognition machine learning model to verify synchronization between mouth movements and audio, preventing presentation attacks by requiring real-time speech verification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If thermal imaging-based facial liveness detection is used, then liveness detection accuracy is improved, but device complexity and cost increase due to additional sensors
Solution Approach 1:
The patent extracts and removes the thermal imaging sensor component from the liveness detection system. Instead of using thermal imaging, the system uses only the existing camera and microphone to capture video and audio data, processing these through machine learning models to achieve liveness detection without requiring additional expensive sensors.
Solution Approach 2:
The patent creates a virtual representation of liveness detection capabilities through machine learning models that analyze audio-visual correlations. Rather than relying on physical thermal sensors, the system uses software-based analysis of standard video and audio inputs to detect liveness, effectively copying the functionality without the hardware overhead.
2Reliability
If rPPG method is used for liveness detection, then liveness detection is achieved, but processing time increases and accuracy decreases
Solution Approach 1:
The patent replaces the complex rPPG signal processing mechanism with a more efficient machine learning-based audio-visual correlation analysis. Instead of using traditional signal processing methods that require extensive computational steps, the system uses trained neural networks to directly analyze the relationship between audio and video data, significantly reducing processing time while maintaining or improving accuracy.
3Reliability
If 3D facial depth analysis is used, then presentation attack detection is improved, but device complexity increases due to additional sensors
Solution Approach 1:
The patent removes the requirement for depth sensors and 3D sensing hardware from the system. Instead of using physical 3D depth analysis, the system extracts depth and spatial information through monocular vision techniques and audio-visual correlation analysis, achieving presentation attack detection without additional sensors.
4Reliability
If speech recognition machine learning model is added to verify audio-visual similarity, then presentation attack prevention is improved, but device complexity increases
Solution Approach 1:
The patent merges the speech recognition functionality with the existing audio processing pipeline. Instead of adding a completely separate verification system, the speech recognition model is integrated into the audio-visual correlation analysis framework, where it works in conjunction with the video analysis model to jointly determine liveness, reducing overall system complexity through unified processing.
Data Source
AI summary
Provided are a method and system for verifying a liveness detection of a user. The method includes: obtaining a video of a user speaking a phrase in response to a question or a randomly generated phrase presented to the user; inputting video data and audio data of the obtained video to a first determination model to obtain a first determination indicative of whether a mouth movement of the user is synchronized with the audio data; inputting, to a second determination model, a first input corresponding to the audio data and a second input corresponding to a predetermined phrase, to obtain a second determination indicative of whether the predetermined phrase is spoken by the user; and determining whether the first determination indicates that the mouth movement is synchronized with the audio data and whether the second determination indicates that the predetermined phrase is spoken by the user.


