Liveness Detection via Audio-Visual Inconsistency Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing biometric verification systems in eKYC procedures are vulnerable to presentation attacks, such as 2D and 3D attacks, and current liveness detection methods like thermal imaging and rPPG are costly, complex, and have low accuracy.

Innovation Solution

A liveness detection verification method and system that uses an audio-visual similarity check based on a random phrase challenge, employing a speech recognition machine learning model to verify synchronization between mouth movements and audio, preventing presentation attacks by requiring real-time speech verification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If thermal imaging-based facial liveness detection is used, then liveness detection accuracy is improved, but device complexity and cost increase due to additional sensors

Engineering Contradiction:
Improveliveness detection accuracyVSAvoidsensor requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes the thermal imaging sensor component from the liveness detection system. Instead of using thermal imaging, the system uses only the existing camera and microphone to capture video and audio data, processing these through machine learning models to achieve liveness detection without requiring additional expensive sensors.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a virtual representation of liveness detection capabilities through machine learning models that analyze audio-visual correlations. Rather than relying on physical thermal sensors, the system uses software-based analysis of standard video and audio inputs to detect liveness, effectively copying the functionality without the hardware overhead.

Inventive Principle:
Principle #26Copying

2Reliability

If rPPG method is used for liveness detection, then liveness detection is achieved, but processing time increases and accuracy decreases

Engineering Contradiction:
Improveliveness detection capabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the complex rPPG signal processing mechanism with a more efficient machine learning-based audio-visual correlation analysis. Instead of using traditional signal processing methods that require extensive computational steps, the system uses trained neural networks to directly analyze the relationship between audio and video data, significantly reducing processing time while maintaining or improving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If 3D facial depth analysis is used, then presentation attack detection is improved, but device complexity increases due to additional sensors

Engineering Contradiction:
Improvepresentation attack detectionVSAvoidsensor requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent removes the requirement for depth sensors and 3D sensing hardware from the system. Instead of using physical 3D depth analysis, the system extracts depth and spatial information through monocular vision techniques and audio-visual correlation analysis, achieving presentation attack detection without additional sensors.

Inventive Principle:
Principle #2Taking out (Extraction)

4Reliability

If speech recognition machine learning model is added to verify audio-visual similarity, then presentation attack prevention is improved, but device complexity increases

Engineering Contradiction:
Improvepresentation attack preventionVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges the speech recognition functionality with the existing audio processing pipeline. Instead of adding a completely separate verification system, the speech recognition model is integrated into the audio-visual correlation analysis framework, where it works in conjunction with the video analysis model to jointly determine liveness, reducing overall system complexity through unified processing.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12039024B2Liveness detection using audio-visual inconsistencies
Publication Date: 2024.07.16 RAKUTEN GROUP INC
  • US12039024B2 patent drawing
  • US12039024B2 patent drawing
  • US12039024B2 patent drawing

AI summary

Provided are a method and system for verifying a liveness detection of a user. The method includes: obtaining a video of a user speaking a phrase in response to a question or a randomly generated phrase presented to the user; inputting video data and audio data of the obtained video to a first determination model to obtain a first determination indicative of whether a mouth movement of the user is synchronized with the audio data; inputting, to a second determination model, a first input corresponding to the audio data and a second input corresponding to a predetermined phrase, to obtain a second determination indicative of whether the predetermined phrase is spoken by the user; and determining whether the first determination indicates that the mouth movement is synchronized with the audio data and whether the second determination indicates that the predetermined phrase is spoken by the user.