Voice Biometric Liveness Detection System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice biometric systems using text-dependent methods are vulnerable to security breaches, such as eavesdropping and voice phishing, as they rely on specific phrases that can be intercepted and replayed by attackers, compromising authentication processes.

Innovation Solution

The implementation of a method that combines text-dependent and text-independent voice biometrics, along with Automatic Speech Recognition, to generate matching scores and a liveness score by comparing voice samples across enrollment and authentication stages, utilizing pseudo-random phrases and intra-session voice variation to verify the authenticity and humanity of the speaker.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text-dependent voice biometric methods are used for speaker recognition, then the authentication process is simplified and user experience is improved, but the system becomes vulnerable to security breaches such as eavesdropping and voice phishing attacks

Engineering Contradiction:
Improveauthentication processVSAvoidsecurity
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent combines text-dependent voice biometric authentication with liveness detection technology into a unified system. The authentication process integrates both the textual content verification (text-dependent) and the speaker's physiological state verification (liveness detection), creating a multi-layered security approach that maintains ease of use while significantly improving security against replay and phishing attacks

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces liveness detection as an intermediary layer between the user and the authentication system. This intermediary verifies the speaker's physiological state (such as detecting natural voice characteristics, breathing patterns, or other biometric indicators) before granting access, thereby preventing unauthorized access from recorded or synthesized voice samples while maintaining the simplicity of text-dependent authentication

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If traditional multi-session enrollment procedures are implemented for voice biometric systems, then the accuracy of speaker recognition is improved, but the time required for enrollment and system setup increases significantly

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidenrollment time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs liveness detection and voice characteristic analysis during the initial enrollment phase, capturing comprehensive biometric data in a single session. By conducting all necessary measurements and creating the voice profile upfront, the system eliminates the need for multiple enrollment sessions while ensuring high accuracy through thorough initial assessment

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous liveness monitoring throughout the authentication process, maintaining verification of the speaker's physiological state from enrollment through authentication. This continuous verification ensures high accuracy without requiring separate enrollment sessions, as the system continuously adapts and refines the voice profile based on ongoing biometric data

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS9484037B2Device, system, and method of liveness detection utilizing voice biometrics
Publication Date: 2016.11.01 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9484037B2 patent drawing
  • US9484037B2 patent drawing
  • US9484037B2 patent drawing

AI summary

Device, system, and method of liveness detection using voice biometrics. For example, a method comprises: generating a first matching score based on a comparison between: (a) a voice-print from a first text-dependent audio sample received at an enrollment stage, and (b) a second text-dependent audio sample received at an authentication stage; generating a second matching score based on a text-independent audio sample; and generating a liveness score by taking into account at least the first matching score and the second matching score.