Identity Authentication via Lip-Reading and Voice Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing identity authentication methods are complex, unreliable, and vulnerable to attacks, particularly facial recognition and voiceprint recognition, which can be bypassed using analog videos or recordings, posing security risks to applications with high security requirements.

Innovation Solution

An identity authentication method that uses a single audio and video stream for user identification, verifying facial features and voiceprints, and ensuring consistency between lip reading and voice to differentiate between real users and fake video records, thereby enhancing security and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If facial recognition and voiceprint recognition are used for identity authentication, then authentication functionality is provided, but the authentication process becomes complex and reliability decreases

Engineering Contradiction:
Improveauthentication reliabilityVSAvoidauthentication process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines facial recognition, voiceprint recognition, and lip-sync verification into a unified authentication system that processes audio and video streams simultaneously. Multiple authentication modalities are merged into a single integrated process that verifies identity through coordinated analysis of facial features, voice characteristics, and lip movements, rather than treating them as separate independent steps.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The authentication system performs multiple functions using a single integrated process: it extracts facial features from video, analyzes voiceprints from audio, verifies lip-sync consistency between audio and video, and cross-validates all three modalities together. This multi-functional approach allows the system to achieve comprehensive verification without requiring separate authentication workflows for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If traditional facial recognition or voiceprint recognition is used, then authentication is performed, but the system is vulnerable to attacks using analog videos or recordings

Engineering Contradiction:
Improveauthentication securityVSAvoiduser operation simplicity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent introduces lip-sync verification as an intermediary layer that mediates between the audio and video authentication streams. By analyzing whether lip movements correspond temporally and spatially to the spoken words, the system creates a verification bridge that detects inconsistencies in replay attacks. This intermediary check adds a layer of security that prevents attackers from successfully using pre-recorded videos or audio files.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The authentication system uses a composite verification approach combining three different biometric modalities (facial features, voiceprints, and lip movements) that work together synergistically. Just as composite materials combine different substances to achieve superior properties, this composite authentication approach combines multiple biometric streams to achieve security levels higher than any single modality alone, making it resistant to attacks that might compromise individual modalities.

Inventive Principle:
Principle #40Composite materials

3Reliability

If multiple independent authentication methods are combined, then security is improved, but the authentication process becomes more complex

Engineering Contradiction:
Improveauthentication securityVSAvoidauthentication process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple authentication modalities into a single unified processing pipeline that simultaneously analyzes audio and video streams. Instead of sequentially executing separate facial recognition, voiceprint recognition, and lip-sync verification processes, the system integrates these functions into one coordinated authentication flow that processes all modalities together, reducing operational complexity while maintaining security.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP3460697B1Identity authentication method and apparatus
Publication Date: 2021.12.08 ADVANCED NEW TECHNOLOGIES CO LTD
  • EP3460697B1 patent drawingFigure 1
  • EP3460697B1 patent drawingFigure 2
  • EP3460697B1 patent drawingFigure 3

AI summary

The present application provides an identity authentication method and apparatus. The method includes obtaining a collected audio and video stream generated by a target object to be authenticated; determining whether lip reading and voice in the audio and video stream are consistent, and if the lip reading and the voice are consistent, using voice content obtained by performing voice recognition on an audio stream in the audio and video stream as an object identifier of the target object; obtaining a model physiological feature corresponding to the object identifier from object registration information, if the pre-stored object registration information includes the object identifier; performing physiological recognition on the audio and video stream to obtain a physiological feature of the target object; and comparing the physiological feature of the target object with the model physiological feature to obtain a comparison result, and if the comparison result satisfies an authentication condition, determining that the target object has been authenticated. The present application improves the efficiency and reliability of identity authentication.