Audio Feedback Video Authentication Against Deep Fakes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current authentication methods for accessing secured resources are cumbersome, requiring multiple identification factors and being vulnerable to video forgery, particularly deep fake attacks, which increase security costs and inconvenience for users.

Innovation Solution

A cloud-based authentication system that uses audio tokens generated from a user's device, where a pre-designated audio signal is played and the reflected audio data is compared to a stored token to verify the user's identity, eliminating the need for traditional credentials and enhancing security against deep fake attacks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional authentication methods using multiple identification factors are used, then security is improved, but user convenience and authentication speed deteriorate

Engineering Contradiction:
Improveauthentication securityVSAvoiduser convenience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent extracts the authentication verification process from traditional multi-factor authentication methods and replaces it with a biometric-based audio analysis system. The system captures audio signals from the user's device, extracts specific acoustic features (such as device handling patterns, background noise characteristics, and voice properties), and compares these against pre-stored biometric templates to verify identity, thereby eliminating the need for users to manually present multiple credentials

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical/manual authentication process (where users physically present ID cards, keys, or remember passwords) with an automated acoustic analysis system. The system uses machine learning algorithms to analyze audio characteristics captured during video calls, automatically verifying user identity without requiring manual intervention or physical credentials

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If video streaming authentication is used, then user convenience is improved, but vulnerability to deep fake attacks increases

Engineering Contradiction:
Improveauthentication convenienceVSAvoidvideo authenticity
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces audio signal analysis as an intermediary verification layer between the video streaming process and authentication decision. Instead of relying solely on visual verification, the system captures audio signals from the user's device, analyzes acoustic features (including device-specific noise patterns, microphone characteristics, and environmental sounds), and uses these audio biomarkers to verify the authenticity of the video feed, making it difficult for deep fake attacks to successfully impersonate the user

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements real-time audio feedback analysis during the video authentication process. The authentication server continuously monitors audio characteristics from the user's device, compares them against pre-stored biometric audio templates, and provides immediate verification feedback. This feedback mechanism allows the system to detect anomalies or inconsistencies in the audio signal that may indicate deep fake manipulation, thereby maintaining authentication reliability while preserving the convenience of video-based verification

Inventive Principle:
Principle #23Feedback

3Reliability

If multiple identification credentials are required, then authentication security is improved, but authentication time and complexity increase

Engineering Contradiction:
Improveauthentication securityVSAvoidauthentication time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a pre-enrollment process where the system captures and stores the user's biometric audio characteristics (device handling patterns, voice properties, background acoustic environment) before actual authentication occurs. During the authentication phase, the system simply compares the newly captured audio features against these pre-stored templates, significantly reducing authentication time while maintaining security through the use of unique biometric identifiers

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This method provides a secure, efficient, and convenient authentication process by verifying user identity through audio feedback, reducing the need for multiple credentials and effectively combating video forgery, thereby enhancing account security and user experience.

Implementation Method 1

a first control signal including a first audio signal... playback of the first audio signal by a speaker... first audio data captured by a microphone

Methodology Applied
Scientific EffectSound wave generation and reflection: Sound

Data Source

PatentUS12259957B1Audio feedback-based video authentication method and system
Publication Date: 2025.03.25 UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)
  • US12259957B1 patent drawing
  • US12259957B1 patent drawing
  • US12259957B1 patent drawing

AI summary

A remote audio signal-based method and system of performing an authentication of video of a person in order to authorize access to a secured resource. An audio security token with a first set of features is generated and stored for each user. During subsequent sessions, the system and method are configured to cause a remote computing device to play a specific audio signal while collecting audio data from a microphone of the same device. The audio data is evaluated to determine if the same features are present as in the first set of features. If the features are present, the system determines the image is authentic and can verify an identity of the person, and can further be configured to automatically grant the person access to one or more services, features, or information for which he or she is authorized.