Voice Biometric Authentication Using Context-Limited Identity Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice biometrics systems require separate enrollment with each vendor system, limiting user experience and adoption across diverse devices and public settings, and face challenges in resource-intensive identification processes and false acceptance/rejection rates.

Innovation Solution

A centralized machine-learning architecture enables seamless user authentication across multiple devices and services by generating biometric and context embeddings, allowing enrollment once and authenticating across disparate systems, while mitigating resource demands and false acceptance/rejection rates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice biometrics systems are deployed privately with single-vendor enrollment, then authentication accuracy is improved, but user experience and adoption across multiple vendors deteriorates

Engineering Contradiction:
Improveauthentication accuracyVSAvoidcross-vendor compatibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces a centralized authentication server as an intermediary between voice devices and users. This server maintains a centralized database of voice prints and handles authentication requests from multiple vendors uniformly. The intermediary architecture allows each vendor to maintain their own authentication system while the centralized server provides cross-vendor interoperability, resolving the contradiction between authentication accuracy and cross-vendor compatibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The centralized authentication server implements universality by serving multiple vendors and device types through a single unified system. The server can authenticate users across different voice devices (smart speakers, telephones, video conferencing systems) without requiring separate enrollment processes for each vendor, thereby achieving both high authentication accuracy and broad adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If centralized authentication system is implemented, then cross-device authentication is enabled, but system complexity increases

Engineering Contradiction:
Improvecross-device authenticationVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system is segmented into distinct functional components: voice devices that capture audio, a centralized authentication server that processes authentication requests, and a database that stores voice prints. This segmentation allows each component to be optimized independently and simplifies the overall architecture by distributing complexity across manageable modules rather than requiring monolithic integration.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The centralized authentication server acts as an intermediary that abstracts the complexity of cross-device authentication from individual vendors and devices. By centralizing the authentication logic in a dedicated server, the patent reduces the complexity burden on each individual device while enabling seamless cross-device authentication functionality.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If comprehensive voice biometric enrollment is performed, then authentication security is improved, but resource consumption and processing time increases

Engineering Contradiction:
Improveauthentication securityVSAvoidprocessing resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary action by capturing and storing voice prints during the enrollment phase before actual authentication is needed. During authentication, the system only needs to compare the incoming voice print against the stored enrollment data, significantly reducing the computational resources required compared to performing comprehensive voice analysis in real-time during authentication.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a copy of the voice print during enrollment and stores it in the database. During authentication, the system compares the incoming voice sample against this pre-created copy rather than performing comprehensive re-analysis. This copying approach maintains high authentication security through accurate voice print comparison while dramatically reducing processing resource consumption during actual authentication events.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12380892B2Limiting identity space for voice biometric authentication
Publication Date: 2025.08.05 PINDROP SECURITY INC
  • US12380892B2 patent drawing
  • US12380892B2 patent drawing
  • US12380892B2 patent drawing

AI summary

Disclosed are systems and methods including computing-processes executing machine-learning architectures extract vectors representing disparate types of data and output predicted identities of users accessing computing services, without express identity assertions, and across multiple computing services, analyzing data from multiple modalities, for various user devices, and agnostic to architectures hosting the disparate computing service. The system invokes the identification operations of the machine-learning architecture, which extracts biometric embeddings from biometric data and context embeddings representing all or most of the types of metadata features analyzed by the system. The context embeddings help identify a subset of potentially matching identities of possible users, which limits the number of biometric-prints the system compares against an inbound biometric embedding for authentication. The types of extracted features originate from multiple modalities, including metadata from data communications, audio signals, and images. In this way, the embodiments apply a multi-modality machine-learning architecture.