Voice Biometrics Authentication Using Codec-Independent Neural Network
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice biometric systems face reduced accuracy due to the use of compressed audio content and variations in audio codecs, leading to inconsistencies in identity verification.
Innovation Solution
A system and method that utilize a neural network to generate generic representations of audio content, which are independent of codec types, allowing for accurate authentication by comparing these representations without the need for decompression, thereby enhancing accuracy and handling mixed codec environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If compressed audio content is used for voice biometric authentication, then data storage and transmission efficiency is improved, but authentication accuracy deteriorates due to information loss from compression
Solution Approach 1:
The patent introduces an intermediary processing step where audio content from different codecs is converted to a universal format before voice print extraction. This intermediary conversion layer eliminates the direct dependency between diverse codec outputs and the voice biometric engine, allowing compressed audio to be used efficiently while maintaining authentication accuracy through standardized processing.
2Adaptability or versatility
If voice prints are stored using different codecs to support various audio sources, then system adaptability is improved, but authentication accuracy deteriorates due to codec sensitivity in VB engines
Solution Approach 1:
The patent applies homogeneity by converting all audio inputs from different codecs into a single universal audio format before processing. This standardization ensures that the voice biometric engine receives consistent, homogeneous input regardless of the source codec, thereby maintaining high authentication accuracy while supporting diverse audio sources through the universal conversion layer.
3Measurement precision
If original uncompressed audio content is used for authentication, then authentication accuracy is improved, but data storage and processing requirements increase
Solution Approach 1:
The patent extracts only the essential voice print features from the audio content after converting to universal format, rather than storing and processing the entire audio stream. This extraction approach maintains authentication accuracy by preserving critical biometric information while significantly reducing data storage and processing requirements by eliminating redundant audio data.
Data Source
AI summary
A system and method for authenticating an identity may include generating a first generic representation representing a stored audio content, generating a second generic representation representing input audio content, and, providing the first and second generic representations to a voice biometrics unit adapted to authenticate an identity based on the first and second generic representations.


