Voice Authentication via Audio Segment Outlier Peak Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital voice data systems are unable to effectively distinguish between human and artificial users, as artificial intelligence and natural language processing engines can mimic human voices, leading to difficulties in authenticating user interactions.
Innovation Solution
A system that captures and compares two audio segments from user interactions, plotting them into time-domain or frequency-domain waveforms, and determining outlier peaks to assign an artificial user probability, which is then displayed on an endpoint device, utilizing machine learning for enhanced authentication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If artificial intelligence and natural language processing engines are used to create audio mimicking human voices, then the capability to interact with entities is improved, but the ability to distinguish between true human customers and artificial users deteriorates
Solution Approach 1:
The patent divides the audio signal into multiple segments and compares corresponding segments from different audio samples. By analyzing local variations within segments rather than treating the entire audio signal as a single unit, the system can detect subtle differences between human and artificial voices that would be imperceptible in aggregate analysis.
Solution Approach 2:
The patent focuses on detecting local variations and outlier peaks within specific segments of audio signals. Rather than analyzing global characteristics, the system examines local quality metrics such as waveform deviations and frequency variations at specific time points, which are more indicative of human vs. artificial voice generation.
2Ease of operation
If traditional voice authentication methods are used, then the simplicity of operation is maintained, but the accuracy of authentication deteriorates due to inability to detect artificial voices
Solution Approach 1:
The system performs self-authentication by automatically comparing audio segments against each other and against stored reference patterns. The algorithm independently identifies outlier peaks and determines authentication results without requiring external intervention or complex manual verification processes, maintaining ease of operation while improving accuracy.
Solution Approach 2:
The system incorporates feedback mechanisms where authentication results are used to refine future comparisons. By continuously learning from authenticated samples and adjusting comparison thresholds, the system improves authentication accuracy over time while maintaining the same simple operational interface.
3Measurement precision
If detailed audio analysis is performed to improve authentication accuracy, then the measurement precision is improved, but the computational resources and time required increase
Solution Approach 1:
The patent applies partial analysis by focusing computational effort only on identifying and comparing outlier peaks rather than analyzing every aspect of the audio signal. By concentrating resources on the most discriminative features (local variations and anomalies), the system achieves high authentication accuracy with reduced computational overhead compared to comprehensive audio analysis.
Solution Approach 2:
By segmenting the audio analysis into discrete comparable units and focusing only on segments with significant deviations, the system reduces the overall computational burden. The segmentation allows parallel processing of multiple segments and enables the system to skip detailed analysis of segments that show minimal variation, thereby improving processing speed.
Data Source
AI summary
Systems, computer program products, and methods are described herein for digital voice data processing and authentication. The present invention is configured to receive a user interaction comprising a digital audio signal and capture a first audio segment and a second audio segment of the digital audio signal. The first audio segment and the second audio segment are plotted into corresponding first and second plots. The first and second plots are compared, wherein comparing comprises subtracting the first plot from the second plot to form a difference plot. A quantity of outlier peaks is determined, then an artificial user probability is assigned to the user interaction, wherein the artificial user probability is low if the quantity of outlier peaks is greater than a predetermined outlier peak threshold. The artificial user probability is then displayed on a user interface of an endpoint device.


