Voice Authentication via Harmonic Standardization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice authentication systems face challenges in efficiently processing and recognizing voices due to high computational resource requirements and variations in voice characteristics over time and environment, particularly in noisy or different volume settings.
Innovation Solution
A method and system for generating voice identification by transforming voice signals into the frequency domain, where the amplitude and frequency of harmonics are standardized, and unnecessary components are filtered out, or by segmenting voice signals into time domains and digitizing specific time portions, reducing computational load and enhancing recognition across varying conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional voice feature extraction methods are used, then voice authentication accuracy is maintained, but computational resource requirements increase significantly
Solution Approach 1:
The patent extracts and focuses only on the most critical voice features - specifically pitch (fundamental frequency) and volume (amplitude) characteristics - rather than computing comprehensive spectral features. This selective extraction of essential parameters reduces computational load while maintaining authentication accuracy by capturing the most distinctive aspects of individual voice patterns.
Solution Approach 2:
The system transforms voice authentication from traditional time-domain or full-spectral analysis to frequency-domain analysis focusing on fundamental frequency (pitch) and amplitude (volume) parameters. By changing the analysis domain and focusing on specific frequency parameters rather than complete spectral decomposition, the system reduces computational complexity while preserving authentication effectiveness.
2Measurement precision
If comprehensive voice characteristics are extracted, then authentication accuracy is improved, but processing time increases
Solution Approach 1:
The patent extracts only the essential pitch and volume characteristics from voice signals, ignoring less critical spectral details. This selective extraction dramatically reduces the number of computations required while maintaining authentication accuracy by focusing on the most distinctive and stable voice features that vary between individuals but remain consistent for the same individual.
Solution Approach 2:
The system performs partial feature extraction by computing only the fundamental frequency and amplitude parameters rather than complete spectral analysis. This partial action approach processes only the necessary minimum set of features required for effective authentication, reducing processing time while maintaining sufficient accuracy for practical applications.
3Measurement precision
If voice signals are processed in detail, then recognition accuracy is maintained, but data representation size increases
Solution Approach 1:
The patent extracts only the essential pitch and volume parameters from complete voice signals, representing complex acoustic information through these two primary characteristics. This extraction reduces data representation size by focusing on the most discriminative features while maintaining recognition accuracy through the use of frequency-domain analysis that captures fundamental voice identity markers.
Solution Approach 2:
The system changes the representation parameters from comprehensive time-domain waveforms or full spectral data to specific frequency-domain parameters (fundamental frequency and amplitude). This parameter transformation reduces data size by representing complex voice signals through a small set of meaningful parameters that capture essential identity information.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A system and method are provided to authenticate a voice in a frequency domain. A voice in the time domain is transformed to a signal in the frequency domain. The first harmonic is set to a predetermined frequency and the other harmonic components are equalized. Similarly, the amplitude of the first harmonic is set to a predetermined amplitude, and the harmonic components are also equalized. The voice signal is then filtered. The amplitudes of each of the harmonic components are then digitized into bits to form at least part of a voice ID. In another system and method, a voice is authenticated in a time domain. The initial rise time, initial fall time, second rise time, second fall time and final oscillation time are digitized into bits to form at least part of a voice ID. The voice IDs are used to authenticate a user's voice.