Speaker Identification via Formant Frequency Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker recognition methods fail to reliably identify speakers from short, low-quality audio records with high noise and distortion, especially when speakers are in different psycho-physiological states or speaking different languages, due to their dependence on identical verbal content, noise sensitivity, and inability to handle varying audio channel conditions.
Innovation Solution
The method selects reference fragments from audio records with formant trajectories of at least three formant frequencies, compares these fragments based on matching formant frequencies, and uses formant vectors calculated over fixed time intervals to ensure noise resistance and reliable identification, even in environments with high noise and distortion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional speaker recognition methods use identical verbal content comparison, then identification accuracy is improved for clear audio records, but the method becomes inapplicable when speakers speak different languages or with different verbal content
Solution Approach 1:
The patent transforms the identification approach by changing the parameters used for comparison from verbal content (words, sentences) to acoustic parameters (formant frequencies F1, F2, F3, F4). This allows comparison of speakers regardless of what they say, enabling cross-language and different verbal content identification while maintaining accuracy through formant frequency matching
Solution Approach 2:
The patent segments the speech signal into reference fragments containing formant trajectories of at least three formant frequencies. By dividing the continuous speech into discrete formant-based segments, the system can identify speakers through acoustic characteristics independent of verbal content, resolving the contradiction between accuracy and versatility
2Duration of action of moving object
If speaker recognition methods process short audio records, then the method becomes applicable for brief speech samples, but identification reliability deteriorates due to insufficient individualizing features
Solution Approach 1:
The patent changes the extraction parameters from requiring multiple speech features (tone, rhythm, intonation) to focusing on formant frequencies (F1, F2, F3, F4) which are present even in short audio fragments. This allows reliable identification from brief recordings by concentrating on the most stable and extractable acoustic parameters
Solution Approach 2:
The patent performs preliminary selection of reference fragments containing formant trajectories before comparison. By pre-processing the audio to identify and isolate fragments with sufficient formant information, the system ensures that even short audio records provide adequate data for reliable identification
3Device complexity
If traditional methods compare speech signals directly, then the process is simple, but the method fails in high noise and distortion environments
Solution Approach 1:
The patent introduces formant frequency analysis as an intermediary layer between the raw speech signal and the comparison process. By extracting formant trajectories (F1, F2, F3, F4) as intermediate features, the system creates noise-resistant representations that filter out harmful environmental factors while preserving speaker-specific characteristics
Solution Approach 2:
The patent replaces direct mechanical signal comparison with acoustic parameter analysis. Instead of comparing raw waveforms directly, the system substitutes this with formant frequency measurement and matching, which are inherently more resistant to noise and distortion while maintaining identification accuracy
4Adaptability or versatility
If speaker identification accounts for different psycho-physiological states, then the method becomes more versatile, but the complexity of accommodating varying speech patterns increases
Solution Approach 1:
The patent changes the basis of comparison from speech content (which varies with psycho-physiological state) to formant frequencies (which remain relatively stable). By focusing on F1, F2, F3, F4 frequencies rather than words or intonation patterns, the system achieves versatility across different speaker states without increasing complexity
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a method for the reliable identification of a speaker based on long or short phonograms, phonograms recorded in different channels with high interference and distortion levels, and phonograms including random speech of narrators under various psycho-physiological conditions and speaking in different languages. The present method can be used in a wide range of applications, including in criminal investigations. The identification of a speaker based on speech phonograms is made by estimating the convergence between a first phonogram of the speaker and a second calibration phonogram. In order to carry out the estimation, the method comprises selecting on the first and second phonograms reference fragments of speech signals containing formant trajectories of at least three formants, comparing the reference fragments in which the values of at least two formant frequencies coincide, estimating the convergence of the compared reference fragments based on the coincidence of the values of the remaining formant frequencies, and determining the convergence of all the phonograms based on the global estimation of the convergence of all the compared reference fragments.