Digital Signal Splitting Using Prosodic Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice biometric authentication methods require strong prior knowledge and are not efficient in utilizing prosodic features for splitting digital signals, leading to less trustworthy authentication results.
Innovation Solution
A method and system that calculate onset value locations in digital signals to identify stress accents, split the signal into prosodic unit candidate sequences, and process these sequences to include only true prosodic units, using a processor and memory to enhance authentication accuracy without relying on speech transcription systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech transcription systems are used to extract phone sequences for speaker verification, then dynamic aspects of speech can be accounted for, but strong prior knowledge is required and system complexity increases
Solution Approach 1:
The patent extracts and utilizes only the necessary prosodic features (stress accents, pitch contours, energy patterns) from speech signals for speaker verification, without requiring full speech transcription systems. This selective extraction of relevant features reduces system complexity while maintaining authentication accuracy by focusing on the most discriminative aspects of speech.
Solution Approach 2:
The patent replaces the mechanical speech transcription system with a more efficient prosodic feature extraction approach. Instead of transcribing full speech content to obtain phone sequences, the system directly extracts prosodic features using signal processing techniques, substituting a complex linguistic processing system with a more streamlined acoustic analysis method.
2Reliability
If unsupervised learning techniques are used to model known phrases, then speaker independent and dependent models can be created, but the method requires processing multiple utterances and complex model estimation
Solution Approach 1:
The patent performs preliminary extraction and storage of prosodic features from enrollment utterances, creating compact representations of stress accents, pitch contours, and energy patterns. This preliminary processing allows the system to quickly compare new utterances against stored prosodic templates without requiring complex real-time model estimation, significantly reducing authentication processing time while maintaining accuracy.
3Reliability
If text-dependent speaker verification is used with speech transcription, then phonetics can be compared against models, but the system requires strong prior knowledge and is less efficient
Solution Approach 1:
The patent extracts and utilizes only the necessary prosodic features (stress accents, pitch contours, energy patterns) from speech signals for speaker verification, without requiring full speech transcription systems. This selective extraction of relevant features reduces system complexity while maintaining authentication accuracy by focusing on the most discriminative aspects of speech.
Solution Approach 2:
The patent replaces the mechanical speech transcription system with a more efficient prosodic feature extraction approach. Instead of transcribing full speech content to obtain phone sequences, the system directly extracts prosodic features using signal processing techniques, substituting a complex linguistic processing system with a more streamlined acoustic analysis method.
Data Source
AI summary
A method for splitting a digital signal using prosodic features included in the signal is provided that includes calculating onset value locations in the signal. The onset values correspond to stress accents in the signal. Moreover, the method includes splitting, using a processor, the signal into a prosodic unit candidate sequence by superimposing the stress accent locations on the signal, and processing the sequence to include only true prosodic units.


