Speech Feature Extraction Using Voice Activity Detection Posteriors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker recognition systems, both ivector-based and DNN-based, rely on equally weighted pooling which fails to accurately represent utterances due to unequal importance of frames, leading to degraded recognition performance.
Innovation Solution
A speech feature extraction apparatus and method that employs voice activity detection to drop non-voice frames and calculate posterior weights for frame-level features, enabling more accurate utterance-level representation by assigning greater importance to frames with higher voice activity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If equally weighted pooling is used to extract utterance-level features from frame-level features, then the system is simple and computationally efficient, but the speaker recognition performance is degraded because all frames are treated equally regardless of their actual information content
Solution Approach 1:
The patent applies local quality by assigning different weights to different frames based on their voice activity detection posteriors. Instead of uniform weighting, each frame receives a weight proportional to its likelihood of containing actual speech, thereby locally optimizing the contribution of each frame to the utterance-level feature representation.
Solution Approach 2:
The patent changes the weighting parameter from uniform (equally weighted) to variable (posterior-based weights). By computing VAD posteriors for each frame and using these as weights in the pooling process, the system dynamically adjusts the importance of each frame based on its acoustic characteristics, improving speaker recognition performance.
2Measurement precision
If all frames are included in the pooling process, then the computation is straightforward, but non-voice frames (silence, noise) dilute the speaker information and reduce recognition accuracy
Solution Approach 1:
The patent performs preliminary voice activity detection and computes posterior probabilities for each frame before the pooling process. This preliminary action identifies and flags non-voice frames, allowing the subsequent pooling operation to either exclude or down-weight these frames, thereby preserving speaker information while maintaining computational efficiency.
Solution Approach 2:
The patent extracts and removes non-voice frames from the pooling process by using VAD posteriors to identify frames that do not contain actual speech. These frames are either excluded entirely or given minimal weight, effectively taking them out of the feature aggregation process to prevent them from diluting the speaker information.
Data Source
AI summary
A speech feature extraction apparatus 100 includes a voice activity detection unit 103 that drops non-voice frames from frames corresponding to an input speech utterance, and calculates a posterior of being voiced for each frame, a voice activity detection process unit 106 calculates a function value as weights in pooling frames to produce an utterance-level feature, from a given a voice activity detection posterior, and an utterance-level feature extraction unit 112 that extracts an utterance-level feature, from the frame on a basis of multiple frame-level features, using the function values.


