Speech Feature Extraction Using Voice Activity Detection Posteriors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker recognition systems, both ivector-based and DNN-based, rely on equally weighted pooling which fails to accurately represent utterances due to unequal importance of frames, leading to degraded recognition performance.

Innovation Solution

A speech feature extraction apparatus and method that employs voice activity detection to drop non-voice frames and calculate posterior weights for frame-level features, enabling more accurate utterance-level representation by assigning greater importance to frames with higher voice activity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If equally weighted pooling is used to extract utterance-level features from frame-level features, then the system is simple and computationally efficient, but the speaker recognition performance is degraded because all frames are treated equally regardless of their actual information content

Engineering Contradiction:
Improvespeaker recognition performanceVSAvoidfeature extraction complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by assigning different weights to different frames based on their voice activity detection posteriors. Instead of uniform weighting, each frame receives a weight proportional to its likelihood of containing actual speech, thereby locally optimizing the contribution of each frame to the utterance-level feature representation.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the weighting parameter from uniform (equally weighted) to variable (posterior-based weights). By computing VAD posteriors for each frame and using these as weights in the pooling process, the system dynamically adjusts the importance of each frame based on its acoustic characteristics, improving speaker recognition performance.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If all frames are included in the pooling process, then the computation is straightforward, but non-voice frames (silence, noise) dilute the speaker information and reduce recognition accuracy

Engineering Contradiction:
Improveutterance representation accuracyVSAvoidprocessing time for frame selection
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary voice activity detection and computes posterior probabilities for each frame before the pooling process. This preliminary action identifies and flags non-voice frames, allowing the subsequent pooling operation to either exclude or down-weight these frames, thereby preserving speaker information while maintaining computational efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts and removes non-voice frames from the pooling process by using VAD posteriors to identify frames that do not contain actual speech. These frames are either excluded entirely or given minimal weight, effectively taking them out of the feature aggregation process to prevent them from diluting the speaker information.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11580967B2Speech feature extraction apparatus, speech feature extraction method, and computer-readable storage medium
Publication Date: 2023.02.14 NEC CORP
  • US11580967B2 patent drawing
  • US11580967B2 patent drawing
  • US11580967B2 patent drawing

AI summary

A speech feature extraction apparatus 100 includes a voice activity detection unit 103 that drops non-voice frames from frames corresponding to an input speech utterance, and calculates a posterior of being voiced for each frame, a voice activity detection process unit 106 calculates a function value as weights in pooling frames to produce an utterance-level feature, from a given a voice activity detection posterior, and an utterance-level feature extraction unit 112 that extracts an utterance-level feature, from the frame on a basis of multiple frame-level features, using the function values.