Voice Activity Segmentation via Self-Adjusting Thresholds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice activity segmentation systems burden users with the need to utter reference speech and struggle with inaccurate parameter updates due to insufficient utterance amounts, leading to inefficiencies in noise robustness.

Innovation Solution

A voice activity segmentation device and method that determines voice-active and voice-inactive segments by comparing feature values with threshold values, with a secondary segmentation step using superimposed reference speech to adjust threshold values based on discrepancy rates, allowing for robust noise resistance without user burden.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a general voice activity segmentation system uses fixed threshold values for determining voice-active and voice-inactive segments, then the system is simple to operate, but the segmentation accuracy is insufficient in noisy environments

Engineering Contradiction:
Improvesegmentation accuracyVSAvoiduser burden
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically updates threshold values using voice-inactive segment data without requiring user participation. The threshold value update means performs self-adjustment by calculating discrepancy rates between segmented results and actual voice activity patterns, eliminating the need for users to utter reference speech while improving segmentation accuracy in noisy environments

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary threshold value updates using voice-inactive segment data before actual voice activity detection. By pre-adjusting threshold values based on environmental noise characteristics captured during inactive periods, the system prepares optimal segmentation parameters in advance, improving accuracy without burdening users during actual speech input

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the system updates parameters using user utterances, then the segmentation accuracy improves, but the user burden increases due to requiring reference speech

Engineering Contradiction:
Improveparameter update accuracyVSAvoiduser burden
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs parameter updates autonomously using voice-inactive segment data from the environment. The threshold value update means calculates discrepancy rates and adjusts parameters without requiring any user utterance, making the system self-configuring while maintaining high accuracy in adapting to noisy environments

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses voice-inactive segment data as an intermediary resource for parameter updates. Instead of directly using user reference speech, it leverages environmental noise periods as intermediate training data to indirectly improve segmentation parameters, achieving accurate adaptation without user burden

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If the system uses insufficient utterance data for parameter updates, then the operation remains simple, but the parameter update accuracy becomes insufficient

Engineering Contradiction:
Improveparameter update accuracyVSAvoidutterance data amount
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system accumulates voice-inactive segment data in advance during periods when no speech is detected. By gathering environmental noise characteristics over time before actual voice activity detection, the system builds sufficient training data reservoir without requiring extensive user utterances, enabling accurate parameter updates with minimal user input

Inventive Principle:
Principle #10Preliminary action

4Reliability

If the system does not update threshold values adaptively, then the operation is simple, but the noise robustness decreases

Engineering Contradiction:
Improvenoise robustnessVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements feedback mechanisms where the threshold value update means continuously monitors discrepancy rates between segmented results and actual voice activity patterns. Based on this feedback, the system automatically adjusts threshold values to improve noise robustness, creating a self-improving system that adapts to changing environmental conditions

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system transitions from static fixed threshold values to dynamic adaptive threshold values that automatically adjust based on environmental noise characteristics. The threshold values become flexible parameters that evolve over time through automatic updates, enhancing noise robustness while maintaining operational simplicity

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9293131B2Voice activity segmentation device, voice activity segmentation method, and voice activity segmentation program
Publication Date: 2016.03.22 NEC CORP
  • US9293131B2 patent drawing
  • US9293131B2 patent drawing
  • US9293131B2 patent drawing

AI summary

Provided is a noise-robust voice activity segmentation device which updates parameters used in the determination of voice-active segments without burdening the user, and also provided are a voice activity segmentation method and a voice activity segmentation program.The voice activity segmentation device comprises: a first voice activity segmentation means for determining a voice-active segment (first voice-active segment) and a voice-inactive segment (first voice-inactive segment) in a time-series of input sound by comparing a threshold value and a feature value of the time-series of the input sound; a second voice activity segmentation means for determining, after a reference speech acquired from a reference speech storage means has been superimposed on a time-series of the first voice-inactive segment, a voice-active segment and a voice-inactive segment in the time-series of the superimposed first voice-inactive segment by comparing the threshold value and a feature value of the time-series of the superimposed first voice-inactive segment; and a threshold value update means for updating the threshold value in such a way that a discrepancy rate between the determination result of the second voice activity segmentation means and a correct segmentation calculated from the reference speech is decreased.