Voice Activity Segmentation via Self-Adjusting Thresholds
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice activity segmentation systems burden users with the need to utter reference speech and struggle with inaccurate parameter updates due to insufficient utterance amounts, leading to inefficiencies in noise robustness.
Innovation Solution
A voice activity segmentation device and method that determines voice-active and voice-inactive segments by comparing feature values with threshold values, with a secondary segmentation step using superimposed reference speech to adjust threshold values based on discrepancy rates, allowing for robust noise resistance without user burden.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a general voice activity segmentation system uses fixed threshold values for determining voice-active and voice-inactive segments, then the system is simple to operate, but the segmentation accuracy is insufficient in noisy environments
Solution Approach 1:
The system automatically updates threshold values using voice-inactive segment data without requiring user participation. The threshold value update means performs self-adjustment by calculating discrepancy rates between segmented results and actual voice activity patterns, eliminating the need for users to utter reference speech while improving segmentation accuracy in noisy environments
Solution Approach 2:
The system performs preliminary threshold value updates using voice-inactive segment data before actual voice activity detection. By pre-adjusting threshold values based on environmental noise characteristics captured during inactive periods, the system prepares optimal segmentation parameters in advance, improving accuracy without burdening users during actual speech input
2Measurement precision
If the system updates parameters using user utterances, then the segmentation accuracy improves, but the user burden increases due to requiring reference speech
Solution Approach 1:
The system performs parameter updates autonomously using voice-inactive segment data from the environment. The threshold value update means calculates discrepancy rates and adjusts parameters without requiring any user utterance, making the system self-configuring while maintaining high accuracy in adapting to noisy environments
Solution Approach 2:
The system uses voice-inactive segment data as an intermediary resource for parameter updates. Instead of directly using user reference speech, it leverages environmental noise periods as intermediate training data to indirectly improve segmentation parameters, achieving accurate adaptation without user burden
3Measurement precision
If the system uses insufficient utterance data for parameter updates, then the operation remains simple, but the parameter update accuracy becomes insufficient
Solution Approach 1:
The system accumulates voice-inactive segment data in advance during periods when no speech is detected. By gathering environmental noise characteristics over time before actual voice activity detection, the system builds sufficient training data reservoir without requiring extensive user utterances, enabling accurate parameter updates with minimal user input
4Reliability
If the system does not update threshold values adaptively, then the operation is simple, but the noise robustness decreases
Solution Approach 1:
The system implements feedback mechanisms where the threshold value update means continuously monitors discrepancy rates between segmented results and actual voice activity patterns. Based on this feedback, the system automatically adjusts threshold values to improve noise robustness, creating a self-improving system that adapts to changing environmental conditions
Solution Approach 2:
The system transitions from static fixed threshold values to dynamic adaptive threshold values that automatically adjust based on environmental noise characteristics. The threshold values become flexible parameters that evolve over time through automatic updates, enhancing noise robustness while maintaining operational simplicity
Data Source
AI summary
Provided is a noise-robust voice activity segmentation device which updates parameters used in the determination of voice-active segments without burdening the user, and also provided are a voice activity segmentation method and a voice activity segmentation program.The voice activity segmentation device comprises: a first voice activity segmentation means for determining a voice-active segment (first voice-active segment) and a voice-inactive segment (first voice-inactive segment) in a time-series of input sound by comparing a threshold value and a feature value of the time-series of the input sound; a second voice activity segmentation means for determining, after a reference speech acquired from a reference speech storage means has been superimposed on a time-series of the first voice-inactive segment, a voice-active segment and a voice-inactive segment in the time-series of the superimposed first voice-inactive segment by comparing the threshold value and a feature value of the time-series of the superimposed first voice-inactive segment; and a threshold value update means for updating the threshold value in such a way that a discrepancy rate between the determination result of the second voice activity segmentation means and a correct segmentation calculated from the reference speech is decreased.


