Voice Activity Detector Dynamic Threshold Adjustment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice activity detection systems face accuracy issues when determining active and non-active voice segments due to unsuitable parameter settings for noise conditions and recording environments.
Innovation Solution
A voice activity detector that adjusts duration thresholds based on labeled segment numbers to improve judgment accuracy by shaping judgment results and updating parameters to minimize differences between calculated and labeled segment counts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If fixed threshold values are used for judgment conditions, then the device complexity is reduced, but the measurement precision deteriorates under different noise and recording conditions
Solution Approach 1:
The patent transforms fixed threshold values into dynamic parameters that are automatically adjusted based on the characteristics of the input signal. The judgment threshold is no longer a static value but adapts to different noise conditions and recording environments, allowing the system to maintain high measurement precision across varying conditions without requiring complex manual parameter configuration.
Solution Approach 2:
The system performs self-adjustment of parameters based on the input signal characteristics. By automatically determining appropriate threshold values from the signal itself, the system eliminates the need for external parameter tuning while maintaining optimal segment determination accuracy for each specific recording condition.
2Measurement precision
If adaptive parameter adjustment is implemented, then the measurement precision improves, but the device complexity increases
Solution Approach 1:
The patent employs feedback mechanisms where the system continuously monitors signal characteristics and adjusts judgment thresholds accordingly. The feedback loop processes signal statistics and automatically modifies parameters to optimize segment determination, achieving high measurement precision through an automated feedback-driven adjustment process rather than complex manual control.
Data Source
AI summary
Judgment result deriving means 74 makes a judgment between active voice and non-active voice every unit time for a time series of voice data in which the number of active voice segments and the number of non-active voice segments are already known as a number of the labeled active voice segment and a number of the labeled non-active voice segment and shapes active voice segments and non-active voice segments as the result of the judgment by comparing the length of each segment during which the voice data is consecutively judged to correspond to active voice by the judgment or the length of each segment during which the voice data is consecutively judged to correspond to non-active voice by the judgment with a duration threshold. Segments number calculating means 75 calculates the number of active voice segments and the number of non-active voice segments. Duration threshold updating means 76 updates the duration threshold so that the difference between the calculated number of active voice segments and the number of the labeled active voice segments decreases or the difference between the calculated number of non-active voice segments and the number of the labeled non-active voice segments decreases.


