Voice Activity Detection Parameter Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems face errors in voice activity detection due to the use of general average parameter values, which can lead to inaccuracies in distinguishing between utterance and non-utterance sections, particularly for users with low or high utterance rates, resulting in inefficiencies and potential loss of user input.
Innovation Solution
An electronic device that identifies user-specific utterance characteristics, such as utterance rate and energy level, to dynamically adjust parameters for voice activity detection, including energy thresholds and hangover times, based on these characteristics to improve detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If general average parameter values are used for voice activity detection, then the system can operate with simple fixed parameters, but detection accuracy deteriorates for users with low or high utterance rates
Solution Approach 1:
The patent implements dynamic parameter adjustment by transitioning from fixed average parameters to user-specific adaptive parameters. The system dynamically modifies voice activity detection parameters based on individually learned utterance characteristics, allowing parameters to adapt in real-time to each user's speech patterns while maintaining system reliability across diverse users with varying utterance rates
Solution Approach 2:
The patent applies parameter changes by modifying voice activity detection parameters according to user-specific utterance characteristics. The system learns individual parameters such as utterance rate, energy levels, and spectral features, then adjusts detection thresholds and timing parameters accordingly to improve accuracy for each user without requiring complex manual configuration
2Measurement precision
If user-specific parameters are learned and applied, then detection accuracy improves for individual users, but the system complexity and processing time increase
Solution Approach 1:
The patent implements preliminary action by performing parameter learning during initial system setup or calibration phases before actual voice recognition tasks. The system pre-learns user-specific utterance characteristics and stores these parameters for rapid deployment, eliminating the need for time-consuming real-time analysis during critical voice recognition operations
Solution Approach 2:
The patent applies this principle by using lightweight, computationally efficient parameter representations that can be quickly processed and updated. The system employs simplified models for capturing utterance characteristics that require minimal processing resources, allowing rapid parameter adaptation without heavy computational overhead or extended learning periods
Data Source
AI summary
An electronic device includes a memory storing one or more instructions; and a processor configured to execute the one or more instructions stored in the memory to receive audio data corresponding to a user's utterance, to identify the user's utterance characteristics based on the received audio data, to determine a parameter for performing voice activity detection, by using the identified user's utterance characteristics, and to perform voice activity detection on the received audio data with respect to the user's utterance by using the determined parameter.


