Voice-Aware Headset Audio With Adaptive Noise-Aware VAD
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice activity detection (VAD) algorithms perform poorly in noisy environments, leading to false detections and difficulty in setting internal parameters for noise reduction, which affects audio quality and VAD performance.
Innovation Solution
A voice aware audio system with a microphone array and noise reduction module that uses beamforming, fractional delay processing, and an improved VAD algorithm to estimate noise levels and adapt noise reduction parameters, allowing users to be aware of external sounds while listening to music without disrupting the audio experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional noise reduction modules are used with fixed internal parameters, then the system can operate in specific noise conditions, but it cannot adapt to different noise levels and categories, causing distortions in silent and quiet environments
Solution Approach 1:
The patent implements dynamic adaptation of noise reduction parameters by continuously estimating noise levels and adjusting the noise reduction module's internal parameters accordingly. The system transitions from static fixed parameters to dynamic adaptive parameters that change based on real-time noise conditions, resolving the contradiction between adaptability and audio quality
Solution Approach 2:
The system changes the operational parameters of the noise reduction module based on estimated noise levels. By monitoring noise characteristics and adjusting parameters such as reduction strength and filtering characteristics, the system adapts to different noise categories without causing distortions in quiet environments
2Object-affected harmful factors
If noise reduction is applied to preprocess speech signals, then noise is reduced, but musical noise appears which misleads the VAD module and creates false detections
Solution Approach 1:
The patent introduces an intermediary noise estimation module that analyzes noise characteristics before they affect the VAD decision process. This intermediary component provides accurate noise level information to the VAD module, allowing it to distinguish between actual speech and musical noise artifacts, thereby maintaining detection reliability
Solution Approach 2:
The system implements feedback mechanisms where the estimated noise level and detection results are continuously monitored and used to adjust the noise reduction parameters. This closed-loop feedback prevents musical noise from misleading the VAD module by dynamically optimizing the noise reduction strength based on actual speech presence
3Measurement precision
If frame lengths of 10-40 ms are used for VAD processing, then speech can be considered statistically stationary, but the system struggles to detect speech in noisy environments with poor detection scores
Solution Approach 1:
The patent extends the analysis from single-frame features to multi-frame temporal patterns. By considering sequences of frames and their temporal relationships, the system maintains the statistical stationarity assumption within each frame while gaining robustness against noise through temporal context, improving detection scores in noisy environments
Data Source
AI summary
A voice aware audio system and a method for a user wearing a headset to be aware of an outer sound environment while listening to music or any other audio source. An adjustable sound awareness zone gives the user the flexibility to avoid hearing far distant voices. The outer sound can be analyzed in a frequency domain to select an oscillating frequency candidate and in a time domain to determine if the oscillating frequency candidate is the signal of interest. If the signal directed to the outer sound is determined to be a signal of interest the outer sound is mixed with audio from the audio source.


