Headphone Conversation Detection Using Dual Voice Activity Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital audio processing techniques for headphones struggle to accurately activate and deactivate transparency mode during conversations in noisy environments, leading to undesirable ambient sounds and increased power consumption and distortion.
Innovation Solution
A conversation detector system that uses a combination of microphone and sensor signals to determine when to activate or deactivate the transparency mode, employing voice activity detectors and machine learning models to isolate and enhance the speech of the other talker, while preventing false triggers and optimizing power usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If transparency mode is activated to reproduce ambient sound during conversations, then the wearer can hear and understand the other person better, but undesirable ambient sounds may be reproduced when the mode is activated at the wrong time
Solution Approach 1:
The system performs preliminary voice activity detection and speaker verification before activating transparency mode. The OVAD and TVAD detectors analyze microphone signals in advance to determine whether a conversation is occurring, and the speaker verification process confirms the identity of the other talker before mode activation, preventing premature or incorrect activation
Solution Approach 2:
The system continuously monitors microphone signals and sensor data to provide feedback on conversation status. The conversation detector uses ongoing voice activity detection and speaker verification results to dynamically adjust transparency mode activation, ensuring the mode remains active only when appropriate and deactivating when conversations end or false conditions are detected
2Reliability
If transparency mode is continuously active to ensure conversation clarity, then the wearer can always hear the other person clearly, but power consumption increases
Solution Approach 1:
Instead of continuous transparency mode operation, the system uses periodic voice activity detection and conversation monitoring. The OVAD and TVAD detectors periodically analyze signals to determine conversation status, and transparency mode is activated only during detected conversation periods, creating an on-demand operation pattern that reduces overall power consumption while maintaining conversation clarity when needed
Solution Approach 2:
The system dynamically adjusts transparency mode activation based on real-time conversation detection results. The mode transitions between active and inactive states according to detected conversation presence, voice activity levels, and speaker verification outcomes, optimizing power consumption by activating the computationally intensive transparency processing only when conversations are detected
3Illumination intensity
If transparency mode is activated during all ambient sound conditions, then the wearer can hear all sounds clearly, but distortion increases when the mode is activated in unsuitable situations
Solution Approach 1:
The system applies different processing qualities to different audio conditions. Voice activity detection and speaker verification are applied selectively to determine when high-quality transparency processing is appropriate. The OVAD and TVAD detectors identify specific local conditions (conversation presence, other talker verification) that trigger enhanced processing, while other ambient sound conditions receive different or reduced processing to avoid distortion
Solution Approach 2:
The system performs preliminary analysis of audio conditions through voice activity detection and speaker verification before applying transparency mode processing. This preliminary action identifies suitable conditions for transparency activation, ensuring that the computationally intensive processing is applied only when it will produce clear audio without distortion, rather than applying it universally to all ambient sound conditions
Data Source
AI summary
A conversation detector processes microphone signals and other sensor signals of a headphone to declare a conversation and configures a filter block to activate a transparency audio signal. It then declares an end to the conversation based on processing one or more of the microphone signals and the other sensor signals, and in response deactivates the transparency audio signal. The conversation detector monitors an idle duration in which an OVAD and a TVAD are both or simultaneously indicating no activity and declares the end to the conversation in response to the idle duration being longer than an idle threshold. Other aspects are also described and claimed.


