Headset Voice Activity Detection via Microphone Phase Difference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing headset technologies struggle to accurately detect a user's voice activity to prevent active noise cancellation or audio output from interfering with conversations.
Innovation Solution
The implementation of a voice activity detection system using a pair of microphones, one inner and one outer, positioned on a headset, which compares the phase difference between the signals from these microphones to determine when a user is speaking.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If active noise cancellation is continuously applied, then noise reduction effectiveness is improved, but it interferes with user conversation
Solution Approach 1:
The noise cancellation system dynamically adjusts its operation based on real-time voice activity detection. When the user speaks, the system pauses or reduces noise cancellation; when the user is silent, it resumes full noise cancellation. This dynamic switching resolves the contradiction by making the system adaptive to conversational states.
Solution Approach 2:
The system uses feedback from voice activity detection to control noise cancellation operation. The inner and outer microphones continuously monitor for user speech, and this feedback signal triggers the pause or resumption of noise cancellation, creating a closed-loop control system that prevents conversation interference while maintaining noise reduction effectiveness.
2Ease of operation
If audio output is continuously played, then user entertainment experience is improved, but it distracts from user conversation
Solution Approach 1:
The audio output system dynamically adjusts playback based on detected voice activity. When the user speaks, audio playback is paused or reduced; when the user is silent, full audio playback resumes. This dynamic behavior ensures entertainment experience is maintained while preventing distraction during conversations.
Solution Approach 2:
The system takes preliminary action by detecting voice activity before conversation interference occurs. By proactively pausing audio output upon detecting user speech, the system prevents distraction rather than reacting after the fact, maintaining both entertainment quality and conversational awareness.
3Device complexity
If single microphone is used for voice detection, then device complexity is reduced, but voice activity detection accuracy is insufficient
Solution Approach 1:
The system uses two microphones with different local qualities and positions: the inner microphone is positioned closer to the user's mouth for capturing voice signals, while the outer microphone is positioned farther away. This spatial differentiation creates a phase difference that enables accurate voice activity detection while maintaining relatively simple device architecture.
Solution Approach 2:
The system transitions from single-point microphone detection to spatial differential detection by using two microphones at different positions. The phase difference dimension added by the second microphone provides additional information for voice activity detection, improving accuracy without significantly increasing device complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution effectively detects voice activity, allowing the headset to adjust noise cancellation and audio output accordingly, thereby enhancing the user's conversational experience by minimizing distractions.
Implementation Method 1
an inner microphone and an outer microphone of a headset
Implementation Method 2
detecting a user's voice activity, according to a phase difference between an inner microphone and an outer microphone of a headset
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A headset that can detect the voice activity of a user includes an inner microphone generating an inner microphone signal; an outer microphone generating an outer microphone signal, wherein the inner microphone and outer microphone are positioned such that, when the headset is worn by a user, the inner microphone is disposed nearer to the user's head; and a voice-activity detector determining a sign of a phase difference between the inner microphone signal and the outer microphone signal and generating a voice activity detection signal representing a user's voice activity when the sign of the phase difference indicates that the outer microphone received an audio signal after the inner microphone received the audio signal.