Headset Voice Activity Detection Using Accelerometer Vibration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Consumer electronic devices using microphone ports or headsets struggle with capturing user speech due to environmental noise interference, degrading voice communication quality.

Innovation Solution

An enhanced headset with earbuds and a microphone array employs an accelerometer to detect vocal chord vibrations and acoustic signals, using a voice activity detector (VAD) to differentiate between voiced and unvoiced speech, and steers beamformers to emphasize user speech and suppress environmental noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speakerphone mode or wired headset is used for hands-free operation, then user convenience is improved, but environmental noise interference increases

Engineering Contradiction:
Improvehands-free operation convenienceVSAvoidenvironmental noise interference
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system segments the speech detection function into multiple independent components: accelerometer-based vibration detection for voiced speech, microphone-based acoustic detection for unvoiced speech, and beamforming for spatial noise rejection. Each component operates independently and their outputs are combined to achieve robust speech detection in noisy environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an accelerometer as an intermediary sensor placed in the earbud to detect vocal chord vibrations directly. This intermediary detection method provides a new pathway for speech detection that is independent of acoustic microphones, thereby avoiding environmental noise contamination while maintaining speech detection capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If microphones are used to capture user speech, then speech detection capability is improved, but speech intelligibility deteriorates due to environmental noise

Engineering Contradiction:
Improvespeech detection capabilityVSAvoidspeech intelligibility
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system merges multiple speech detection pathways: accelerometer-based vibration detection, microphone-based acoustic detection, and beamforming-based spatial filtering. By combining these complementary methods, the system achieves both high speech detection precision and maintains speech intelligibility through noise suppression.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system dynamically switches between different detection modes based on speech characteristics. For voiced speech, it relies on accelerometer detection; for unvoiced speech, it uses microphone detection with high-pass filtering. This dynamic adaptation optimizes speech detection accuracy across different speech types while maintaining intelligibility.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If accelerometer signals are used to detect voiced speech, then voiced speech detection accuracy is improved, but unvoiced speech detection capability is lost

Engineering Contradiction:
Improvevoiced speech detection accuracyVSAvoidunvoiced speech detection capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system implements a universal speech detection framework that handles both voiced and unvoiced speech through multiple pathways. The accelerometer serves voiced speech detection while microphones with appropriate processing handle unvoiced speech, creating a multi-functional detection system that adapts to different speech types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically selects the appropriate detection method based on speech type. Voiced speech triggers accelerometer-based detection, while unvoiced speech activates microphone-based detection with high-pass filtering. This dynamic switching ensures both detection accuracy for voiced speech and capability for unvoiced speech.

Inventive Principle:
Principle #15Dynamics

4Measurement precision

If beamformers are steered to emphasize user speech, then speech signal quality is improved, but system complexity increases

Engineering Contradiction:
Improvespeech signal qualityVSAvoidbeamforming system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary speech activity detection using the accelerometer before activating complex beamforming operations. This preliminary detection allows the system to prepare beamformer steering in advance, reducing computational complexity by only performing full beamforming when speech is detected, rather than continuously.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Effectively isolates user speech from environmental noise, enhancing voice communication quality by accurately detecting voice activity and adaptively steering beamformers to minimize noise contamination.

Implementation Method 1

the accelerometer may detect speech caused by the vibrations of the user's vocal chords

Methodology Applied
Scientific EffectVibration: Vibration

Implementation Method 2

data output by a sensor detecting movement that is included in the pair of earbuds

Methodology Applied
Scientific EffectAccelerometer: Accelerometer

Data Source

PatentUS9438985B2System and method of detecting a user's voice activity using an accelerometer
Publication Date: 2016.09.06 APPLE INC
  • US9438985B2 patent drawing
  • US9438985B2 patent drawing
  • US9438985B2 patent drawing

AI summary

A method of detecting a user's voice activity in a headset with a microphone array is described herein. The method starts with a voice activity detector (VAD) generating a VAD output based on acoustic signals received from microphones included in a pair of earbuds and the microphone array included on a headset wire and data output by an accelerometer that is included in the pair of earbuds. A noise suppressor may then receive the acoustic signals from the microphone array and the VAD output and suppress the noise included in the acoustic signals received from the microphone array based on the VAD output. The method may also include steering one or more beamformers based on the VAD output. Other embodiments are also described.