Accelerometer Earphone Voice Activity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Consumer electronic devices, such as mobile phones and computers, face challenges in voice communication due to environmental noise interference when using speakerphone or headset modes, rendering user speech unintelligible.

Innovation Solution

The use of an accelerometer in earbuds or earphones to detect vocal chord vibrations combined with microphone signals from a headset or mobile device, enabling a voice activity detector (VAD) to differentiate between voiced and unvoiced speech, and a noise suppressor to filter out environmental noise by steering beamformers to emphasize user speech and deemphasize background noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speakerphone mode or wired headset is used to receive speech, then hands-free operation is enabled, but environmental noise such as secondary speakers and background noises degrades voice communication quality

Engineering Contradiction:
Improvehands-free operationVSAvoidvoice communication quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent divides the audio signal processing into multiple independent components: accelerometer-based vibration detection, microphone-based acoustic detection, beamforming for spatial filtering, and noise suppression modules. Each component processes specific aspects of the signal separately before combining results, allowing targeted optimization of speech detection while filtering environmental noise.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an accelerometer as an intermediary sensor that detects bone conduction vibrations from the user's vocal cords. This intermediary measurement pathway provides a reference signal that is largely immune to environmental noise, which then serves as a basis for enhancing the microphone signals through adaptive processing and beamforming.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If microphone port or headset is used to capture speech, then voice communication is enabled, but environmental noise renders user speech unintelligible

Engineering Contradiction:
Improvespeech capture capabilityVSAvoidenvironmental noise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent replaces reliance on purely acoustic mechanical systems (microphones) with a hybrid system that incorporates inertial sensing (accelerometers). The accelerometer detects mechanical vibrations from bone conduction, providing a non-acoustic measurement pathway that is insensitive to airborne environmental noise, thereby substituting the vulnerable acoustic-only approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent applies different processing strategies to different signal components: the accelerometer signal undergoes vibration analysis for voiced speech detection, while microphone signals receive beamforming and noise suppression processing. Each signal pathway is optimized for its specific characteristics, with the accelerometer pathway focusing on low-frequency vibration patterns and the microphone pathway handling full-band acoustic signals.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If accelerometer signals are used to detect voiced speech, then detection accuracy is improved, but device complexity increases due to multiple sensors and processing requirements

Engineering Contradiction:
Improvevoiced speech detection accuracyVSAvoidmultiple sensors and processing components
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the accelerometer serve multiple functions: it detects voiced speech through vibration analysis, provides a reference signal for adaptive noise cancellation, and enables speaker verification. The microphone array also performs dual roles in capturing acoustic signals and providing spatial information for beamforming. This multi-functionality reduces the need for additional dedicated components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the processing of accelerometer and microphone signals into a unified framework where both signal types are processed simultaneously through adaptive algorithms. The beamforming operation combines spatial filtering of microphone signals with vibration-based speech detection from the accelerometer, creating an integrated processing pipeline that reduces overall system complexity compared to separate independent systems.

Inventive Principle:
Principle #5Merging (Combining)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution effectively enhances voice communication quality by accurately detecting user speech and suppressing environmental noise, improving the intelligibility of voice communications in noisy environments.

Implementation Method 1

the accelerometer may detect speech caused by the vibrations of the user's vocal chords

Methodology Applied
Scientific EffectVibration: Vibration

Data Source

PatentUS9313572B2System and method of detecting a user's voice activity using an accelerometer
Publication Date: 2016.04.12 APPLE INC
  • US9313572B2 patent drawing
  • US9313572B2 patent drawing
  • US9313572B2 patent drawing

AI summary

A method of detecting a user's voice activity in a mobile device is described herein. The method starts with a voice activity detector (VAD) generating a VAD output based on (i) acoustic signals received from microphones included in the mobile device and (ii) data output by an inertial sensor that is included in an earphone portion of the mobile device. The inertial sensor may detect vibration of the user's vocal chords modulated by the user's vocal tract based on vibrations in bones and tissue of the user's head. A noise suppressor may then receive the acoustic signals from the microphones and the VAD output and suppress the noise included in the acoustic signals received from the microphones based on the VAD output. The method may also include steering one or more beamformers based on the VAD output. Other embodiments are also described.