Wearable Audio Ducking for Speech Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users of wearable audio devices face inconvenience when transitioning from a private audio experience to real-world interactions, as they must manually adjust volume or pause music to hear conversations, due to the lack of automatic ducking functionality that distinguishes user speech from ambient noise.

Innovation Solution

A wearable device equipped with microphones and processing capabilities that detect ambient noise, differentiate user speech from ambient speech, and automatically duck audio playback when user speech is detected, continuing the ducking based on ambient speech, thereby reducing the need for manual volume adjustments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If noise-cancelling functionality is provided to detect and analyze ambient noise, then the user's private audio experience is improved, but the user's ability to hear conversations in real-world interactions degrades

Engineering Contradiction:
Improveprivate audio experienceVSAvoidability to hear conversations
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system dynamically adjusts the audio playback based on real-time detection of user speech. When user speech is detected, the audio playback is automatically attenuated; when user speech is not detected, the audio playback continues at normal volume. This dynamic adaptation resolves the contradiction by making the system responsive to the user's immediate needs rather than static.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses microphones to continuously monitor ambient noise and detect user speech, then feeds this information back to the audio processing system. Based on this feedback, the system automatically adjusts audio playback in real-time, allowing the user to transition smoothly between private listening and conversation modes without manual intervention.

Inventive Principle:
Principle #23Feedback

2Reliability

If manual volume control is required to hear conversations, then the user can adjust audio levels, but the transition between private listening and real-world interactions becomes repetitive and cumbersome

Engineering Contradiction:
Improveaudio level adjustmentVSAvoidtransition convenience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system performs self-adjustment of audio playback based on automatic detection of user speech. Instead of requiring the user to manually manipulate volume controls, the system monitors ambient noise, detects speech patterns, and automatically attenuates audio playback when needed, then restores it when conversation ends. This eliminates the repetitive manual operations.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If automatic ducking is implemented based on ambient noise detection, then the transition between private listening and conversation is seamless, but the system cannot distinguish between user speech and ambient speech

Engineering Contradiction:
Improvetransition smoothnessVSAvoidspeech detection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system uses multiple microphones positioned at different locations to capture ambient noise with different spatial characteristics. By analyzing the temporal and spatial patterns of the audio signals from multiple microphones, the system can distinguish between distant ambient speech and nearby user speech, enabling more precise detection.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The audio signal processing is divided into multiple analysis stages: initial ambient noise detection, speech pattern recognition, and user speech differentiation. The system segments the audio analysis process to first detect any speech presence, then further analyze the characteristics to determine whether it is user speech or ambient speech, improving overall accuracy.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10325614B2Voice-based realtime audio attenuation
Publication Date: 2019.06.18 GOOGLE LLC
  • US10325614B2 patent drawing
  • US10325614B2 patent drawing
  • US10325614B2 patent drawing

AI summary

An example implementation may involve driving an audio output module of a wearable device with a first audio signal and then receiving, via at least one microphone of wearable device, a second audio signal comprising first ambient noise. The device may determine that the first ambient noise is indicative of user speech and responsively duck the first audio signal. While the first audio signal is ducked, the device may detect, in a subsequent portion of the second audio signal, second ambient noise, and determine that the second ambient noise is indicative of ambient speech. Responsive to the determination that the second ambient noise is indicative of ambient speech, the device may continue the ducking of the first audio signal.