Wearable Self-Voice Separation for Natural Listen-Through Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Wearable devices struggle to provide a natural listening experience due to the inability to seamlessly distinguish between self-voice signals and external signals, resulting in distorted audio input when using the listen-through feature, as both types of signals have different distortion patterns and are not effectively separated by existing microphones.

Innovation Solution

The implementation of a multi-microphone speech generative network and beamforming techniques to separate self-voice signals from external signals, applying distinct filters to each and remixing them to generate a natural-sounding output audio signal, while also utilizing the audio zoom feature to focus sound pickup in a desired direction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If microphones pick up external audio signals including background noise and self-voice signals, then the device can capture all audio input, but the different distortion patterns from self-voice and external signals result in unnatural sounding audio output

Engineering Contradiction:
Improveaudio signal handling capabilityVSAvoidaudio output quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the audio signal processing by separating self-voice signals from external background signals using a speech generative network. The system divides the mixed audio input into distinct signal components, applying different processing paths to each type of signal to resolve the distortion pattern conflict and achieve natural-sounding output.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by using different filtering approaches for different signal types. Self-voice signals receive one type of processing while external background signals receive different processing, allowing each signal type to be optimized for its specific characteristics and distortion patterns.

Inventive Principle:
Principle #3Local quality

2Device complexity

If the device applies the same filtering to all audio signals, then the processing is simple, but self-voice and external signals have different distortion patterns that require different filtering approaches

Engineering Contradiction:
Improvesignal processing complexityVSAvoidsignal separation accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The speech generative network segments the audio processing task by identifying and separating self-voice signals from external signals. This segmentation enables the system to apply appropriate filtering to each signal type, improving separation accuracy while managing complexity through automated signal classification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes processing parameters dynamically based on signal type. The system adjusts filtering parameters and processing characteristics according to whether the detected signal is self-voice or external background, allowing optimal processing for each signal category without requiring manual configuration.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12069425B2Separation of self-voice signal from a background signal using a speech generative network on a wearable device
Publication Date: 2024.08.20 QUALCOMM INC
  • US12069425B2 patent drawing
  • US12069425B2 patent drawing
  • US12069425B2 patent drawing

AI summary

A wearable device may include a processor configured to detect a self-voice signal, based on one or more transducers. The processor may be configured to separate the self-voice signal from a background signal in an external audio signal based on using a multi-microphone speech generative network. The processor may also be configured to apply a first filter to an external audio signal, detected by at least one external microphone on the wearable device, during a listen through operation based on an activation of the audio zoom feature to generate a first listen-through signal that includes the external audio signal. The processor may be configured to produce an output audio signal that is based on at least the first listen-through signal that includes the external signal, and is based on the detected self-voice signal.