Ear-Worn Audio Spatial Focusing for Multi-Speaker Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional ear-worn devices face challenges in noise reduction, particularly in scenarios with multiple speakers, due to warped beamforming patterns caused by the wearer's head and torso, limited sound reduction capabilities, and performance issues in reverberant environments, with beamforming being more effective for high-frequency sounds and adding noise in quiet environments.
Innovation Solution
Implementing neural networks for spatial focusing in ear-worn devices, which apply different weights to audio signals based on sound source locations, using multiple microphones to distinguish between target and interfering speakers, and independently controlling background noise and interfering speech volumes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If conventional beamforming is used to reduce noise and interfering speakers, then sound reduction from certain directions is improved, but the beamforming pattern becomes warped due to interference from the head, torso, and ear
Solution Approach 1:
The patent replaces conventional mechanical/acoustic beamforming with a neural network-based spatial focusing system. The neural network learns optimal filtering operations from training data, substituting the rigid mathematical beamforming patterns with adaptive, data-driven filters that are not constrained by the warped acoustic patterns caused by head and torso interference.
Solution Approach 2:
The patent changes the approach from fixed beamforming patterns to dynamic, learnable filtering parameters. The neural network adjusts its internal parameters (weights and biases) based on training data, allowing the system to adapt to different acoustic environments and speaker configurations rather than relying on predetermined beamforming patterns that become warped.
2Measurement precision
If conventional beamforming is used, then high-frequency sound localization is improved, but low-frequency sound reduction is limited
Solution Approach 1:
The neural network is trained to handle multiple frequency ranges simultaneously by using spectrogram inputs that capture both low and high-frequency information. The network learns to apply appropriate filtering at different frequencies through its trained weights, overcoming the inherent limitation of conventional beamforming that works better for high frequencies.
3Measurement precision
If conventional beamforming is used in reverberant environments, then direct sound focusing is improved, but indirect reverberant paths from the front are not attenuated
Solution Approach 1:
The neural network substitutes the directional-based beamforming approach with a data-driven approach that learns to identify and filter reverberant components. By training on data that includes reverberant environments, the network learns temporal and spectral patterns characteristic of reverberation and can attenuate them even when they arrive from the front direction.
4Object-affected harmful factors
If conventional beamforming is used in quiet environments, then noise reduction is improved, but additional noise is added to the output
Solution Approach 1:
The neural network applies filtering operations that are adapted to the specific acoustic conditions. In quiet environments, the trained network can modulate its filtering strength to avoid excessive processing that would introduce artifacts, whereas conventional beamforming applies fixed patterns regardless of environmental conditions.
Data Source
AI summary
An ear-worn device includes two or more microphones and noise reduction circuitry including neural network circuitry. The neural network circuitry is configured to: receive multiple audio signals wherein at least two of the multiple audio signals each originate from a different one of the two or more microphones and/or at least one of the multiple audio signals is a beamformed audio signal originating from the two or more microphones; and implement one or more neural network layers trained to perform background noise modification and spatial focusing based on the multiple audio signals, such that the neural network circuitry generates, based on the multiple audio signals, one or more neural network outputs. The noise reduction circuitry is configured to output, based on the one or more neural network outputs, an output audio signal comprising a background noise-modified and spatially-focused version of a first audio signal of the multiple audio signals.


