Multi-Microphone Ear-Worn Audio Processing with Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional hearing aids face challenges in effectively reducing noise and spatially focusing audio signals, particularly in scenarios with interfering speakers, due to limitations in beamforming patterns and neural networks that cannot accurately determine sound direction with a single microphone.
Innovation Solution
An ear-worn device employs multiple microphones and neural networks to process frequency-domain audio signals, applying spatial focusing and noise reduction by using a single neural network trained on both beamformed and non-beamformed signals to enhance audio signals based on their direction of arrival.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional beamforming patterns are used with a single microphone, then device complexity is reduced, but measurement precision of sound direction deteriorates
Solution Approach 1:
The patent combines multiple microphones (at least two) to form a microphone array that works together to determine sound direction. By merging the capabilities of multiple microphones, the system achieves accurate spatial focusing and noise reduction while maintaining reasonable device complexity through integrated processing circuitry.
Solution Approach 2:
The processing circuitry acts as an intermediary that receives signals from multiple microphones, performs beamforming operations, and generates enhanced audio outputs. This intermediary component enables the system to process multiple microphone inputs effectively without requiring complex external processing equipment.
2Reliability
If multiple frequency-domain input signals are processed through a single neural network, then noise reduction and spatial focusing are improved, but processing time increases
Solution Approach 1:
The patent performs Short-Time Fourier Transformation (STFT) to convert time-domain signals to frequency-domain signals before neural network processing. This preliminary transformation enables the neural network to operate on frequency components separately, improving noise reduction effectiveness while the efficient architecture maintains acceptable processing speeds for real-time applications.
Solution Approach 2:
The system transforms audio signals from the time domain to the frequency domain, changing the parameter space in which processing occurs. This parameter change allows the neural network to more effectively identify and suppress noise components across different frequency bands while maintaining temporal coherence through the STFT framework.
3Measurement precision
If beamforming circuitry is used to generate frequency-domain beamformed signals, then spatial focusing is improved, but device complexity increases
Solution Approach 1:
The processing circuitry is designed to perform multiple functions: it conducts STFT on microphone inputs, applies beamforming algorithms to generate frequency-domain beamformed signals, and feeds these signals to the neural network. This multi-functional design achieves high spatial focusing accuracy while containing device complexity by integrating all processing functions into a unified circuit architecture.
Data Source
AI summary
An ear-worn device may include two or more microphones configured to generate time-domain audio signals, each of the two or more microphones configured to generate one of the time-domain audio signals; processing circuitry comprising analog processing circuitry, digital processing circuitry, beamforming circuitry, and short-time Fourier transformation (STFT) circuitry, the processing circuitry configured to generate, from the time-domain audio signals, one or more frequency-domain non-beamformed audio signals and one or more frequency-domain beamformed signals; and enhancement circuitry comprising neural network circuitry configured to receive multiple frequency-domain input audio signals originating from the one or more frequency-domain non-beamformed audio signals and the one or more frequency-domain beamformed signals, and implement a single neural network trained to generate, based on the multiple frequency-domain input audio signals, a noise-reduced and spatially-focused output audio signal or an output for generating a noise-reduced and spatially-focused output audio signal.


