Ear-Worn Multi-Mic Neural Audio Focusing for Speaker Isolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional hearing aids face challenges in effectively reducing noise, especially in scenarios where a wearer is listening to one speaker while there are other interfering speakers, due to limitations in beamforming patterns and noise reduction techniques, which can be warped by the wearer's anatomy and perform better on high-frequency sounds than low-frequency sounds.

Innovation Solution

The development of neural networks that utilize multiple microphones on an ear-worn device to perform spatial focusing by applying different gains based on sound source locations, derived from timing differences, and incorporating inertial measurement units for head movement tracking, enabling improved noise reduction and spatial focusing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional beamforming patterns are used for noise reduction, then high-frequency sound processing is improved, but low-frequency sound processing deteriorates and the patterns are warped by wearer anatomy

Engineering Contradiction:
Improvenoise reduction performanceVSAvoidfrequency range coverage
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms the beamforming approach from spatial domain to temporal domain by using autocorrelation and cross-correlation calculations on time-delayed microphone signals. This parameter change allows the system to achieve frequency-independent noise reduction that works effectively across both high and low frequencies, overcoming the limitations of conventional spatial beamforming patterns that are warped by anatomy and frequency-dependent

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical/spatial beamforming system with a signal-processing-based system using autocorrelation and cross-correlation algorithms. This substitution eliminates the need for fixed geometric patterns that are distorted by wearer anatomy, allowing the noise reduction to adapt to different frequencies and anatomical variations through computational methods rather than physical constraints

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If multiple microphones are used for spatial focusing, then speaker isolation is improved, but device complexity increases

Engineering Contradiction:
Improvespeaker isolationVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the noise reduction problem into distinct computational components: autocorrelation calculation for each microphone signal, cross-correlation calculation between microphone pairs, and combination of these correlations to determine noise presence. This segmentation allows the complex task of spatial focusing with multiple microphones to be broken down into manageable, efficient computational steps that can be processed in real-time

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary autocorrelation and cross-correlation calculations on the microphone signals before making the final noise reduction decision. By pre-processing the signals to extract correlation information, the system prepares the data in advance, making the subsequent speaker isolation and noise reduction steps more efficient and reducing overall processing complexity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12581250B2Ear-worn device with neural network for noise reduction and/or spatial focusing using multiple input audio signals
Publication Date: 2026.03.17 FORTELL RESEARCH INC
  • US12581250B2 patent drawing
  • US12581250B2 patent drawing
  • US12581250B2 patent drawing

AI summary

An ear-worn device may include two or more microphones configured to generate time-domain audio signals, each of the two or more microphones configured to generate one of the time-domain audio signals; processing circuitry comprising analog processing circuitry, digital processing circuitry, beamforming circuitry, and short-time Fourier transformation (STFT) circuitry, the processing circuitry configured to generate, from the time-domain audio signals, one or more frequency-domain non-beamformed audio signals and one or more frequency-domain beamformed signals; and enhancement circuitry comprising neural network circuitry configured to receive multiple frequency-domain input audio signals originating from the one or more frequency-domain non-beamformed audio signals and the one or more frequency-domain beamformed signals, and implement a single neural network trained to generate, based on the multiple frequency-domain input audio signals, a noise-reduced and spatially-focused output audio signal or an output for generating a noise-reduced and spatially-focused output audio signal.