In-Ear Voice Capture Using Generative Audio Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Recognizing speech from a person's mouth using in-ear microphones is challenging due to the complexity of noisy systems, as conventional signal processing struggles with noise and low-pass filtering effects, leading to poor quality speech recording.

Innovation Solution

A method involving deep learning and generative modeling that processes signals from both in-ear and external microphones to isolate and enhance speech, using a combination of noise cancellation and audio super-resolution techniques to generate noise-free audible signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If in-ear microphone is used for speech capture, then proximity to sound source is improved, but noise and low-pass filtering effects increase

Engineering Contradiction:
Improvespeech capture accuracyVSAvoidnoise and low-pass filtering
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces an external microphone as an intermediary device to capture the speech signal before it enters the ear canal. This external microphone serves as a mediator that records the speech in a cleaner environment, which then compensates for the degradation caused by the in-ear microphone's low-pass filtering and noise. The system combines signals from both microphones, using the external microphone's cleaner capture to offset the in-ear microphone's limitations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If conventional signal processing is used, then system complexity is reduced, but speech recognition accuracy deteriorates

Engineering Contradiction:
Improvesignal processing complexityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent replaces conventional mechanical signal processing methods with deep learning-based neural network processing. Instead of using traditional filtering and enhancement algorithms, the system employs trained neural networks that can automatically learn and adapt to various speech patterns, noise types, and ear canal characteristics. This substitution enables more accurate speech recognition while handling the complex interactions between multiple microphones and noise sources.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If multiple microphones are used, then speech signal quality is improved, but system complexity increases

Engineering Contradiction:
Improvespeech signal qualityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the functionality of multiple microphones (in-ear and external) into a unified processing system. Rather than treating them as separate components requiring independent processing, the system combines their outputs and uses a joint deep learning model that processes both signals simultaneously. This merging approach leverages the complementary strengths of each microphone while avoiding the complexity of managing them as separate systems.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10685663B2Enabling in-ear voice capture using deep learning
Publication Date: 2020.06.16 NOKIA TECHNOLOGIES OY
  • US10685663B2 patent drawing
  • US10685663B2 patent drawing
  • US10685663B2 patent drawing

AI summary

A method includes accessing, by at least one processing device, an audible signal including at least one in-ear microphone audible signal and at least one external microphone audible signal and at least one noise signal; training a generative network to generate an enhanced external microphone signal from an in-ear microphone signal based on the at least one in-ear microphone audible signal and the at least one external microphone audible signal; and outputting the generative network.