In-Ear Voice Capture Using Generative Audio Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Recognizing speech from a person's mouth using in-ear microphones is challenging due to the complexity of noisy systems, as conventional signal processing struggles with noise and low-pass filtering effects, leading to poor quality speech recording.
Innovation Solution
A method involving deep learning and generative modeling that processes signals from both in-ear and external microphones to isolate and enhance speech, using a combination of noise cancellation and audio super-resolution techniques to generate noise-free audible signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If in-ear microphone is used for speech capture, then proximity to sound source is improved, but noise and low-pass filtering effects increase
Solution Approach 1:
The patent introduces an external microphone as an intermediary device to capture the speech signal before it enters the ear canal. This external microphone serves as a mediator that records the speech in a cleaner environment, which then compensates for the degradation caused by the in-ear microphone's low-pass filtering and noise. The system combines signals from both microphones, using the external microphone's cleaner capture to offset the in-ear microphone's limitations.
2Device complexity
If conventional signal processing is used, then system complexity is reduced, but speech recognition accuracy deteriorates
Solution Approach 1:
The patent replaces conventional mechanical signal processing methods with deep learning-based neural network processing. Instead of using traditional filtering and enhancement algorithms, the system employs trained neural networks that can automatically learn and adapt to various speech patterns, noise types, and ear canal characteristics. This substitution enables more accurate speech recognition while handling the complex interactions between multiple microphones and noise sources.
3Measurement precision
If multiple microphones are used, then speech signal quality is improved, but system complexity increases
Solution Approach 1:
The patent merges the functionality of multiple microphones (in-ear and external) into a unified processing system. Rather than treating them as separate components requiring independent processing, the system combines their outputs and uses a joint deep learning model that processes both signals simultaneously. This merging approach leverages the complementary strengths of each microphone while avoiding the complexity of managing them as separate systems.
Data Source
AI summary
A method includes accessing, by at least one processing device, an audible signal including at least one in-ear microphone audible signal and at least one external microphone audible signal and at least one noise signal; training a generative network to generate an enhanced external microphone signal from an in-ear microphone signal based on the at least one in-ear microphone audible signal and the at least one external microphone audible signal; and outputting the generative network.


