Close Signal Capture for Speech Recognition in Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in noisy environments and when the speaker must remain silent, as traditional audio microphones fail to provide a clean signal, and advanced noise filtering or lipreading techniques are not effective in all situations.

Innovation Solution

The use of close signal capture techniques, such as microphones, cameras, and electrographic sensors placed near the speaker's head, to supplement or replace traditional microphones, capturing and processing speech information directly from the speaker's mouth and muscle movements, which can reject environmental noise and improve speech recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional audio microphones are used for speech recognition, then the system can capture speech in various environments, but the audio signal becomes corrupted by environmental noise and ambiguous in noisy conditions

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidenvironmental noise
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent transitions from traditional audio-only capture to multi-dimensional sensing by incorporating visual (camera), tactile (accelerometer, gyroscope), and electrographic (EMG sensors) dimensions. This allows the system to capture speech-related information through multiple modalities, reducing dependence on audio quality and improving reliability in noisy environments.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces intermediate processing layers including noise filtering algorithms, signal processing pipelines, and machine learning models that act as intermediaries between the raw sensor data and speech recognition. These intermediaries clean and refine the input signals before they reach the recognition engine, effectively reducing environmental noise impact.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If advanced noise filtering technology is applied to separate speech from background noise, then speech clarity may be improved, but the technique is useless when the speaker must remain silent and fails to use other available information

Engineering Contradiction:
Improvespeech signal clarityVSAvoidapplicability in silent conditions
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal speech recognition system that functions across multiple conditions (noisy environments, silent conditions, whispered speech) by incorporating multiple sensor types. The same device can switch between or combine audio, visual, and tactile modalities depending on the situation, making it universally applicable where traditional audio-only systems fail.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces the mechanical/audio-based speech capture system with alternative physical sensing mechanisms. Instead of relying solely on acoustic waves, the system uses electromagnetic sensors (cameras, EMG sensors, accelerometers) to detect speech-related physical phenomena, enabling speech recognition in conditions where audio capture is impossible or ineffective.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If automated lipreading is used in teleconferencing systems, then speech guidance may be provided, but the technique requires a teleconferencing system environment and lacks integration with personal speech recognition systems

Engineering Contradiction:
Improvespeech information recoveryVSAvoidsystem integration requirements
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges multiple previously separate technologies (lipreading, muscle movement detection, audio processing) into a single integrated personal device. By combining camera, EMG sensors, accelerometer, gyroscope, and audio microphone in one portable system, the patent eliminates the need for separate teleconferencing infrastructure while achieving comprehensive speech information recovery.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a self-contained speech recognition system that does not depend on external teleconferencing infrastructure. The device uses its own integrated sensors and processing capabilities to capture and interpret speech information independently, making the technology accessible in any environment without requiring specialized rooms or systems.

Inventive Principle:
Principle #25Self-service

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances speech recognition accuracy by reducing environmental noise and allowing whispering or silent speech recognition, even in noisy environments, by using adaptive learning and joint processing of close and traditional audio signals to produce a more reliable output.

Implementation Method 1

The device includes an accelerometer, a gyroscope, a close microphone, a non-visible camera, and an electrographic probe

Methodology Applied
Scientific EffectAccelerometer detection: Accelerometer

Implementation Method 2

The device includes an accelerometer, a gyroscope, a close microphone, a non-visible camera, and an electrographic probe

Methodology Applied
Scientific EffectGyroscope detection: Gyroscope

Implementation Method 3

The device includes an accelerometer, a gyroscope, a close microphone, a non-visible camera, and an electrographic probe

Methodology Applied
Scientific EffectElectrographic detection: Electromagnetic Induction

Data Source

PatentUS11373653B2Portable speech recognition and assistance using non-audio or distorted-audio techniques
Publication Date: 2022.06.28 EPSTEIN JOSEPH ALAN
  • US11373653B2 patent drawing
  • US11373653B2 patent drawing
  • US11373653B2 patent drawing

AI summary

A method and system for detecting speech using close sensor applications, according to some embodiments. In some embodiments, a close microphone is applied to detect sounds with higher muscle or bone transmission components. In some embodiments, a close camera is applied that collects visual information and motion that is correlated with the potential phonemes for such positions and motion. In some embodiments, myography is performed, to detect muscle movement. In an earbud form factor embodiment, processing of different channels of close information is performed to improve the accuracy of the recognition.