Microphone Array Noise Separation for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional electronic devices struggle to accurately separate human speech from noise in noisy environments due to limitations in microphone design and signal processing, leading to poor signal-to-noise ratios and reduced battery life from high power consumption.

Innovation Solution

A method and apparatus using a geometrical array of microphones that employ both cardioid and beamforming signal processing techniques to separate unwanted noise from audible signals across the full speech frequency range, enhancing signal-to-noise ratios and reducing power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If conventional filtering methods are used to reduce noise, then noise reduction is achieved, but speech components are also suppressed along with noise, reducing signal-to-noise ratio

Engineering Contradiction:
ImprovenoiseVSAvoidsignal-to-noise ratio
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent segments the audio signal processing into multiple frequency bands using filter banks. Each band is processed independently to extract speech components while suppressing noise. This segmentation allows selective enhancement of speech frequencies without suppressing important speech components that conventional full-band filtering would remove.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing strategies to different frequency regions. Speech-enhanced filter banks are designed with varying Q-factors and center frequencies tailored to speech characteristics in specific bands. This local optimization ensures that noise reduction is applied where appropriate while preserving speech components in critical frequency regions.

Inventive Principle:
Principle #3Local quality

2Loss of information

If voice recognition systems process full speech frequency range (100-8000 Hz), then more useful data is captured, but power consumption increases reducing battery life

Engineering Contradiction:
Improvespeech informationVSAvoidpower consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent implements partial processing by applying full speech enhancement only to frequency bands containing speech components, while using simpler processing or noise-only suppression in bands without speech. This partial action approach captures essential speech information across the full frequency range while reducing computational load and power consumption compared to processing all bands equally.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system periodically analyzes the input signal to detect presence of speech components in different frequency bands, and dynamically adjusts processing intensity accordingly. During periods when speech is detected, full enhancement is applied; during noise-only periods, reduced processing is used, creating a periodic action pattern that optimizes power consumption while maintaining speech quality.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS10366700B2Device for acquiring and processing audible input
Publication Date: 2019.07.30 LOGITECH EUROPE SA
  • US10366700B2 patent drawing
  • US10366700B2 patent drawing
  • US10366700B2 patent drawing

AI summary

Embodiments of the disclosure generally include a method and apparatus for receiving and separating unwanted external noise from an audible input received from an audible source using an audible signal processing system that contains a plurality of audible signal sensing devices that are arranged and configured to detect an audible signal that is received from any position or angle within three dimensional space. The audible signal processing system is configured to analyze the received audible signals using a first signal processing technique that is able to separate unwanted low frequency range noise from the received audible signal and a second signal processing technique that is able to separate unwanted higher frequency range noise from the received audible signal. The audible signal processing system can then combine the signals processed by the first and second signal processing techniques to form a desired audible signal that has a high signal-to-noise ratio throughout the full speech range.