Microphone Array Source Location Estimation for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems in noisy environments struggle to accurately distinguish between intended and unintended voice signals, particularly in video game applications, due to unreliable voice volume estimation and sub-optimal performance of far-field microphones for close-talk scenarios.

Innovation Solution

A method using two or more microphones to estimate the distance and direction of a sound source by comparing volume and time delay properties of signals, allowing for reliable identification of the intended voice signal and rejection of background noise, employing a sound source discriminator with modules for voice segment detection, source location estimation, and decision-making to enable precise voice recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If far-field microphones are used for speech recognition, then the coverage area is increased, but the performance for close-talk scenarios deteriorates

Engineering Contradiction:
Improvecoverage areaVSAvoidclose-talk performance
Core Design Contradiction:
Area of stationary objectVSReliability

Solution Approach 1:

The patent segments the microphone array into multiple groups, with some microphones optimized for far-field capture and others for close-talk scenarios. This segmentation allows the system to handle both distant and nearby speech sources effectively by processing signals from different microphone groups according to their respective optimization characteristics.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatial dimensionality by using an array of microphones arranged in specific geometries (linear, planar, or volumetric configurations). By leveraging the spatial distribution of microphones and calculating arrival direction and distance in three-dimensional space, the system can distinguish between far-field and close-talk sources, resolving the contradiction between coverage area and close-talk performance.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If voice volume is used for source distance estimation, then the estimation process is simplified, but the reliability of distance estimation deteriorates due to unknown real voice volume

Engineering Contradiction:
Improveestimation process complexityVSAvoiddistance estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary approach by using the array of microphones to measure the arrival direction and calculate distance based on signal propagation characteristics rather than relying solely on voice volume. This intermediary measurement method (using multiple microphones as mediators) provides more reliable distance estimation by accounting for the unknown real voice volume through comparative analysis of signals from multiple microphones.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the simple volume-based estimation mechanism with a more sophisticated acoustic field analysis mechanism. Instead of relying on the mechanical simplicity of volume measurement, the system uses the physical principles of sound wave propagation, arrival time differences, and intensity variations across multiple microphones to calculate distance, thereby improving measurement precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If a single microphone is used for voice detection, then the device complexity is reduced, but the ability to distinguish between intended and unintended voice signals deteriorates in noisy environments

Engineering Contradiction:
Improvemicrophone configurationVSAvoidvoice signal discrimination
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the voice detection function across multiple microphones arranged in an array, with each microphone capturing a portion of the acoustic field. By segmenting the detection task and combining results from multiple microphones with different spatial positions, the system achieves reliable voice signal discrimination while maintaining manageable device complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent makes the microphone array universally applicable to multiple functions: far-field speech recognition, close-talk detection, voice signal discrimination, and spatial localization. The same array configuration and processing algorithms serve multiple purposes, allowing the system to distinguish between intended and unintended voice signals while maintaining cost-effectiveness and avoiding the need for separate specialized hardware for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach effectively filters out unwanted voice signals, improving the accuracy of voice recognition in noisy environments and enhancing the reliability of voice-activated commands in applications like video games by accurately determining the source of speech, thereby reducing false triggers and improving user interaction.

Implementation Method 1

estimates a distance and direction to a speaker based on a comparison of a volume and a time delay of the signals from the two or more microphones

Methodology Applied
Scientific EffectTime delay of sound signals: Speed of Sound

Data Source

PatentUS8442833B2Speech processing with source location estimation using signals from two or more microphones
Publication Date: 2013.05.14 SONY INTERACTIVE ENTERTAINMENT LLC
  • US8442833B2 patent drawing
  • US8442833B2 patent drawing
  • US8442833B2 patent drawing

AI summary

Computer implemented speech processing is disclosed. First and second voice segments are extracted from first and second microphone signals originating from first and second microphones. The first and second voice segments correspond to a voice sound originating from a common source. An estimated source location is generated based on a relative energy of the first and second voice segments and/or a correlation of the first and second voice segments. A determination whether the voice segment is desired or undesired may be made based on the estimated source location.