Acoustic Speaker Localization via Reflected Sound Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The localization of speakers in communication systems is challenging, leading to difficulties in detecting speech signals and introducing noise, which affects the intelligibility and signal-to-noise ratio in electronic communications, especially in audio and video conferences.

Innovation Solution

A method using a microphone array to detect sound reflections from a loudspeaker, determining the speaker's direction and distance through signal processing, including beamforming and echo compensation, allowing for accurate localization before the speaker begins uttering, thereby improving signal clarity and reducing background noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speaker localization is not implemented, then the system is simpler, but the speech signal detection becomes difficult and noise corruption increases

Engineering Contradiction:
Improvespeech signal detection reliabilityVSAvoidlocalization system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary speaker localization by detecting sound reflections before speech recognition occurs. The microphone array captures reflected sounds from the speaker's body, and the processor determines speaker position in advance, allowing the system to prepare appropriate beamforming weights and noise suppression parameters before the actual speech signal needs to be processed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses sound reflections as an intermediary to indirectly detect speaker position. Instead of directly measuring speaker location, the system detects reflections of sound waves from the speaker's body off surrounding surfaces, which serve as a mediator to infer speaker position without requiring direct line-of-sight or additional sensors.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If speaker localization is implemented, then the signal-to-noise ratio improves, but the system complexity increases

Engineering Contradiction:
Improvenoise interference levelVSAvoidsignal processing complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The system segments the audio signal processing into distinct components: first detecting sound reflections separately from direct speech, then processing localization information separately from speech recognition, and finally applying separate beamforming and noise suppression treatments. This segmentation allows each component to be optimized independently while working together to reduce noise interference.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces physical acoustic filtering mechanisms with computational signal processing methods. Instead of using physical barriers or acoustic chambers to filter noise, the system uses digital signal processing techniques including beamforming, echo cancellation, and adaptive filtering to achieve noise suppression based on localized speaker position.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If the microphone array detects reflected sound, then speaker localization accuracy improves, but the detection of direct speech may be compromised

Engineering Contradiction:
Improvespeaker position accuracyVSAvoiddirect speech signal quality
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system uses feedback by continuously monitoring the detected sound field, comparing expected reflection patterns with actual measurements, and adjusting localization estimates accordingly. The processed localization information feeds back into the beamforming and speech enhancement algorithms, which in turn improve the quality of detected speech signals, creating a positive feedback loop that resolves the apparent contradiction.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs periodic updates of speaker localization based on continuously detected sound reflections, rather than attempting to simultaneously process all acoustic information. By periodically refreshing localization data and using it to adjust speech detection parameters, the system maintains accurate position tracking while preserving direct speech signal quality through adaptive processing.

Inventive Principle:
Principle #19Periodic action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This method enhances the ability to detect voices and relevant audio signals with high clarity, eliminating irrelevant noise and improving the signal-to-noise ratio, enabling reliable speech recognition and communication systems to function effectively.

Implementation Method 1

The loudspeaker emits a sound

Methodology Applied
Scientific EffectSound wave propagation: Sound

Implementation Method 2

the sound reflected by the speaker

Methodology Applied
Scientific EffectAcoustic reflection: Reflection

Implementation Method 3

The microphone array detects the sound reflected by the speaker and converts the sound into a microphone signal

Methodology Applied
Scientific EffectAcoustic detection: Sound

Data Source

PatentUS9338549B2Acoustic localization of a speaker
Publication Date: 2016.05.10 NUANCE COMMUNICATIONS INC
  • US9338549B2 patent drawing
  • US9338549B2 patent drawing
  • US9338549B2 patent drawing

AI summary

A system locates a speaker in a room containing a loudspeaker and a microphone array. The loudspeaker transmits a sound that is partly reflected by a speaker. The microphone array detects the reflected sound and converts the sound into a microphone array, the speaker's distance from the microphone array, or both, based on the characteristics of the microphone signals.