Depth-Aware Speech Recognition Using Scene Depth Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face limitations in noise cancellation and recognition accuracy due to the inability to differentiate between speech and noise from the same direction, and decreased performance with increasing distance from the microphone, leading to suboptimal performance in voice-controlled applications.

Innovation Solution

Integration of a depth camera with a microphone array to compute a scene depth map, segment the speaker from the background, and determine distance, allowing the speech recognition engine to select a distance-specific acoustic model, apply weighting factors to suppress noise and boost the signal, and provide feedback for improved performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a depth camera is integrated with a microphone array to compute scene depth map and determine speaker distance, then speech recognition accuracy is improved by suppressing noise and adapting to varying distances, but device complexity increases due to additional sensors and processing requirements

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines a depth camera with a microphone array into an integrated system. The depth camera captures scene depth maps while the microphone array captures audio signals, and both are processed together to achieve distance-aware speech recognition and noise suppression, resolving the contradiction by merging sensing and processing functions.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The depth camera serves as an intermediary that provides distance information to the speech recognition system. This intermediate depth data enables the system to adapt acoustic models and suppress noise based on speaker distance, improving recognition accuracy without requiring the microphone array itself to perform complex distance estimation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If beam forming is used to suppress noise from different directions, then noise cancellation performance is improved, but the system fails when speech and noise come from the same direction

Engineering Contradiction:
Improvenoise cancellation performanceVSAvoidadaptability to same-direction noise
Core Design Contradiction:
Object-affected harmful factorsVSAdaptability or versatility

Solution Approach 1:

The patent introduces depth information as an additional dimension beyond traditional directional beam forming. By incorporating the depth axis from the depth camera, the system can distinguish speech from noise even when they originate from the same direction but at different distances, extending the effectiveness of noise cancellation to scenarios where conventional beam forming fails.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system dynamically changes the parameter of distance based on depth map analysis. By adjusting the acoustic model and beam forming parameters according to the determined speaker distance, the system adapts to varying noise conditions and maintains effectiveness even when noise and speech share the same direction, thus resolving the adaptability limitation.

Inventive Principle:
Principle #35Parameter changes

3Area of stationary object

If the speaker is at a greater distance from the microphone array, then the coverage area is increased, but the speech signal amplitude decreases and recognition accuracy deteriorates

Engineering Contradiction:
Improvecoverage areaVSAvoidrecognition accuracy
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The patent implements dynamic adaptation of the speech recognition system based on real-time distance measurement from the depth camera. As the speaker's distance changes, the system dynamically adjusts acoustic models and processing parameters to maintain optimal recognition accuracy across varying distances and coverage areas.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The depth camera provides continuous feedback on speaker distance to the speech recognition system. This feedback loop enables the system to monitor and adjust its performance based on the determined distance, compensating for signal amplitude loss at greater distances and maintaining recognition accuracy across the full coverage area.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9847082B2System for modifying speech recognition and beamforming using a depth image
Publication Date: 2017.12.19 RESIDEO USA LLC
  • US9847082B2 patent drawing
  • US9847082B2 patent drawing
  • US9847082B2 patent drawing

AI summary

A system includes a speech recognition processor, a depth sensor coupled to the speech recognition processor, and an array of microphones coupled to the speech recognition processor. The depth sensor is operable to calculate a distance and a direction from the array of microphones to a source of audio data. The speech recognition processor is operable to select an acoustic model as a function of the distance and the direction from the array of microphones to the source of audio data. The speech recognition processor is operable to apply the distance measure in the microphone array beam formation so as to boost portions of the signals originating from the source of audio data and to suppress portions of the signals resulting from noise.