Audio Signal Processing Using Visual Speaker Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio systems face challenges in accurately acquiring and processing speech signals due to reverberations, noise interference, and mislocalization issues in environments like vehicles and conference rooms, where sound reflections and multiple sound sources complicate the detection and processing of speech.

Innovation Solution

An audio system that uses a camera to estimate the distance and orientation of a speaker relative to a microphone, adjusting processing parameters such as gain and filter aggressiveness based on visual data to improve speech signal quality and reduce reverberation effects, while also identifying and prioritizing the correct sound source.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If visual information from a camera is used to estimate speaker distance and orientation, then speech signal quality is improved by reducing reverberation and noise interference, but device complexity increases due to integration of camera and image analysis components

Engineering Contradiction:
Improvespeech signal qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines the camera system and image analysis functionality with the existing audio processing system into an integrated audio-visual speech enhancement system. The image analyzer processes visual data from the camera to estimate speaker parameters, which are then used to control audio signal processing operations, merging two previously separate systems into a unified architecture that leverages both visual and audio information streams.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an image analyzer as an intermediary component that bridges the camera and the audio signal processor. This intermediary analyzes visual information to extract speaker distance and orientation estimates, which are then fed as control parameters to the audio processing system. This intermediary layer enables the system to make informed audio processing decisions based on visual data without requiring direct integration between all components.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If processing parameters are dynamically adjusted based on visual feedback, then reverberation and noise interference are reduced, but processing time increases due to real-time image analysis requirements

Engineering Contradiction:
Improvespeech signal qualityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs image analysis and speaker parameter estimation in advance of the actual speech processing operation. By analyzing visual information to determine speaker distance and orientation beforehand, the system can pre-configure appropriate audio processing parameters such as gain values and filter settings before speech signals arrive, eliminating the need for real-time adjustments during critical speech processing intervals.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements periodic updates of speaker parameter estimates based on visual feedback rather than continuous real-time adjustment. The system periodically re-evaluates speaker position and distance using image analysis, updating audio processing parameters at intervals that balance processing quality with computational efficiency. This periodic approach reduces the continuous processing burden while maintaining adequate adaptation to changing speaker positions.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentEP2766901B1Speech signal enhancement using visual information
Publication Date: 2016.09.21 NUANCE COMMUNICATIONS INC
  • EP2766901B1 patent drawingFigure 1
  • EP2766901B1 patent drawingFigure 2
  • EP2766901B1 patent drawingFigure 3

AI summary

Visual information is used to alter or set an operating parameter of an audio signal processor, other than a beamformer. A digital camera captures visual information about a scene that includes a human speaker and/or a listener. The visual information is analyzed to ascertain information about acoustics of a room. A distance between the speaker and a microphone may be estimated, and this distance estimate may be used to adjust an overall gain of the system. Distances among, and locations of, the speaker, the listener, the microphone, a loudspeaker and/or a sound- reflecting surface may be estimated. These estimates may be used to estimate reverberations within the room and adjust aggressiveness of an anti-reverberation filter, based on an estimated ratio of direct to indirect (reverberated) sound energy expected to reach the microphone. In addition, orientation of the speaker or the listener, relative to the microphone or the loudspeaker, can also be estimated, and this estimate may be used to adjust frequency-dependent filter weights to compensate for uneven frequency propagation of acoustic signals from a mouth, or to a human ear, about a human head.