Audio-Visual Sound Enhancement via Gaze Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio systems fail to effectively isolate and enhance specific sound sources in noisy environments, such as multiple human speakers, due to indiscriminate amplification and inability to differentiate between sources of the same type.

Innovation Solution

A computer-implemented method that acquires image and sensor data to determine a user's gaze direction and focus, processes audio signals to identify the source of interest, and enhances that audio signal relative to others, using a combination of audio and visual features for precise sound separation and output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If personal sound amplification products amplify all received sounds indiscriminately, then the user can hear sounds from all sources, but the user cannot focus on specific desirable sounds among many sound sources

Engineering Contradiction:
Improvequantity of sounds amplifiedVSAvoidprecision of sound source selection
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the mixed audio signal into multiple independent sound source signals by analyzing visual information from multiple cameras. Each camera captures a specific sound source, and the system separates these signals spatially and temporally, allowing selective enhancement of individual sources rather than amplifying all sounds uniformly.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces visual information from cameras as an intermediary to identify and locate sound sources. The visual data serves as a mediator that guides the audio processing system to select which sound sources to enhance, resolving the contradiction between amplifying multiple sounds and focusing on specific ones.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If sound isolation devices separate sounds by type, then the user can hear desired types of sounds, but the user cannot differentiate between multiple sources of the same type

Engineering Contradiction:
Improveprecision of sound type separationVSAvoidquantity of distinguishable sound sources
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies segmentation by assigning different spatial regions to different cameras, where each camera is responsible for capturing a specific sound source. This spatial segmentation allows the system to distinguish between multiple sources of the same sound type (e.g., multiple speakers) by their location, not just by sound type.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a spatial dimension to sound source differentiation by using multiple cameras positioned at different locations. Instead of relying solely on audio characteristics, the system uses spatial positioning information from visual data to distinguish between multiple sources of the same sound type.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If audio processing systems require the source and user to remain stationary, then the system can maintain precise sound enhancement, but the system lacks adaptability to moving sources or users

Engineering Contradiction:
Improveprecision of sound enhancementVSAvoidadaptability to motion
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamics by continuously tracking the positions of both the user and sound sources using visual information from cameras. The system dynamically adjusts the audio enhancement parameters based on real-time position data, allowing precise sound enhancement to be maintained even when sources or users are moving.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent uses feedback from visual tracking systems to continuously monitor and adjust the audio processing parameters. The position information from cameras provides feedback that enables the system to adapt to motion in real-time, maintaining precision without requiring stationarity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11482238B2Audio-visual sound enhancement
Publication Date: 2022.10.25 HARMAN INT IND INC
  • US11482238B2 patent drawing
  • US11482238B2 patent drawing
  • US11482238B2 patent drawing

AI summary

Embodiments of the present disclosure sets forth a computer-implemented method comprising acquiring image information associated with an environment, acquiring, from one or more sensors, sensor data associated with a gaze of a user, determining a source of interest based on the image information and the sensor data, processing a set of audio signals associated with the environment based on the image information to identify an audio signal associated with the source of interest, enhancing the audio signal associated with the source of interest relative to other audio signals in the set of audio signals, and outputting the enhanced audio signal associated with the source of interest to the user.