Audio-Visual Speech Separation in Noisy Environments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio-visual speech separation technologies fail to effectively isolate speech signals in noisy environments with overlapping audio and background noise, particularly in settings where multiple speakers are present, and often require specific speaker visibility for accurate separation.

Innovation Solution

A system utilizing a speech separation engine that processes real-time video and audio inputs using neural networks to generate isolated speech signals for each speaker, employing joint audio-visual features and spectrogram masks to separate speech from background noise and other speakers, even when speakers are not in the current camera view, and allows for user preference and automatic speech recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional audio-visual speech separation is used, then speech separation can be achieved, but it fails to effectively isolate speech signals in noisy environments with overlapping audio and background noise

Engineering Contradiction:
Improvespeech separation effectivenessVSAvoidbackground noise and overlapping audio
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent combines audio and visual inputs into a unified processing system that uses joint audio-visual features to improve speech separation. The system merges microphone array captures with camera video feeds, processing both modalities simultaneously through a neural network that leverages the complementary information to achieve more reliable speech isolation in noisy environments.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary processing layer using neural networks that take both audio and visual inputs and produce enhanced speech separation outputs. This intermediary system processes joint audio-visual features through spectrogram masks to separate speech from background noise and overlapping audio, achieving improved reliability without direct physical separation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If speaker visibility is required for accurate separation, then processing can be simplified, but it limits functionality in crowded settings where speakers may not be visible

Engineering Contradiction:
Improveprocessing requirementsVSAvoidfunctionality in crowded settings
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal speech separation system that functions whether speakers are visible or not. The system can process audio from microphone arrays and video from cameras, and the neural network is trained to handle various scenarios including visible speakers, invisible speakers, and crowded environments. This multi-functional approach maintains accurate speech separation across different conditions without requiring speaker visibility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Speed

If real-time processing is implemented, then responsiveness is improved, but processing delay may occur in complex environments

Engineering Contradiction:
Improveresponse timeVSAvoidprocessing delay
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent implements preliminary processing by pre-computing spectrogram masks and preparing neural network models for rapid inference. The system pre-processes audio and visual inputs into standardized representations that can be quickly processed in real-time. This preliminary preparation reduces computational overhead during actual speech separation, minimizing processing delay while maintaining real-time responsiveness even in complex environments.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240428816A1Audio-visual hearing aid
Publication Date: 2024.12.26 GOOGLE LLC
  • US20240428816A1 patent drawing
  • US20240428816A1 patent drawing
  • US20240428816A1 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for audio-visual speech separation. A method includes: receiving, by a user device, a first indication of one or more first speakers visible in a current view recorded by a camera of the user device, in response, generating a respective isolated speech signal for each of the one or more first speakers that isolates speech of the first speaker in the current view and sending the isolated speech signals for each of the one or more first speakers to a listening device operatively coupled to the user device, receiving, by the user device, a second indication of one or more second speakers visible in the current view recorded by the camera of the user device, and in response generating and sending a respective isolated speech signal for each of the one or more second speakers to the listening device.