Audio-Visual Neural Network for Speech Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio-only methods for speech separation struggle to effectively isolate specific human voices from similar gender voices and background noise, especially in noisy environments, due to their inability to leverage visual information.

Innovation Solution

The use of visual information from face and mouth movements in video signals to enhance speech separation, employing an audio-visual end-to-end neural network model that integrates visual data with audio inputs to generate a spectrogram of enhanced speech, improving separation performance by exploiting visual features alongside audio features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio-only methods are used for speech separation, then the system complexity is low, but the speech separation performance deteriorates in noisy environments and when separating similar gender voices

Engineering Contradiction:
Improvespeech separation performanceVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges audio signal processing with visual information processing by integrating a video analysis module that processes both audio and video inputs simultaneously. The audio signal and video frames are fed into a unified neural network model that learns joint audio-visual representations, enabling the system to leverage complementary information from both modalities to improve speech separation performance while maintaining a cohesive system architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary video analysis module that extracts visual features from video frames and combines them with audio features. This intermediary component processes visual information about speaker characteristics, facial expressions, and lip movements, then integrates these features with audio spectral information to enhance speech separation capability without directly modifying the core audio processing pipeline.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If visual information is integrated with audio inputs, then speech separation performance improves, but the device complexity increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the processing pipeline into distinct functional modules: a video analysis module that extracts visual features from video frames, an audio processing module that processes audio signals, and a fusion module that combines both feature sets. This segmentation allows each module to be optimized independently and facilitates modular implementation, reducing overall system complexity while maintaining high speech intelligibility through coordinated multi-modal processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10777215B2Method and system for enhancing a speech signal of a human speaker in a video using visual information
Publication Date: 2020.09.15 YISSUM RESEARCH DEVELOPMENT COMPANY OF THE HEBREW UNIVERSITY OF JERUSALEM LTD
  • US10777215B2 patent drawing
  • US10777215B2 patent drawing
  • US10777215B2 patent drawing

AI summary

A method and system for enhancing a speech signal is provided herein. The method may include the following steps: obtaining an original video, wherein the original video includes a sequence of original input images showing a face of at least one human speaker, and an original soundtrack synchronized with said sequence of images; and processing, using a computer processor, the original video, to yield an enhanced speech signal of said at least one human speaker, by detecting sounds that are acoustically unrelated to the speech of the at least one human speaker, based on visual data derived from the sequence of original input images.