Visual Speech Separation Network for Unknown Speakers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech separation technologies based on audio and video convergence have poor generalization capability for unknown speakers, leading to low precision and user experience issues in real-time applications.

Innovation Solution

A speech separation method that uses a combination of audio and video information to extract visual semantic features from facial motions, which are then input into a visual speech separation network to accurately separate the user's speech from environmental noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech separation technology based on face representation is used, then speech separation can be performed using visual cues, but generalization capability for unknown speakers deteriorates

Engineering Contradiction:
Improvegeneralization capabilityVSAvoidspeech separation precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the speech separation task into two independent components: visual feature extraction from face videos and audio feature extraction from mixed speech. By processing visual and audio streams separately and then fusing their features, the system achieves better generalization to unknown speakers while maintaining high separation precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a temporal dimension by using video sequences instead of static images. By extracting visual features from temporal sequences of face videos, the system captures dynamic facial motion information that improves generalization capability across different speakers while maintaining separation accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If deep learning algorithms are used for speech separation, then separation capability is improved, but computational delay increases

Engineering Contradiction:
Improvespeech separation accuracyVSAvoidspeech separation delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the deep learning processing into separate visual and audio branches that can be computed independently and in parallel. This segmentation reduces the overall computational delay while maintaining high separation accuracy through feature fusion.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts only the most relevant visual semantic features from face videos rather than processing all visual information. This partial action approach reduces computational complexity and delay while retaining sufficient information for accurate speech separation.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12334092B2Speech separation method, electronic device, chip, and computer- readable storage medium
Publication Date: 2025.06.17 HUAWEI TECH CO LTD
  • US12334092B2 patent drawing
  • US12334092B2 patent drawing
  • US12334092B2 patent drawing

AI summary

A speech separation method is provided, and relates to the field of speech. The method includes: obtaining, in a speaking process of a user, audio information including a user speech and video information including a user face; coding the audio information to obtain a mixed acoustic feature; extracting a visual semantic feature of the user from the video information; inputting the mixed acoustic feature and the visual semantic feature into a preset visual speech separation network to obtain an acoustic feature of the user; and decoding the acoustic feature of the user to obtain a speech signal of the user. An electronic device, a chip, and a computer-readable storage medium are provided.