Visual Speech Separation Network for Unknown Speakers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech separation technologies based on audio and video convergence have poor generalization capability for unknown speakers, leading to low precision and user experience issues in real-time applications.
Innovation Solution
A speech separation method that uses a combination of audio and video information to extract visual semantic features from facial motions, which are then input into a visual speech separation network to accurately separate the user's speech from environmental noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech separation technology based on face representation is used, then speech separation can be performed using visual cues, but generalization capability for unknown speakers deteriorates
Solution Approach 1:
The patent segments the speech separation task into two independent components: visual feature extraction from face videos and audio feature extraction from mixed speech. By processing visual and audio streams separately and then fusing their features, the system achieves better generalization to unknown speakers while maintaining high separation precision.
Solution Approach 2:
The patent introduces a temporal dimension by using video sequences instead of static images. By extracting visual features from temporal sequences of face videos, the system captures dynamic facial motion information that improves generalization capability across different speakers while maintaining separation accuracy.
2Measurement precision
If deep learning algorithms are used for speech separation, then separation capability is improved, but computational delay increases
Solution Approach 1:
The patent divides the deep learning processing into separate visual and audio branches that can be computed independently and in parallel. This segmentation reduces the overall computational delay while maintaining high separation accuracy through feature fusion.
Solution Approach 2:
The patent extracts only the most relevant visual semantic features from face videos rather than processing all visual information. This partial action approach reduces computational complexity and delay while retaining sufficient information for accurate speech separation.
Data Source
AI summary
A speech separation method is provided, and relates to the field of speech. The method includes: obtaining, in a speaking process of a user, audio information including a user speech and video information including a user face; coding the audio information to obtain a mixed acoustic feature; extracting a visual semantic feature of the user from the video information; inputting the mixed acoustic feature and the visual semantic feature into a preset visual speech separation network to obtain an acoustic feature of the user; and decoding the acoustic feature of the user to obtain a speech signal of the user. An electronic device, a chip, and a computer-readable storage medium are provided.


