Speech Signal Processing With Audio-Visual Auxiliary Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional target speaker extraction techniques face challenges in accurately extracting audio signals of a speaker of interest from mixed audio signals due to issues with speaker characteristics, such as similar voice characteristics or varying quality of visual cues from speaker movements, leading to reduced extraction accuracy.
Innovation Solution
An audio signal processing apparatus utilizing multiple neural networks to convert audio and video features into auxiliary features, combined through an attention mechanism, and a training apparatus for parameter optimization to enhance accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speaker characteristics in audio signal are utilized for target speaker extraction, then extraction accuracy is improved when speakers have distinct voice characteristics, but extraction accuracy deteriorates when speakers have similar voice characteristics
Solution Approach 1:
The patent combines audio-based auxiliary features and video-based auxiliary features into a unified feature representation. The audio auxiliary features capture speaker characteristics while the video auxiliary features capture visual speaker characteristics, and their combination enables the system to leverage both modalities for more accurate and robust target speaker extraction, particularly when speakers have similar voice characteristics
Solution Approach 2:
The patent introduces video information as an intermediary modality to supplement audio-based speaker extraction. When audio features alone are insufficient (e.g., similar voice characteristics), the video auxiliary features serve as an additional clue to distinguish between speakers, thereby improving extraction accuracy in challenging scenarios
2Stability of the object's composition
If pre-recorded audio signal of target speaker is input to auxiliary neural network, then auxiliary features can be extracted with stable quality, but the system cannot adapt to varying speaker characteristics in real-time scenarios
Solution Approach 1:
The auxiliary neural networks are designed to process multiple types of input signals (pre-recorded audio, real-time audio, video) and generate auxiliary features applicable to various extraction scenarios. This multi-functionality enables the system to maintain stable feature extraction quality while adapting to different input conditions and speaker characteristics
3Adaptability or versatility
If video of target speaker is input to auxiliary neural network, then the system can handle speakers with similar voices, but extraction accuracy deteriorates when video quality varies due to speaker movement or occlusion
Solution Approach 1:
The system dynamically adjusts the reliance on video auxiliary features based on their quality. When video quality is high and provides useful speaker characteristics, the video features are heavily weighted. When video quality deteriorates due to movement or occlusion, the system automatically reduces reliance on video features and increases reliance on audio features, thereby maintaining stable extraction accuracy across varying conditions
4Device complexity
If single auxiliary neural network is used, then device complexity is reduced, but the system cannot simultaneously process multiple types of auxiliary information effectively
Solution Approach 1:
The auxiliary feature extraction system is segmented into multiple specialized neural networks: one for processing audio-based auxiliary information and another for processing video-based auxiliary information. Each network is optimized for its specific modality, enabling effective processing of multiple auxiliary information types while maintaining reasonable system complexity through modular architecture
Data Source
AI summary
An audio signal processing apparatus (10) includes a first auxiliary feature conversion unit (12) and a second auxiliary feature conversion unit (13) that convert a plurality of signals relating to processing of an audio signal of a target speaker into a plurality of auxiliary features for the plurality of signals using a plurality of auxiliary neural networks corresponding to the plurality of signals, and an audio signal processing unit (11) that estimates information regarding an audio signal of the target speaker included in a mixed audio signal using a main neural network based on an input feature of the mixed audio signal and the plurality of auxiliary features, wherein the plurality of signals relating to processing of the audio signal of the target speaker are two or more pieces of information of different modalities.


