Audio Signal Processing for Robust Speech Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Supervised learning-based deep models for speech separation have poor robustness and generalization when processing audio signals not labeled during training, leading to reduced accuracy in scenarios other than the training scenario, especially in single-channel speech separation tasks.

Innovation Solution

An audio signal processing method that performs embedding processing, generalized feature extraction, and collaborative iterative training using a teacher and student model based on unlabeled sample mixed signals to obtain an encoder network and abstractor network, enabling robust and generalizable audio signal processing across various scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If supervised learning-based deep models are used for speech separation, then speech separation can be performed in training scenarios, but robustness and generalization performance deteriorate when processing audio signals not labeled during training

Engineering Contradiction:
Improverobustness and generalizationVSAvoidperformance in scenarios other than training scenario
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs self-supervised learning by automatically generating supervision signals from the mixed audio itself through masking and prediction tasks, eliminating the need for external labeled data. The model learns to predict masked portions of the embedding using contextual information, enabling it to adapt to unseen scenarios without manual labeling while maintaining robust speech separation performance

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary embedding extraction and masking operations on the mixed audio signal before the actual speech separation task. By pre-processing the audio through embedding extraction and creating masked versions for training, the model is prepared to handle various scenarios without requiring scenario-specific labeled data, thus improving generalization

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual labeling of training data is performed to improve model accuracy, then speech separation accuracy improves, but labor costs and time consumption increase

Engineering Contradiction:
Improvespeech separation accuracyVSAvoidtime for manual labeling
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system eliminates manual labeling by implementing self-supervised learning where the model generates its own training signals. Through automatic masking of portions of the embedding and requiring the model to predict the masked content, the system creates supervision signals autonomously from the unlabelled mixed audio, achieving high accuracy without any manual annotation effort

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system introduces an embedding layer as an intermediary between the raw mixed audio and the speech separation task. This embedding serves as a intermediate representation that can be automatically masked and used for self-supervised training, bridging the gap between unlabelled audio and the supervision needed for accurate speech separation without manual labeling

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4006901B1Audio signal processing method and apparatus, electronic device, and storage medium
Publication Date: 2024.12.04 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP4006901B1 patent drawingFigure 1~2
  • EP4006901B1 patent drawingFigure 3~4
  • EP4006901B1 patent drawingFigure 5~6

AI summary

An audio signal processing method and apparatus, an electronic device, and a storage medium, belonging to the technical field of signal processing. Embedding processing is performed on a mixed audio signal to obtain an embedding feature of the mixed audio signal, and generalization feature extraction is performed on the embedding feature to extract a generalization feature of a target component in the mixed audio signal. The generalization feature of the target component has good generalization capability and expression capability, and can be well applied to different scenarios, thereby improving robustness and generalization during the audio signal processing, and improving the accuracy of audio signal processing.