一种基于自适应聚类扩散的变源语音分离方法及系统
The adaptive clustering diffusion model solves the problems of unknown speaker number and arrangement ambiguity in speech separation, improves the processing efficiency of long sequences, and achieves high-efficiency speech separation results, which are suitable for complex acoustic environments and mobile devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGSU ZHIHENG INFORMATION TECH SERVICES CO LTD
- Filing Date
- 2025-08-19
- Publication Date
- 2026-07-17
AI Technical Summary
Existing speech separation techniques face challenges in handling unknown numbers of speakers, permutation ambiguity, and long sequence efficiency, especially due to the fixed output structure of diffusion models, imperfect handling of permutation ambiguity, and high computational cost of long sequence modeling.
An adaptive clustering diffusion model is adopted. By adding a normalization layer and a Mamba module to the U-Net model and combining it with a dynamic adaptive clustering module, the separation of a variable number of speakers is achieved. Speech reconstruction is performed using an adaptive clustering denoiser and a reconstruction decoder. The model is trained by combining score matching loss and reconstruction loss.
It achieves flexible and high-quality separation of an unknown number of speakers, eliminates the ambiguity of arrangement in multi-speaker scenarios, improves the efficiency and separation performance of long speech sequence processing, and is suitable for complex acoustic environments and mobile devices.
Smart Images

Figure CN121122308B_ABST
Abstract
Citation Information
Patent Citations
Audio source separation
CN105989851A
Multi-modal sentiment analysis method and system under uncertain content deficiency
CN120123993A