一种基于自适应聚类扩散的变源语音分离方法及系统

The adaptive clustering diffusion model solves the problems of unknown speaker number and arrangement ambiguity in speech separation, improves the processing efficiency of long sequences, and achieves high-efficiency speech separation results, which are suitable for complex acoustic environments and mobile devices.

CN121122308BActive Publication Date: 2026-07-17JIANGSU ZHIHENG INFORMATION TECH SERVICES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU ZHIHENG INFORMATION TECH SERVICES CO LTD
Filing Date
2025-08-19
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing speech separation techniques face challenges in handling unknown numbers of speakers, permutation ambiguity, and long sequence efficiency, especially due to the fixed output structure of diffusion models, imperfect handling of permutation ambiguity, and high computational cost of long sequence modeling.

Method used

An adaptive clustering diffusion model is adopted. By adding a normalization layer and a Mamba module to the U-Net model and combining it with a dynamic adaptive clustering module, the separation of a variable number of speakers is achieved. Speech reconstruction is performed using an adaptive clustering denoiser and a reconstruction decoder. The model is trained by combining score matching loss and reconstruction loss.

Benefits of technology

It achieves flexible and high-quality separation of an unknown number of speakers, eliminates the ambiguity of arrangement in multi-speaker scenarios, improves the efficiency and separation performance of long speech sequence processing, and is suitable for complex acoustic environments and mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122308B_ABST
    Figure CN121122308B_ABST
Patent Text Reader

Abstract

本发明提供了一种基于自适应聚类扩散的变源语音分离方法及系统,属于语音分离技术领域。其方法包括:获取混合语音信号,生成与混合语音信号嵌入维度相同的噪声源嵌入;将混合语音信号和噪声源嵌入输入预先构建的自适应聚类扩散模型,输出分离出的干净语音信号;其中,自适应聚类扩散模型的构建包括:在获取的U‑Net模型的编码器、解码器中的卷积层和激活层之间添加归一化层,并在U‑Net模型的编码器和解码器后均添加Mamba模块,得到自适应聚类去噪器;在自适应聚类去噪器后依次添加动态自适应聚类模块和重构解码器,得到构建好的自适应聚类扩散模型。本发明能够适应说话者数量、解决排列模糊性、兼顾长序列处理效率。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Audio source separation

    CN105989851A

  • Multi-modal sentiment analysis method and system under uncertain content deficiency

    CN120123993A