一种声学环境感知引导的自适应语音增强方法及装置

By employing an acoustic environment-aware adaptive speech enhancement method, which combines the Transformer architecture and the U-Net diffusion generation model, the problem of poor speech enhancement performance in multi-channel far-field environments is solved, achieving efficient speech enhancement and recognition under different acoustic environments.

CN121789701BActive Publication Date: 2026-07-17UNIV OF SCI & TECH BEIJING

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH BEIJING
Filing Date
2025-12-17
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In multi-channel far-field environments, existing speech enhancement methods suffer from the problem that beamformers are sensitive to sound source orientation errors and reverberation, leading to target speech distortion and poor speech enhancement effects.

Method used

An acoustic environment-aware adaptive speech enhancement method is adopted. This method performs time-frequency transformation on the multi-channel speech signals received by the microphone array, performs spatial sampling using a fixed beamforming filter, and combines an acoustic environment-aware network based on the Transformer architecture with a U-Net diffusion generation model to extract acoustic environment coding vectors and form joint features for speech enhancement.

Benefits of technology

It significantly improves the speech enhancement algorithm's performance in different acoustic environments, reduces speech distortion, and enhances speech intelligibility and speech recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789701B_ABST
    Figure CN121789701B_ABST
Patent Text Reader

Abstract

本发明公开一种声学环境感知引导的自适应语音增强方法及装置,涉及语音信号处理技术领域。方法包括:对麦克风阵列接收的多通道语音信号进行时频变换,得到复数频谱;对复数频谱进行空间采样,通过一组指向不同空间方向的固定波束成形滤波器,得到空间‑时频三维特征张量;将空间‑时频三维特征张量输入至基于Transformer架构的声学环境感知网络,提取声学环境感知网络的最后一层网络输出作为声学环境编码向量;将复数频谱与声学环境编码向量进行拼接,输入至语音增强网络,得到初步增强语音频谱;将初步增强语音频谱与声学环境编码向量进行拼接,输入至U‑Net扩散生成模型,生成最终的增强语音频谱。采用本发明,可以显著提高语音增强效果。
Need to check novelty before this filing date? Find Prior Art