Robustness / performance improvements for deep learning-based speech enhancement against artifacts and distortion.

A multi-stage deep learning framework with a separator and improver addresses speech enhancement artifacts and distortions, enhancing speech quality by using autoencoders and neural networks to refine denoised signals.

JP7863597B2Active Publication Date: 2026-05-21DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2024-09-25
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing deep learning-based speech enhancement techniques suffer from artifacts and distortions, particularly when dealing with different acoustic environments, leading to suboptimal performance.

Method used

A multi-stage deep learning-based framework comprising a separator and an improver, where the separator performs an initial denoising step, and the improver reduces distortions and artifacts introduced by the separator, utilizing autoencoders, recurrent neural networks, and generative models to refine the output.

Benefits of technology

The framework effectively reduces artifacts and distortions, improving the perceptual quality of speech signals by enhancing the speech component and suppressing noise, even in challenging acoustic conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007863597000004
    Figure 0007863597000004
  • Figure 0007863597000005
    Figure 0007863597000005
  • Figure 0007863597000006
    Figure 0007863597000006
Patent Text Reader

Abstract

To provide a method for processing audio signals, as well as a corresponding device, a computer program, and computer-readable storage media.SOLUTION: A method 1000 includes the steps of: applying emphasis to a first component of an audio signal and / or applying suppression to a second component of the audio signal with respect to the first component S1010; and modifying an output of the step S1010 by applying a deep learning based model to the output to remove artifacts and / or distortions introduced into the audio signal in the step S1010, and perceptually improving the first component of the audio signal S1020.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art