Modified Diffusion Model for Speech Restoration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech enhancement methods fail to effectively restore degraded speech quality due to non-linear and lossy compression processes, leading to reduced intelligibility and suitability for applications like automatic speech recognition and speaker identification.

Innovation Solution

A modified diffusion model incorporating a deep CNN upsampler is trained to invert lossy transformations and restore missing information in speech signals, using a diffusion-based vocoder to generate estimated original speech from degraded mel-spectrum samples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If speech compression algorithms are used to reduce sampling rate and compress input, then file size and transmission efficiency are improved, but speech quality and intelligibility deteriorate

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidspeech quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The patent replaces traditional mechanical signal processing methods (compression algorithms, filtering) with a neural network-based diffusion model. The model learns to reconstruct high-quality speech waveforms from compressed representations by modeling the probability distribution of speech signals, thereby avoiding the quality loss inherent in deterministic compression methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of speech representation by using mel-spectrogram features as input to the diffusion model rather than raw waveforms. This parameter transformation allows the model to capture essential speech characteristics in a compressed form while reconstructing high-fidelity waveforms through the diffusion process.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If clipping of speech is applied to compress signal, then data transmission is improved, but high frequency content is introduced with negative impact on quality

Engineering Contradiction:
Improvedata transmissionVSAvoidspeech quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent introduces a neural network-based diffusion model as an intermediary between the compressed speech representation and the reconstructed waveform. This intermediary learns to smoothly transition from compressed representations to high-quality waveforms, avoiding the abrupt artifacts introduced by traditional clipping methods.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Object-affected harmful factors

If traditional speech enhancement methods are used to remove background noise, then noise reduction is improved, but non-linear compression artifacts are introduced that reduce intelligibility

Engineering Contradiction:
Improvenoise reductionVSAvoidintelligibility
Core Design Contradiction:
Object-affected harmful factorsVSManufacturing precision

Solution Approach 1:

The patent replaces non-linear traditional speech enhancement methods with a probabilistic diffusion model that models the underlying distribution of clean speech. This approach naturally handles noise removal while preserving speech intelligibility by learning the statistical characteristics of clean speech rather than applying deterministic filtering operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11978466B2Systems, methods, and apparatuses for restoring degraded speech via a modified diffusion model
Publication Date: 2024.05.07 THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA
  • US11978466B2 patent drawing
  • US11978466B2 patent drawing
  • US11978466B2 patent drawing

AI summary

Systems, methods, and apparatuses to restore degraded speech via a modified diffusion model are described. An exemplary system is specially configured to train a diffusion-based vocoder containing an upsampler, based on pairing original speech x and degraded speech mel-spectrum mT samples; train a deep convoluted neural network (CNN) upsampler based on a mean absolute error loss to match the estimated original speech {circumflex over (x)}′ outputted by the diffusion-based vocoder by extracting the upsampler, generating a reference conditioner, and generating a weighted altered conditioner cT<sub2>n</sub2>′. The system further optimizes speech quality to invert non-linear transformation and estimate lost data by feeding the degraded mel-spectrum mT through the CNN upsampler and feeding the degraded mel-spectrum mT through the diffusion-based vocoder. The system then generates estimated original speech {circumflex over (x)}′ based on the corresponding degraded speech mel-spectrum mT. Other related embodiments are described.