Bandwidth-Extended Speech via Discriminative DNN Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing bandwidth extension methods for telephony networks, such as those using deep neural networks, often result in over-smoothing, where vowels are extended too strongly and fricatives are not extended enough, leading to degraded voice quality due to the inability to correctly predict high-frequency formants across different speakers.

Innovation Solution

Incorporating additional discriminative terms into the cost function, such as the average power ratio (APR) between phoneme classes, to force the deep neural network to maintain better separability of fricatives and vowels, thereby improving the quality of bandwidth-extended speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If standard mean squared error cost function is used for training the deep neural network, then the network can be trained to extend narrow-band speech to wideband speech, but over-smoothing occurs where vowels are extended too strongly and fricatives are not extended enough, degrading voice quality

Engineering Contradiction:
Improvebandwidth extension accuracyVSAvoidvoice quality
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent changes the training parameter by introducing a discriminative cost function that incorporates phoneme class separability constraints. This modifies the optimization landscape to prevent over-smoothing by penalizing predictions that collapse distinct phoneme classes, thereby maintaining voice quality while achieving accurate bandwidth extension.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback by using phoneme class labels during training to provide discriminative guidance. The cost function incorporates feedback about predicted phoneme class separability, adjusting the network weights to maintain distinct characteristics of different phoneme classes, thus preventing the over-smoothing effect that degrades voice quality.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If the deep neural network is trained to predict high-frequency formants, then bandwidth extension can be achieved, but the network fails to correctly predict high-frequency formants across different speakers, leading to degraded voice quality

Engineering Contradiction:
Improvebandwidth extension capabilityVSAvoidformant prediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the training objective by adding phoneme class separability as a discriminative constraint. This modifies the parameter optimization to preserve speaker-specific and phoneme-specific characteristics in the predicted high-frequency formants, enabling accurate formant prediction across different speakers while maintaining adaptability for bandwidth extension.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies local quality by treating different phoneme classes differently during training. The discriminative cost function applies localized constraints to preserve the specific spectral characteristics of each phoneme class, ensuring that high-frequency formants are predicted accurately for each class rather than applying a uniform smoothing effect across all inputs.

Inventive Principle:
Principle #3Local quality

3Quantity of substance

If conventional bandwidth extension is applied, then narrow-band speech can be extended to wideband speech, but the separation between phoneme classes is lost due to over-smoothing

Engineering Contradiction:
Improvespectral bandwidthVSAvoidphoneme class separability
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent changes the training parameters by incorporating phoneme class separability constraints into the cost function. This modification preserves the separation between phoneme classes during the bandwidth extension process, preventing information loss while achieving the desired spectral bandwidth expansion from narrow-band to wideband speech.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses phoneme class labels as feedback during training to maintain separability. The discriminative cost function provides feedback about the predicted separability of different phoneme classes, adjusting the network to preserve this information while extending the spectral bandwidth, thus preventing the loss of phoneme class distinctions.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3785189B1System and method for discriminative training of regression deep neural networks
Publication Date: 2023.12.13 CERENCE OPERATING CO
  • EP3785189B1 patent drawingFigure 1
  • EP3785189B1 patent drawingFigure 2
  • EP3785189B1 patent drawingFigure 3

AI summary

A method, computer program product, and computer system for transforming, by a computing device, a speech signal into a speech signal representation. A regression deep neural network may be trained with a cost function to minimize a mean squared error between actual values of the speech signal representation and estimated values of the speech signal representation, wherein the cost function may include one or more discriminative terms. Bandwidth of the speech signal may be extended by extending the speech signal representation of the speech signal using the regression deep neural