Bandwidth-Extended Speech via Discriminative DNN Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing bandwidth extension methods for telephony networks, such as those using deep neural networks, often result in over-smoothing, where vowels are extended too strongly and fricatives are not extended enough, leading to degraded voice quality due to the inability to correctly predict high-frequency formants across different speakers.
Innovation Solution
Incorporating additional discriminative terms into the cost function, such as the average power ratio (APR) between phoneme classes, to force the deep neural network to maintain better separability of fricatives and vowels, thereby improving the quality of bandwidth-extended speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If standard mean squared error cost function is used for training the deep neural network, then the network can be trained to extend narrow-band speech to wideband speech, but over-smoothing occurs where vowels are extended too strongly and fricatives are not extended enough, degrading voice quality
Solution Approach 1:
The patent changes the training parameter by introducing a discriminative cost function that incorporates phoneme class separability constraints. This modifies the optimization landscape to prevent over-smoothing by penalizing predictions that collapse distinct phoneme classes, thereby maintaining voice quality while achieving accurate bandwidth extension.
Solution Approach 2:
The patent implements feedback by using phoneme class labels during training to provide discriminative guidance. The cost function incorporates feedback about predicted phoneme class separability, adjusting the network weights to maintain distinct characteristics of different phoneme classes, thus preventing the over-smoothing effect that degrades voice quality.
2Adaptability or versatility
If the deep neural network is trained to predict high-frequency formants, then bandwidth extension can be achieved, but the network fails to correctly predict high-frequency formants across different speakers, leading to degraded voice quality
Solution Approach 1:
The patent changes the training objective by adding phoneme class separability as a discriminative constraint. This modifies the parameter optimization to preserve speaker-specific and phoneme-specific characteristics in the predicted high-frequency formants, enabling accurate formant prediction across different speakers while maintaining adaptability for bandwidth extension.
Solution Approach 2:
The patent applies local quality by treating different phoneme classes differently during training. The discriminative cost function applies localized constraints to preserve the specific spectral characteristics of each phoneme class, ensuring that high-frequency formants are predicted accurately for each class rather than applying a uniform smoothing effect across all inputs.
3Quantity of substance
If conventional bandwidth extension is applied, then narrow-band speech can be extended to wideband speech, but the separation between phoneme classes is lost due to over-smoothing
Solution Approach 1:
The patent changes the training parameters by incorporating phoneme class separability constraints into the cost function. This modification preserves the separation between phoneme classes during the bandwidth extension process, preventing information loss while achieving the desired spectral bandwidth expansion from narrow-band to wideband speech.
Solution Approach 2:
The patent uses phoneme class labels as feedback during training to maintain separability. The discriminative cost function provides feedback about the predicted separability of different phoneme classes, adjusting the network to preserve this information while extending the spectral bandwidth, thus preventing the loss of phoneme class distinctions.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method, computer program product, and computer system for transforming, by a computing device, a speech signal into a speech signal representation. A regression deep neural network may be trained with a cost function to minimize a mean squared error between actual values of the speech signal representation and estimated values of the speech signal representation, wherein the cost function may include one or more discriminative terms. Bandwidth of the speech signal may be extended by extending the speech signal representation of the speech signal using the regression deep neural