Neural Speech Style Extraction for Noise-Robust Pitch Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis systems struggle with accurate pitch tracking, particularly in noisy environments and real-time applications, due to the lack of reliable pitch annotations and the inflexibility to control style-related elements like pitch, voicing, and emotion, leading to inaccurate and unreliable results.
Innovation Solution
A neural network-based system that separates input speech into content and style components, using a deconstructor module to extract pitch and other style elements without explicit labels, and a reconstructor module to reconstruct the speech with preserved pitch, enhancing robustness against noise and degradation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional pitch tracking methods (autocorrelation or harmonic product spectrogram) are used, then the system structure is simple and easy to implement, but the pitch tracking accuracy deteriorates in noisy environments and real-time scenarios
Solution Approach 1:
The patent replaces traditional mechanical signal processing methods (autocorrelation and harmonic product spectrogram) with a neural network-based system. The neural network learns to extract pitch and style information directly from audio signals, substituting complex mathematical computations with trained model inference that adapts to noisy conditions and provides robust pitch tracking in real-time scenarios.
Solution Approach 2:
The patent transforms the approach by changing from explicit pitch parameter estimation to latent space representation learning. The neural network maps audio signals to a compressed latent representation that captures both content and style information, allowing for more accurate pitch tracking without requiring explicit pitch annotations or complex mathematical processing.
2Measurement precision
If manual or semi-automatic pitch annotation is performed to create training data, then the training data quality is high, but the time and computational resources required increase significantly
Solution Approach 1:
The patent employs self-supervised learning where the neural network learns to extract pitch and style information by analyzing the audio signals themselves without requiring external annotations. The model uses the audio data as both input and teacher signal, allowing it to automatically learn accurate pitch representations without manual labeling or complex data preparation processes.
Solution Approach 2:
The patent uses unsupervised learning techniques where the neural network creates its own internal representations of pitch and style by copying and analyzing the temporal patterns and spectral characteristics directly from the audio signals. This eliminates the need for human annotators to create training data while maintaining high accuracy in pitch extraction.
3Adaptability or versatility
If traditional pitch tracking is used for real-time applications, then the computational requirements are low, but the system lacks flexibility to control style-related elements like pitch, voicing, and emotion
Solution Approach 1:
The patent segments the audio representation into distinct content and style components in the latent space. The neural network separates pitch information, voicing characteristics, and emotional cues into independent dimensions, allowing for selective manipulation of style elements while maintaining content integrity. This segmentation enables fine-grained control over synthesized speech characteristics.
Solution Approach 2:
The patent implements a dynamic system where the neural network can adaptively adjust style parameters in real-time based on input conditions. The model dynamically learns to modulate pitch, voicing, and emotional expressions by analyzing contextual information and adjusting its latent representations accordingly, providing flexible style control that responds to varying speech scenarios.
Data Source
AI summary
The disclosed technology relates to methods, speech processing systems, and non-transitory computer readable media for style extraction in speech synthesis. In some examples, one or more content elements and one or more non-content elements are extracted from input audio data obtained via an audio interface and corresponding to input speech. The one or more non-content elements comprise style elements comprising at least an input pitch. A trained autoencoder is applied to encode the input pitch in a latent representation comprising a low-dimensional vector and combine the one or more content elements and the one or more non-content elements based on the low-dimensional vector to generate a new representation of the input speech. Output audio data is then generated and provided based on the new representation of the input speech. The output audio data comprises a pitch-consistent reconstruction of the input speech.


