Controllable Diffusion Speech Model Prosody Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing diffusion-based voice conversion systems lack controllability over prosodic features such as intonation, stress, and speaking rate, relying on a single speaker embedding that spans the entire time and frequency domain.
Innovation Solution
A controllable diffusion-based speech generative model is introduced, which includes a prosody conversion engine that extracts and converts prosody features from source speech to match those of target speech, using a prosody embedding generated from target speech prosody data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single speaker embedding is used for voice conversion, then the system complexity is reduced, but the controllability over prosodic features deteriorates
Solution Approach 1:
The patent segments the single speaker embedding into multiple independent embedding vectors, each representing specific prosodic features (intonation, stress, speaking rate). This segmentation allows independent control of each prosodic dimension while maintaining manageable system complexity through modular architecture.
Solution Approach 2:
The patent transitions from a single-dimensional speaker embedding to a multi-dimensional embedding space where each dimension corresponds to a specific prosodic feature. This dimensional expansion enables fine-grained control over different aspects of speech prosody independently.
2Speed
If prosody features are converted using a single embedding, then the processing speed is maintained, but the precision of prosodic feature control deteriorates
Solution Approach 1:
The patent applies local quality by assigning specific prosodic control capabilities to specific embedding vectors. Each embedding vector is responsible for particular prosodic features (e.g., one for intonation, another for speaking rate), enabling precise local control without requiring global reprocessing of all features.
3Adaptability or versatility
If frame-level intonation control is implemented, then the controllability over prosodic features is improved, but the device complexity increases
Solution Approach 1:
The patent implements dynamic control by enabling frame-level adjustment of intonation parameters. The system can adapt prosodic features on a frame-by-frame basis, providing real-time controllability while using efficient neural network architectures to manage computational complexity.
Data Source
AI summary
Systems and techniques described herein relate to a diffusion-based model for generating converted speech from a source speech based on target speech. For example, a device may extract first prosody data from input data and may generate a content embedding based on the input data. The device may extract second prosody data from target speech, generate a speaker embedding from the target speech, and generate a prosody embedding from the second prosody data. The device may generate, based on the first prosody data and the prosody embedding, converted prosody data. The device may then generate a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding.


