Controllable Diffusion Speech Model Prosody Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing diffusion-based voice conversion systems lack controllability over prosodic features such as intonation, stress, and speaking rate, relying on a single speaker embedding that spans the entire time and frequency domain.

Innovation Solution

A controllable diffusion-based speech generative model is introduced, which includes a prosody conversion engine that extracts and converts prosody features from source speech to match those of target speech, using a prosody embedding generated from target speech prosody data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single speaker embedding is used for voice conversion, then the system complexity is reduced, but the controllability over prosodic features deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidcontrollability over prosodic features
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments the single speaker embedding into multiple independent embedding vectors, each representing specific prosodic features (intonation, stress, speaking rate). This segmentation allows independent control of each prosodic dimension while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a single-dimensional speaker embedding to a multi-dimensional embedding space where each dimension corresponds to a specific prosodic feature. This dimensional expansion enables fine-grained control over different aspects of speech prosody independently.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Speed

If prosody features are converted using a single embedding, then the processing speed is maintained, but the precision of prosodic feature control deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidprecision of prosodic feature control
Core Design Contradiction:
SpeedVSManufacturing precision

Solution Approach 1:

The patent applies local quality by assigning specific prosodic control capabilities to specific embedding vectors. Each embedding vector is responsible for particular prosodic features (e.g., one for intonation, another for speaking rate), enabling precise local control without requiring global reprocessing of all features.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If frame-level intonation control is implemented, then the controllability over prosodic features is improved, but the device complexity increases

Engineering Contradiction:
Improvecontrollability over prosodic featuresVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic control by enabling frame-level adjustment of intonation parameters. The system can adapt prosodic features on a frame-by-frame basis, providing real-time controllability while using efficient neural network architectures to manage computational complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250078810A1Controllable diffusion-based speech generative model
Publication Date: 2025.03.06 QUALCOMM INC
  • US20250078810A1 patent drawing
  • US20250078810A1 patent drawing
  • US20250078810A1 patent drawing

AI summary

Systems and techniques described herein relate to a diffusion-based model for generating converted speech from a source speech based on target speech. For example, a device may extract first prosody data from input data and may generate a content embedding based on the input data. The device may extract second prosody data from target speech, generate a speaker embedding from the target speech, and generate a prosody embedding from the second prosody data. The device may generate, based on the first prosody data and the prosody embedding, converted prosody data. The device may then generate a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding.