Latent Representation Voice Customization via Normalizing Flows

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech processing systems lack effective methods for voice customization and anonymization, particularly in the context of synthetic speech generation, which limits their ability to efficiently modify voice characteristics during compression and transmission.

Innovation Solution

The system employs normalizing flows and machine learning models to transform speech data into a latent representation, allowing for voice customization by modifying speech attributes before or after decompression, enabling changes in voice characteristics such as age, gender, and accent, and anonymizing the source voice.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speech data is compressed for transmission and storage, then transmission efficiency and storage capacity are improved, but voice characteristics and speaker identity are lost or altered

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidvoice characteristics
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The speech data is segmented into two separate components: acoustic features (capturing voice characteristics and speaker identity) and linguistic features (capturing speech content). This segmentation allows independent processing of each component, enabling lossless compression of linguistic features while preserving acoustic features for voice customization and anonymization operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A neural network model acts as an intermediary that reconstructs speech data from compressed linguistic features and selected acoustic features. This intermediary enables the system to recover or transform voice characteristics after compression by controlling which acoustic features are preserved or modified during reconstruction.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If voice customization is implemented by processing speech data before compression, then voice modification flexibility is improved, but processing complexity and computational resources increase

Engineering Contradiction:
Improvevoice modification flexibilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary extraction of acoustic features from speech data before compression. This preliminary action separates voice characteristics from speech content, allowing subsequent voice customization or anonymization to be achieved by selecting or modifying acoustic features without requiring complex processing of the entire speech signal during transmission or storage.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Voice customization is achieved by changing parameters in the acoustic feature representation rather than manipulating the raw speech signal. The neural network model allows selective modification of acoustic parameters (such as speaker identity, pitch, or timbre) while keeping linguistic content intact, reducing processing complexity compared to traditional speech manipulation methods.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If speech data is compressed with high compression ratios, then transmission bandwidth and storage requirements are reduced, but speech quality and voice fidelity deteriorate

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidspeech quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

By segmenting speech into acoustic and linguistic features, the system can apply different compression strategies to each component. Linguistic features can be heavily compressed with minimal quality loss, while acoustic features are preserved at higher fidelity or selectively modified, achieving high overall compression ratios while maintaining speech quality and enabling voice customization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the representation parameters of speech data from raw waveforms to feature vectors extracted by neural networks. This parameter transformation enables more efficient compression while preserving essential speech quality attributes, as the feature representation captures the most important characteristics with fewer data points.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250014567A1Voice customization for synthetic speech generation
Publication Date: 2025.01.09 AMAZON TECH INC
  • US20250014567A1 patent drawing
  • US20250014567A1 patent drawing
  • US20250014567A1 patent drawing

AI summary

Voice customization is an application of voice synthesis that involves synthesizing speech having certain voice characteristics, and/or modifying the voice characteristics of human speech. Certain techniques for voice customization may be used in conjunction with compressing speech for storage and/or transmission. For example, speech may be received at a first device and transformed into a latent representation and/or compressed for storage and/or transmission to a second device. The system may use normalizing flows to transform the source audio to a latent representation having a desired variable distribution, and to transform the latent representation back into audio data. A flow model may be conditioned using first speech attributes when transforming the source audio, and an inverse flow model may use second speech attributes when transforming the latent representation back into audio data. The first and/or second speech attributes may be modified to alter voice characteristics of the transmitted speech.