Voice Conversion Using Spectral Warping and Unit Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice conversion technologies face challenges in achieving both high fidelity and naturalness of human speech, with existing methods often resulting in unnatural-sounding synthesized speech due to inherent limitations in spectral conversion and text-to-speech systems.
Innovation Solution
A novel voice conversion method that combines spectral conversion technologies like frequency warping with unit selection from text-to-speech systems, using the converted source speech as a target for unit selection to replace parts of the spectrum, thereby enhancing similarity and naturalness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spectral conversion methods (codebook mapping or GMM) are used to convert source speaker speech to target speaker speech, then similarity to target speaker is improved, but speech quality degrades severely
Solution Approach 1:
The patent combines frequency warping method with unit selection from TTS systems to create a hybrid voice conversion approach. The frequency warping provides spectral transformation for similarity, while unit selection provides natural speech units for quality maintenance, resolving the contradiction between similarity and quality degradation
Solution Approach 2:
The patent changes the spectral parameters through frequency warping to transform the source speaker's spectrum to match the target speaker's characteristics. By adjusting frequency parameters and selecting appropriate speech units, the system achieves both similarity transformation and quality preservation
2Manufacturing precision
If frequency warping is used to convert speech spectra, then speech quality is maintained better than codebook mapping, but detailed spectral differences between speakers remain perceptible
Solution Approach 1:
The patent segments the speech signal into discrete units (words or phrases) and applies unit selection from TTS systems to these segments. This segmentation allows selective replacement of spectral content while maintaining overall speech quality and naturalness
Solution Approach 2:
The patent uses an intermediary approach by introducing unit selection as a mediator between frequency warping and the final speech output. The unit selection process acts as an intermediate step that refines the spectral transformation and eliminates perceptible differences
3Measurement precision
If traditional TTS systems are used to synthesize speech, then speaker identity is preserved, but naturalness of human speech is lost due to unnatural-sounding synthesized speech
Solution Approach 1:
The patent copies natural speech units from the target speaker's corpus and selects them to replace synthesized speech segments. By copying actual natural speech units rather than generating them through TTS synthesis, the system preserves both speaker identity and naturalness
Data Source
AI summary
A method, system and computer program product for voice conversion. The method includes performing speech analysis on the speech of a source speaker to achieve speech information; performing spectral conversion based on said speech information, to at least achieve a first spectrum similar to the speech of a target speaker; performing unit selection on the speech of said target speaker at least using said first spectrum as a target; replacing at least part of said first spectrum with the spectrum of the selected target speaker's speech unit; and performing speech reconstruction at least based on the replaced spectrum.


