Voice Morphing via Intermediary Standard Speaker
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice morphing technologies face challenges in achieving high-quality voice conversion while maintaining similarity to the original speaker's voice, often resulting in impaired sound quality, especially when the original and target voices have significant differences.
Innovation Solution
The method involves selecting a standard speaker from a TTS database whose voice is similar to the original speaker, performing voice synthesis, and then applying voice morphing to accurately mimic the original voice, ensuring the converted voice is closer to the original voice features while maintaining similarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If frequency warping method is used for voice conversion, then the voice similarity to target speaker is improved, but the sound quality is impaired when original and target voices differ significantly
Solution Approach 1:
The patent introduces an intermediary standard speaker whose voice characteristics lie between the original speaker and target speaker. The voice conversion process proceeds in two stages: first from original speaker to standard speaker, then from standard speaker to target speaker. This intermediary approach reduces the magnitude of frequency warping at each stage, thereby maintaining sound quality while achieving voice similarity to the target speaker.
Solution Approach 2:
The patent segments the voice conversion process into multiple intermediate steps rather than performing direct conversion. By dividing the conversion path into original speaker → standard speaker → target speaker, the system applies frequency warping in smaller increments, reducing cumulative quality degradation while still achieving the desired voice transformation.
2Device complexity
If direct voice morphing is applied from original speaker to target speaker, then the conversion process is simple, but the sound quality impairment increases rapidly when voices differ far
Solution Approach 1:
The standard speaker serves as a mediator that simplifies the conversion process by providing a intermediate representation. Rather than directly converting between dissimilar voices, the system uses the standard speaker as a bridge, making each conversion step more manageable and maintaining higher sound quality throughout the process.
Solution Approach 2:
The patent performs preliminary voice synthesis using the standard speaker before applying frequency warping to match the target speaker. This preliminary action creates a base voice that is already partially transformed, reducing the magnitude of subsequent warping operations and minimizing quality impairment.
3Manufacturing precision
If standard speaker selection is performed from TTS database, then the voice synthesis quality is improved, but additional processing steps are required
Solution Approach 1:
The system performs preliminary selection of a standard speaker from the TTS database based on voice characteristic matching. This preliminary action prepares an optimal intermediate representation before the actual voice conversion, ensuring that the subsequent frequency warping operations start from the best possible base and achieve higher overall quality.
Solution Approach 2:
The patent implements feedback mechanisms to evaluate and select the most suitable standard speaker from the TTS database by comparing voice characteristics. This feedback-driven selection ensures that the intermediary speaker optimally bridges the gap between original and target voices, justifying the additional processing steps through improved conversion quality.
Data Source
AI summary
The invention proposes a method and apparatus for significantly improving the quality of voice morphing and guaranteeing the similarity of converted voice. The invention sets several standard speakers in a TTS database, and selects the voices of different standard speakers for speech synthesis according to different roles, wherein the voice of the selected standard speaker is similar to the original role to a certain extent. Then the invention further performs voice morphing on the standard voice similar to the original voice to a certain extent, in order to accurately mimic the voice of the original speaker, so as to make the converted voice closer to the original voice features while guaranteeing the similarity.


