Prompt-Guided Speech Synthesis for Accent-Filtered Voice Personalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech generation technologies struggle to generate personalized speech that accurately captures the desired speech features of a target object while avoiding undesired characteristics such as accent, requiring improved methods for fine-tuned customization.
Innovation Solution
A speech synthesis method that constructs input sequences using placeholders or speech feature representations within a sequence template, trained with a target model to generate target speech content, allowing for precise control over speech attributes like timbre and prosody.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech features are extracted from reference speech using machine learning models, then personalized speech generation is achieved, but undesired characteristics such as accent are retained
Solution Approach 1:
The patent extracts and separates desired speech features (timbre, prosody) from undesired features (accent) through the sequence template structure. The placeholder mechanism allows selective extraction of specific speech characteristics while filtering out unwanted elements like accent, enabling personalized speech generation without carrying over undesirable traits.
Solution Approach 2:
The sequence template applies different quality requirements to different parts of the speech generation process. By specifying particular speech features (timbre, prosody) in certain positions while excluding others (accent) in the placeholder segments, the system achieves local quality control over which speech characteristics are preserved and which are filtered.
2Adaptability or versatility
If speech features are extracted from prompt speech content, then speech customization is achieved, but control over specific speech attributes is reduced
Solution Approach 1:
The patent segments the speech feature extraction process into distinct components through the sequence template structure. Different placeholders can represent different speech attributes (timbre, prosody, pitch), allowing independent control and customization of each attribute while maintaining overall speech coherence.
Solution Approach 2:
The system enables precise control over speech attributes by allowing parameter-specific customization in the sequence template. Users can adjust individual speech parameters (timbre, prosody, pitch) independently through the placeholder mechanism, achieving manufacturing precision in speech attribute control while maintaining adaptability.
Data Source
AI summary
Embodiments of the disclosure relate to speech synthesis. A method provided herein includes: constructing, based on target text and prompt speech content, an input sequence corresponding to a sequence template, wherein the sequence template includes a placeholder, and a sequence segment, in the input sequence, corresponding to the placeholder is: preset content independent of the prompt speech content, or a speech feature representation generated based on the prompt speech content; and processing the input sequence with a target model to generate target speech content corresponding to the target text, wherein the target model is trained with a set of training sequences constructed based on the sequence template, the set of training sequences corresponds to a set of training speech content, and the set of training sequences is constructed by replacing the placeholder with the preset content or a training speech feature representation corresponding to respective training speech content.


