Prompt-Guided Speech Synthesis for Accent-Filtered Voice Personalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech generation technologies struggle to generate personalized speech that accurately captures the desired speech features of a target object while avoiding undesired characteristics such as accent, requiring improved methods for fine-tuned customization.

Innovation Solution

A speech synthesis method that constructs input sequences using placeholders or speech feature representations within a sequence template, trained with a target model to generate target speech content, allowing for precise control over speech attributes like timbre and prosody.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech features are extracted from reference speech using machine learning models, then personalized speech generation is achieved, but undesired characteristics such as accent are retained

Engineering Contradiction:
Improvepersonalized speech generationVSAvoidundesired characteristics like accent
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts and separates desired speech features (timbre, prosody) from undesired features (accent) through the sequence template structure. The placeholder mechanism allows selective extraction of specific speech characteristics while filtering out unwanted elements like accent, enabling personalized speech generation without carrying over undesirable traits.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The sequence template applies different quality requirements to different parts of the speech generation process. By specifying particular speech features (timbre, prosody) in certain positions while excluding others (accent) in the placeholder segments, the system achieves local quality control over which speech characteristics are preserved and which are filtered.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If speech features are extracted from prompt speech content, then speech customization is achieved, but control over specific speech attributes is reduced

Engineering Contradiction:
Improvespeech customizationVSAvoidcontrol over speech attributes
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the speech feature extraction process into distinct components through the sequence template structure. Different placeholders can represent different speech attributes (timbre, prosody, pitch), allowing independent control and customization of each attribute while maintaining overall speech coherence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system enables precise control over speech attributes by allowing parameter-specific customization in the sequence template. Users can adjust individual speech parameters (timbre, prosody, pitch) independently through the placeholder mechanism, achieving manufacturing precision in speech attribute control while maintaining adaptability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250356841A1Speech synthesis
Publication Date: 2025.11.20 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250356841A1 patent drawing
  • US20250356841A1 patent drawing
  • US20250356841A1 patent drawing

AI summary

Embodiments of the disclosure relate to speech synthesis. A method provided herein includes: constructing, based on target text and prompt speech content, an input sequence corresponding to a sequence template, wherein the sequence template includes a placeholder, and a sequence segment, in the input sequence, corresponding to the placeholder is: preset content independent of the prompt speech content, or a speech feature representation generated based on the prompt speech content; and processing the input sequence with a target model to generate target speech content corresponding to the target text, wherein the target model is trained with a set of training sequences constructed based on the sequence template, the set of training sequences corresponds to a set of training speech content, and the set of training sequences is constructed by replacing the placeholder with the preset content or a training speech feature representation corresponding to respective training speech content.