Speech Synthesis Expression Control Using Contrastive Text Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech synthesis systems fail to accurately describe the expression state of generated speech, resulting in a poor presentation effect.

Innovation Solution

A method involving a contrastive learning module to obtain a reference description feature, construct a target description feature, and generate target speech content based on an input phoneme sequence to control the expression state.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech synthesis systems are used, then speech content can be generated, but the expression state cannot be accurately described resulting in poor presentation effect

Engineering Contradiction:
Improveexpression state description accuracyVSAvoidpresentation effect
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent segments the speech synthesis process into distinct feature extraction and expression control components. The reference description feature is separated from the target description feature, allowing independent optimization of each component's contribution to expression state accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation by introducing expression state features as additional dimensions in the speech synthesis process. This transforms the synthesis from basic phoneme generation to controlled expression generation through parameter expansion.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If expression state control is added to speech synthesis, then presentation effect is enhanced, but system complexity increases

Engineering Contradiction:
Improvepresentation effectVSAvoidsystem structure
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent creates a universal framework where the reference description feature can be extracted from any speech content and applied to control expression states across different synthesis scenarios. This multi-functional approach handles various expression requirements through a unified mechanism.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces description features as intermediary representations between input text and output speech. These features act as mediators that carry expression state information through the synthesis process without requiring direct complex interactions between all system components.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250356840A1Method, apparatus, device, and storage medium for speech synthesis
Publication Date: 2025.11.20 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250356840A1 patent drawing
  • US20250356840A1 patent drawing
  • US20250356840A1 patent drawing

AI summary

A method, an apparatus, a device, and a storage medium for speech synthesis are provided. A reference description feature corresponding to prompted speech content is obtained, the reference description feature includes a text encoding representation determined by processing the prompted speech content with a contrastive learning module, and the text encoding representation describes a first expression state of the prompted speech content. Based on the reference description feature, a target description feature for indicating a target expression state is constructed. Target speech content corresponding to the target expression state is generated based on an input phoneme sequence including the target description feature.