A controllable speech generation method and system based on emotion trajectory reasoning

By employing a controllable speech generation method based on emotion trajectory reasoning, and utilizing Markov processes and diffusion models for multi-stage emotion modeling and acoustic feature control, this method solves the problem of unnatural multi-emotion speech generation in existing technologies, achieves automated and refined control of speech generation, and improves the naturalness and consistency of speech.

CN122135697APending Publication Date: 2026-06-02SUN YAT SEN UNIV +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUN YAT SEN UNIV
Filing Date
2026-04-14
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing emotional speech generation technologies are ill-suited to scenarios with continuous evolution of multiple emotions, and cannot achieve natural emotional transitions or fine-grained control of multi-stage emotions within speech. This results in abrupt and unnatural transitions in the generated speech during emotional shifts.

Method used

A controllable speech generation method based on emotion trajectory reasoning is adopted. The potential emotional state is modeled by Markov process, the location of emotion transfer is detected, a multi-stage emotion trajectory structure is generated, and the acoustic features are finely controlled by diffusion model. Combined with multi-dimensional evaluation and reinforcement learning optimization, the speech generation process is automated and the transition is smooth.

Benefits of technology

It achieves automatic reasoning and boundary segmentation of multi-stage emotional speech, improves the naturalness and consistency of speech generation, avoids emotional abrupt changes and acoustic breaks, enhances the coherence and realism of speech, and improves the degree of matching user commands and the quality of generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135697A_ABST
    Figure CN122135697A_ABST
Patent Text Reader

Abstract

This invention discloses a controllable speech generation method and system based on emotion trajectory reasoning. The method includes: an instruction understanding and emotion reasoning step: performing joint semantic parsing on the input speech text content and user instruction information, performing semantic reasoning on the explicit and implicit emotional needs in the user instruction information, constructing a multi-stage emotion trajectory structure of the target speech, and generating prosodic control factors and pronunciation control information corresponding to each emotion stage to form stage-level speech control factors; a controllable speech generation step: inputting the stage-level speech control factors and the speech text content into a conditional speech generation model, generating multiple candidate speech under the constraints of the stage-level speech control factors; and an optimization and verification step: performing multi-dimensional comprehensive evaluation on the multiple candidate speech, optimizing the generation process based on the evaluation results, and selecting the optimal candidate speech as the final output speech.
Need to check novelty before this filing date? Find Prior Art