A controllable speech generation method and system based on emotion trajectory reasoning
By employing a controllable speech generation method based on emotion trajectory reasoning, and utilizing Markov processes and diffusion models for multi-stage emotion modeling and acoustic feature control, this method solves the problem of unnatural multi-emotion speech generation in existing technologies, achieves automated and refined control of speech generation, and improves the naturalness and consistency of speech.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUN YAT SEN UNIV
- Filing Date
- 2026-04-14
- Publication Date
- 2026-06-02
AI Technical Summary
Existing emotional speech generation technologies are ill-suited to scenarios with continuous evolution of multiple emotions, and cannot achieve natural emotional transitions or fine-grained control of multi-stage emotions within speech. This results in abrupt and unnatural transitions in the generated speech during emotional shifts.
A controllable speech generation method based on emotion trajectory reasoning is adopted. The potential emotional state is modeled by Markov process, the location of emotion transfer is detected, a multi-stage emotion trajectory structure is generated, and the acoustic features are finely controlled by diffusion model. Combined with multi-dimensional evaluation and reinforcement learning optimization, the speech generation process is automated and the transition is smooth.
It achieves automatic reasoning and boundary segmentation of multi-stage emotional speech, improves the naturalness and consistency of speech generation, avoids emotional abrupt changes and acoustic breaks, enhances the coherence and realism of speech, and improves the degree of matching user commands and the quality of generation.
Smart Images

Figure CN122135697A_ABST