Emotion Control in Speech Synthesis via Acoustic Feature Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis methods produce emotionless audio due to the lack of emotional sound libraries, which are difficult and inefficient to create, resulting in weak expressive force in synthesized speech.
Innovation Solution
A method that acquires text and specifies an emotion type, determines corresponding acoustic features, and inputs these into a pre-trained speech synthesis model trained without the specified emotion type, generating target audio with the desired emotional characteristics using existing corpora.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If an emotional sound library is created to enhance the expressive force of synthesized audio, then the emotion and expressiveness of the audio is improved, but the workload and time required for recording staff increases significantly
Solution Approach 1:
The patent uses acoustic features extracted from emotional audio samples as templates or copies to guide the synthesis process. Instead of requiring recording staff to manually create emotional sound libraries, the system copies acoustic characteristics (pitch, energy, duration) from existing emotional audio to generate new emotional speech, dramatically reducing recording workload while maintaining emotional expressiveness
Solution Approach 2:
The patent replaces the mechanical process of manual emotional audio recording with an automated computational system. The speech synthesis model automatically generates emotional audio by controlling acoustic features through algorithms, substituting the mechanical recording process with an automated information processing system that eliminates the need for extensive manual recording work
2Reliability
If a sound library with emotion is created to make synthesized audio more expressive, then the emotion of the audio is improved, but the workload for recording staff increases heavily
Solution Approach 1:
The patent extracts essential emotional characteristics from audio samples in the form of acoustic features (pitch contour, energy distribution, duration patterns). By taking out only the critical emotional elements rather than requiring complete emotional sound libraries, the system reduces the complexity of data preparation while preserving emotional expressiveness in the synthesized output
Solution Approach 2:
The patent controls emotional expression by adjusting acoustic parameters (fundamental frequency, energy, duration) generated by the speech synthesis model. By changing these parameters according to acoustic features derived from emotional samples, the system achieves emotional audio output without requiring complex emotional sound libraries, thereby reducing workload complexity
3Reliability
If emotional sound libraries are manually created to enhance audio expressiveness, then the expressiveness of synthesized speech is improved, but the time required for creation increases significantly
Solution Approach 1:
The patent performs preliminary extraction of acoustic features from emotional audio samples before the actual synthesis process. By pre-processing and storing acoustic characteristics (pitch, energy, duration patterns) in a structured format, the system prepares emotional expressions in advance, enabling rapid generation of emotional speech during synthesis without time-consuming manual recording during the actual creation phase
Solution Approach 2:
The system copies acoustic feature patterns from pre-processed emotional samples to guide the speech synthesis model. This copying mechanism allows rapid generation of emotional speech by reusing extracted acoustic characteristics, dramatically reducing the time required to create emotional audio compared to manual recording while maintaining high expressiveness
Data Source
AI summary
The present disclosure relates to a speech synthesis method, apparatus, readable medium and electronic device, which relates to the technical field of electronic information processing. The method comprises: acquiring a text to be synthesized and a specified emotion type (101), determining specified acoustic features corresponding to the specified emotion type (102), and inputting the text to be synthesized and the specified acoustic features into a pre-trained speech synthesis model, to acquire a target audio with the specified emotion type corresponding to the text to be synthesized which is output by the speech synthesis model (102). The acoustic features of the target audio match with the specified acoustic features, and the speech synthesis model is trained from a corpus without the specified emotion type.


