Virtual Teacher Voice Synthesis Using Emotion Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual teacher systems face limitations in generating simulated voices with dynamic emotion features and voice styles, leading to fixed timbres, low correlation with real teachers' voices, and inadequate rapid synthesis capabilities, which hinders their application in educational settings.
Innovation Solution
A method and terminal for generating simulated voices involve collecting real voice samples, converting them into text sequences, constructing emotion and tone training sets, and using a lexical item emotion model to extract emotion features and voice style features, enabling the synthesis of voices with emotional changes and tone features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing voice synthesis methods are used to obtain real person voices through natural language processing and training, then the timbre can be obtained, but the system complexity and cost increase significantly, making it difficult for users to replace voices with other teachers' voices
Solution Approach 1:
The patent uses voiceprint extraction and copying technology to capture the timbre characteristics of real teachers' voices. By extracting voiceprint features from original voice samples and storing them in a database, the system can replicate and reproduce specific teachers' voice characteristics without requiring complex retraining processes, enabling easy voice replacement while maintaining system simplicity
Solution Approach 2:
The patent extracts the essential voiceprint features (timbre characteristics) from complete voice samples. By separating and storing only the critical timbre information in a structured database format, the system eliminates the need to process entire voice datasets during synthesis, reducing computational complexity while preserving the ability to replace voices
2Adaptability or versatility
If star or idol voices are used as samples to enhance affinity, then the emotional connection may be improved, but the correlation with real teachers' voices decreases, making it difficult to awaken the sense of presence in learning
Solution Approach 1:
The patent changes the selection parameter from celebrity voices to actual teacher voices. By modifying the source material parameter to use real teachers' voice samples instead of stars' or idols' voices, the system maintains emotional connection while significantly improving voice correlation accuracy and sense of presence in educational contexts
Solution Approach 2:
The patent copies the actual voice characteristics of real teachers rather than using celebrity voices. By extracting and reproducing the authentic timbre and speech patterns of teachers who actually teach the students, the system preserves both emotional connection and high correlation accuracy, creating a more immersive learning experience
3Productivity
If cloud-edge-end architecture is used to enable rapid voice synthesis, then the synthesis speed can be improved, but ensuring consistency of synthesized voice styles and emotional features with real people requires complex coordination
Solution Approach 1:
The patent segments the voice synthesis system into distinct functional modules: voiceprint extraction, feature storage in database, and synthesis execution. This segmentation allows each component to operate independently and efficiently, enabling rapid synthesis through cloud-edge-end architecture while maintaining voice style consistency through standardized feature interfaces
Solution Approach 2:
The patent introduces a voiceprint database as an intermediary layer between the cloud training system and edge synthesis devices. This intermediary stores standardized voice features and enables consistent voice style reproduction across different devices and platforms, ensuring fidelity to the original teacher's voice while enabling rapid local synthesis
Data Source
AI summary
Disclosed are a method and a terminal for generating simulated voices of virtual teachers. Real voice samples of teachers are collected and converted into text sequences, and a text emotion polarity training set and a text tone training set are constructed according to the text sequences; a lexical item emotion model is constructed based on lexical items in the text sequences and is trained by using the emotion polarity training set, and word vectors, an emotion polarity vector, and a weight parameter are obtained by training; and the similarity between the word vector and the emotion polarity vector is calculated, and emotion features are extracted according to a similarity calculation result, a conditional vocoder is constructed according to the voice styles and emotion features to generate new voices with emotion changes. The method and the terminal contribute to satisfying the application requirements of high-quality virtual teachers.


