Neural Network Training for Emotive Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing communication devices for individuals with motor neurone disease (MND) or other speech impairments are slow and lack the emotional and personal nuances of human speech, making them inadequate for spontaneous conversation.
Innovation Solution
A method of training a neural network to generate conversational replies using a dataset of stored phrases, with optional user-specific customization, emotional categorization, and integration with speech-to-text and text-to-speech engines to provide personalized and emotionally contextually relevant responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If gaze-tracking systems with text-to-speech are used to assist communication, then speech capability is restored, but conversation speed becomes extremely slow
Solution Approach 1:
The system pre-loads and prepares multiple potential response options before they are needed. When a user receives an input, the neural network has already generated several candidate replies based on the conversation context, allowing for rapid selection and output without real-time processing delays
Solution Approach 2:
The system uses neural networks to generate synthetic conversational responses that copy and emulate natural human speech patterns, emotional tones, and contextual appropriateness. Multiple response variants are generated and stored, allowing the system to quickly select and output the most suitable reply without requiring slow, deliberate typing
2Productivity
If predictive text and word prediction are used to speed up communication, then typing speed improves, but subtle verbal cues and emotional expression are lost
Solution Approach 1:
The system transforms the output by adjusting multiple parameters including emotional tone, speech rhythm, volume, and stylistic characteristics. The neural network generates responses that match the emotional context of the conversation, and the text-to-speech engine applies appropriate vocal parameters to preserve subtle verbal cues and emotional expression while maintaining fast communication speed
3Ease of operation
If existing communication devices are used, then basic communication is enabled, but spontaneous conversation interaction becomes very difficult
Solution Approach 1:
The system continuously monitors the conversation flow, user selections, and contextual information to provide real-time feedback for improving response generation. The neural network learns from each interaction and adjusts its predictions, while the system provides feedback loops that refine response timing, emotional appropriateness, and contextual relevance, enabling natural spontaneous conversation
Solution Approach 2:
The system dynamically adapts its response generation based on conversation context, user preferences, and emotional tone. Rather than using static pre-programmed responses, the neural network continuously adjusts its predictions to match the evolving conversation dynamics, allowing for natural spontaneous interaction while maintaining ease of use
Data Source
AI summary
A method of training a neural network to generate conversational replies, the method including: providing a first dataset of stored phrases linked to form a plurality of conversational sequences; training the neural network to generate responses to input phrases using the first dataset; and using the trained neural network to generate a list of conversational replies in response to conversational inputs.


