Voice Grafting Machine Learning for Natural Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice rehabilitation methods for patients with impaired or missing phonation, such as those on mechanical ventilation or post-laryngectomy, often result in distorted and unnatural speech, and are cumbersome or invasive, failing to restore natural-sounding speech effectively.
Innovation Solution
The technique involves measuring the time-varying physical configuration of the vocal tract to numerically synthesize a speech waveform in real time, using machine learning algorithms to 'graft' the patient's articulation onto a healthy phonation source, outputting the waveform as an acoustic wave via a loudspeaker, allowing for natural-sounding speech restoration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If current voice rehabilitation methods (tracheoesophageal puncture, esophageal speech, electrolarynx) are used, then patients can produce speech output, but the speech sounds distorted and unnatural
Solution Approach 1:
The patent replaces mechanical voice restoration systems (electrolarynx, tracheoesophageal puncture) with a computational system that uses machine learning to synthesize natural-sounding speech. The system measures vocal tract configuration and uses algorithms to generate speech waveforms, substituting mechanical vibration methods with digital signal processing and neural network-based synthesis.
Solution Approach 2:
The patent changes the fundamental parameters of voice restoration by transitioning from mechanical vibration frequency and amplitude control to machine learning-based waveform synthesis. The system measures vocal tract parameters (shape, volume, configuration) and uses these to generate speech waveforms with natural spectral characteristics, fundamentally changing how speech parameters are controlled and synthesized.
2Reliability
If invasive procedures (tracheoesophageal puncture, surgery) are used for voice restoration, then speech capability is restored, but the methods are cumbersome and invasive
Solution Approach 1:
The patent replaces invasive surgical procedures with a non-invasive measurement and computational synthesis system. Instead of creating physical openings in the trachea or esophagus, the system uses external sensors to measure vocal tract configuration and computational algorithms to generate speech, eliminating surgical intervention entirely.
Solution Approach 2:
The patent introduces an intermediary computational system that bridges the gap between vocal tract measurement and speech output. Rather than directly using mechanical or surgical methods to produce speech, the system uses machine learning algorithms as an intermediary to translate vocal tract measurements into natural-sounding speech waveforms.
3Reliability
If real-time speech synthesis is implemented, then natural-sounding speech is restored, but computational complexity and processing requirements increase
Solution Approach 1:
The patent applies preliminary action by pre-training machine learning models with large datasets of speech and vocal tract measurements before deployment. The system performs offline training and model optimization, so that during real-time operation, the pre-trained models can rapidly synthesize speech without requiring complex real-time computations, reducing operational complexity.
Solution Approach 2:
The patent implements dynamics by using adaptive machine learning models that can adjust to individual patient characteristics. The system dynamically adapts to each patient's unique vocal tract configuration and speech patterns through personalized training data, allowing complex synthesis to be optimized for each user's specific needs rather than using a fixed complex system for all users.
Data Source
AI summary
A process labeled “voice grafting” can be understood in terms of the source-filter model of speech production as follows: For a patient who has partially or completely lost the ability to phonate, but retained at least a partial ability to articulate, the techniques described herein computationally “graft” the patient's time varying filter function, i.e. articulation, onto a source function, i.e. phonation, which is based on the speech output of one or more healthy speakers, in order to synthesize natural sounding speech in real time.


