Dysarthria Speech Intelligibility System Using Automatic Corpus Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice conversion systems for dysarthria patients face challenges in achieving complete alignment of speech signals, leading to suboptimal voice conversion quality and requiring manual alignment, which is time-consuming and labor-intensive.
Innovation Solution
A system that includes a speech disordering module to automatically generate a synchronous corpus from paired reference and patient corpora, eliminating the need for conventional corpus alignment technologies, thereby improving the training and conversion qualities of the voice conversion model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If conventional corpus alignment technologies (DTW, PSOLA) are used to align patient corpus with reference corpus, then alignment quality is improved, but the process requires manual intervention and is time-consuming
Solution Approach 1:
The system employs automatic alignment algorithms that self-correct timing offsets between patient and reference corpora without manual intervention. The alignment module automatically detects and adjusts temporal misalignments, allowing the system to serve itself in the alignment process rather than requiring external manual operation.
Solution Approach 2:
The patent replaces manual mechanical alignment operations with automated computational algorithms. Instead of human operators manually adjusting and aligning corpus data, the system uses computer-based alignment technologies that automatically process and synchronize the corpora, substituting mechanical human labor with automated computational mechanisms.
2Manufacturing precision
If speech alignment technologies are used to pre-process training corpus, then training quality is improved, but complete alignment is difficult to achieve due to unclear patient speech
Solution Approach 1:
The system incorporates feedback mechanisms where the alignment module continuously monitors the alignment quality between patient and reference corpora. By analyzing the degree of alignment and detecting misaligned portions, the system adjusts its processing to achieve complete synchronization, ensuring that training data is properly aligned before being used for model training.
Solution Approach 2:
The patent performs alignment pre-processing before the actual voice conversion training. By aligning the patient corpus with the reference corpus in advance and generating synchronized training data beforehand, the system ensures that complete and accurate alignment is achieved prior to training, preventing propagation of alignment errors throughout the training process.
3Manufacturing precision
If manual alignment operation is performed to achieve complete synchronization, then alignment accuracy is improved, but manpower cost increases
Solution Approach 1:
The system employs automatic alignment algorithms that self-correct timing offsets between patient and reference corpora without manual intervention. The alignment module automatically detects and adjusts temporal misalignments, allowing the system to serve itself in the alignment process rather than requiring external manual operation.
Solution Approach 2:
The patent replaces manual mechanical alignment operations with automated computational algorithms. Instead of human operators manually adjusting and aligning corpus data, the system uses computer-based alignment technologies that automatically process and synchronize the corpora, substituting mechanical human labor with automated computational mechanisms.
Data Source
AI summary
A system for improving dysarthria speech intelligibility and method thereof, are provided. In the system, user only needs to provides a set of paired corpus including a reference corpus and a patient corpus, and a speech disordering module can automatically generate a new corpus completely synchronous with the reference corpus, and the new corpus can be used as a training corpus for training a dysarthria voice conversion model. The present invention does not need to use a conventional corpus alignment technology or a manual manner to perform pre-processing on the training corpus, so that manpower cost and time cost can be reduced, and synchronization of the training corpus can be ensured, thereby improving both training and conversion qualities of the voice conversion model.


