Synchronous Dysarthria Voice Corpus Generation via Consonant Signal Replacement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice conversion technologies for dysarthria patients face challenges in aligning reference and patient corpora due to unclear voices, leading to suboptimal alignment and resulting in poor voice conversion quality, characterized by noise and popping in converted voices, which requires high manpower and time for manual alignment.

Innovation Solution

A device and method for generating synchronous corpora using a phoneme database, syllable detector, and voice synthesizer that replaces dysarthria consonant signals with normal consonant signals based on script data and phoneme data, utilizing techniques like autocorrelation or deep neural networks to detect signal positions and improve alignment, thereby enhancing voice conversion quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional alignment technologies (DTW or PSOLA) are used to align reference corpus to patient corpus, then the alignment process can be automated, but the alignment quality is poor due to unclear dysarthria voices, resulting in incomplete alignment and requiring high manual effort for correction

Engineering Contradiction:
Improvealignment automationVSAvoidalignment quality
Core Design Contradiction:
Extent of automationVSManufacturing precision

Solution Approach 1:

The patent introduces a phoneme-level analysis as an intermediary step between the reference corpus and patient corpus. By detecting consonant signal positions at the phoneme level and using them as anchors for frame alignment, the system creates a mediating structure that bridges the gap between automated processing and high alignment quality, eliminating the need for manual correction

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the speech signal into phoneme-level units, specifically identifying consonant signal positions within each phoneme. This segmentation allows for precise localization of alignment points, transforming the continuous alignment problem into a discrete set of anchor points that guide the overall alignment process, thereby achieving both automation and high precision

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If manual alignment is performed to achieve high alignment quality, then the alignment precision is improved, but the time cost and manpower requirements increase significantly

Engineering Contradiction:
Improvealignment qualityVSAvoidalignment time cost
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs self-alignment by automatically detecting consonant signal positions in the dysarthria patient's speech and using these positions to guide the alignment of reference corpus frames to patient corpus frames. The algorithm independently identifies alignment anchors and completes the synchronization process without requiring manual intervention, achieving both high precision and efficiency

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary detection of consonant signal positions before the main alignment process. By pre-identifying the positions of consonant signals in the patient corpus and preparing corresponding reference frames in advance, the system sets up the alignment structure beforehand, enabling the subsequent alignment to proceed quickly and accurately without manual effort

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If DNN is used to learn relationships between aligned frames, then the voice conversion model can capture fine-grained temporal relationships, but slight offsets in alignment cause incorrect learning and result in popping or noise in converted voices

Engineering Contradiction:
Improvetemporal relationship precisionVSAvoidvoice conversion quality
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies different quality requirements to different parts of the alignment process. At the phoneme level, it achieves precise consonant signal position detection to establish accurate anchor points. At the frame level, it uses these anchors to guide alignment with appropriate tolerance. This localized precision strategy ensures that the DNN learns correct temporal relationships without being sensitive to minor frame-level variations, preventing popping and noise in the output

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11222650B2Device and method for generating synchronous corpus
Publication Date: 2022.01.11 NATIONAL CHUNG CHENG UNIV
  • US11222650B2 patent drawing
  • US11222650B2 patent drawing
  • US11222650B2 patent drawing

AI summary

A device and a method for generating synchronous corpus is disclosed. Firstly, script data and a dysarthria voice signal having a dysarthria consonant signal are received and the position of the dysarthria consonant signal is detected, wherein the script data have text corresponding to the dysarthria voice signal. Then, normal phoneme data corresponding to the text are searched and the text is converted into a normal voice signal based on the normal phoneme data corresponding to the text. The dysarthria consonant signal is replaced with the normal consonant signal based on the positions of the normal consonant signal and the dysarthria consonant signal, thereby synchronously converting the dysarthria voice signal into a synthesized voice signal. The synthesized voice signal and the dysarthria voice signal are provided to train a voice conversion model, retain the timbre of the dysarthria voices and improve the communication situations.