Personalized synthesis and recognition enhancement of dysarthric speech

By using a speech synthesis model for articulation disorders and employing long-range dependent feature encoding and non-stationary feature encoding modules to generate personalized synthesized speech, the problem of poor recognition performance in existing technologies has been solved, achieving more efficient speech synthesis and recognition and improving the communication ability of patients with articulation disorders.

CN120412540BActive Publication Date: 2026-04-07TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing automatic speech recognition systems are not performing well in recognizing speech disorders, mainly due to the scarcity of speech data for speech disorders, the large differences in individual patient characteristics, the inability of existing data augmentation methods to effectively preserve individual characteristics, and insufficient research on end-to-end model synthesis.

Method used

A speech synthesis model for articulation disorders is adopted, including a long-range dependency feature encoding module, a non-stationary feature encoding module, and a decoding module. The long-range dependency feature encoding module captures long-range dependency features, and the non-stationary feature encoding module introduces random Gaussian noise. Combined with noise processing in the training and testing phases, synthesized speech that conforms to personalized features is generated.

Benefits of technology

It improves the ability of speech synthesis models for articulation disorders to extract personalized features and improves speech synthesis performance, reduces the word error rate of recognition models, provides more accurate speech synthesis and recognition assistance tools, and enhances the communication ability of patients with articulation disorders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412540B_ABST
    Figure CN120412540B_ABST
Patent Text Reader

Abstract

This invention discloses a personalized synthesis and recognition enhancement method for speech disorders. The speech disorder synthesis model includes a long-range dependent feature encoding module, a non-stationary feature encoding module, and a decoding module. The input of the speech disorder synthesis model includes samples, and the output includes synthesized speech disorder speech. The samples are speech disorder text sequences. The input of the long-range dependent feature encoding module includes samples, and the output is an alignment vector z. The input of the non-stationary feature encoding module includes the alignment vector z, and the output is the final embedding representation. The input of the decoding module is the final embedding representation, and the output is the synthesized speech disorder speech. The speech disorder synthesis model of this invention improves the ability to extract personalized features of speech disorder speech, enhances speech synthesis performance, and improves the fine-grained expression of speech disorder speech features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech synthesis and recognition for articulation disorders, and specifically relates to a personalized synthesis and recognition enhancement method for speech with articulation disorders. Background Technology

[0002] Current research indicates that there are over 20 million people with articulation disorders in China. Articulation disorders negatively impact patients' quality of life. Those with moderate to severe articulation disorders, due to their inability to effectively control the muscles involved in articulation, experience unclear pronunciation, fatigue while speaking, excessively fast or slow speech rate, and variations in intonation and volume, leading to decreased speech clarity and comprehensibility. Therefore, communicating with people with articulation disorders presents a significant challenge for those without prior experience. Furthermore, some cases of articulation disorders are caused by underlying medical conditions. As these conditions worsen, the resulting impairment in speech ability can threaten patients' ability to live independently. Therefore, researching the individualized speech characteristics and speech recognition abilities of people with articulation disorders is of great practical significance in alleviating the suffering of these over 20 million people in my country and improving their quality of life.

[0003] To date, Automatic Speech Recognition (ASR) systems developed based on the standard speech paradigm have made remarkable progress, but they have not achieved ideal recognition results in recognizing pathological speech such as speech disorders. The main reasons for this are as follows:

[0004] (1) Dysarthria is a neurological disorder that causes dysarthria, usually accompanied by muscle weakness and fatigue, which poses a challenge to long-term collection of speech samples, making it relatively difficult to obtain speech samples of dysarthria.

[0005] (2) Currently, a large number of open-source healthy speech datasets have established a good ecosystem, but there are very few publicly available speech datasets for speech disorders.

[0006] (3) Due to the different pathogenesis of patients with articulation disorders, there are significant differences in the pronunciation of patients with articulation disorders as speakers. This makes the changes in the acoustic space of articulation disorder speech greater and more complex compared with normal speech.

[0007] Therefore, addressing the scarcity of speech data for dysarthric speech is one of the key factors in improving the performance of dysarthric speech recognition (DSR) systems. Existing research employs data augmentation methods to enhance the capabilities of dysarthric speech recognition systems. Data augmentation methods simulate dysarthric speech by modifying healthy speech and use it as augmented training data, thereby improving the performance of dysarthric speech recognition systems. However, data augmentation methods often sacrifice the personalized features of speech.

[0008] In contrast, methods based on dysarthric speech synthesis (DSS) can expand datasets without limitation, enhance the performance of speech recognition systems for articulation disorders, and effectively preserve the speaker's individual characteristics. Speech synthesis includes two-stage models and end-to-end models. Two-stage models are difficult to apply to speech synthesis for articulation disorders due to prediction errors caused by speech irregularities and poor generalization ability of vocoders. End-to-end models, on the other hand, generate speech directly from text without relying on intermediate features. Studies have shown that adding pause mechanisms and personalized features such as pitch, energy, and duration during the synthesis of speech disorders using end-to-end methods can significantly improve the performance of speech recognition systems for articulation disorders. However, research on end-to-end methods in considering personalized synthesis of speech disorders remains insufficient. Summary of the Invention

[0009] To address the shortcomings of existing technologies, the purpose of this invention is to provide a speech synthesis model for articulation disorders.

[0010] Another objective of this invention is to provide a training method for a speech synthesis model with articulation disorders.

[0011] Another objective of this invention is to provide a personalized synthesis method for speech disorders.

[0012] Another object of the present invention is to provide an application of a speech synthesis model for articulation disorders in reducing the word error rate of a speech recognition model for articulation disorders.

[0013] This invention is achieved through the following technical solution.

[0014] A speech synthesis model for articulation disorders includes: a long-range dependent feature encoding module, a non-stationary feature encoding module, and a decoding module connected in sequence. The input of the speech synthesis model for articulation disorders includes samples, and the output includes synthesized speech with articulation disorders. The samples are sequences of text with articulation disorders.

[0015] The long-range dependency feature encoding module takes samples as input and outputs an alignment vector z. This module includes a phoneme encoding module, an affine transformation module, a monotonic alignment module, and a long-range dependency duration prediction module. The phoneme encoding module takes samples as input and outputs a long-range dependency feature vector h. The affine transformation module takes the long-range dependency feature vector h as input and outputs a long-range dependency feature vector. The long-term dependency duration prediction module includes: a streaming module;

[0016] When training a speech synthesis model for articulation disorders, the input to the monotonic alignment module includes long-range dependent feature vectors. The speaker-level representation z corresponding to the sample s The output is the alignment vector z and the audio duration feature vector d; the input of the long-range dependent duration prediction module is the long-range dependent feature vector h and the audio duration feature vector d, and the output is random noise N. d ;

[0017] When testing the speech synthesis model for articulation disorders, the input to the long-range dependent duration prediction module is random noise N. d The output is the audio duration feature vector d″; the input of the monotonic alignment module includes the audio duration feature vector d″ and the long-range dependency feature vector. The output is an aligned vector z;

[0018] The input to the non-stationary feature encoding module includes the alignment vector z, and the output is the final embedding representation. The non-stationary feature encoding module includes: a noise perturbation module, a feature extraction module, and a feature fusion module; the noise perturbation module takes an alignment vector z as input and outputs vectors s1 and s2; the feature extraction module takes vectors s1 and s2 as input and outputs non-stationary feature vectors. Non-stationary eigenvectors The input to the feature fusion module is a non-stationary feature vector. Non-steady eigenvectors And align the vector z, and output the final embedding representation.

[0019] The input to the decoding module is the final embedded representation. The output is a synthesized speech with articulation disorders.

[0020] In the above technical solution, random noise N d Used to calculate the loss function value during training of a speech synthesis model with articulation disorders, random noise N d It is also used as input to the long-term dependent duration prediction module during testing of speech synthesis models for articulation disorders.

[0021] In the above technical solution, the noise perturbation module is used to randomly generate Gaussian noise ∈1 and Gaussian noise ∈2, and concatenates Gaussian noise ∈1 and Gaussian noise ∈2 with the alignment vector z respectively. The output of the noise perturbation module is vector s1 = [∈1; z] and vector s2 = [∈2; z], where ∈1 and ∈2 ∈ R. |z| .

[0022] In the above technical solution, the feature extraction module includes: seven fully connected layers and a max pooling layer connected sequentially. The first fully connected layer maps vectors s1 and s2 to D1 dimensions. The second to sixth fully connected layers perform linear mapping and make the output of the sixth fully connected layer D1 dimension. The seventh fully connected layer maps the output of the sixth fully connected layer from D1 dimension to D2 dimension. A ReLU activation function is connected after the first, second, fourth, and sixth fully connected layers. The max pooling layer is used for downsampling and outputting non-stationary feature vectors. Non-stationary eigenvectors

[0023] In the above technical solution, the feature fusion module is used to perform feature fusion, and the fusion formula is as follows:

[0024] A personalized synthesis method for speech with articulation disorders includes:

[0025] The test set samples and the random noise N obtained during the training phase of the speech synthesis model for articulation disorders are used to... d The input is fed into the trained speech synthesis model for speech synthesis of speech disorders, and the synthesized speech of speech disorders corresponding to the sample is obtained.

[0026] In the above technical solution, during the testing phase, the long-range dependency feature encoding module of the trained articulation disorder speech synthesis model includes only the stream module in the long-range dependency duration prediction module.

[0027] The training method for obtaining the trained speech synthesis model with articulation disorders in the above technical solution includes:

[0028] Each patient with articulation disorder was asked to speak based on a text sequence of articulation disorders. Microphones were used to collect speech audio data from each patient in different contexts, resulting in approximately one hour of articulation disorder speech sequences for each patient. Each patient's text sequence was used as a sample, and the corresponding articulation disorder speech sequence was used as the ground truth. The ground truth was then used as input to a speaker encoding module, whose output was the speaker frame-level representation z corresponding to that sample. s ; All samples, along with their corresponding ground truth values ​​and speaker frame-level representations z sCreate a dataset D, and then divide dataset D into a training set and a test set in an 8:2 ratio.

[0029] The training set samples and their corresponding speaker frame-level representations z s The input is fed into the speech synthesis model for 120,000 training steps. During each training step, the speech synthesis model outputs synthesized speech with articulation disorders. The total loss function value is calculated based on the ground truth sample and the output synthesized speech with articulation disorders. The parameters of the speech synthesis model with articulation disorders are updated using the AdamW optimizer. The speech synthesis model with articulation disorders is obtained when the total loss function value is minimized. The ground truth sample is the speech sequence with articulation disorders.

[0030] In the training method for obtaining a trained speech synthesis model with articulation disorders, the inputs to the speech synthesis model during the training phase include: samples and speaker frame-level representation z. s The output consists of synthesized speech with articulation disorders and random noise N. d ;

[0031] The long-range dependency duration prediction module during the training phase includes: an audio duration feature extraction module, a phoneme duration feature extraction module, a weighting module, and a streaming module. Both the audio duration feature extraction module and the phoneme duration feature extraction module include a sequentially connected one-dimensional convolutional layer and a Mamba layer. The audio duration feature extraction module extracts the duration of phonemes, with input being an audio duration feature vector `d` and output being a duration feature vector `d′`. The phoneme duration feature extraction module extracts the duration of phonemes, with input being a long-range dependency feature vector `h` and outputting a duration feature vector `h′`. The duration feature vectors `d′` and `h′` are used as input to the weighting module, which performs a weighting operation to obtain... The input to the stream module is The audio duration feature vector d and the output are random noise N. d .

[0032] The application of a speech synthesis model for articulation disorders in reducing the word error rate of speech recognition models for articulation disorders includes:

[0033] All texts from the AiShell-3 text corpus are input into the trained speech synthesis model for speech disorders to synthesize speech disorders, resulting in synthesized speech disorders corresponding to each text segment. All synthesized speech disorders are used as speech samples, and all speech samples and their corresponding texts constitute the corpus.

[0034] The Whisper pre-trained model was used as the speech recognition model for articulation disorders. The model was trained using all speech samples in the corpus and all ground truth values ​​of the training set. The trained speech recognition model was then used to perform speech recognition on the ground truth values ​​of the test set, so that it could recognize the text for each ground truth value.

[0035] The present invention has the following advantages due to the adoption of the above technical solutions:

[0036] 1. The speech synthesis model for articulation disorders of the present invention learns the personalized features of the speech of patients with articulation disorders and generates synthesized speech with articulation disorders that is more consistent with the characteristics of speech with articulation disorders.

[0037] 2. The long-range dependency feature encoding module in the speech synthesis model for articulation disorders of the present invention effectively captures long-range dependency features in the samples and integrates the speaker frame-level representation z. s The alignment vector z is obtained, which improves the ability of the speech synthesis model for speech disorders to extract personalized features of speech disorders and improves the speech synthesis performance.

[0038] 3. The non-stationary feature encoding module in the speech synthesis model for articulation disorders of the present invention introduces random Gaussian noise ∈1 and random Gaussian noise ∈2. The non-stationary feature encoding module automatically captures non-stationary features, thereby improving the speech synthesis model for articulation disorders in terms of its ability to express speech features in a refined manner and its speech synthesis performance.

[0039] 4. The speech recognition and enhancement method of this invention combines a trained speech synthesis model and a speech recognition model for speech disorders, achieving both the synthesis and recognition of speech disorders. This improves the accuracy of speech synthesis and recognition, providing an effective auxiliary tool for communication for patients with speech disorders.

[0040] 5. The personalized synthesis method and recognition enhancement method for speech disorders of the present invention are used to assist patients with speech disorders in communication and speech rehabilitation training, thereby improving the quality of life of individuals with speech disorders. Attached Figure Description

[0041] Figure 1 This is a schematic diagram of the speech synthesis model for articulation disorders of the present invention, (a) is the training phase, and (b) is the testing phase;

[0042] Figure 2 This is a schematic diagram of the structure of the long-range dependency feature encoding module of the present invention;

[0043] Figure 3This is a schematic diagram of the long-range dependent duration prediction module of the present invention, (a) is the training phase, and (b) is the testing phase;

[0044] Figure 4 This is a schematic diagram of the non-steady-state feature encoding module of the present invention. Detailed Implementation

[0045] The personalized synthesis and recognition enhancement method for speech disorders of the present invention will be described in detail below with reference to the accompanying drawings.

[0046] Example 1

[0047] like Figure 1 As shown, a speech synthesis model for articulation disorders includes: a long-range dependent feature encoding module, a non-stationary feature encoding module, and a decoding module connected in sequence. The input of the speech synthesis model for articulation disorders includes samples, and the output of the speech synthesis model for articulation disorders includes synthesized speech with articulation disorders, wherein the samples are sequences of text with articulation disorders.

[0048] like Figure 2 As shown, the input of the long-range dependent feature encoding module includes samples, and the output includes the alignment vector z.

[0049] The long-range dependency feature encoding module includes: a phoneme encoding module, an affine transformation module, a monotonic alignment module, and a long-range dependency duration prediction module. The phoneme encoding module (Dao T, Gu A. Transformers are SSMs: Generalized models and efficient algorithms through structured state spaceduality[J]. arXiv preprint arXiv:2405.21060,2024.) takes samples as input and outputs long-range dependency feature vector h. The phoneme encoding module includes: four attention layers and two Mamba layers connected in sequence. The phoneme encoding module is used to extract the long-range dependency feature vector h of phonemes from the samples.

[0050] The affine transformation module takes a long-range dependent eigenvector h as input and outputs a long-range dependent eigenvector h. The affine transformation module includes a one-dimensional convolutional layer with a kernel size of 1 and a stride of 1.

[0051] The long-range dependency duration prediction module includes: the streaming module (Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C] / / International Conference on Machine Learning.PMLR,2021:5530-5540.);

[0052] When training a speech synthesis model for articulation disorders, the input to the monotonic alignment module (Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C] / / International Conference on Machine Learning.PMLR,2021:5530-5540.) includes long-range dependent feature vectors. The speaker-level representation z corresponding to the sample s The output of the monotonic alignment module is the alignment vector z and the audio duration feature vector d; the input of the long-range dependent duration prediction module is the long-range dependent feature vector h and the audio duration feature vector d, and the output of the long-range dependent duration prediction module is random noise N. d Random noise N d Used to calculate the loss function value and as input to the long-range dependent duration prediction module during testing of the articulation disorder speech synthesis model, such as Figure 3 As shown in (a);

[0053] When testing the speech synthesis model for articulation disorders, the input to the long-range dependent duration prediction module is random noise N. d The output of the long-range dependent duration prediction module is the audio duration feature vector d″, such as Figure 3 As shown in (b); the input to the monotonic alignment module includes the audio duration feature vector d″ and the long-range dependency feature vector. The output of the monotonic alignment module is the alignment vector z;

[0054] like Figure 4 As shown, the input of the non-stationary feature encoding module includes the alignment vector z, and the output is the final embedding representation. The non-stationary feature encoding module is used to extract non-stationary features from the alignment vector z; the non-stationary feature encoding module includes: a noise perturbation module, a feature extraction module, and a feature fusion module connected in sequence;

[0055] The noise perturbation module takes an alignment vector z as input and randomly generates Gaussian noise ∈1 and Gaussian noise ∈2. It then concatenates Gaussian noise ∈1 and Gaussian noise ∈2 with the alignment vector z. The output of the noise perturbation module is vectors s1 = [∈1; z] and s2 = [∈2; z], where ∈1 and ∈2 ∈ R. |z| ;

[0056] The feature extraction module takes vectors s1 and s2 as inputs and outputs non-stationary feature vectors. Non-stationary eigenvectors The feature extraction module includes seven fully connected layers and a max-pooling layer connected sequentially. The first fully connected layer maps vectors s1 and s2 to D1 dimensions. The second through sixth fully connected layers perform linear mapping and make the output of the sixth fully connected layer D1 dimension. The seventh fully connected layer maps the output of the sixth fully connected layer from D1 dimension to D2 dimension. ReLU activation functions are connected after the first, second, fourth, and sixth fully connected layers. The max-pooling layer downsamples and outputs the non-stationary feature vector. Non-stationary eigenvectors

[0057] The input to the feature fusion module is a non-stationary feature vector. Non-steady eigenvectors The alignment vector z and the output of the feature fusion module are the final embedding representations. The feature fusion module is used to perform feature fusion, and the fusion formula is as follows:

[0058] The input to the decoding module (Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C] / / International Conference on Machine Learning.PMLR,2021:5530-5540.) is the final embedded representation. The output of the decoding module is synthesized speech with articulation disorders. The decoding module is used to embed the final representation. Decode to the waveform domain and output synthesized speech with articulation disorders.

[0059] In this embodiment, D1 = 512, D2 = 32.

[0060] Example 2

[0061] A method for obtaining dataset D includes:

[0062] S1, each patient with articulation disorder pronounces according to the text sequence of articulation disorder, and uses a microphone to collect the speech audio data of each patient with articulation disorder in different situations. Each patient with articulation disorder obtains about 1 hour of speech sequence of articulation disorder.

[0063] S2, taking the articulation disorder text sequence of each patient with articulation disorder as a sample, taking the corresponding articulation disorder speech sequence as the sample ground value, and taking the sample ground value as the input of the speaker encoding module (Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C] / / International Conference on Machine Learning.PMLR,2021:5530-5540.), the output of which is the speaker frame-level representation z corresponding to the sample. s ;

[0064] The speaker encoding module is used to extract features from the sample ground truth. Specifically, the speaker encoding module converts the sample ground truth from a time-domain signal to a frequency-domain signal using Fast Fourier Transform (FFT), and extracts the spectral features within each window using Short Time Fourier Transform (STFT), thereby obtaining the speaker frame-level representation z of the sample ground truth. s ;

[0065] S3, which includes all samples and their corresponding sample ground values ​​and speaker frame-level representations z. s Create a dataset D, and then divide dataset D into a training set and a test set in an 8:2 ratio.

[0066] In step S1, the sampling frequency is set to be greater than or equal to 16kHz when collecting speech audio data to ensure that the quality of the collected speech audio data is high. Each sample in dataset D is a speech disorder text sequence, and the true value of each sample is a speech disorder speech sequence.

[0067] In this embodiment, the sampling frequency is set to 16kHz, the number of samples in dataset D is 24, the window size of the speaker coding module's Fast Fourier Transform (FFT) is 1024, the window size of the Short Time Fourier Transform (SFT) is 1024, and the step size of both the FFT and SFT windows is 256.

[0068] Example 3

[0069] A training method for a speech synthesis model with articulation disorders includes: processing samples from the training set in Example 2 and the corresponding speaker frame-level representations z of the samples. s The input was fed into the speech synthesis model with articulation disorders in Example 1 for 120,000 training steps. During each training step, the speech synthesis model outputs synthesized speech with articulation disorders, and the total loss function value is calculated based on the sample ground truth and the output synthesized speech with articulation disorders. The parameters of the speech synthesis model with articulation disorders are updated using the AdamW optimizer (Loshchilov I. Decoupled weight decay regularization[J].arXiv preprint arXiv:1711.05101,2017.). The trained speech synthesis model with articulation disorders is obtained when the total loss function value is minimized.

[0070] Among them, such as Figure 1 As shown in (a), the inputs to the articulation disorder speech synthesis model during the training phase include: samples and speaker frame-level representations z. s The output consists of synthesized speech with articulation disorders and random noise N. d ;

[0071] The long-range dependency duration prediction module during the training phase includes: an audio duration feature extraction module, a phoneme duration feature extraction module, a weighting module, and a streaming module. Both the audio duration feature extraction module and the phoneme duration feature extraction module include a sequentially connected one-dimensional convolutional layer and a Mamba layer. The audio duration feature extraction module extracts the duration of phonemes, with input being an audio duration feature vector `d` and output being a duration feature vector `d′`. The phoneme duration feature extraction module extracts the duration of phonemes, with input being a long-range dependency feature vector `h` and outputting a duration feature vector `h′`. The duration feature vectors `d′` and `h′` are used as input to the weighting module, which performs a weighting operation to obtain... The input to the stream module is The audio duration feature vector d and the output are random noise N. d ;

[0072] The total loss function is:

[0073] in, The loss function for predicting long-range dependent durations is calculated as follows:

[0074]

[0075] In the formula, u is a random variable and 0≤u≤1, v is a random variable and Random variables u and v satisfy the posterior distribution To Prior distribution of the sample, To v-sampling prior distribution, It is a sequence of positive real numbers. For the expected calculation, h is the long-range dependent feature vector;

[0076] To reconstruct the loss function, the calculation formula is as follows:

[0077]

[0078] In the formula, x represents the spectrum of synthesized speech with articulation disorders corresponding to a batch of samples. mel The spectrum is composed of the true values ​​of a batch of samples. The subscript 1 in the formula represents the L1 loss (absolute error loss).

[0079] To optimize the loss function, the calculation formula is as follows:

[0080]

[0081] In the formula, ε is a constant, ε∈(0, 10). -12 ), ε is used to prevent the denominator from being 0;

[0082] The KL divergence loss function is used (Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C] / / International Conference on Machine Learning.PMLR,2021:5530-5540.).

[0083] loss function The loss function is used to calculate the loss of synthesized articulation disordered speech synthesized by decoder G in the decoding module. This is used to calculate the loss between the synthesized articulated speech synthesized by the discriminator D′ and the ground truth sample in the decoding module; loss function. and loss function The calculation formula can be found in: Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C] / / International Conference on Machine Learning.PMLR,2021:5530-5540;

[0084] The feature mapping loss function is (Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C] / / International Conference on Machine Learning.PMLR,2021:5530-5540.).

[0085] In each training round, the number of samples in each batch of the input training set is 16, the weight decay coefficient λ = 0.01, and the initial learning rate is 2 × 10⁻⁶. -4 The learning rate in each training round is 0.999. 1 / 8 The factors decay, and the hyperparameters β1 and β2 in the AdamW optimizer are used to control the exponential decay rate of the first-order and second-order momentum estimates during the optimization process. The hyperparameters β1 = 0.8 and β2 = 0.99.

[0086] Example 4

[0087] A personalized synthesis method for speech with articulation disorders includes the following steps:

[0088] The samples in the test set and the random noise N obtained during the training phase are combined. d The input is fed into the articulation disorder speech synthesis model trained in Example 3 to synthesize the articulation disorder speech, and the synthesized articulation disorder speech corresponding to each sample in the test set is obtained.

[0089] Among them, such as Figure 1 (b) and Figure 3 As shown in (b), in the long-range dependency feature encoding module of the articulation disorder speech synthesis model trained during the testing phase, the long-range dependency duration prediction module only includes the streaming module, whose input is random noise N. d The output is the predicted audio duration feature vector d″, and the input of the monotonic alignment module is the long-range dependent feature vector. The predicted audio duration feature vector d″ is used, and the output is the alignment vector z.

[0090] By synthesizing speech samples with pronunciation styles consistent with those of patients with articulation disorders using a trained speech synthesis model, the diversity and richness of the corpus can be improved.

[0091] Example 5

[0092] A personalized synthesis method for speech disorders is basically the same as that in Example 4, with the only difference being: all Mamba layers in the phoneme encoding module, audio duration feature extraction module and phoneme duration feature extraction module of the speech disorder synthesis model are deleted during training, and the Mamba layer of the phoneme encoding module is deleted during testing.

[0093] Example 6

[0094] A personalized synthesis method for speech disorders is basically the same as that in Example 4, with the only difference being that the non-steady-state feature encoding module in the speech synthesis model for speech disorders is removed, and the alignment vector z output by the long-range dependent feature encoding module is used as the input of the decoding module.

[0095] The synthesized articulation disorder speech samples from Examples 4, 5, and 6 were evaluated using the Mean Opinion Score (MOS), and the results are as follows:

[0096] Table 1

[0097] Example MOS Example 5 3.58 Example 6 3.26 Example 4 4.00

[0098] As shown in Table 1, the Mamba layer and non-stationary coding module in the speech synthesis model for articulation disorders in Example 4 play a significant role in improving the quality of synthesized speech, and both are indispensable.

[0099] The Mean Opinion Score (MOS) was used to evaluate the synthesized articulation disorder speech of Example 4, Tacotron2+HiFi-GAN, Tacotron2+WaveGlow, GlowTTS+HiFi-GAN, GlowTTS+waveGlow, and VITS. The results are shown in Table 2. As can be seen from Table 2, the synthesized articulation disorder speech of the present invention has high quality.

[0100] Table 2

[0101] method MOS Tacotron2+HiFi-GAN 1.43 Tacotron2+WaveGlow 1.05 GlowTTS+HiFI-GAN 2.33 GlowTTS+waveGlow 1.35 VITS 3.40 Example 4 4.00

[0102] In the methods in Table 2, the model before the "+" sign is used as an acoustic model to convert the text into a Mel spectrum, and the model after the "+" sign synthesizes the final articulation disorder speech.

[0103] Among them, Tacotron2, see: Shen J, Pang R, Weiss RJ, et al.Natural ttssynthesis by conditioning wavenet on mel spectrogram predictions[C] / / 2018IEEEinternational conference on acoustics, speech and signal processing(ICASSP).IEEE, 2018:4779-4783;

[0104] HiFi-GAN see: Kong J, Kim J, Bae J.Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis[J]. Advances inneural information processing systems, 2020,33:17022-17033;

[0105] WaveGlow see: Prenger R, Valle R, Catanzaro B. Waveglow: A flow-based generative network for speech synthesis [C] / / ICASSP 2019-2019IEEEInternational Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019: 3617-3621;

[0106] GlowTTS see: Kim J, Kim S, Kong J, et al. Glow-tts: A generative flow fortext-to-speech via monotonic alignment search[J]. Advances in NeuralInformation Processing Systems, 2020,33:8067-8077;

[0107] VITS see: Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech[C] / / InternationalConference on Machine Learning.PMLR, 2021:5530-5540.

[0108] Example 7

[0109] Based on Example 4, a method for enhancing the recognition of speech with articulation disorders includes:

[0110] All texts from the AiShell-3 text corpus are input into the trained speech synthesis model for speech disorders in Example 3 to synthesize speech disorders, resulting in synthesized speech disorders corresponding to each text segment (each synthesized speech disorder segment is 6 hours long). All synthesized speech disorders are used as speech samples, and all speech samples and their corresponding texts constitute the corpus.

[0111] The Whisper pre-trained model (Radford A, Kim JW, Xu T, et al. Robust speech recognition via large-scale weak supervision[C] / / International conference on machine learning.PMLR,2023:28492-28518.) was used as the speech recognition model for articulation disorders. The model was trained using all speech samples in the corpus and all ground truth values ​​of samples in the training set of Example 2. The trained speech recognition model was then used to perform speech recognition on the trained model using ground truth values ​​from the test set of Example 2, so that it could recognize the recognized text for each ground truth value.

[0112] Example 8

[0113] It is basically the same as Example 7, except that the speech recognition model for speech disorders is directly used to perform speech recognition on the ground truth values ​​of the samples in dataset D to obtain the recognized text.

[0114] Example 9

[0115] It is basically the same as Example 7, except that the duration of each synthesized speech disorder speech segment is Y hours.

[0116] The recognized texts from Examples 7, 8, and 9 were evaluated using the character error rate (CER) metric, and the results are as follows:

[0117] Table 3

[0118] Example CER Example 8 86.51% Example 7 31.40%

[0119] As shown in Table 3, in Example 7, the trained speech synthesis model for articulation disorders was used to synthesize speech and obtain a corpus. The speech recognition model for articulation disorders was then trained using the corpus, resulting in a significant improvement in the performance of the speech recognition model for articulation disorders and a significant reduction in the word error rate.

[0120] Table 4

[0121] Example Y CER Example 9 1 86.51% Example 9 2 72.19% Example 9 3 59.32% Example 9 4 68.35% Example 9 5 49.06% Example 7 6 31.40%

[0122] As shown in Table 4, as the duration of Y increases, the performance of the obtained speech recognition model for articulation disorders is significantly improved, and the word error rate is significantly reduced.

[0123] This invention combines an articulation disorder speech synthesis model with an articulation disorder speech recognition model, realizing the integration of speech synthesis and speech recognition. It builds a language-friendly communication bridge for patients with articulation disorders, providing them with better speech recognition and synthesis services.

[0124] The present invention has been described above by way of example. It should be noted that any simple modifications, alterations or other equivalent substitutions that can be made by those skilled in the art without creative effort without departing from the core of the present invention fall within the protection scope of the present invention.

Claims

1. A speech synthesis model for articulation disorders, characterized in that, include: The system comprises a long-range dependency feature encoding module, a non-stationary feature encoding module, and a decoding module. The long-range dependency feature encoding module includes: a phoneme encoding module, an affine transformation module, a monotonic alignment module, and a long-range dependency duration prediction module. The phoneme encoding module takes samples as input and outputs long-range dependency feature vectors. The samples are text sequences containing articulation disorders; the input to the affine transformation module is a long-range dependent feature vector. The output is a long-range dependent feature vector. The long-range dependency duration prediction module includes: a streaming module; When training a speech synthesis model for articulation disorders, the input to the monotonic alignment module includes long-range dependent feature vectors. Speaker-level representation corresponding to the sample The output is an alignment vector. and audio duration feature vector The input to the long-range dependency duration prediction module is the long-range dependency feature vector. and audio duration feature vector The output is random noise. ; When testing the speech synthesis model for articulation disorders, the input to the long-range dependent duration prediction module is random noise. The output is an audio duration feature vector. The input to the monotonic alignment module includes the audio duration feature vector. and long-range dependent feature vectors The output is an alignment vector. ; The input to the non-stationary feature encoding module includes the alignment vector. The output is the final embedded representation. ; The input to the decoding module is the final embedded representation. The output is a synthesized speech with articulation disorders.

2. The speech synthesis model for articulation disorders according to claim 1, characterized in that, Random noise Used to calculate the loss function value during the training of speech synthesis models with articulation disorders, random noise. It is also used as input to the long-term dependent duration prediction module during testing of speech synthesis models for articulation disorders.

3. The speech synthesis model for articulation disorders according to claim 1, characterized in that, The non-stationary feature encoding module includes: a noise perturbation module, a feature extraction module, and a feature fusion module; the input to the noise perturbation module is the alignment vector. The output is a vector. sum vector The input to the feature extraction module is a vector. sum vector The output is a non-stationary eigenvector. Non-stationary eigenvectors The input to the feature fusion module is a non-stationary feature vector. Non-stationary eigenvectors and alignment vector The output is the final embedded representation. .

4. The speech synthesis model for articulation disorders according to claim 3, characterized in that, The noise disturbance module is used to randomly generate Gaussian noise. and Gaussian noise and Gaussian noise and Gaussian noise respectively with the alignment vector The noise perturbation module outputs a vector after splicing. sum vector , and .

5. The speech synthesis model for articulation disorders according to claim 3, characterized in that, The feature extraction module includes seven fully connected layers and a max pooling layer connected sequentially. The first fully connected layer in the seven layers is used to process the vector... sum vector Mapped to The second to sixth fully connected layers are used for linear mapping and make the output of the sixth fully connected layer... The seventh fully connected layer is used to transfer the output of the sixth fully connected layer from... Dimension mapping to In this architecture, the first, second, fourth, and sixth fully connected layers are all followed by ReLU activation functions. Max pooling layers are used for downsampling and outputting non-stationary feature vectors. Non-stationary eigenvectors .

6. The speech synthesis model for articulation disorders according to claim 3, characterized in that, The feature fusion module is used to perform feature fusion, and the fusion formula is as follows: .

7. A personalized synthesis method for speech with articulation disorders, characterized in that, include: The test set samples and random noise acquired during the training phase of the speech synthesis model for speech disorders were used. The input is fed into the trained speech synthesis model for speech disorders according to claim 1 to synthesize speech disorders, and the synthesized speech disorders corresponding to the sample are obtained.

8. The personalized synthesis method according to claim 7, characterized in that, The training method for obtaining a trained speech synthesis model with articulation disorders includes: processing the training set samples and the corresponding speaker frame-level representations. The input is fed into the speech synthesis model for articulation disorders for training. At each training step, the speech synthesis model outputs synthesized speech with articulation disorders. The total loss function value is calculated based on the ground truth sample and the output synthesized speech with articulation disorders. The speech synthesis model with articulation disorders is obtained when the total loss function value is minimized. The ground truth sample is the speech sequence with articulation disorders.

9. The personalized synthesis method according to claim 8, characterized in that, The long-range dependency duration prediction module during the training phase includes: an audio duration feature extraction module, a phoneme duration feature extraction module, a weighting module, and a streaming module. The audio duration feature extraction module is used to extract the duration of phonemes, and its input is the audio duration feature vector. The output is a duration feature vector. The phoneme duration feature extraction module is used to extract the duration of phonemes. The input to the phoneme duration feature extraction module is a long-range dependent feature vector. The output duration feature vector is ;The duration feature vector and duration feature vector As input to the weighting module, the weighting module performs a weighting operation to obtain... The input of the stream module is and audio duration feature vector The output is random noise. .

Citation Information

Patent Citations

  • Double-flow voice conversion method, device and equipment and storage medium

    CN113436608A

  • Obstacle speech recognition and reconstruction method based on speech data retrieval enhancement technology

    CN119207481A