A feature enhanced dysarthric speech processing method
Patent Information
- Application Number
- CN202211156788.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-09-22
Smart Images

Figure CN115565551B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of speech signal processing, and particularly relates to a dysarthria speech processing method with feature enhancement. BACKGROUND
[0002] Dysarthria is a speech and movement disorder caused by damage to the central nervous system. This speech disorder affects the individual's vocal tract and vocal cord phonation, thereby affecting the patient's language expression ability and speech intelligibility, which will have a very serious impact on the patient's daily communication. With the development of signal processing technology, some speech signal processing methods are often used for the study of pathological speech. Meanwhile, due to the rapid development of machine learning and deep learning, some problems in the medical field can be solved by combining medical treatment with technology through the cross-fusion of signal processing technology.
[0003] It is worth noting that features play an important role in the field of machine learning, as they represent the objects to be studied. Finding features that are more representative and better represent the characteristics of the subjects is of great significance, which will greatly improve the performance of model classification and recognition. For the study of pathological speech, some features are often used to represent the information of the patients. Common features include fundamental frequency, jitter, shimmer, harmonic-to-noise ratio (HNR), mel-frequency cepstral coefficient (MFCC), linear prediction coefficient (LPC), etc. At present, many researchers at home and abroad are committed to this field. Joshy and Rajan used deep learning algorithms such as DNN, CNN, and LSTM on MFCCs features to detect the severity of dysarthria in TORGO and UA Speech datasets, and proposed a new method based on gated neural network (GNN) to improve the integration of acoustic and pitch features. They also studied the Bayesian estimation of GNN to further improve its robustness. Juliette Millet and Neil Zeghidour used the original waveform features of the TORGO dataset to detect the severity of dysarthria. Abner Hernandez et al. used rhythm indicator features to detect dysarthria on QoLT Korean and TORGO datasets. N.P. Narendra and Paavo Alku used raw speech signals and glottal flow waveforms to detect the severity of dysarthria on TORGO and UA Speech datasets. Siddhartha Prakash input data into a model based on convolutional neural network and trained the model output through a softmax classifier. Krishnagurbelli and Anil Kumar Vuppala proposed using a set of auditory perceptual operators to enhance the perceptual enhancement single-frequency cepstral coefficient (PESFCC) features of dysarthria speech. Kadi used a set of prosodic features selected by LDA / SVM to detect the severity of dysarthria, and achieved an accuracy of 93% using the Nemours database. Ina Kodrasi proposed spectral temporal sparsity features, using Gini index as a robust feature to identify dysarthria speech, with a result of 83.3%. Nida Sae Jong and Pornchai Phukpattaranont used dimensionality reduction of six features to detect dysarthria speech. N.P. Narendra and Paavo Alku used acoustic and glottal features extracted from TORGO and UA Speech datasets to detect dysarthria language. Liu Shansong's research uses visual features to identify dysarthria. Although there are many features to represent pathological speech, the pronunciation of dysarthria patients is significantly different from that of normal people, and the vocal tract and vocal cords contain a wealth of information. Many commonly used features can only reflect single information.For example, Mel frequency cepstral coefficients (MFCC), linear prediction coefficients (LPC) and formants are mainly manifestations of vocal tract information; while the fundamental frequency reflects vocal cord information. Therefore, it is necessary to further explore representative feature parameters of dysarthria speech to more effectively represent dysarthria speech and improve the recognition rate of the model. SUMMARY
[0004] The dysarthria speech processing method with feature enhancement provided by the present application overcomes the problem of low recognition rate caused by ignoring the irregularity and non-stationarity of pathological speech in the prior art, and improves the recognition rate of dysarthria speech.
[0005] To solve the above technical problems, the technical scheme adopted by the present application is as follows: a dysarthria speech processing method with feature enhancement, comprising the following steps:
[0006] S1, performing fast Fourier transform on the original signal to calculate the frequency spectrum signal thereof;
[0007] S2, performing empirical mode decomposition on the frequency spectrum signal to obtain each intrinsic mode decomposition component;
[0008] S3, calculating the power spectrum density of the first m intrinsic mode decomposition components to obtain a power spectrum feature vector with a dimension of m;
[0009] S4, performing fast Walsh-Hadamard transform on each of the first m intrinsic mode decomposition components to obtain Walsh transform coefficients, and then extracting statistical features of each Walsh transform coefficient to obtain a statistical feature vector with a dimension of m x a; wherein m and a are integers, and a represents the number of statistical features;
[0010] S5, combining the power spectrum feature vector and the statistical feature vector to obtain a combined feature vector with a dimension of m x (a+1).
[0011] In the step S4, the statistical features include mean, standard deviation, maximum value, minimum value and variance.
[0012] In the step S3, the calculation method of the power spectrum density is as follows:
[0013] S301, segmenting the corresponding intrinsic mode decomposition component to obtain a plurality of segmented signals;
[0014] S302, performing windowing preprocessing and fast Fourier transform on each segmented signal to obtain a corresponding periodogram;
[0015] S303, calculating the power spectrum according to the periodogram corresponding to each segmented signal.
[0016] The calculation formula of the periodogram is as follows:
[0017]
[0018] wherein S k represents the kth segment signal, k = 1, 2, …, T, T represents the total number of segment signals, P k (f) represents the periodogram of the kth segment signal, ω[n] is a window function, f represents the frequency of the signal, n represents the current sample data point, n = 0, 1, 2, …, M, M represents the number of samples in the segment signal S k ;
[0019] The calculation formula of the power spectral density is:
[0020]
[0021] wherein P welch represents the power spectral density.
[0022] The calculation formula of the Walsh transform coefficient is:
[0023]
[0024] wherein FWT(u) represents the Walsh transform coefficient, IMF i (f) represents the i th intrinsic mode decomposition component, b t (u) is the value of the t+1th bit of the binary number of u, b (j-1-t) (u) represents the value of the j-tth bit, u represents the current u th data point, N represents the number of sampling points, j represents the number of intrinsic mode decomposition components, and t is an integer.
[0025] The value of m is 5.
[0026] In the step S1, the calculation formula of the spectrum signal is:
[0027]
[0028]
[0029] wherein s(t) represents the original signal, t represents time, f represents frequency, and L represents the length of the original signal s(t).
[0030] Compared with the prior art, the present application has the following beneficial effects:
[0031] The application provides a dysarthria speech processing method with enhanced features, which processes speech based on empirical mode decomposition (EMD) and fast Fourier transform (FFT) (referred to as FEMD), fully retains irregular and non-stationary information in pathological speech, so as to retain more sufficient pathological information for subsequent features. In the method, the speech is subjected to empirical mode decomposition to obtain intrinsic mode functions (IMFs), because the intrinsic mode functions are basic components of the speech signal. Secondly, considering the irregular shape and non-stationarity of the pathological speech, adaptive analysis of local features can better depict the features of the speech signal; meanwhile, IMFi(t) contains all the formant frequency information, vocal cord and vocal tract information closely related to phonation. Therefore, compared with normal people, the speech information and frequency obtained after FEMD processing can better represent the information in dysarthria speech, especially the pathological information carried by the patients, and the power spectrum vector formed by the power spectrum density can better represent the features of dysarthria speech.
[0032] In addition, the application also combines the fast Walsh-Hadamard transform (FWHT) method to transform the intrinsic mode functions obtained by FEMD decomposition, which is referred to as WHFEMD, then extracts statistical features of each Walsh transform coefficient, can obtain the pathological individual uniqueness in each speech, that is, the statistical feature vector, and the combined feature vector obtained by combining the power spectrum vector, can further optimize the feature extraction accuracy of dysarthria speech, and facilitate subsequent classification and recognition of dysarthria patients with different disease courses. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 A flowchart of a feature-enhanced dysarthria speech processing method provided by the embodiment of the application is shown in the figure.
[0034] Figure 2 A signal diagram in the application is shown in the figure, wherein (a) is the FFT spectrum of a dysarthria speech, (b) is the first five intrinsic mode decomposition components IMF obtained after decomposition, and (c) is the Walsh transform coefficient obtained by performing fast Walsh-Hadamard transform on the intrinsic mode decomposition components respectively.
[0035] Figure 3 A principle diagram of a comparative test FEMD feature extraction method and a FEMD_PSD feature extraction method is shown in the figure.
[0036] Figure 4 A principle diagram of a comparative test WHFEMD feature extraction method and a feature extraction method of the embodiment of the application is shown in the figure. DETAILED DESCRIPTION
[0037] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0038] As shown in Figure 1 , the embodiment of the present application provides a feature-enhanced dysarthria speech processing method, comprising the following steps:
[0039] S1, performing fast Fourier transform on the original signal to calculate a frequency spectrum signal.
[0040] In the step S1, the calculation formula of the frequency spectrum signal is:
[0041]
[0042]
[0043] Wherein, s(t) represents the original signal, t represents time, f represents frequency, and L represents the length of the original signal s(t).
[0044] S2, performing empirical mode decomposition (EMD) on the frequency spectrum signal to obtain each intrinsic mode decomposition component.
[0045] Wherein, when the decomposition is to the jth intrinsic mode decomposition component IMF j (f), the corresponding residual r j (f) has become a monotonic function or a constant, and the intrinsic mode function cannot be decomposed, the entire EMD process ends. The definition is as follows:
[0046]
[0047] Wherein, IMF i (f) represents the ith intrinsic mode decomposition component, i=1, 2, 3,..., j. r j (f) is a residual, generally representing the average trend of a signal, which is a constant sequence or a monotonic sequence, and j represents the number of intrinsic mode decomposition components.
[0048] S3, calculating the power spectrum density PSD of the first m intrinsic mode decomposition components to obtain a power spectrum feature vector with a dimension of m.
[0049] In the embodiment, the value of m is 5, that is, the first 5 intrinsic mode decomposition components IMF are taken to calculate the power spectrum feature vector, and the expression is:
[0050] Xi = IMF i (f), i = 1, 2,..., 5; (4)
[0051] Specifically, in the step S3 of the embodiment, the calculation method of the power spectral density is:
[0052] S301, segmenting the corresponding eigenmode decomposition component to obtain a plurality of segmented signals. In the embodiment, each eigenmode decomposition component is segmented, assuming that the sample number of the original signal is N, the sample number of the eigenmode decomposition component is also N, which is evenly divided into T segments, each segment has M samples, so N = M*T. Specifically, the segment number T can be determined according to the frequency domain resolution, specifically, T = f / N.
[0053] S302, respectively windowing and pre-processing and fast Fourier transforming each segmented signal to obtain a periodogram corresponding to each segmented signal.
[0054] The calculation formula of the periodogram is:
[0055]
[0056] Wherein, S k represents the kth segmented signal, k = 1, 2,..., T, T represents the number of segmented signals, P k (f) represents the periodogram of the kth segmented signal, ω[n] is the window function, f represents the frequency of the signal, n represents the current sample data point, n = 0, 1, 2,..., M, M represents the sample number in the segmented signal S k ;
[0057] S303, calculating the power spectrum according to the periodogram corresponding to each segmented signal. The calculation formula of the power spectral density is:
[0058]
[0059] Wherein, P welch represents the power spectral density.
[0060] In the embodiment, the Welch function is used to estimate the power spectrum of the periodogram method. Welch periodogram method is a method for estimating the modified periodogram power spectral density. The power spectral density is used to represent the speech signal, which can carry more characteristics of dysarthric speech.
[0061] S4, respectively performing fast Walsh-Hadamard transform on the first m eigenmode decomposition components to obtain Walsh transform coefficients, and then extracting statistical features of each Walsh transform coefficient to obtain a statistical feature vector with a dimension of mx a; wherein m and a are integers, and a represents the number of statistical features.
[0062] The calculation formula of the Walsh transform coefficient is:
[0063]
[0064] Wherein, FWT(u) represents the Walsh transform coefficient, b t (u) is the value of the t+1th bit of the binary number of u, b (j-1-t) (u) represents the value of the j-tth bit, u represents the current uth data point, N represents the number of sampling points, N is generally an integer power of 2 for easy calculation, j represents the number of intrinsic mode decomposition components, and t is an integer.
[0065] Specifically, in the step S4, considering that the speech of patients and normal people has great difference in amplitude and time change, the statistical characteristics selected in the embodiment include mean, standard deviation, maximum value, minimum value, and variance.
[0066] Mean: X meani =∑X i / N; (8)
[0067] Standard deviation: X stdi =sqrt(∑((x i -x mean ) 2 ) / N); (9)
[0068] Maximum value: X max i =max(X i ); (10)
[0069] Minimum value: X min i =min(X i ); (11)
[0070] Variance: X var i =∑[(x i -x mean ) 2 ] / (N-1); (12)
[0071] Wherein, N represents the number of sample points of the current IMF i (f).
[0072] S5, combine the power spectrum feature vector and the statistical feature vector to obtain a combined feature vector with a dimension of m x (a+1).
[0073] As Figure 2As shown, (a) is an FFT spectrum of the dysarthric speech, (b) is the first five intrinsic mode decomposition components IMFs obtained after decomposition, and (c) is the Walsh transform coefficients obtained by performing fast Walsh-Hadamard transform on the intrinsic mode decomposition components respectively. It can be observed that the spectra of the IMFs represent different speech parts. IMF1 carries almost all the energy of the signal, and the remaining spectrum of the IMFs is reduced by nearly 2 times as the IMF order increases. It can be seen from the Walsh transform coefficients that most of the signal energy is usually in the first sequence. That is, the fast Walsh-Hadamard transform compresses the energy by ordering, and the first coefficient usually contains most of the signal energy.
[0074] To verify the effectiveness of the algorithm, the TORGO database and the UA Speech dataset were selected, and the random forest and bagging algorithms were used for classification and recognition, and compared with a variety of classical feature extraction algorithms. The specific results are shown in Tables 1-2. Among them, RF represents the random forest algorithm, and BG represents the bagging algorithm.
[0075] Table 1 Accuracy rate of TORGO database
[0076]
[0077] Table 2 Accuracy rate of UA Speech database
[0078]
[0079] As shown in Table 1 and Table 2, the principle of the FEMD feature extraction method and the FEMD_PSD feature extraction method is shown in the schematic diagram as shown in Figure 3 Figure 4 The principle of the WHFEMD feature extraction method and the feature extraction method of the embodiment of the application is shown in the schematic diagram as shown in
[0080] Specifically, the FEMD feature extraction method refers to taking 5 statistical features of the first 5 intrinsic mode decomposition components obtained based on Fourier transform (FFT) and empirical mode decomposition (EMD), i.e., 25 parameters, as extraction features. The statistical feature sequence of each intrinsic mode decomposition component IMF is composed into a new feature, and then there are:
[0081] X IMFi = (X meani ,X stdi , X max i , X min i , X var i ), i = 1, 2, …, 5; (13)
[0082] wherein X IMFi represents a statistical feature set of the i-th intrinsic mode decomposition component IMF.
[0083] The statistical feature set of each intrinsic mode decomposition component IMF can constitute a vector Xstat, namely:
[0084] X stat .=(X IMF1 ,X IMF2 ,X IMF3 ,X IMF4 ,X IMF5 ) T ; (14)
[0085] The vector Xstat represents the statistical feature set of each intrinsic mode decomposition component IMF of the original data.
[0086] FEMD_PSD refers to taking the power spectrum density PSD of the first 5 intrinsic mode decomposition components obtained based on Fourier transform (FFT) and empirical mode decomposition (EMD) as the extracted features of the signal, and WHFEMD feature extraction method refers to performing fast Walsh-Hadamard transform (FWHT) on the first 5 intrinsic mode decomposition components obtained based on Fourier transform (FFT) and empirical mode decomposition (EMD), and then calculating 5 statistical features of 5 Walsh transform coefficients obtained by the transformation, and taking 25 parameters as the extracted features.
[0087] As can be seen from Table 1 and Table 2, the speech processing method of the embodiment of the application has an algorithm performance improved by 4-15% compared with the results of directly extracting LPC, PSD or even the classical MFCC feature, which proves that the application has good robustness in correctly classifying and identifying different pathological degrees.
[0088] In summary, the application overcomes the general performance of the traditional feature extraction method, processes the original signal from Fourier transform, empirical mode decomposition and fast Walsh-Hadamard transform, extracts feature functions respectively and fuses them to improve the overall feature performance, and thus improves the accuracy of dysarthria speech recognition, provides an objective and effective method for the diagnosis of dysarthria and the improvement of the intelligibility of dysarthria patients, and helps speech therapists to better understand the dysarthric speech disorder of patients and guide them to develop more scientific and effective rehabilitation programs.
[0089] It should be finally pointed out that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; the speech processing method of the present application is not limited to dysarthric speech, and is also applicable to the study of the pronunciation movement of other articulatory organs, such as glottal data. The protection scope of the present application is not limited to the above embodiments. Although the present application is described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A feature enhanced dysarthria speech processing method, characterized by, The method comprises the following steps: S1, performing fast Fourier transform on the original signal to obtain a frequency spectrum signal; S2, performing empirical mode decomposition on the frequency spectrum signal to obtain each intrinsic mode decomposition component; S3, calculating the power spectrum density of the first m intrinsic mode decomposition components to obtain a power spectrum feature vector with a dimension of m; S4, performing fast Walsh-Hadamard transform on each of the first m intrinsic mode decomposition components to obtain a Walsh transform coefficient, and then extracting statistical features of each Walsh transform coefficient to obtain a statistical feature vector with a dimension of m*a; wherein m and a are integers, and a represents the number of statistical features; S5, combining the power spectrum feature vector and the statistical feature vector to obtain a combined feature vector with a dimension of m*(a+1).
2. The method of dysarthric speech processing with feature enhancement as claimed in claim 1, wherein, In the step S4, the statistical features include mean, standard deviation, maximum value, minimum value, and variance.
3. The method of dysarthric speech processing with feature enhancement as claimed in claim 1, wherein, In the step S3, the power spectrum density is calculated by: S301, segmenting the corresponding intrinsic mode decomposition component to obtain a plurality of segmented signals; S302, performing windowing preprocessing and fast Fourier transform on each segmented signal to obtain a corresponding periodogram; S303, calculating the power spectrum according to the periodogram corresponding to each segmented signal.
4. The method of dysarthric speech processing with feature enhancement as claimed in claim 3, wherein, The calculation formula of the periodogram is: where S k represents the kth segment signal, k = 1, 2, …, T, T represents the total number of segment signals, P k (f) represents the periodogram of the kth segment signal, ω[n] is a window function, f represents the frequency of the signal, n represents the current sample data point, n = 0, 1, 2, …, M, M represents the number of samples in the segment signal S k . The calculation formula of the power spectrum density is: where P welch represents the power spectral density.
5. The method of dysarthric speech processing with feature enhancement as claimed in claim 1, wherein, The calculation formula of the Walsh transform coefficient is: where FWT(u) represents a Walsh transform coefficient, IMF i (f) represents the i-th intrinsic mode decomposition component, b t (u) is the value of the t+1 bit of the binary number of u, b (j-1-t) (u) represents the value of the j-t bit, u represents the current u-th data point, N represents the number of sampling points, j represents the number of intrinsic mode decomposition components, and t is an integer.
6. The method of dysarthria speech processing with feature enhancement of claim 1, wherein, The value of m is 5.
7. The method of dysarthria speech processing with feature enhancement of claim 1, wherein, In the step S1, the calculation formula of the frequency spectrum signal is: wherein s(t) represents the original signal, t represents time, f represents frequency, and L represents the length of the original signal s(t).