Parkinson's disease recognition system for natural voice
High characterization features are extracted through the QR-FOLDA method, which solves the redundant information and operation dependency problems of continuous speech recognition in the prior art, and achieves more efficient PD diagnosis and remote monitoring.
Patent Information
- Application Number
- CN202510572157.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing speech recognition technology mainly relies on fixed text speech in the diagnosis of Parkinson's disease, making it difficult to effectively recognize redundant information in continuous speech, and subject operation dependence and cognitive differences limit their promotion and application.
The rapid orthogonal linear discriminant analysis (QR-FOLDA) method based on QR decomposition is adopted to extract high characterization features, eliminate redundant features, and improve the discrimination ability of features. It is suitable for the Parkinson's disease recognition system of natural speech.
It improves the performance of continuous voice data in PD diagnosis, reduces the impact of noise and individual differences, enhances the reliability and generalization capabilities of the diagnostic system, and is suitable for early detection and telemedicine scenarios.
Smart Images

Figure CN120279953A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical technology, and more specifically, to a Parkinson's disease recognition system for natural speech. Background Art
[0002] Parkinson's disease (PD) is a common neurodegenerative disease that mainly affects the elderly population, and the risk of getting the disease increases significantly with age. PD is irreversible and incurable, and can only be effectively delayed by early diagnosis and timely treatment to improve the quality of life of patients. The early symptoms of PD include movement disorders, tremors, stiffness, and speech disorders, etc. Among them, dysarthria can be identified by computer in acoustic analysis at the initial stage of PD onset. The appearance of speech disorders is usually manifested as hoarseness, tremors, roughness of the voice, and changes in the fluency and rhythm of speech. Therefore, speech recognition technology has become a potential early diagnosis solution, and because of its non-invasive, convenient, and low-cost advantages, it can provide convenient remote monitoring and diagnosis means for patients.
[0003] However, current research on speech recognition of PD mainly focuses on fixed-text speech, and uses feature extraction methods such as Relief, PCA, LDA, etc. to classify and recognize speech. Although these methods can improve the classification accuracy of PD speech, reduce the high aliasing and small sample problems in the data, and alleviate the overfitting phenomenon of the model through feature selection or transformation, their application in continuous PD speech is often unsatisfactory. This is closely related to the fact that such methods cannot remove redundant information in continuous PD speech.
[0004] In addition, most existing PD speech data collection methods rely on professionals to guide subjects to read fixed texts. Such an operation method requires subjects to repeat fixed processes, which is likely to bring boredom to subjects in daily monitoring. At the same time, limited by the differences in the cognitive abilities of subjects, some subjects are difficult to fully understand and correctly follow the instructions of the guiding personnel to complete the operation of speech recognition of PD. These have severely restricted the popularization and application of speech recognition of PD. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a Parkinson's disease recognition system for natural speech (hereinafter referred to as "PD speech recognition system"). The system collects speech data that simulates patients speaking in a natural state, which can comprehensively reflect the daily performance of patients' articulation conditions, thus truly reflecting the possible speech abnormalities of patients in daily life, and is helpful for discovering early potential symptoms of PD and the daily monitoring of the condition of PD patients.
[0006] The present invention replaces the feature extraction method in the original PD speech recognition system with the proposed Quick Orthogonal Linear Discriminant Analysis based on QR decomposition (QR-FOLDA) method, aiming to eliminate redundant features through orthogonal mapping, improve the discriminative ability of the extracted features, and thus improve the performance of continuous speech data in PD diagnosis. Continuous speech recognition is of great significance in Parkinson's disease diagnosis. Compared with fixed text data, natural speech can capture the real performance of patients in daily communication, especially in the expression of speech emotions and dynamic features, which is more representative. This method reduces the dependence of subjects on repetitive tasks and overcomes operation biases caused by cognitive differences, providing a more friendly basis for the popularization of speech diagnosis. By combining the proposed feature extraction technology (QR-FOLDA), natural speech analysis not only improves the stability of feature extraction, but also effectively reduces performance fluctuations caused by data noise and individual differences, significantly enhancing the reliability and generalization ability of the diagnostic system. This technological innovation opens up new directions for the early detection of PD, cross-language applications, and dynamic monitoring in telemedicine scenarios, and is expected to promote the advancement of speech diagnosis technology towards broader practical applications.
[0007] A Parkinson's disease recognition system for natural speech, comprising: a data collector, a feature extraction module, a classifier module, and an output module, wherein,
[0008] The data collector is used to acquire continuous speech data of PD patients and provide label information related to the Parkinson's disease condition. By training and testing the PD speech data in this dataset, the aim is to identify Parkinson's disease using the speech features in the form of natural conversations.
[0009] The feature extraction module uses the Quick Orthogonal Linear Discriminant Analysis (QR-FOLDA) method to extract high-representation features from the PD speech data. This method enhances the distinguishability of the features, improves the expression ability of the feature space, enhances the generalization ability of the model and the feature discrimination effect, and provides effective information for subsequent analysis and training.
[0010] The classifier module trains a classifier model for label prediction based on the high-representation features extracted in the feature extraction module.
[0011] The output module is used to output the label prediction results of the classifier module for the PD natural speech data.
[0012] Optionally, when performing high-representation feature extraction on the feature extraction module, the aim is to find a set of projection vectors W, which increases the intra-class compactness between samples of the same class and the distance between samples of different classes, and maintains the orthogonality between the projection vectors. Its objective function can be expressed as:
[0013]
[0014] Among them, W represents the mapping matrix, and S w represents the within-class scatter matrix, and S b represents the between-class scatter matrix, and Tr(·) represents the trace of a matrix.
[0015] Optionally, the within-class scatter matrix S w represents the distribution of samples in the same class and is used to measure the similarity and dispersion of samples within a class. Suppose there are C classes of samples, and there are n c samples in class c, and X c is the sample set of class c, and the mean vector of class c is μ c , then the within-class scatter matrix S w is calculated as follows:
[0016]
[0017] where x is a sample in class c, and μ c is the mean vector of class c. The within-class scatter matrix describes the similarity and dispersion among samples, and the purpose is to minimize the within-class scatter during classification.
[0018] Optionally, the between-class scatter matrix S b represents the dispersion among the means of each class and is used to measure the differences between different classes. Suppose μ is the overall mean vector of all data, the mean of class c is μ c , and there are n c samples in class c. The between-class scatter matrix S b is calculated as follows:
[0019]
[0020] where μ is the mean vector of all data, μ c is the mean vector of class c, and n c is the mean vector of class c.
[0021] Optionally, by calculating the within-class scatter matrix S w and the between-class scatter matrix S b , the generalized eigenvalue problem is calculated by singular value decomposition and is expressed by the formula:
[0022] S b υ = λS w υ,
[0023] where υ is the eigenvector and λ is the eigenvalue. By solving the eigenvalue problem, the eigenvalue λ and the corresponding eigenvector υ are obtained, and the eigenvectors corresponding to the top k largest eigenvalues are selected to form the projection matrix V = [υ1, υ2,..., υ k , and V is an n featuresA matrix of ×k, representing the k most important feature components in the data.
[0024] Optionally, to further enhance the discriminability of features and improve the expressive power of the feature space, QR-FOLDA introduces a random matrix Q on the basis of OLDA. The introduction of the random matrix helps to break certain structures in the original feature space. The matrix VQ (VQ = V×Q) contains the potential directions for feature extraction. The dimension of the random matrix Q is k×k, and its elements follow the standard normal distribution.
[0025] Furthermore, perform QR decomposition on the matrix VQ to obtain an orthogonal matrix W and an upper triangular matrix R, which can be expressed by the formula:
[0026] QR(VQ) = WR,
[0027] where W is an orthogonal matrix (i.e., W T W = I), which is the eigenvector to be found, and R is an upper triangular matrix that contains the weights of the linear transformation. QR decomposition ensures the orthogonality between eigenvectors, prevents the problem of multicollinearity, and thus improves the stability of the model.
[0028] Finally, the original data X will be projected and mapped to a new low-dimensional space through the projection matrix W, retaining the discriminative features in the data. The mapping process can be expressed as:
[0029] X projected = XW,
[0030] where X represents the original data, with a shape of n samples ×n features , W is the orthogonal matrix obtained from QR decomposition, with a shape of n features ×k, and X projected is the data after dimensionality reduction, with a shape of n samples ×k.
[0031] Optionally, before the high-representation feature extraction by the feature extraction module, data preprocessing is also performed on the natural speech data.
[0032] Optionally, the preprocessing includes low-pass filtering of the speech signal, speech segment clipping, etc.
[0033] Optionally, the preprocessing also includes feature extraction operations on the speech signal data for training and testing. The extracted speech feature signals can include spectral features, acoustic features, Mel cepstral coefficient features, etc.
[0034] Optionally, the preprocessing also includes normalization / standardization processing of the speech feature signals, etc.
[0035] The beneficial effects of the present invention are:
[0036] For the training and testing tasks of continuous PD voice data, by designing a feature extraction and classification model for natural continuous speech, the QR-FOLDA method is used to extract high-representation feature data from PD data, focusing on enhancing the intra-class compactness of similar samples and expanding the discrimination of inter-class features, while maintaining the orthogonality of the projection vectors, thereby improving the classification effect of the model on PD voice data. Compared with the existing Parkinson's voice feature extraction algorithms, the present invention pays more attention to optimizing and enhancing the discriminative ability of data through fast orthogonal characteristics, avoiding the complexity of processing data transformation, and the algorithm has a more obvious improvement in classification accuracy.
[0037] Other advantages, objectives and features of the present invention will to some extent be described in the subsequent specification, and to some extent will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] To better present the objectives, technical solutions and advantages of the present invention, the present invention will be further described in detail below in conjunction with the drawings, where:
[0039] Figure 1 is a schematic structural diagram of a PD voice recognition system shown in an embodiment of the present application;
[0040] Figure 2 is a schematic flowchart of extracting high-representation features with high inter-class separability and high intra-class compactness from PD voice data shown in an embodiment of the present application;
[0041] Figure 3 is a schematic flowchart of a PD voice recognition method shown in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The preferred embodiments of the present invention will be described in detail below with reference to the drawings. It should be understood that the preferred embodiments are only for explaining the present invention, rather than limiting the protection scope of the present invention.
[0043] One of the early symptoms of Parkinson's disease (PD) is dysarthria. Because voice signals are rich in information and the acquisition process is non-invasive and harmless, they have become one of the main means for PD diagnosis. However, in order to fully obtain the dysarthria information of PD, researchers design characters or text fragments that can better reflect PD dysarthria according to local language characteristics for subjects to read, and often require subjects to keep pronouncing continuously to ensure that the collected voice signals can fully reflect PD lesions. In addition, for the collected voice signals, professional personnel are needed to detect the voice quality to ensure the voice quality. This has brought great limitations to the popularization and application of voice recognition for PD. To address these challenges, this application proposes a PD classification system based on continuous speech recognition. For continuous speech data, this system effectively reduces the within-class scatter and increases the between-class distance through a designed QR-FOLDA feature extraction module, thereby improving the discriminability and expressiveness of features. This method focuses on the target data itself, further improving the classification prediction accuracy of the model for PD voice data and assisting in the diagnosis effect of Parkinson's disease.
[0044] As Figure 1 shown, the system includes a data collector, a feature extraction module, a classifier module, and an output module.
[0045] The data collector is used to obtain PD continuous speech data, and this data set contains label information representing the condition of Parkinson's disease. The obtained PD voice data will be used for the training and testing of the model, with the aim of identifying Parkinson's disease through voice data in the form of natural conversations.
[0046] The feature extraction module aims to extract high-representation features from PD voice data using the proposed fast orthogonal linear discriminant analysis based on QR decomposition (QR-FOLDA) method. This method enhances the discriminability of features, improves the expressiveness of the feature space, enhances the generalization ability of the model and the feature discrimination effect, and provides effective information for subsequent analysis and training.
[0047] The classifier module trains a classifier model for label prediction based on the high-representation features extracted in the feature extraction module.
[0048] The output module is used to output the label prediction results of the classifier module for PD natural voice data.
[0049] Figure 2 is a schematic flowchart for extracting high-representation features with high between-class separability and high within-class compactness in PD voice data shown according to an embodiment of the present application;
[0050] First, at S1, continuous PD voice data in a simulated natural state can be obtained. The continuous PD voice data may include, but is not limited to, simulated conversations, reading articles aloud, etc., and should include corresponding PD tag information marked under the guidance of professional medical staff.
[0051] Next, at S2, based on the continuous PD voice data, preprocessing operations should be performed on it. The goal of the preprocessing operations is to initially remove the noise in the continuous PD voice data and ensure the efficient extraction of highly representative features. In some embodiments, the preprocessing operations should not be limited to extracting spectral features, acoustic features, and Mel cepstral coefficient features of the continuous PD voice data.
[0052] Next, at S3, the designed feature extraction algorithm can be used to extract features from the preprocessed continuous PD voice data that are highly representative of PD and have high between-class separability and within-class compactness.
[0053] Figure 3 It is a schematic flowchart of the PD voice recognition method shown in an embodiment of the application.
[0054] Before recognizing the test PD voice data, the classifier module needs to be trained first. The classifier module can be trained from the highly representative features in the PD voice data and the labeled PD voice data.
[0055] First, through the preprocessing of continuous speech, the analog speech signal needs to be converted into a digital signal through sampling and quantization techniques, and spectral subtraction is used for noise reduction to remove background noise. Then, meaningful features are extracted from the clean speech spectrum for subsequent speech analysis and recognition tasks. The preprocessing mainly extracts the MFCC features of the speech signal by simulating the characteristics of the human auditory system, which is a method widely used in speech processing and audio feature extraction.
[0056] The MFCC features obtained after preprocessing also need to be further subjected to the extraction of highly representative features of PD voice data. To fully eliminate the redundancy between features, an orthogonal mapping method is introduced to achieve the purpose of eliminating redundancy. Linear discriminant analysis (LDA) is a classic feature mapping method, and its orthogonal mapping method - orthogonal LDA (OLDA) aims to find a set of projection vectors (W) that, while increasing the within-class compactness between samples of the same class and the distance between samples of different classes, maintain the orthogonality between the projection vectors. The calculation formula of feature mapping can be expressed as Formula 1:
[0057]
[0058] where represents the within-class scatter matrix, Denotes the between-class scatter matrix. Denotes the center of the samples in class c, Denotes the sample center. There are n samples in class C c samples, X c is the sample set of class C, and the mean vector of class C is μ c , and Tr(·) denotes the trace of a matrix.
[0059] However, OLDA mainly uses an iterative method to solve the projection vector, resulting in low solving efficiency. Therefore, this application proposes a fast orthogonal linear discriminant analysis method based on QR decomposition (denoted as QR-FOLDA), aiming to quickly and orthogonally remove the redundancy between features by means of QR matrix decomposition. Specifically, first calculate the within-class scatter matrix S w and the between-class scatter matrix S b , and calculate the generalized eigenvalue problem through singular value decomposition, which can be expressed as Equation 2:
[0060] S b υ = λS w υ, (Equation 2)
[0061] where υ is the eigenvector and λ is the eigenvalue. By solving the eigenvalue problem, the eigenvalue λ and the corresponding eigenvector υ are obtained, and the eigenvectors corresponding to the first k largest eigenvalues are selected to form the projection matrix V = [υ1, υ2,..., υ k , and V is an n features ×k matrix, representing the most important k feature directions in the data.
[0062] To further enhance the discrimination of features and reduce the feature dependence under different conditions, QR-FOLDA introduces a random matrix Q on the basis of OLDA. The introduction of the random matrix helps to break the potential dependence relationship between features, thereby improving the model's ability to identify different classes. The matrix VQ (VQ = V×Q) contains the potential directions for feature extraction. The dimension of the random matrix Q is k×k, and its elements follow the standard normal distribution. Perform QR decomposition on the matrix VQ to obtain the orthogonal matrix W and the upper triangular matrix R, which can be expressed as Equation 3:
[0063] QR(VQ) = WR, (Equation 3) where W is an orthogonal matrix (i.e., W T W = I), which is the eigenvector to be found, and R is the upper triangular matrix, containing the weights of the linear transformation. QR decomposition ensures the orthogonality between eigenvectors, prevents the problem of multicollinearity, and thus improves the stability of the model.
[0064] Finally, the original data X will be projected and dimension-reduced into a new low-dimensional space through the projection matrix W, retaining the discriminative features in the data, and its mapping process can be expressed as Equation 4:
[0065] X projected = XW (Equation 4)
[0066] where X represents the original data, with a shape of n samples ×n features , W is an orthogonal matrix obtained from QR decomposition, with a shape of n features ×k, and X projected is the data after dimensionality reduction, with a shape of n samples ×k.
[0067] From the above derivation process, it can be seen that for QR decomposition, by multiplying the original data matrix with the projection matrix, the orthogonal matrix W is obtained, which is the desired projection matrix.
[0068] After the classifier module is trained, the label prediction for the tested PD voice data can be performed.
[0069] To verify the effectiveness of this system, two datasets are adopted in this implementation, which are respectively defined as MDVR-KCL (Mobile Device Voice Recordings at King's College London) and the Italian voice dataset. These two datasets can be freely obtained at https: / / zenodo.org / records / 2867216 and https: / / hf-mirror.com / datasets / birgermoell / Italian_Parkinsons_Voice_and_Spe websites respectively.
[0070] MDVR-KCL (KCL): This dataset contains voice recordings from 16 Parkinson's patients and 21 healthy controls. The recording content includes reading passages and spontaneous conversations on the phone, and the recordings are segmented by speaker identity through speaker separation technology, which is particularly suitable for analyzing spontaneous conversations. The data collection process includes the following steps: (i) After the subject relaxes, call the test executor; (ii) Read the specified article aloud; (iii) According to the subject's physical condition, require them to read the article completely; (iv) The test executor has a spontaneous conversation with the subject, randomly asking about topics such as scenic spots and local transportation; (v) The test executor ends the phone call with a farewell. The annotation of this dataset includes the Hoehn and Yahr staging (H&Y), the scores of part 5 of UPDRS II and part 18 of UPDRSIII, providing detailed Parkinson's disease-related information for each participant.
[0071] Italian: This dataset consists of three types. First is the Young Healthy Controls (YHC) dataset, which contains the voice data of 15 healthy individuals aged 19 - 29, including 13 males and 2 females, mainly from Bari and Brindisi in the Apulia region of Italy. Then is the Healthy Elderly Controls (HEC) dataset, which contains the voice data of 22 healthy individuals aged 60 - 77, including 10 males and 12 females, all from Bari in the Apulia region of Italy. Finally is the Parkinson's Disease (PD) dataset, which contains 28 individuals aged 40 - 80, including 19 males and 9 females, with 27 from Bari in the Apulia region of Italy and 1 from Venice. All patients received anti - Parkinson's treatment. The conditions and content of the experiment were carried out in a quiet, anechoic room, and the participants were kept at a distance of 15 - 25 cm from the microphone. The experimental room was warm (about 22 °C), and a short friendly conversation was carried out before the experiment to ensure that the participants were relaxed.
[0072] Table 1 Brief information of the dataset
[0073]
[0074] In this example, the classifier used is a neural network. Combining the features generated by the feature extraction module, it outputs the classification probability after multiple layers of processing, providing a reliable classification result for Parkinson's disease voice diagnosis.
[0075] To comprehensively evaluate the performance of the proposed system, this paper selects accuracy (Acc), and uses its mean (Mean) and standard deviation (Std) as the key indicators for evaluating the system. These indicators are all calculated based on the confusion matrix, which is a tool for recording the comparison between the prediction results of the classification model and the actual situation, and is widely used in the performance evaluation of supervised learning algorithms. In a binary classification problem, the confusion matrix includes four core elements: True Positive (TP), which is the number of patients correctly identified by the model; True Negative (TN), which is the number of healthy people correctly identified by the model; False Positive (FP), which is the number of healthy people misjudged as patients by the model; and False Negative (FN), which is the number of patients misjudged as healthy people by the model. These elements constitute the basic data for evaluating the model performance, and the confusion matrix is shown in Table 2.
[0076] Table 2 Confusion matrix for binary classification problems
[0077]
[0078] Accuracy (Acc): The proportion of the number of correctly classified samples to the total number of samples, expressed by the formula:
[0079]
[0080] Mean: It represents the average result of all data points numerically. It represents the average accuracy of multiple experiments and can be expressed as Equation 5:
[0081]
[0082] where x i is the i-th data point and n is the total number of data points.
[0083] Standard Deviation (Std): It represents the degree of dispersion between data points and the mean, reflecting the stability of the results, and can be expressed as Equation 6:
[0084]
[0085] where x i is the i-th data point, Mean is the data mean, and n is the total number of data points.
[0086] Through these metrics, the diagnostic performance of the model can be comprehensively evaluated from different dimensions. It can not only measure the model's ability to identify patient samples (positive) but also examine its ability to distinguish healthy samples (negative). The comprehensive consideration of these metrics can verify the effectiveness of the proposed method in PD diagnosis and its generalization performance.
[0087] The simulation platform for this embodiment is Python, and the system-related parameter settings are as follows: the random sampling rate is 0.7, the optimizer is Adam, the learning rate η is set to 0.001, and the number of iterations is set to 100.
[0088] To verify the reliability of this system, the system first optimizes the data with MFCC features. In this example, the dataset is randomly divided into a training set (80%) and a test set (20%). Each subject's sample only appears in the training set or the test set, effectively avoiding the possibility of data overlap during the training process. All experiments are repeated eight times to eliminate the influence of contingency on the experimental results.
[0089] To further demonstrate the effectiveness of the system algorithm (QR-FOLDA), some typical dimensionality reduction algorithms are compared with the method designed in this system. These feature extraction algorithms for comparison include PCA, KPCA, OLDA, LPP, AE, and the classifier used is a neural network. To eliminate the influence of contingency, all experiments are repeated ten times, and the experimental results are shown in Table 3.
[0090] Table 3 Experimental Results of Classical Feature Extraction Algorithms and System Algorithms (%)
[0091]
[0092] As can be seen from Table 3, the classification performance of QR-FOLDA is significantly better than other feature extraction methods. Specifically, i) on the KCL dataset, the classification accuracy of QR-FOLDA is 76.62%, which is significantly higher than other methods. Compared with the sub-optimal OLDA (70.20%), it is improved by 6.42, and is far ahead of the PCA method by 24.40%. On the Italian dataset, QR-FOLDA leads other feature extraction methods with an accuracy of 78.42%, better than PCA (67.88%) and the sub-optimal method KPCA (70.24%). ii) The standard deviation of the accuracy of QR-FOLDA ranks 5th in both tasks. However, compared with the method with the second-best accuracy, it lags behind by only 3.06% (3.25%) on the KCL (Italian) dataset. Although QR-FOLDA lags behind the method with the best standard deviation, LPP (OLDA), by 6.99% (5.59%) on the KCL (Italian) dataset, QR-FOLDA has an obvious advantage of 15.02% (20.66%) in terms of accuracy. This indicates that the proposed QR-FOLDA algorithm can effectively extract high-representation features of PD continuous speech, thus improving the recognition accuracy of the system for PD continuous speech. Although the standard deviation of the accuracy of QR-FOLDA is not very good, this slight deficiency is completely acceptable.
[0093] Due to its excellent performance on complex and non-linear data, the QR-FOLDA algorithm of the present invention can be regarded as an efficient feature extraction method. Compared with other classical methods, QR-FOLDA can better capture the differences between classes and suppress the influence of noise. In addition, QR-FOLDA not only has obvious advantages in scenarios with complex data characteristics, but also shows good robustness and wide applicability in various tasks, and is a leading PD continuous speech feature extraction algorithm.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A Parkinson's disease recognition system based on continuous speech, characterized in that, Including: A data collector, a feature extraction module, a classifier module, and an output module. Among them, The data collector is used to obtain PD continuous speech data, and this data set contains label information representing the Parkinson's disease condition. The obtained PD speech data will be used for the training and testing of the model, with the aim of identifying Parkinson's disease through speech data in the form of natural conversations. The feature extraction module aims to extract high-representative features from PD speech data by using the proposed fast orthogonal linear discriminant analysis (QR-FOLDA) method based on QR decomposition. This method enhances the discriminability of features, improves the expression ability of the feature space, enhances the generalization ability of the model and the feature discrimination effect, and provides effective information for subsequent analysis and training. The classifier module trains a classifier model to make label predictions based on the high-representative features extracted in the feature extraction module. The output module is used to output the label prediction results of the classifier module for PD continuous speech data.
2. The Parkinson's disease recognition system for natural speech according to claim 1, wherein The extraction of high-representative features from PD speech data includes increasing the within-class compactness between samples of the same class and the distance between samples of different classes, and maintaining the orthogonality between projection vectors.
3. The Parkinson's disease recognition system for natural speech according to claim 2, characterized in that, $W$ represents the mapping matrix, and $S$ w is denoted as the within-class scatter matrix, and $S$ b is denoted as the between-class scatter matrix, and $\text{Tr}(\cdot)$ represents the trace of a matrix. The objective function can be formulated as follows:
4. The Parkinson's disease recognition system for natural speech according to claim 3, wherein Assume that for a C classification task, there are n samples in class c, X is the sample set of class c, and the mean vector of class c is μ; c The within-class scatter matrix S describes the similarity and discreteness among samples, with the aim of minimizing the within-class scatter. The between-class scatter matrix S represents the degree of discreteness among samples of different classes, used to measure the differences between different classes, and should be as large as possible; among them, c is the within-class scatter matrix; c ; w The within-class scatter matrix S describes the similarity and discreteness among samples, with the aim of minimizing the within-class scatter. The between-class scatter matrix S represents the degree of discreteness among samples of different classes, used to measure the differences between different classes, and should be as large as possible; among them, b represents the degree of discreteness among samples of different classes, used to measure the differences between different classes, and should be as large as possible; among them, represents the within-class scatter matrix; is represented as the between-class scatter matrix.
5. The Parkinson's disease recognition system for natural speech according to claim 4, characterized in that, Using the within-class scatter matrix S w and the between-class scatter matrix S b , the generalized eigenvalue problem is calculated by singular value decomposition, which is expressed by the formula as follows: S b υ = λS w υ, Among them, υ is the feature vector and λ is the eigenvalue.
6. The Parkinson's disease recognition system for natural speech according to claim 5, characterized in that, Solve the eigenvalue problem to obtain the eigenvalues λ and the corresponding eigenvectors υ, and select the eigenvectors corresponding to the first k largest eigenvalues to form the projection matrix V = [υ1, υ2,..., υ k , where V is an n features × k matrix, representing the k most important feature directions in the data. To further enhance the discrimination of features and improve the expressive power of the feature space, the proposed Quick Orthogonal Linear Discriminant Analysis (QR-FOLDA) method based on QR decomposition introduces an orthogonal constraint on the basis of Linear Discriminant Analysis (LDA), making the features extracted by it independent of each other and having stronger discriminative ability. Its solution is obtained by introducing a random full-rank matrix Q on the basis of the solution V of LDA and performing QR matrix decomposition on the matrix VQ (VQ = V × Q). Expressed by the formula as: QR(VQ) = WR, Among them, the matrix VQ contains the latent directions for feature extraction. The random matrix Q has a shape of k×k, W is an orthogonal matrix (i.e., W T W = I), R is an upper triangular matrix that contains the weights of the linear transformation. The QR decomposition ensures the orthogonality between the eigenvectors.
7. The Parkinson's disease recognition system for natural speech according to claim 6, wherein, The original data X is dimensionally reduced and mapped to a new low-dimensional space through the projection matrix W, and the discriminative features in the data are retained. Its The dimensional reduction process can be expressed by the formula: X projected = XW, Among them, X is the original data with a shape of n samples ×n features , W is an orthogonal matrix obtained from QR decomposition with a shape of n features ×k, and X projected is the data after dimensionality reduction with a shape of n samples ×k.
8. The Parkinson's disease recognition system for natural speech according to claim 7, characterized in that Before the feature extraction module performs high-representative feature extraction, it is also necessary to perform data preprocessing on the natural speech data.
9. The Parkinson's disease recognition system for natural speech according to claim 8, characterized in that, The data preprocessing includes operations such as low-pass filtering, speech segment clipping, speech feature extraction, and standardization of speech features on the speech data.
10. The Parkinson's disease recognition system for natural speech according to claim 9, characterized in that, The extracted speech signal features may include spectral features, acoustic features, Mel-frequency cepstral coefficient (MFCC) features, etc.
11. The Parkinson's disease recognition system for natural speech according to claim 10, characterized in that, The data collector can be any recording device or online sound collection device with a certain sampling accuracy.
12. The Parkinson's disease recognition system for natural speech according to claim 11, characterized in that, The classifier used in the classifier module is not limited and can be any one of neural networks, support vector machines, random forests, extreme learning machines, or other classifiers.