Parkinson prediction method and device based on voice signal
By constructing a multi-classifier ensemble model and optimizing using Bayes' theorem, and dynamically selecting classifiers for speech signal analysis, the accuracy and efficiency issues of early prediction of Parkinson's disease are solved, providing personalized and efficient Parkinson's disease risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies struggle to accurately and quickly predict the onset and severity of Parkinson's disease in its early stages, and machine learning methods exhibit significant performance variations across different population groups, impacting accuracy and efficiency.
We employ a classifier ensemble model that includes multiple classifiers such as decision trees, random forests, and neural networks. Through speech signal acquisition and feature extraction, we dynamically select the most suitable classifier for personalized prediction, construct a machine learning model, and optimize the classifier selection using Bayes' theorem.
It achieves highly accurate and rapid prediction of Parkinson's disease risk, and is suitable for self-analysis by the general population to detect high risk early and seek medical attention in a timely manner.
Smart Images

Figure CN114863911B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical data acquisition, and particularly relates to a Parkinson prediction method and device based on a voice signal. BACKGROUND
[0002] With the aging of society, there are more and more old people, and the probability of old people suffering from Parkinson's disease (PD) will increase. At present, there is no reversible treatment for Parkinson's disease. Although drugs can significantly alleviate the symptoms of the disease, it is difficult to predict the occurrence and severity of the disease in the early stage with a simple method. Not only is the evaluation time period long, but the evaluation process is also complex, which affects the accuracy. On the other hand, it is not convenient to evaluate patients with Parkinson's disease tendency in the hospital, which consumes medical resources and increases medical costs. Therefore, it is necessary to develop an intelligent portable device that can conveniently screen at any time and any place. Only when the risk of Parkinson's disease tendency is high, the patient can go to the hospital for diagnosis and treatment.
[0003] Some symptoms of Parkinson's disease will be shown in the sound before the occurrence of the disease. At present, there are some machine learning prediction methods based on voice. The problem is that machine learning requires a large number of training samples, and there are few labeled samples for PD tendency prediction. Because the labeled samples require professional knowledge and a large amount of manpower and time.
[0004] Existing machine learning methods rely on inertia-based modeling, making them prone to misclassifying test samples. In reality, humans dynamically adapt their methods based on the current test samples, rather than using the same method for all test samples. Because different people have unique vocal systems, and even the same person's pronunciation can vary at different times, the training speech samples used for machine learning may differ significantly from the test subject's speech during prediction. This is because they may come from different individuals, leading to significant performance differences in machine learning methods across different populations and resulting in poor prediction performance. While many excellent speech acquisition and feature extraction methods exist, they are still affected by many unpredictable factors, such as the subject's gender, highly variable acoustic environments, and the subject's physical condition and characteristics. Furthermore, the methods used for acquiring and measuring training speech may differ from those used during testing, and these methods exhibit varying robustness to the aforementioned influences. Thus, for the same test sample, these methods have different predictive capabilities. Each prediction method corresponds to a classifier, and many classifiers possess different capabilities and complementarities. In our experiments, we found that a classifier might work well for some test samples, but frequently err for others. In particular, when two classifiers are used to classify test samples, their classification abilities may be completely opposite. Therefore, it is reasonable to select different classifiers to achieve personalized prediction based on the specific circumstances of the test subject. To this end, this invention proposes a personalized prediction method and portable device for Parkinson's disease risk based on speech signals. For different test subjects, it first analyzes the test subject's situation and selects the most suitable classifier algorithm, and then uses the selected classifier algorithm to predict the test subject's risk of developing the disease. Summary of the Invention
[0005] The purpose of this invention is to provide a Parkinson's disease prediction method and device based on speech signals, which is not only highly accurate but also fast.
[0006] To achieve the above objectives, in a first aspect, the present invention provides a Parkinson's disease prediction method based on speech signals, comprising the following steps: acquiring the speech signal of a test subject and extracting the speech feature vector of the speech signal; constructing a classifier set model containing at least three different candidate classifiers, training the classifier set model to obtain a machine learning model, wherein the machine learning model selects a suitable classifier from the candidate classifiers according to the input speech feature vector; inputting the speech feature vector into the trained machine learning model and predicting the probability that the test subject has Parkinson's disease.
[0007] Preferably, the steps of acquiring the test subject's speech signal and extracting the speech feature vector of the speech signal include: acquiring the test subject's voice in a quiet environment to obtain a speech signal; preprocessing the acquired speech signal; and extracting the speech feature vector from the preprocessed speech signal.
[0008] Preferably, the at least three different candidate classifiers include decision trees, random trees, and neural networks.
[0009] Preferably, a classifier ensemble model containing at least three different candidate classifiers is constructed. The classifier ensemble model is trained to obtain a machine learning model. The machine learning model selects a suitable classifier from the candidate classifiers based on the input speech feature vectors. This step includes: assigning a class label to each speech feature vector, where the class label value is 1 if the test subject of the speech signal is a Parkinson's patient, and -1 otherwise, constructing a speech dataset SD for training; training each candidate classifier using the speech dataset SD to obtain candidate classifier models to construct the classifier ensemble model; selecting a suitable candidate classifier model for each speech feature vector in the speech dataset SD, labeling this speech feature vector with the suitable candidate classifier model, and thus constructing a new classifier model dataset CD; training a machine learning method using the classifier model dataset CD to obtain a machine learning model; and using the trained machine learning model to classify the speech feature vectors of the input speech signal to obtain the labels of the candidate classifier models for that speech feature vector.
[0010] Preferably, the machine learning method includes support vector machine, random forest, and neural network.
[0011] Preferably, for each speech feature vector in the speech dataset SD, selecting a suitable candidate classifier model and labeling the speech feature vector with the suitable candidate classifier model to construct a new classifier model dataset CD includes: dividing the speech dataset SD into k-fold cross-partitions, selecting 1 fold sequentially, and using the remaining (k-1) folds as the training set, with the entire dataset SD as the test set; selecting a classifier model from the classifier set, training it on the training set, and classifying it on the test set; obtaining the average classification accuracy of this classifier model for speech feature vector samples after k tests; calculating the selection probability of this classifier model for each speech feature vector according to Bayes' theorem; and selecting the classifier model with the highest selection probability as its label for each speech feature vector to construct a new classifier model dataset CD.
[0012] Preferably, the step of inputting the speech feature vector into the trained machine learning model and predicting the probability that the test subject has Parkinson's disease includes: using the machine learning model to classify the speech feature vector of the input speech signal, and obtaining the risk probability that the speech feature vector belongs to Parkinson's disease and does not belong to Parkinson's disease.
[0013] Preferably, after the step of inputting the speech feature vector into the trained machine learning model and predicting the probability of the test subject having Parkinson's disease, the method includes: reminding the test subject to seek medical diagnosis when the risk probability is greater than a set threshold.
[0014] Secondly, the present invention also provides a Parkinson's disease prediction device based on speech signals, the device comprising a speech signal acquisition device, a display touch screen, a processor, and software code for the speech signal-based Parkinson's disease prediction method of the first aspect; the display touch screen is used to provide feedback on the probability of the test subject having Parkinson's disease.
[0015] Preferably, the device is a smartphone or tablet computer.
[0016] Compared with existing technologies, this invention uses at least three different risk prediction classifiers as candidate classifiers, which is not only highly accurate but also fast. This method allows ordinary people to analyze themselves at any time and seek medical diagnosis and treatment as early as possible if they find that they have a tendency to Parkinson's disease or have a high risk. Attached Figure Description
[0017] Figure 1 This is a flowchart of a Parkinson's disease prediction method based on speech signals, according to an embodiment of the present invention. Detailed Implementation
[0018] To illustrate the technical content, structural features, and effects of the present invention in detail, the following description is provided in conjunction with the embodiments and accompanying drawings.
[0019] This invention provides a Parkinson's disease prediction method based on speech signals, which includes the following steps:
[0020] S1. Collect the test subject's speech signal and extract the speech feature vector of the speech signal;
[0021] S2. Construct a classifier set model containing at least three different candidate classifiers, train the classifier set model to obtain a machine learning model, and the machine learning model selects a suitable classifier from the candidate classifiers based on the input speech feature vector.
[0022] S3. Input the speech feature vector into the trained machine learning model and predict the probability of the test subject having Parkinson's disease.
[0023] This invention employs at least three different risk prediction classifiers as candidate classifiers, achieving both high accuracy and speed. This method allows the general public to easily analyze their own risk profiles and seek early diagnosis and treatment if they identify a predisposition to Parkinson's disease or a higher risk.
[0024] In embodiments of the present invention, such as Figure 1 As shown, step S1, which involves acquiring the test subject's speech signal and extracting the speech feature vector, includes:
[0025] S11. The method for collecting the test subject's speech signal includes a doctor-patient dialogue, the patient reading a designated passage and pronouncing it; when selecting pronunciation, vowels are chosen because different sounds are formed through different mechanisms. Speech data collection can be performed using various devices. To meet the needs of portable devices, the devices used for collection are typically tablets and smartphones. The information collected from the test subject includes: speech signal, test subject number, whether diagnosed with Parkinson's disease, whether there are other diseases causing speech impairment, duration of illness, UPDRS (motor), UPDRS (holistic), and the date and time of collection. The test subject's voice is collected in a quiet environment to obtain the speech signal; specifically, in this embodiment, the fable "The North Wind and the Sun," commonly used in phonetics research, is selected. The test subject reads the test material aloud in Cantonese or Mandarin at a natural speaking speed and appropriate loudness. Before the formal test, the test subject can read silently to familiarize themselves with the short text. The collection process is completed in a quiet environment, with environmental noise controlled below 45dB. Using a smartphone, the audio of patients reading text is collected and linked to their health records, recording the test subject's ID card, name, whether they have been previously diagnosed with Parkinson's disease, whether they have other diseases that cause speech impairment, the duration of the illness, UPDRS (motor), UPDRS (holistic), and the date and time of collection.
[0026] S12. Preprocess the acquired speech signal; specifically, preprocessing the speech signal includes format conversion, sampling frequency conversion, pre-emphasis, windowing and framing, removal of silent parts, separation of voiced data (vocal cord vibration) and unvoiced data (vocal cord non-vibration), data standardization, and outlier removal to improve the patient's voice quality. For example, XAduioPro tool can be used for preprocessing.
[0027] S13. Extract speech feature vectors from the preprocessed speech signal. Specifically, two methods are used for feature extraction. The first method involves manual extraction, including commonly used amplitude parameters, impulse parameters, frequency parameters, vocalization parameters, pitch parameters, and harmony parameters. The second method uses deep learning, first converting the speech signal into image data, transforming each speech signal into the time-frequency domain to preserve the time and frequency information of the data, and then extracting features according to image processing methods. This embodiment of the invention uses the openSMILE speech feature extraction tool, which has wide applications in speech recognition, emotion computing, music information retrieval, and other fields. This implementation uses openSMILE to extract the following features: Frame Energy, Frame Intensity / Loudness (approximation), Critical Band Spectrum (Mel / Bark / Octave, triangular masking filters), Mel- / Bark-Frequency-Cepstral Coefficients (MFCC), Auditory Spectrum, Loudness approximated from auditory spectra, Linear Predictive Coefficients (LPC), Line Spectral Pairs (LSP), Fundamental Frequency (via ACF / Cepstrum method and via Subharmonic-Summation (SHS)), Probability of Voicing from ACF and SHS spectrumpeaks, Voice-Quality: Jitter and Shimmer, Formal frequencies, and Bandwidths (resonant frequency and bandwidth), Zero-and Mean-Crossing rate, Psychoacoustic sharpness, and spectral harmonicity.
[0028] In this embodiment of the invention, the at least three different candidate classifiers are decision trees, random forests, and neural networks. Using different candidate classifiers can complement each other and maintain diversity. In some other embodiments, other classifiers may also be selected.
[0029] In this embodiment of the invention, step S2 involves constructing a classifier set model containing at least three different candidate classifiers, training the classifier set model to obtain a machine learning model, and the machine learning model selecting a suitable classifier from the candidate classifiers based on the input speech feature vector.
[0030] S21. Assign a category label to each speech feature vector. When the test subject of the speech signal becomes a Parkinson's patient, the category label value is 1, otherwise the category label value is -1. Construct a speech dataset SD for training in this way.
[0031] S22. Train each candidate classifier using the speech dataset SD to obtain candidate classifier models and construct a classifier set model.
[0032] S23. For each speech feature vector in the speech dataset SD, select a suitable candidate classifier model, label the speech feature vector with the suitable candidate classifier model, and then construct a new classifier model dataset CD.
[0033] S24. A machine learning method is trained using the classifier model dataset CD to obtain a machine learning model. Specifically, the machine learning method includes support vector machine, random forest and neural network. In this embodiment of the invention, the machine learning method selected is support vector machine.
[0034] S25. Use the machine learning model obtained through training to classify the speech feature vector of the input speech signal, and obtain the label of the candidate classifier model of the speech feature vector, that is, the classifier model most suitable for classifying the speech feature vector.
[0035] In this embodiment of the invention, step S23 involves selecting a suitable candidate classifier model for each speech feature vector in the speech dataset SD, labeling the speech feature vector with the suitable candidate classifier model, and then constructing a new classifier model dataset CD, which includes:
[0036] S231. Divide the speech dataset SD into k-fold cross-partitions, select 1 fold in sequence, and use the remaining (k-1) folds as the training set. The entire dataset SD is the test set. Specifically, k is a parameter, for example, a value of 10.
[0037] S232. Select a classifier model from the classifier set, train it on the training set, and classify it on the test set.
[0038] S233. After k tests, the average classification accuracy of this classifier model for speech feature vector samples is obtained.
[0039] S234. According to Bayes' theorem, calculate the probability of selecting this classifier model for each speech feature vector.
[0040] S235. For each speech feature vector, select the classifier model with the highest selection probability as its label to construct a new classifier model dataset CD.
[0041] In this embodiment of the invention, step S3, which involves inputting the speech feature vector into a trained machine learning model and predicting the probability that the test subject has Parkinson's disease, includes:
[0042] S31. The machine learning model is trained to classify the speech feature vector of the input speech signal, and the risk probability of the speech feature vector belonging to Parkinson's disease and not belonging to Parkinson's disease is obtained. After the speech feature vector of the speech signal is input into the machine learning model, the machine learning model will select an appropriate classifier model and use the selected classifier model to predict the probability of the test subject having Parkinson's disease.
[0043] The embodiment of the present invention further includes, after step S3, step S4: when the risk probability is greater than a set threshold, reminding the test subject to seek medical diagnosis. The threshold for the risk probability of having Parkinson's disease can be set to 50%. When the test result shows a risk probability of Parkinson's disease greater than or equal to 50%, the test subject is reminded to go to the hospital for diagnosis and treatment.
[0044] This invention also provides a Parkinson's disease prediction device based on voice signals. The device includes a voice signal acquisition device, a display touch screen, a processor, a speaker, and software code containing the above-described Parkinson's disease prediction method based on voice signals. The display touch screen is used to provide feedback on the probability of the test subject having Parkinson's disease.
[0045] In this embodiment of the invention, the device is a portable device such as a smartphone or tablet. When the risk probability is greater than a given threshold, the result is displayed on the touchscreen to remind the test subject to go to the hospital for diagnosis and treatment.
[0046] The Parkinson's disease prediction method and portable device based on speech signals in this invention not only have high accuracy but also high speed. The portable device allows people to analyze data anytime, anywhere, promptly detect disease risks, and seek early medical diagnosis and treatment.
[0047] The above-disclosed examples are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, any equivalent variations made in accordance with the claims of the present invention shall still fall within the scope of the present invention.
Claims
1. A Parkinson's disease prediction method based on speech signals, characterized in that, Includes the following steps: Collect the test subject's speech signal and extract the speech feature vector of the speech signal; A classifier set model containing at least three different candidate classifiers is constructed, the classifier set model is trained to obtain a machine learning model, and the machine learning model selects a suitable classifier from the candidate classifiers based on the input speech feature vector. The at least three different candidate classifiers include decision tree, random forest and neural network. The speech feature vectors are input into a trained machine learning model to predict the probability of a test subject having Parkinson's disease. The steps of constructing a classifier set model containing at least three different candidate classifiers, training the classifier set model to obtain a machine learning model, and selecting a suitable classifier from the candidate classifiers based on the input speech feature vector include: Each speech feature vector is assigned a class label. When the test subject of the speech signal is a Parkinson's patient, the class label value is 1, otherwise the class label value is -1. This is used to construct a speech dataset SD for training. Each candidate classifier is trained using the speech dataset SD to obtain candidate classifier models and construct a classifier set model. For each speech feature vector in the speech dataset SD, select a suitable candidate classifier model, label the speech feature vector with the suitable candidate classifier model, and then construct a new classifier model dataset CD. A machine learning method is trained using the classifier model dataset CD to obtain a machine learning model. The machine learning model is trained to classify the speech feature vector of the input speech signal, and the labels of the candidate classifier models for the speech feature vector are obtained.
2. The Parkinson's disease prediction method based on speech signals as described in claim 1, characterized in that, The steps for collecting the test subject's speech signal and extracting the speech feature vector include: In a quiet environment, the test subject's voice was collected to obtain speech signals; Preprocess the collected speech signals; Speech feature vectors are extracted from the preprocessed speech signal.
3. The Parkinson's disease prediction method based on speech signals as described in claim 1, characterized in that, The machine learning methods include support vector machines, random forests, and neural networks.
4. The Parkinson's disease prediction method based on speech signals as described in claim 1, characterized in that, For each speech feature vector in the speech dataset SD, a suitable candidate classifier model is selected, and the speech feature vector is labeled with the suitable candidate classifier model to construct a new classifier model. The dataset CD includes: The speech dataset SD is divided into k-fold cross-partitions. One fold is selected in turn, and the remaining (k-1) folds are used as the training set. The entire dataset SD is used as the test set. Select a classifier model from the classifier set, train it on the training set, and classify it on the test set; After k tests, the average classification accuracy of this classifier model for speech feature vector samples was obtained. According to Bayes' theorem, calculate the probability of selecting this classifier model for each speech feature vector; For each speech feature vector, the classifier model with the highest selection probability is selected as its label to construct a new classifier model dataset CD.
5. The Parkinson's disease prediction method based on speech signals as described in claim 1, characterized in that, The steps involved in inputting speech feature vectors into a trained machine learning model and predicting the probability of a test subject having Parkinson's disease include: The machine learning model is trained to classify the speech feature vector of the input speech signal, and the risk probability of the speech feature vector belonging to Parkinson's disease and not belonging to Parkinson's disease is obtained.
6. The Parkinson's disease prediction method based on speech signals as described in claim 1, characterized in that, After the steps of inputting the speech feature vectors into the trained machine learning model and predicting the probability of a test subject having Parkinson's disease include: When the probability of risk exceeds the set threshold, the test subject is reminded to seek medical diagnosis.
7. A Parkinson's disease prediction device based on speech signals, characterized in that, The device includes a voice signal acquisition device, a display touch screen, a processor, and software code containing the voice signal-based Parkinson's disease prediction method according to any one of claims 1-6; the display touch screen is used to provide feedback on the probability that the test subject has Parkinson's disease.
8. The speech signal-based Parkinson's disease prediction device as described in claim 7, characterized in that, The device is a smartphone or tablet.
Citation Information
Patent Citations
Detection method and system for pathological voice
CN103730130A
Speech emotion recognition method based on integrated deep belief network
CN106297825A