Parkinson's disease and idiopathic tremor identification model and device based on video
Through the video-based identification model of Parkinson's disease and idiopathic tremor, the hand key points are extracted and classified using the RTMPose-L and PatchTST models, the problems of insufficient accuracy and wearable devices for identification of Parkinson's disease and idiopathic tremor in the prior art are solved, and high-precision distinction and remote identification are achieved.
Patent Information
- Application Number
- CN202510355830.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has strong subjectivity, time-consuming and complex problems when distinguishing Parkinson's disease from idiopathic tremor, especially when clinical characteristics are similar, the existing video-based methods are insufficient in the accuracy of the wearable devices, and the cost and complex operation are limited by wearable devices.
The video-based Parkinson's disease and idiopathic tremor identification model is adopted, including the pose estimation module, the data processing module and the classification module. The RTMPose-L model is used to extract the coordinates of the hand key point, combined with the PatchTST time series classification model, and feature extraction and classification are performed through the Transformer encoder.
It has achieved high-precision distinction between Parkinson's disease and idiopathic tremor, especially suitable for elderly patients with scarce medical resources or mobility difficulties, and provides objective and remote identification methods, which significantly improves the classification accuracy.
Smart Images

Figure CN120298768A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the application of artificial intelligence in Parkinson's disease and essential tremor, and particularly relates to a video-based discrimination model for Parkinson's disease and essential tremor and its device. Background Art
[0002] With the rapid growth of the global elderly population, the prevalence of Parkinson's disease (PD) and essential tremor (ET) has increased sharply [1]. PD is the most common neurodegenerative disease [2]. The prevalence of Parkinson's disease in the elderly population over 65 years old in China is about 1.7%, and over 4% in those over 80 years old [3]. ET is the movement disorder disease with the highest incidence, and its incidence also increases with age, reaching 5.79% in the elderly population over 65 years old [4].
[0003] The disease courses and prognoses of these two diseases are different. Therefore, accurate diagnosis is crucial for predicting the development of the disease and formulating long-term management plans. However, the current diagnosis of the two diseases mainly relies on face-to-face assessment by neurologists using clinical rating scales. This method is highly subjective, time-consuming, and complex [5]. Since PD and ET have similar clinical features, such as tremors and gait abnormalities, and there is a certain overlap in symptoms, the differential diagnosis of the two diseases is relatively difficult. A past research report showed that about one-third of PD patients were misdiagnosed as ET [6], resulting in these patients failing to receive appropriate treatment in a timely manner [7].
[0004] To overcome the deficiencies of traditional diagnostic methods, intelligent diagnostic methods based on video and wearable sensor technologies have emerged [8]. Such methods use artificial intelligence to analyze the movement data of patients, thereby extracting the kinematic characteristics of limb movements and achieving objective and accurate diagnosis and evaluation of movement disorder diseases. The National Institute for Health and Care Excellence in the UK officially recommended several artificial intelligence-based technologies as options for PD monitoring in its 2023 guidelines [9]. The release of this guideline marks the official institution's gradual recognition of the clinical value of these technologies. However, wearable technologies also have certain limitations, such as high cost, complex operation, and poor comfort.
[0005] Currently, there is no shortage of research on the diagnosis and differentiation of PD and ET based on accelerometers [8]. For example, Hathaliya et al.
[10] (2022) used accelerometer data to distinguish between PD and ET patients based on a CNN model; Ma et al.
[11] (2022) adopted a machine learning method to quantify the tremor severity of ET patients through accelerometer data. Compared with methods using wearable devices such as accelerometers, video-based tremor assessment has advantages such as non-contact and remote operation.
[0006] Kovalenko et al.
[12] (2021) developed a machine learning-based video analysis method for differentiating PD and ET. The researchers used a Logitech camera to record videos of 15 movement tasks (resolution 640×480, frame rate 30Hz) of 42 PD patients and 13 ET patients, including general movements, fine coordination movements, resting tremors, and clinical assessment tasks, etc. After the video recording, the doctor marked the videos, adjusted the pixels to 368×368, and then input them into the OpenPose library to extract key points, and calculated the relative speed and acceleration. The researchers further used the Savitzky-Golay filter for data denoising, and combined time and frequency domain analysis to screen out the features with the largest amount of information, and adopted dimensionality reduction technology to simplify the data. Finally, various machine learning models such as Logistic Regression, XGBoost, Random Forest, and Support Vector Machine (SVM) were used for classification. Among them, the Random Forest model performed the best, with a classification accuracy of 0.77.
[0007] Hayashida et al.
[13] (2023) used a GoPro camera or the built-in camera of a smartphone to record videos of hand tremors of 19 PD patients and 6 ET patients (resolution 1920×1080, frame rate 30Hz or 60Hz) for the assessment of tremor severity and the differential diagnosis between PD and ET. During video recording, the camera was fixed at a position about 30 cm from the hand, and a green felt was placed behind the hand to extract the hand area. After experts manually selected and labeled the video frames most relevant to the motor symptoms, the captured video data was trimmed to remove the extra frames outside the selected area, and the obtained trimmed data was used as the input video of the system. Subsequently, the input video was segmented into an image sequence, the hand area was separated from the green screen background using image processing methods, and the images were converted to the HSV color space. The hand area was extracted by manually setting the hue range of the background. The time series of hand movements was output by calculating the displacement of the hand area in consecutive binary images, and statistical features were further extracted from the time series. Finally, a logistic regression model was used to classify the severity of tremors and the disease types. The research results showed that the accuracy rate of classifying PD and ET diseases was only 0.56. This result indicates that this study failed to achieve a high accuracy rate in distinguishing PD and ET at the practical application level and still needs further improvement. Summary of the Invention
[0008] In view of the above problems, the present application aims to propose a video-based differential diagnosis model and device for Parkinson's disease and essential tremor.
[0009] The video-based differential diagnosis model for Parkinson's disease and essential tremor of the present application includes: a pose estimation module, a data processing module, and a classification module;
[0010] The pose estimation module includes a pre-trained RTMPose-L model; the pose estimation module performs full-body pose estimation on each frame of the three upper limb movement videos of the subject, extracts the wrist and five finger key point coordinates of the target hand, and forms a multi-group absolute coordinate sequence of hand key points;
[0011] The data processing module converts the absolute coordinate sequence output by the pose estimation module into a relative coordinate sequence, and calculates a statistical feature sequence of speed, acceleration, amplitude, frequency, and entropy based on the relative coordinate sequence;
[0012] The classification module includes a PatchTST time series classification model; the pose estimation module divides the relative coordinate sequence and the statistical feature sequence as its input into multiple patches of a predetermined length, then extracts features from each patch through the encoder of the Transformer and outputs patch predictions, and finally flattens the output patch predictions into a one-dimensional vector and classifies Parkinson's disease and essential tremor through a fully connected layer.
[0013] Preferably, the three upper limb movements include finger-to-nose, palm-turning, and finger-palm flexion and extension movements of both hands.
[0014] Preferably, the absolute coordinate sequences of the multiple groups of hand key points are eight groups of hand key point coordinate sequences, and the dimension of each group of hand key point coordinate sequences is 6, respectively representing the coordinates of the unilateral wrist and the ends of five fingers.
[0015] Preferably, for each hand key point, when the absolute coordinate sequence is converted into a relative coordinate sequence, starting from the second frame, the absolute coordinates of the hand key point in each frame of the image are respectively subtracted from the absolute coordinates of the hand key point in the first frame of the image to obtain the relative coordinates of the hand key point corresponding to this frame of the image, and the relative coordinates represent displacement; the relative coordinates of the hand key point corresponding to each frame of the image except the first frame form a relative coordinate sequence.
[0016] Preferably, there is no background restriction when recording the three upper limb movement videos of the subject.
[0017] The video-based discrimination device for Parkinson's disease and essential tremor of the present application includes: a pose estimation unit, a data processing unit, and a classification unit;
[0018] The pose estimation unit is a computing unit running the RTMPose-L model; the input of the pose estimation unit is the three upper limb movement videos of the subject, and the output is the absolute coordinate sequences of multiple groups of hand key points; the pose estimation unit performs full-body pose estimation on each frame of the image in the three upper limb movement videos of the subject, extracts the wrist and finger key point coordinates of the target hand, and forms the absolute coordinate sequences of multiple groups of hand key points;
[0019] The data processing unit is used to convert the absolute coordinate sequences output by the pose estimation unit into relative coordinate sequences, and use the relative coordinate sequences to calculate the statistical feature sequences of speed, acceleration, amplitude, frequency, and entropy;
[0020] The classification unit is a computing unit running the PatchTST model; the input of the classification unit is the relative coordinate sequences and statistical feature sequences output by the data processing unit, and the output is the classification result of Parkinson's disease and essential tremor; the classification unit divides the input relative coordinate sequences and statistical feature sequences into multiple patches of a predetermined length, then extracts features from each patch through the encoder of the Transformer and outputs patch predictions, and finally flattens the output patch predictions into a one-dimensional vector and classifies Parkinson's disease and essential tremor through a fully connected layer.
[0021] The video-based discrimination model and device for Parkinson's disease and essential tremor of the present application can achieve high-precision judgment only through an ordinary camera compared with wearable devices, and is especially suitable for areas with scarce medical resources or elderly patients with inconvenient mobility, providing a new idea for the objective and remote discrimination of movement disorder diseases. Description of the Drawings
[0022] Figure 1 Schematic diagrams of three upper limb movements included in the present application.
[0023] Figure 2 Architecture diagram of the video-based discrimination model for Parkinson's disease and essential tremor of the present application.
[0024] Figure 3 Schematic diagram of human key points extracted using the RTMPose-L model.
[0025] Figure 4 Confusion matrix and ROC curve for classifying PD and ET with three upper limb movement tasks. Detailed Implementation Modes
[0026] Research Subjects
[0027] This study has obtained the approval of the Ethics Committee of Chinese PLA General Hospital (S2018-021-00 / 01), and all subjects have signed the informed consent form. Subject screening was conducted in the outpatient department of the PLA General Hospital, and the collection of experimental data started on November 19, 2021. Subjects need to meet the inclusion criteria: 1. Diagnosed by the Department of Neurology, ineffective or intolerant after taking medications; 2. Have the ability and willingness to participate in all research visits; 3. Have no cognitive impairment and can communicate feelings with the doctor during treatment; 4. Have no severe anxiety or depression; have no serious cardiovascular or cerebrovascular diseases. For PD patients, they meet the UK Brain Bank PD clinical diagnostic criteria, including bradykinesia, at least one additional symptom (muscle rigidity, resting tremor or postural instability), and three or more diagnostic criteria (such as unilateral onset, presence of resting tremor, progressive disease, etc.). For ET patients, they meet the ET diagnostic criteria proposed by the American Academy of Neurology and the International Essential Tremor Foundation, including core diagnostic criteria (such as action tremor of both hands and forearms, without other neurological signs, etc.) and secondary diagnostic criteria (such as disease duration exceeding 3 years, family history, etc.). As of January 16, 2024, a total of 63 ET patients and 14 PD patients were included, and the age distribution of the participants was between 30 and 80 years old. The detailed information of the subjects is shown in Table 1.
[0028] In this study, we first designed a motor task for video-based discrimination of PD and ET. We referred to the consensus statement on tremor classification written by the International Parkinson and Movement Disorder Society Tremor Working Group
[15] , and selected the Unified Parkinson's Disease Rating Scale (UPDRS)
[16] , which is most commonly used in clinical assessment of PD, and the Fahn-Tolosa-Marin Tremor Rating Scale (FTM-TRS)
[17] , which is most commonly used in assessment of ET, as the basis for task design. By summarizing the motor tasks related to tremor assessment in the third part (motor function assessment) of the UPDRS scale and the FTM-TRS scale, we finally selected three upper limb motor tasks included in both scales for video acquisition, namely: bilateral action tremor (finger-to-nose), bilateral forearm circumduction (palm flipping), and bilateral palm movement (finger-palm flexion and extension). As Figure 1 shown.
[0029] The video data of the subjects were collected under the guidance of a neurologist. The subjects sat on a chair in a comfortable position and performed the above three motor tasks under the guidance of a neurologist according to the criteria of the FTM-CRST scale. The collection process was recorded in full by the rear camera of a Xiaomi tablet. Finally, a total of 1,136 videos (video resolution: 1920×1080, frame rate: 30fps, format: mp4) were collected from the above three tasks. The data collection scope of this study is wider than that of previous studies (such as Kovalenko et al.
[12] , Hayashida et al.
[13] ).
[0030] Table 1 Subject Information
[0031]
[0032] Software Environment and Research Framework
[0033] An AMD Ryzen 9 5950X CPU and an NVIDIA GeForce RTX 3090 GPU were used for data processing, model training and testing. PyTorch 2.5.1 and CUDA 12.4 were used for training and inference of the deep learning model, and the human pose estimation task was completed based on the MMPose 1.2.0 framework.
[0034] In this study, a human pose estimation model was used to extract the key point coordinates of the wrists and fingers of PD and ET patients in the movement task videos, generating key point sequences and statistical features of hand movements. Subsequently, a time series prediction and classification model based on the Transformer architecture was used to classify and analyze these key point sequences, thus realizing the differential diagnosis of PD and ET based on videos. The model architecture of this study is as Figure 2 shown.
[0035] Extraction of Hand Key Point Coordinate Sequences and Movement Features
[0036] Considering the clinical application scenarios and the movement assessment based on mobile devices or remote servers, it is necessary to achieve high efficiency and reliable speed and accuracy in hand key point extraction. We investigated the pose estimation performance of various human pose estimation methods (PaddleDetection
[18] , AlphaPose
[19] , MMPose
[20] ) combined with different detectors (YOLOv3
[21] , Faster-RCNN
[22] , and RTMDet
[23] ). These methods were all trained on the 2D human pose estimation dataset COCO
[24] . After comparison, we found that the RTMPose (Real-Time Multi-Person Pose Estimation based on MMPose) model
[25] based on the MMPose framework performed best in balancing model performance and complexity and showed strong detection robustness.
[0037] We selected the optimized RTMPose-L model with the best performance and applied it to the movement task videos of patients for human pose estimation, and extracted the human key point coordinates of each frame of the video. The specific pre-trained model used was set as rtmpose-l_8xb32-270e_coco-wholebody-384×288. The model weights were obtained by training on the COCO whole body key point dataset with an input image size of 384×288 and a batch size of 32 on 8 GPUs for 270 epochs. We directly used the RTMPose-L model to load the above pre-trained model weights to process the patient videos.
[0038] For the three selected evaluation movement tasks - bilateral action tremor (finger-to-nose), bilateral palm movement (finger-palm flexion and extension movement), and bilateral forearm circumduction movement (palm turning), with the left and right hands as the evaluation objects respectively, full body pose estimation was performed on each frame of the video and the key point coordinates of the wrists and fingers of the target hand were extracted, forming a total of eight groups of hand key point coordinate sequences. The dimension of each coordinate sequence is 6 (representing the coordinates of the wrist and five fingers respectively). The output results are asFigure 3 As shown, where A. Bilateral action tremor (finger-nose test); B. Forearm rotation movement (palm flipping); C. Bilateral palm movement (finger-palm flexion and extension movement). Next, for each group of coordinate sequences, we convert the absolute coordinates output by the RTMPose-L model into relative coordinates, that is, the coordinates of each hand key point minus its coordinates in the first frame, to obtain relative coordinates representing displacement. Using the obtained relative coordinate sequences, we calculate five statistical features on the coordinate sequences of velocity, acceleration, amplitude, frequency, and entropy, and save the coordinate sequences and the statistical features obtained from the coordinate sequences as the input for the next feature extraction and analysis.
[0039] Classification of Hand Key Point Coordinate Sequences
[0040] We use the PatchTST (Patch Time Series Transformer) model
[26] to classify the hand key point coordinate sequences. PatchTST is a time series prediction and classification model based on the Transformer architecture. By dividing the time series data into "patch blocks" and using the self-attention mechanism of the Transformer, it can effectively reduce the dimension of the input sequence while capturing long-term dependencies and complex patterns in the sequence, thus showing excellent performance in time series prediction and classification tasks. The PatchTST model for classification first divides the input sequence into multiple patch blocks of length P, then extracts features from each input patch through the encoder of the Transformer and outputs patch predictions. Finally, the output results are flattened into a one-dimensional vector and classified into PD and ET through a fully connected layer.
[0041] We adopt five-fold cross-validation to train and evaluate a PatchTST model for each evaluation task and the coordinate sequences of each hand. Specifically, we group all hand key point coordinate sequences by patient number. The coordinate sequences in each group come from the same patient. Then these groups are divided into five mutually exclusive subsets. Each time, four of these subsets are used as the training set, and the remaining one subset is used as the validation set. This is repeated five times, and each time a different validation set is used for training and testing.
[0042] During model training, we used binary cross-entropy loss function as the loss function, adopted the Adam optimizer for model training, set the learning rate to 0.0001, the batch size to 64, and the number of training epochs to 200. During the training process, we used a learning rate scheduler (lr_scheduler.ReduceLROnPlateau) to dynamically adjust the learning rate according to the validation loss. Set the learning rate scaling value to 0.5. When the validation loss has not decreased for 10 consecutive epochs, the learning rate will be dynamically reduced to half of the original. At the same time, we also adopted an early stopping strategy, monitored according to the validation set loss value. When the validation set loss value has not decreased for 20 consecutive epochs, stop training and save the model weights with the lowest validation set loss.
[0043] Baseline experiment design
[0044] To further verify the effectiveness of the proposed method in this study and conduct a comparative analysis with related studies, we designed a baseline experiment. To evaluate the classification performance of different models on the same dataset, thus providing a comparison benchmark for the proposed method in this study. We selected the main model of this study, the time series classification model PatchTST based on the Transformer architecture. Related studies used logistic regression, XGBoost, random forest, and support vector machine models, as well as the commonly used long short-term memory network model (Long Short-Term Memory, LSTM) and Informer model.
[0045] In the experiment, we selected three feature combinations for the experiment: the hand key point sequence based on x and y coordinates, which directly reflects the position change of the hand in the video. And statistical features (based on speed, acceleration, amplitude, frequency, and entropy) that more comprehensively describe the dynamic characteristics of hand movement, including speed change, acceleration change, movement amplitude, frequency component, and signal complexity. And the feature combination of the hand key point sequence of x and y coordinates combined with statistical features. Similar to the main experiment, after extracting the hand key point coordinate sequence, we performed normalization to eliminate the influence of scale and position. Using the five-fold cross-validation method, different validation sets were used for training and testing each time. And calculated evaluation metrics such as accuracy (Accuracy, ACC), area under the curve (Area Under Curve, AUC), precision, recall, and F1 value (F1Score) to evaluate the classification performance of the model.
[0046] Results
[0047] 1. Model classification performance
[0048] In this study, the classification performance of different feature combinations and models in three upper limb movement tasks was evaluated through five-fold cross-validation. The evaluation results of the classification performance of the models in three upper limb movement tasks (both hands palm movement, both hands forearm rotation movement, and both hands action tremor) are shown in Tables 2-4, and Table 5 shows the overall ranking results of the classification performance of the models with three feature combinations: hand key point coordinates, statistical features, and the combination of key point coordinates and statistical features. The results show that the PatchTST model based on the Transformer architecture performs optimally when fusing the hand key point coordinate sequence and statistical movement features (Table 5). Especially in the task of both hands action tremor (finger-to-nose), it is significantly better than other models (Table 2).
[0049] In the task of both hands action tremor (finger-to-nose), the average ACC, AUC, and F1 values of the PatchTST model under the combination of fusing the hand key point coordinate sequence and statistical movement features reached 0.917, 0.957, and 0.916 respectively, and its performance improved by more than 12% compared with traditional machine learning models (such as Random Forest with ACC = 0.779) and time series models (such as LSTM with ACC = 0.777). In the forearm rotation movement (palm turning) task, the average ACC of the PatchTST model under the feature combination of the hand key point coordinate sequence was 0.798, and the AUC was 0.837. Although it was slightly lower than other tasks, it was still better than Informer (average ACC = 0.800) and XGBoost (average ACC = 0.801). In the both hands palm movement (finger-palm flexion and extension movement) task, the average ACC of the PatchTST model under the combination of fusing the hand key point coordinate sequence and statistical movement features was 0.873, and the AUC was 0.911, which was about 0.5%-1.2% higher than that of the single feature combination.
[0050] Comparing with the baseline experiment, it was found that traditional machine learning models (such as Logistic Regression, SVM) were limited in performance when dealing with high-dimensional time series data, and LSTM performed the worst due to gradient problems (average ACC = 0.695). In addition, the models relying only on statistical features had low robustness, especially in complex action tasks, the AUC value decreased significantly. For example, in the palm movement task, LSTM (average AUC = 0.623).
[0051] Through the analysis of the model performance results of the three movement tasks, it was found that in the task of both hands action tremor (finger-to-nose), because the action was simple and the tremor characteristics were significant, the overall performance of the model was the best, and the average F1 value of the PatchTST under the combination of the hand key point coordinate sequence and statistical movement features reached 0.895.
[0052] The influence of feature combinations on model performance was found through analysis. The sequence of hand key points based on x and y coordinates directly reflects the spatio-temporal dynamics of hand movement trajectories, but lacks a quantitative description of movement characteristics. Under this feature, the average ACC of PatchTST in the finger-to-nose task reaches 0.897, higher than that of LSTM (0.777) and random forest (0.779), indicating that it can effectively capture long-term dependencies in the sequence. The dynamic indicators of statistical motion features (speed, acceleration, amplitude, frequency, and entropy), but their information density is low. For example, the average ACC of LSTM under statistical features in the palm-turning task is only 0.636, much lower than its average of 0.791 under the combination of key point coordinate sequences and statistical motion features, indicating that a single statistical feature is difficult to characterize complex motion patterns. When the hand key point coordinate sequence is combined with statistical motion features, the spatio-temporal trajectory and dynamic characteristics are fused, and the classification effect of the model is the best.
[0053] Table 2 Classification performance evaluation results of the model in the bilateral action tremor (finger-to-nose) task
[0054]
[0055] Table 3 Classification performance evaluation results of the model in the forearm rotation movement (palm-turning) task
[0056]
[0057]
[0058] Table 4 Classification performance evaluation results of the model in the bilateral palm movement (finger-palm flexion and extension movement) task
[0059]
[0060] Table 5 Overall classification performance evaluation results of the model
[0061]
[0062] 2. Confusion matrix and ROC curve
[0063] We plotted the confusion matrix (Cnfusion Matrix) and ROC (Receiver Operating Characteristic) curve of the PatchTST model when fusing the hand key point coordinate sequence and statistical motion features in three motion tasks (see Figure 4)。We observe the confusion matrix to deeply analyze the patterns of classification errors. The results show that the model performs excellently in the task of bilateral action tremor (finger-to-nose test). The confusion matrix shows a high classification accuracy and fewer misclassified samples. The AUC value of the ROC curve is close to 1, indicating good discrimination ability at different thresholds. The AUC values of each fold line are relatively high, suggesting that the model performs stably on different data subsets. In the task of bilateral forearm circumduction (palms-up movement), the confusion matrix shows 50 false positives and 40 false negatives. Compared with the action tremor task, the classification performance of this task is slightly inferior, with a higher number of false positives. The AUC values of different folds of the ROC curve range from 0.764 to 0.934, and there is some fluctuation in the AUC values of each fold line. The model's performance on different data subsets is not very stable, but still reaches an average AUC of 0.816. The model performs well in the task of bilateral palm movement (finger-palm flexion and extension movement), with a reasonable distribution of predicted and actual values and a high classification accuracy.
[0064] Overall, the PatchTST model demonstrates excellent classification performance in differentiating PD and ET, especially in the action tremor task, with high accuracy and good robustness.
[0065] This study proposes an intelligent discrimination method for PD and ET based on video analysis, which realizes the efficient discrimination of PD and ET through human pose estimation and deep learning techniques. The research results show that the proposed PatchTST model performs excellently in three upper limb movement tasks. Especially in the task of bilateral action tremor (finger-to-nose test), it achieves an average accuracy of 0.917 and an AUC of 0.957 for five folds, significantly outperforming the existing research methods based on video or wearable sensors (the accuracy of 0.77 by Kovalenko et al.
[12] ).
[0066] This study combines the RTMPose-L human pose estimation model with the Transformer-based PatchTST time series classification model for the first time to discriminate PD and ET. Compared with traditional machine learning models (such as random forest, SVM) and time series models such as LSTM, PatchTST effectively captures the long-term dependencies of hand movement trajectories through the self-attention mechanism, and combines key point coordinates with statistical features (such as speed, acceleration, frequency, etc.), significantly improving the classification performance. In addition, the data coverage (1136 video segments) and task design (three standardized upper limb movements) in this study are more comprehensive than previous studies, enhancing the generalization ability of the model. For example, Hayashida et al.
[13] only used hand tremor videos, and the classification accuracy was only 0.56, while this study achieved higher robustness through multi-task feature fusion.
[0067] When directly inputting video images into a deep learning model, the model may mistakenly learn irrelevant factors such as the background, lighting, or clothing as features. To prevent the model from learning irrelevant features, Hayashida et al.
[13] required patients to record videos in front of a green screen background, with a demanding recording angle and a complex process. Our research provides an efficient and reliable method for hand movement assessment, which only extracts the key point sequences of the wrist and fingers in each frame of the image, overcoming the influences of the shooting scene, environment, individual differences of patients, and clothing, and has strong generalization and practicality.
[0068] References
[0069] [1]Feigin V L, Vos T, Nichols E, et al. The global burden of neurological disorders: translating evidence into policy[J]. The Lancet Neurology, 2020, 19(3): 255 - 265.
[0070] [2]Kalia L V, Lang D A E. Parkinson’s disease[J]. Lancet, 2015, 386(9996): 896 - 912.
[0071] [3]Zhu J, Cui Y, Zhang J, et al. Temporal trends in the prevalence of Parkinson’s disease from 1980 to 2023: a systematic review and meta - analysis[J]. The Lancet Healthy Longevity, 2024, 5(7): e464 - e479.
[0072] [4]Louis E D, Ottman R, Allen Hauser W. How common is the most common adult movement disorder?Estimates of the prevalence of essential tremor throughout the world[J]. Movement Disorders, 1998, 13(1): 5 - 10.
[0073] [5] Jankovic J, Tan E K. Parkinson’s disease: Etiopathogenesis and treatment[J]. Journal of Neurology Neurosurgery & Psychiatry, 2020, 91(8): jnnp - 2019 - 322338.
[0074] [6] Jain S, Lo S E, Louis E D. Common misdiagnosis of a common neurological disorder - How are we misdiagnosing essential tremor?[J]. Archives of Neurology, 2006, 63(8): 1100 - 1104.
[0075] [7] Moon S, Song H J, Sharma V D, et al. Classification of Parkinson’s disease and essential tremor based on balance and gait characteristics from wearable motion sensors via machine learning techniques: a data - driven approach[J]. Journal of Neuroengineering and Rehabilitation, 2020, 17(1): 125.
[0076] [8] Peng Y, Ma C, Li M, et al. Intelligent devices for assessing essential tremor: a comprehensive review[J]. Journal of Neurology, 2024[2024 - 06 - 07].
[0077] [9] Overview|Devices for remote monitoring of Parkinson’s disease|Guidance |NICE[M]. NICE, 2023. https: / / www.nice.org.uk / guidance / DG51.
[0078]
[10] Hathaliya J J, Modi H, Gupta R, et al. Parkinson and essential tremor classification to identify the patient?s risk based on tremor severity[J]. Computers & Electrical Engineering, 2022, 101: 107946.
[0079]
[11] Ma C, Li D, Pan L, et al. Quantitative assessment of essential tremor based on machine learning methods using wearable device[J]. Biomedical Signal Processing and Control, 2022, 71: 103244.
[0080]
[12] Kovalenko E, Talitckii A, Anikina A, et al. Distinguishing Between Parkinson’s Disease and Essential Tremor Through Video Analytics Using Machine Learning: A Pilot Study[J]. Ieee Sensors Journal, 2021, 21(10): 11916 - 11925.
[0081]
[13] Hayashida T, Sugiyama T, Sakai K, et al. Validation and Discussion of Severity Evaluation and Disease Classification Using Tremor Video[J]. ELECTRONICS, 2023, 12(7): 1674.
[0082]
[14] Park K W, Mirian M S, McKeown M J. Artificial intelligence-based video monitoring of movement disorders in the elderly: a review on current and future landscapes[J]. SINGAPORE MEDICAL JOURNAL, 2024, 65(3): 141-149.
[0083]
[15] Bhatia K P, Bain P, Bajaj N, et al. Consensus Statement on the classification of tremors. from the task force on tremor of the International Parkinson and Movement Disorder Society: IPMDS Task Force on Tremor Consensus Statement[J]. Movement Disorders, 2018, 33(1): 75-87.
[0084]
[16] Goetz C G, Tilley B C, Shaftman S R, et al. Movement Disorder Society-sponsored revision of the Unified Parkinson’s Disease Rating Scale (MDS-UPDRS): Scale presentation and clinimetric testing results[J]. Movement Disorders, 2008, 23(15): 2129-2170.
[0085]
[17] Fahn S, Tolosa E, Marin C. Clinical Rating Scale for Tremor[J]. Journal of Neurology, Neurosurgery & Psychiatry, 1988[2023-06-07].
[0086]
[18] PaddlePaddle Authors.PaddleDetection:Object Detection toolkit based on PaddlePaddle[EB / OL].[2024-12-29].https: / / github.com / PaddlePaddle / PaddleDetection.
[0087]
[19] Fang H S,Li J,Tang H,et al.AlphaPose:Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-Time[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2023,45(6):7157-7173.
[0088]
[20] MMPose Contributors.Openmmlab pose estimation toolbox and benchmark[EB / OL].(2020)[2024-12-29].https: / / github.com / openmmlab / mmpose.
[0089]
[21] Redmon J,Farhadi A.YOLOv3:An Incremental Improvement[J].arXiv e-prints,2018[2024-12-29].
[0090]
[22] Ren S,He K,Girshick R,et al.Faster R-CNN:Towards Real-Time Object Detection with Region Proposal Networks[J].IEEE Transactions on Pattern Analysis & Machine Intelligence,2017,39(6):1137-1149.
[0091]
[23] Liu Z,Ning J,Cao Y,et al.Video Swin Transformer[J].2021[2024-12-29].
[0092]
[24] Lin T Y,Maire M,Belongie S,et al.Microsoft COCO:Common Objects inContext[C] / / European Conference on Computer Vision.2014.
[0093]
[25] Jiang T,Lu P,Zhang L,et al.RTMPose:Real-Time Multi-Person PoseEstimation based on MMPose[M].arXiv,2023[2024-12-28].DOI:10.48550 / arXiv.2303.07399.
[0094]
[26] Nie Y,Nguyen N H,Sinthong P,et al.A Time Series is Worth 64Words:Long-term Forecasting with Transformers[M].arXiv,2023[2025-01-01].DOI:10.48550 / arXiv.2211.14730。
Claims
1. A video-based discrimination model for Parkinson's disease and essential tremor, comprising: Pose Estimation module, data processing module, classification module; The pose estimation module includes a pre-trained RTMPose-L model; The pose estimation module performs full-body pose estimation frame by frame on the three upper limb movement videos of the subject, extracts the wrist and five finger key point coordinates of the target hand, and forms a multi-group absolute coordinate sequence of hand key points; The data processing module converts the absolute coordinate sequence output by the pose estimation module into a relative coordinate sequence, and calculates the statistical feature sequences of speed, acceleration, amplitude, frequency, and entropy based on the relative coordinate sequence; The classification module includes a PatchTST time series classification model; the pose estimation module divides the relative coordinate sequence and the statistical feature sequence, which are used as its input, into multiple patches of a predetermined length, then extracts features from each patch through the encoder of the Transformer and outputs patch predictions, and finally flattens the output patch predictions into a one-dimensional vector, and classifies Parkinson's disease and essential tremor through a fully connected layer.
2. The discrimination model according to claim 1, wherein: The three upper limb movements include finger-to-nose, palm-turning, and double-finger palm flexion and extension movements.
3. The discrimination model according to claim 1, wherein: The multi-group absolute coordinate sequence of hand key points is eight groups of hand key point coordinate sequences, and the dimension of each group of hand key point coordinate sequences is 6, respectively representing the coordinates of the unilateral wrist and the five finger tips.
4. The discrimination model according to claim 1, wherein: For each hand key point, when the absolute coordinate sequence is converted into a relative coordinate sequence, starting from the second frame, the absolute coordinates of the hand key point in each frame image are respectively subtracted from the absolute coordinates of the hand key point in the first frame image to obtain the relative coordinates of the hand key point corresponding to this frame image, and the relative coordinates represent displacement; the relative coordinates of the hand key point corresponding to each frame image except the first frame form a relative coordinate sequence.
5. The discrimination model according to claim 1, wherein: There is no background restriction when recording the three upper limb movement videos of the subject.
6. A video-based discrimination device for Parkinson's disease and essential tremor, comprising: Pose Estimation unit, data processing unit, classification unit; The pose estimation unit is a computing unit running the RTMPose-L model; The input of the pose estimation unit is the three upper limb movement videos of the subject, and the output is a multi-group absolute coordinate sequence of hand key points; the pose estimation unit performs full-body pose estimation on each frame image in the three upper limb movement videos of the subject, extracts the wrist and finger key point coordinates of the target hand, and forms a multi-group absolute coordinate sequence of hand key points; The data processing unit is used to convert the absolute coordinate sequence output by the pose estimation unit into a relative coordinate sequence, and use the relative coordinate sequence to calculate the statistical feature sequences of speed, acceleration, amplitude, frequency, and entropy; The classification unit is a computing unit running the PatchTST model; the input of the classification unit is the relative coordinate sequence and the statistical feature sequence output by the data processing unit, and the output is the classification result of Parkinson's disease and essential tremor; The classification unit divides the input relative coordinate sequence and statistical feature sequence into multiple patches of a predetermined length, then extracts features from each patch through the encoder of the Transformer and outputs patch predictions. Finally, the output patch predictions are flattened into a one-dimensional vector, and classification between Parkinson's disease and essential tremor is performed through a fully connected layer.
7. The discrimination device according to claim 6, wherein: The three upper limb movements are finger-to-nose, palm-turning, and finger-palm flexion and extension movements of both hands.
8. The discrimination device according to claim 6, wherein: The absolute coordinate sequences of the multiple groups of hand key points are eight groups of hand key point coordinate sequences, and the dimension of each group of hand key point coordinate sequences is 6, respectively representing the coordinates of the wrist and the ends of five fingers.
9. The discrimination device according to claim 6, wherein: For each hand key point, when the coordinate conversion and sequence calculation unit converts the absolute coordinate sequence into a relative coordinate sequence, starting from the second frame, the absolute coordinates of the hand key points in each frame of the image are respectively subtracted from the absolute coordinates of the hand key point in the first frame of the image to obtain the relative coordinates of the hand key point corresponding to this frame of the image, and the relative coordinates represent displacement; The relative coordinates of the hand key point corresponding to each frame of the image except the first frame constitute a relative coordinate sequence.