Tremor symptom multitask evaluation model for contactless video-based fine full-body pose estimation and apparatus thereof
By employing contactless video assessment technology, HRTNet-Dark and Transformer modules are used to extract key points throughout the body and learn tremor characteristics. This solves the problems of large measurement errors and high costs in existing technologies, and achieves highly accurate tremor level assessment and multi-task assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GENERAL HOSPITAL OF PLA
- Filing Date
- 2023-03-27
- Publication Date
- 2026-04-28
AI Technical Summary
Existing tremor symptom assessment techniques mostly rely on contact sensors or visual markers, resulting in large measurement errors, high costs, and an inability to effectively assess whole-body posture. Furthermore, existing video-based methods cannot accurately capture fine-grained information about tremors or perform multi-task assessments.
A non-contact, video-based fine-grained full-body pose estimation model is adopted. Video is captured by a camera, and the HRTNet-Dark module is used to extract key point coordinates and eliminate jitter. The Transformer module is combined to learn motion features, and the TDT module is used to estimate the tremor level, thereby realizing full-body pose estimation and tremor detection.
It achieves highly accurate tremor level assessment, reduces quantification errors and shaking problems, and provides a non-contact, low-cost whole-body posture assessment tool suitable for multi-task tremor assessment.
Smart Images

Figure CN116129530B_ABST
Abstract
Description
Technical Field
[0001] This application relates to tremor symptom assessment technology, and more particularly to a non-contact, video-based, fine-grained whole-body posture estimation multi-task assessment model and apparatus for tremor symptoms. Background Technology
[0002] Essential tremor (ET) is one of the most common movement disorders, reportedly affecting more than 60 million people worldwide. A prospective study showed that patients with ET have a more than fourfold increased risk of developing Parkinson's disease (PD). The main manifestation of ET is intention tremor, and clinical history taking and scale testing are important references for diagnosing disease severity. However, in the clinic examination room, ET can be confused with PD and dystonia tremor. Clinicians assess tremor symptoms primarily based on visual features, such as the Clinical Tremor Scoring Scale (CRST), which is the gold standard for clinical scoring; tremor severity increases with amplitude and frequency. Therefore, clinicians' scoring of symptoms is subjective, especially for grades 2 and 3. For example, one report indicated that in the diagnosis of 48 PD patients, the accuracy rate of scoring by six physicians was only 58.9-78.6%.
[0003] To provide objective tremor scores, previous techniques and methods have largely employed wearable sensors to record motion information at specific sites. These techniques and methods offer a non-invasive, real-time, cost-effective, and objective solution to support standard clinical assessments by neurologists. However, these sensors inevitably require physical contact with the patient, and the weight of the sensors and devices, due to the need to strap them on, inevitably hinders the user's tremor activity, potentially leading to significant measurement errors. Furthermore, these inertial sensors require regular calibration to avoid accumulating further errors, which also increases the difficulty of using these techniques or methods.
[0004] Furthermore, some vision-based technologies or methods require unique markers as keypoints for calibration to assist posture estimation models in accurately detecting key points on the human body. These markers, which come into contact with the limbs, can also affect the fine tremor movements of patients and increase the cost of clinical examinations. Recently proposed methods use depth cameras to capture hand movements in PD patients and extract kinematic features, which are then input into machine learning models for classification. While these abstract, high-dimensional features are helpful for model decision-making, some information is inevitably lost. This is because multi-sensor signals are compressed into single-valued features, and these dense, multi-dimensional signals are often crucial details for capturing the severity of fine-grained tremors. In addition, devices such as depth cameras and visual markers are expensive and not readily available. Moreover, most existing technologies or methods focus only on hand posture analysis; in fact, tremor information from other parts of the body can provide rich diagnostic reference information for clinicians in clinical scale task assessments. In summary, these shortcomings limit the application scenarios and accuracy of tremor symptom assessment.
[0005] While existing techniques and methods for tremor assessment using remote cameras have improved the accuracy and generalization of predictions, they are often limited by experimental paradigms, dataset quality, and the fitting ability of the prediction algorithms themselves. Furthermore, there are currently no quantitative studies assessing the severity of ET tremor using multiple laboratory testing tasks (e.g., those supporting simultaneous analysis of resting tremor and postural tremor tasks). In fact, managing movement disorders such as ET requires timely monitoring of disease progression, and non-contact, low-cost, and high-precision visual assessment methods show great promise. Summary of the Invention
[0006] In view of the above problems, this application aims to propose a non-contact, video-based, fine-grained whole-body posture estimation multi-task assessment model and device for tremor symptoms, which obtains the subject's tremor level by capturing video of a subject performing a specific action in any one of the multi-tasks using a camera.
[0007] The contactless, video-based, fine-grained whole-body posture estimation multi-task assessment model for tremor symptoms in this application includes: a posture key point estimation module and a tremor level estimation module.
[0008] The pose key point estimation module is used for pose estimation of the subject, extracting a sequence of key point coordinates reflecting the subject's tremor from video frames of the subject performing any of the multi-tasks;
[0009] The motion feature sequence is obtained through the keypoint coordinate sequence; the motion feature sequence is the sequence of amplitude, velocity, acceleration, frequency, or entropy of key points obtained through the keypoint coordinate sequence.
[0010] The tremor level estimation module estimates the subject's tremor level based on a sequence of motion characteristics.
[0011] Preferably, the pose keypoint estimation module eliminates image jitter between consecutive video frames.
[0012] Preferably, the pose keypoint estimation module is the HRTNet-Dark module;
[0013] Continuous video frames are processed by the Faster R-CNN human detector to obtain human boundary maps;
[0014] The human body boundary map is input into the HRTNet-Dark module to obtain a sequence of key point coordinates for the entire body of the subject, which is then precisely located.
[0015] Preferably, the HRTNet-Dark module includes: a first DARK module, an HRNet model, a Transformer module, and a second DARK module;
[0016] The first DARK module is used to encode the human body boundary map to obtain a reduced-resolution human body boundary map.
[0017] The HRNet model is used to obtain heatmaps of key points from a human body boundary map with self-reduced resolution.
[0018] The Transformer module is used to eliminate jitter in the heatmap of key points and send the processing results to the second DARK module.
[0019] The second DARK module corresponds to the first DARK module and is used for decoding to reduce quantization errors during encoding and decoding, thereby obtaining a stable sequence of keypoint coordinates.
[0020] Preferably, the tremor level estimation module is a TDT module;
[0021] The motion feature sequence is cut into equal-length segments by a sliding window and then input into the TDT module, which estimates the subject's tremor level.
[0022] The present application discloses a contactless, video-based, fine-grained whole-body posture estimation multi-task assessment device for tremor symptoms, which includes a computing unit for running a posture key point estimation module and a tremor level estimation module.
[0023] The pose key point estimation module is used for pose estimation of the subject, extracting a sequence of key point coordinates reflecting the subject's tremor from video frames of the subject performing any of the multi-tasks;
[0024] The motion feature sequence is obtained through the keypoint coordinate sequence; the motion feature sequence is the sequence of amplitude, velocity, acceleration, frequency, or entropy of key points obtained through the keypoint coordinate sequence.
[0025] The tremor level estimation module calculates the subject's tremor level based on the motion characteristic sequence.
[0026] This application presents a contactless, video-based, fine-grained whole-body pose estimation model and apparatus for multi-task assessment of tremor symptoms, effectively fusing complementary spatial-temporal information from multiple body parts. The model refines and quantifies tremor severity for various tasks through whole-body pose estimation and sequence learning. The fine-grained whole-body pose estimation model (HRTNet-Dark) in this application reduces quantization errors during encoding and decoding, avoiding jitter issues arising from continuous video frame prediction. The Transformer-based tremor detection algorithm (TDT) in this application learns tremor patterns from kinematic feature sequences, thereby achieving high-accuracy tremor level classification performance. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the pipeline for the contactless, video-based, fine-grained whole-body posture estimation multi-task assessment model for tremor symptoms proposed in this application.
[0028] Figure 2 The HRTNet-Dark model is used for the contactless, video-based, fine-grained whole-body posture estimation multi-task assessment model for tremor symptoms proposed in this application.
[0029] Figure 3 This is a schematic diagram of a top-down approach for full-body pose estimation.
[0030] Figure 4 A schematic diagram of HRTNet-Dark for the contactless, video-based, fine-grained whole-body posture estimation multi-task assessment model for tremor symptoms proposed in this application.
[0031] Figure 5 This is a schematic diagram of the Transformer (TDT) pipeline for tremor detection, which is a non-contact, video-based, fine-grained whole-body posture estimation tremor symptom multi-task assessment model proposed in this application.
[0032] Figure 6 This is the confusion matrix of the TDT model proposed in this application in resting tremor and posture tremor tasks;
[0033] Figure 7 The ROC curves of the TDT model proposed in this application are shown in resting tremor and postural tremor tasks.
[0034] Figure 8The image shows the SHAP summary obtained by the optimal XGBoost algorithm for predicting resting tremor and postural tremor, respectively. Detailed Implementation
[0035] The present application will now be described in detail with reference to the accompanying drawings.
[0036] (1) Video data collection. The entire laboratory examination was supervised by a neurologist with expertise in movement disorders, and video data (CMOS camera, 48MP, 1920*1080 HD, 60 frames / second) was recorded to support independent scoring by three neurologists (mutually blinded). During the laboratory examination, the patient completed resting tremor and postural tremor tasks as required.
[0037] Two experts jointly score the patients' action tremors. In cases of discrepancies, a third expert reviews the video recordings to make the final determination. The specific scoring process is as follows: 1) Two neurologists score the severity of the patient's action tremor by watching video footage and observing the remaining handwriting, obtaining two mutually blinded rating sheets; 2) A data analysis engineer tallies the consistent scores from the two experts. For actions with inconsistent scores, the video footage and remaining handwriting are used by another experienced neurologist to make the final decision until a reliable score is obtained. This experimental design also avoids training errors caused by systematic bias, making the deep learning model more reliable.
[0038] The data acquisition unit uses only cameras. A database of video examinations can be built based on the data acquisition unit, including video recordings and patient information from doctor consultations and examinations (including demographic information, past medical history, and scale tests).
[0039] The patient performed two tasks according to the CRST: 1) a resting tremor task, with hands resting on the armrests of a chair; and 2) a postural task, with arms extended in front of the chest and wrists slightly extended. A total of 122 videos of these two tasks were collected during the experiment.
[0040] The collected videos were submitted to a committee of neurologists for evaluation. Subjective visual diagnoses are often influenced by factors such as the physician's clinical experience level and fatigue. Therefore, to ensure the objectivity of the scoring, scores from two independent physicians were compared. For discrepancies, a third experienced neurologist was consulted for arbitration to obtain a final consensus score. Research has shown that the five severity levels defined in the CRST increase the difficulty of the committee's scoring. For example, research by Guo et al. indicated that the accuracy rate of six physicians was only 58.9-78.6%. To avoid this, the scoring in this application was modified to three levels: normal (corresponding to 0 on the original scale), mild (corresponding to 1-3 on the original scale), and severe tremor (corresponding to 4 on the original scale).
[0041] (2) The overall structure of this application. Figure 1 This is a schematic diagram of the pipeline for the contactless, video-based, fine-grained whole-body posture estimation multi-task assessment model for tremor symptoms proposed in this application.
[0042] First, we define the problem description of this application. The input to this application is a collected video frame sequence V = {f1, f2, ..., f...} M The goal of this invention is to estimate the precise keypoints of the entire body in each video frame. We define the CRST task pose as P = {P1, P2, ..., P...} T},in It is a key point in the τth video frame. It is the two-dimensional position of the key point k.
[0043] The framework proposed by the method of this invention accepts input. and output Here it is given by formula (1):
[0044]
[0045] in, θ is the classifier function, and θ is the hyperparameter of the neural network. Furthermore, and y i ∈{0,1,2} is a pose sequence with N keypoint features and the i-th frame of T frames. th The severity label for each video.
[0046] This automated pipeline consists of two stages. The first stage is designed as a refined full-body pose estimation algorithm, extracting key points from video frames. To filter out anomalous jitter in the pose estimator between video frames (…),… Figure 2(b) This invention proposes a Transformer-based HRNet-DARK algorithm. These extracted keypoints describe the changes of limb joints over time and contain rich spatiotemporal information. Therefore, in the second stage, this invention uses the coordinate sequence of these keypoints to train a Transformer (TDT) classification model for tremor detection to quantify the severity of a patient's tremor.
[0047] (3) Refined Full-Body Pose Estimation Algorithm. Compared to bottom-up methods that directly extract key points from the human body, top-down methods are better able to handle scale differences. For example... Figure 3 As shown, the top-down approach has the advantage of magnifying human body details by cropping and adjusting the bounding boxes for each individual. Especially when assessing tremors, it is necessary to pay attention to detailed features, such as the tremors in the finger areas. Therefore, the method proposed in this invention offers a top-down approach with higher accuracy. On the other hand, prior to the release of the COCO whole-body dataset, most algorithms, such as OpenPose, attempted to aggregate five separately trained pose estimation subnetworks, thus complicating training and inference.
[0048] On the other hand, although ZoomNet was the first method to aggregate all subnetworks for full-body pose estimation, allowing for end-to-end training and effectively handling scale differences across different body parts, the most challenging aspect of vision-based tremor symptom diagnosis is the problem of keypoint jitter. For example... Figure 2 As shown in (b), keypoint detection algorithms exhibit significant jitter in continuous frame prediction due to rare movements or occlusion of movements. Current technologies and methods generally offer solutions in two main areas: low-pass filters and learning-based models. Low-pass filters, such as exponential moving average filters and Kalman filters, often face difficult trade-offs because the jitter in pose estimation models is non-uniform. Excessively high filter strength can delay output results and even over-smooth the details of limb tremors. On the other hand, learning-based methods generally utilize spatial-temporal models to optimize prediction accuracy and temporal stability for each frame, such as the Skeleton and ST-A2J models. However, due to the long-term jitter and tremor in continuous frames of ET patients, the extracted spatial / temporal features are inherently unreliable, leading to a bottleneck in spatial-temporal optimization.
[0049] This invention proposes a novel, refined method for full-body pose estimation. For example... Figure 3 and Figure 4As shown, we obtain stable keypoint predictions by dynamically capturing time-varying global features through the insertion of Transformer-based encoder blocks into HRNet. This invention names this HRTNet. Furthermore, we combine HRTNet with a distributed-aware coordinate representation (DARK) of keypoints to reduce quantization errors during encoding and decoding. Specifically, given a set of video frames, we use an existing Faster R-CNN human detector to obtain human boundaries (…). Figure 3 Then, we use HRTNet-Dark to precisely locate key points throughout the body. Figure 4 The following section will detail the specific definition of the refined whole-body pose estimation algorithm.
[0050] First, we consider improving coordinate decoding in heatmaps. In standard coordinate decoding methods, the location of key points is predicted as:
[0051]
[0052] Where ||ps|| defines the Euclidean norm of the difference between the primary peak p and the secondary peak s in heatmap t. α represents the rate of resolution reduction. To obtain the accurate location of sub-pixels, DARK assumes that the predicted heatmap follows a two-dimensional Gaussian distribution, and the logarithm of the predicted heatmap is expressed as:
[0053]
[0054] Where x is the pixel position, Σ represents the diagonal matrix, and our goal is to estimate the Gaussian mean μ corresponding to the keypoint positions. Here, a quadratic term of the Taylor series is used to approximate the activation.
[0055]
[0056] in Represents the Hessian matrix evaluated at point p. Therefore, we can obtain μ by calculating the peak derivative of the heatmap:
[0057]
[0058] Furthermore, DARK is used to generate precise coordinate codes of a Gaussian distribution by quantizing pre-quantized coordinates. Figure 4 This can generate higher resolution heatmaps during the learning process.
[0059] On the other hand, motion blur or sudden changes in the original video may cause frequent jitter of the detected keypoints. Figure 2 (b)). Meanwhile, traditional smoothing algorithms filter out the details of tremors in ET patients. For example... Figure 4 As shown, our proposed Transformer-inspired HRTNet-Dark addresses this issue by refining the jitter of keypoints based on their spatial-temporal context. Considering that traditional sinusoidal embeddings lose the translational equivalence property of convolutional layers, we add relative height and width information to the standard multi-head self-attention (MHSA) module, using two-dimensional relative position encoding. Specifically, the method of this invention applies pixel pairs of keypoints a = (ax, a... y ) and b = (bx, b y Encode the pairs of attention pairs before Softmax:
[0060]
[0061] Where q a and k b These represent the query and key-value vectors for a and b, respectively. and It is a learnable embedding with relative width and height. Therefore, fine-grained self-attention can be represented as:
[0062]
[0063] in It is a relative logarithmic matrix along the width and height dimensions. Represents raw data This is a low-dimensional embedding form. This downsampling transformation avoids the inefficient and redundant nature of pixel pairs in the original Transformer. Instead of simply integrating multi-head sub-attention modules onto the feature maps of the CNN backbone, we apply Transformer modules to each layer of the pose keypoint detection network to collect long-term dependencies from multiple scales.
[0064] (4) Extraction and selection of kinematic features.
[0065] Based on the above detailed full-body pose estimation stage, we obtain the pose sequence of key points for performing the CRST task. For the video frame sequence V = {f1, f2, ..., f...} T Our goal is to accurately estimate the whole-body keypoints in each image. We define a CRST task pose sequence of T video frames. in This represents the keypoints in the τ-th video frame; each keypoint contains Cartesian coordinates in the x and y directions. Furthermore, K represents the number of keypoints per motion, and C represents the extracted motion feature vector. By default, there are 133 keypoints throughout the body, and the original motion features are the coordinates of the keypoints, therefore C = 2. We can further define kinematic features, such as amplitude, velocity, acceleration, frequency, and entropy. Specifically, amplitude can be derived as:
[0066]
[0067] in, and These represent the x and y coordinates of the k-th keypoint corresponding to the video frame, respectively. Similarly, velocity and acceleration can also be calculated using difference operations:
[0068]
[0069]
[0070] All features are detailed in Table 1, including the name, abbreviation, and dimension of each kinematic feature. We use a feature selection method based on Recursive Feature Elimination (RFE) to eliminate redundant features and improve the classification performance of tremor detection. The base learner is XGBoost, and the sampled SHapley Additive Explanation (SHAP) method is used to quantify the importance of features for tremor quantification. The SHAP value is calculated as follows:
[0071]
[0072] in, Indicates the included feature set The predicted BP value, Indicates the feature set that is not included. The predicted BP value.
[0073] Table 1
[0074] Motion features extracted in this study
[0075]
[0076]
[0077] Note: This excludes the six key points on both feet.
[0078] (5) Transformer-based tremor detection algorithm (TDT).
[0079] For a set of sequence-label pairs extracted from M videos in and y i ∈{0,1,2} is the i-th th We will design a sequence learning model based on the pose sequence and tremor severity labels of T frames in a video. To detect the severity of tremor. Specifically, this invention proposes a Transformer (TDT) model for tremor detection to quantify these multivariate feature sequences. Figure 5 This is a schematic diagram of TDT. The core of TDT lies in the Transformer encoder. For simplicity, we will use each training sample... Represented as a multivariate time series of T frames and K×C different pose features. Each feature vector It is normalized and then projected into a d-dimensional vector space through a one-dimensional convolutional layer with a kernel size of (K×C):
[0080]
[0081] Then, we use the learnable positional encoding E pos For the input pose vector Encode:
[0082]
[0083] We did not use the original deterministic and sinusoidal coding methods here because we found that the positional embeddings of adjacent pose motion features had high similarity. Finally, we used batch normalization (BN) instead of layer normalization. There are two main reasons for this: 1) Unlike word vectors, pose motion sequences may contain numerically anomalous samples; 2) Pose motion sequences have the same length, which avoids the problem of unequal signal lengths caused by BN. For the tremor quantization task, we chose cross-entropy loss as the objective function for learning the model parameters.
[0084]
[0085] Where σ(W) o ·z τ +b o () represents the probability of predicting the value as j. In the final vector representation... A linear layer with parameters is used on top of that.
[0086] (6) Evaluation Metrics of the Invention. We used a paradigm-segmented dataset among patients to ensure that the model did not encounter data from the same individuals during training and testing. To compare the stability of model performance, all experimental results were presented using five-fold cross-validation. We used mean precision (AP) and mean recall (AR) to evaluate pose estimation performance. To quantify jitter error between video frames, we chose to use distance error (DE), velocity error (VE), acceleration error (AE), and keypoint pixel precision (KPA), where DE was calculated using Euclidean distance. For the assessment of tremor severity classification, we used accuracy (ACC), precision (PRE), recall (REC), specificity (SPE), F1 score, receiver operating characteristic (ROC) curve, and area under the ROC curve (AUC) scores.
[0087] Specifically, the dataset collected in this invention is described as follows:
[0088] To improve the generalization ability of the proposed HRTNet-Dark, we pre-trained it on the publicly available COCO-WholeBody dataset and then fine-tuned it on our collected video ET dataset. COCO-WholeBody contains 200,000 full-body images annotated in the wild, each with 133 keypoints. Compared to separately annotated datasets, COCO-WholeBody is more conducive to training full-body pose estimation models, mainly because: 1) Wild images enhance the model's generalization ability to natural scenes. 2) Previous datasets only contained specific parts of the body; training separately on these datasets may introduce biases caused by variations in lighting, scale, etc. Figure 5 As shown, for the tremor detection dataset, we used a sliding window with a window length of 3s and a step size of 2s to increase the data sample size. The sample distribution of the datasets in the two tremor tasks is shown in Table 2.
[0089] Table 2
[0090] Sample distribution of the dataset
[0091]
[0092] Specifically, the deployment of the model in this invention is as follows:
[0093] HRTNet-Dark and TDT were implemented in PyTorch using the Adam optimizer on an NVIDIA RTX 3090. For training, the initial learning rate was 1e-4, decaying exponentially at a rate of 0.95. We used a pre-trained model provided by mmPose as the initial weights for the baseline method.
[0094] For ease of comparison, the method of this invention employs state-of-the-art technologies or methods for each task and is evaluated using the same experimental setup:
[0095] ① For full-body pose estimation, we selected the first four models provided by mmPose as baseline methods:
[0096] 1) ResNet-152. Composed of residual blocks, it achieved state-of-the-art (SOTA) results on the COCO2017 challenge dataset in 2018.
[0097] 2) HRNet. It maintains high-resolution representations throughout the learning process and achieved state-of-the-art (SOTA) results in 2019.
[0098] 3) ViPNAS. It simultaneously searches for single-frame network structure and inter-frame temporal connections, achieving a better trade-off between accuracy and efficiency.
[0099] 4) TCFormer. It can dynamically adjust the size and shape of visual tokens through progressive clustering, better preserving image details. It achieved state-of-the-art results on COCO-WholeBody in 2022.
[0100] ② For tremor detection, we chose the state-of-the-art sequence learning model as the baseline method.
[0101] 1) GRU-FCN. Its classification performance outperforms state-of-the-art (SOTA) on many time-series datasets in 2018.
[0102] 2) TCN. It consists of causal convolutions, dilated convolutions, and residual connection layers. It performs better than traditional recurrent neural networks.
[0103] 3) Inceptiontime. Inspired by the Inception-v4 architecture. It has low computational complexity and is scalable.
[0104] 4) GatedTabTransformer. This is the latest Transformer-based model, which incorporates a gating unit into the TabTransformer and performs best in binary classification tasks.
[0105] (7) To verify the effectiveness of the proposed model, we conducted extensive experiments: 1) Proof of the effectiveness of attitude estimation; 2) Proof of the effectiveness of tremor detection; 3) The utility of key point features in tremor quantification; 4) The utility of kinematic features in tremor quantification; 5) Ablation studies (to verify the utility of the attention mechanism); and 6) Comparison with state-of-the-art methods.
[0106] The specific experimental results for verification are as follows:
[0107] (1) Proof of the effectiveness of attitude estimation
[0108] We designed experiments to verify the performance of the proposed HRTNet-Dark model and compared it with some state-of-the-art pose estimation methods. The experimental results are shown in Table 3. Compared with methods used for single-frame pose detection, Table 3 shows the effectiveness of HRTNet in whole-body pose estimation, with the highest performance in AP and AR for hands, face, limbs, and the whole body. HRTNet inserts a Transformer-based encoder block into HRNet, which is an attention mechanism optimized for video frames. This novel module helps the network capture globally dependent features, thereby improving the PRE and REC metrics for keypoint recognition. Based on the prediction results of HRTNet, the method of this invention further embeds the DARK module to reduce the quantization error generated during encoding and decoding, i.e., the HRTNet-Dark model, which improves the AP and AR metrics for hands, face, and the whole body. However, HRTNet-Dark shows a slight decrease in prediction for limbs. We believe that the detection of limbs is relatively complex in the whole-body dataset collected in the method of this invention, possibly due to the similarity between the gowns of hospitalized patients and the background of the examination room. In addition, neurologists pay more attention to the degree of hand tremor in clinical visual assessment. Therefore, the method of this invention still recommends choosing HRTNet-Dark as the pose estimation model.
[0109] Table 3
[0110] Performance comparison of the HRTNet-Dark model in this application with several state-of-the-art pose evaluation methods
[0111]
[0112] 2) Proof of the effectiveness of tremor detection
[0113] We compared the classification performance of the proposed TDT model with other state-of-the-art (SOTA) sequence learning classifiers on resting tremor and postural tremor tasks. The six evaluation metrics for tremor detection are shown in Tables 4 and 5. It can be seen that the proposed TDT model achieved F1 scores of 83.0% and 92.4% on the resting tremor and postural tremor tasks, respectively, which are higher than the results of other SOTA models.
[0114] Table 4
[0115] Comparison of the classification performance of the proposed TDT model and the SOTA method in resting tremor assessment
[0116]
[0117]
[0118] Table 5
[0119] A comparison of the classification performance of the proposed TDT model and the SOTA method for posture tremor assessment.
[0120]
[0121] In addition, we Figure 6 and Figure 7 The confusion matrix and ROC curve of the TDT model's classification results in resting tremor and postural tremor tasks are visualized. It can be seen that the TDT model classifies ACC severity 100% in the resting tremor task. Figure 6 (a)). Furthermore, the results of the five-fold cross-validation showed that the AUC at the normal level was 96.7% ( Figure 7 (a) Since ET patients are less likely to exhibit tremor symptoms at rest, normal samples account for 83.6% of all samples. This leads to class imbalance, with only 80.4% of PREs identifying tremor severity in the resting tremor task. Nevertheless, the proposed TDT model has the best classification performance compared to other methods.
[0122] Compared to the resting tremor task, ET patients are more likely to elicit tremor symptoms by performing specific movements in the postural tremor task. Therefore, all classifiers performed better in the postural tremor task. In particular, the TDT model achieved 100% ( Figure 6 (b) and 96.7% Figure 7 (b)) The classification ACC and AUC are at normal levels. Furthermore, Figure 7 The ROC curves show that the TDT model has better stability in the postural tremor task. Clinically, the postural tremor task can more effectively identify ET (postural tremor) compared to other types of diseases, such as PD (postural tremor).
[0123] (3) The utility of key point features in tremor quantification
[0124] The method of this invention posits that keypoints throughout the body influence the assessment of tremor symptoms. For example, as hand tremors worsen, low-frequency vibrations of the head or legs may also be involved. Therefore, it is necessary to discuss the importance of keypoints to provide more accurate classification performance. We combined keypoints from three parts and used REF to select keypoints for each part. The test results for two tasks are shown in Table 6. It can be seen that keypoint features are often equally important in different tremor tasks. Specifically, with the baseline performance of the classification model using only hand keypoint features, fusing keypoint features from other parts can effectively improve classification performance. However, using keypoint features from the whole body can lead to excessively high dimensionality of the input data for the model, resulting in a slight loss in classification performance. The results show that limb keypoints provide a greater performance improvement than facial keypoints. Compared to the model that only uses hand keypoints, the F1 score using limb keypoint features improved by 9.7% and 8.3% in the resting tremor and postural tremor tasks, respectively.
[0125] Table 6
[0126] Comparison of classification performance of TDT models proposed with different keypoint features
[0127]
[0128] (4) The utility of kinematic characteristics in tremor quantification
[0129] The discussion of keypoint features in the previous experiment reveals the need to consider the trade-off between the number of features and classification performance. For example... Figure 8 As shown, we use a SHAP plot to illustrate the influence of the top 15 important kinematic features on the predicted tremor severity. SHAP values are calculated in the REF XGBoost model. The calculated SHAP values are arranged in descending order of F1 score, with colors representing feature values. A wider vertical axis indicates a larger cluster of samples, while a larger SHAP value on the horizontal axis indicates a positive effect on the estimated tremor level. It can be seen that tremor velocity in the limbs and hands are the two most influential features in the resting tremor and postural tremor tasks, respectively. Furthermore, keypoint features of the hand are the most important in both tasks, accounting for 46.6%. Describing tremor velocity and amplitude is also effective; the selected keypoints are generally the easiest parts to identify.
[0130] (5) Ablation studies (to verify the effectiveness of the attention mechanism)
[0131] As mentioned earlier, jitter errors occurring at keypoints during video frame prediction can interfere with the performance of tremor assessment. To quantify the smoothing effect of the proposed refined whole-body pose estimation method, we designed an ablation study comparing it with three commonly used filters. As shown in Table 7, HRNet-Dark achieves the best performance with our proposed Transformer-based encoder block, reducing DE by 21.19% and AE by 48.15% compared to the original pose estimation results. Saritzky-Golay shows the best performance among the three filters, but it still lags behind HRTNet-Dark.
[0132] Table 7
[0133] Performance comparison of Transformer-based encoders and commonly used filters
[0134]
[0135] Note that w / indicates adding the module to the model.
[0136] (6) Comparison with existing advanced technologies and methods
[0137] To our knowledge, there are currently no relevant technologies or methods focused on video-based whole-body posture assessment to quantify tremor severity in ET patients. However, several recent studies have explored assessing PD symptoms through specific tasks. Considering the similarity of research methods, we summarize these state-of-the-art (SOTA) works in Table 8. It can be seen that most works do not consider issues such as keypoint jitter, false detections, and omissions that occur in video frame prediction. They directly call existing posture estimation models, which may be a potential reason for low classification accuracy. For example, Williams et al. directly used the computer vision software DeepLabCut. They only labeled keypoints for 3% of the video frames, and the remaining video frames were predicted using the ResNet-50 model with the coordinates of the keypoints. This approach easily leads to a large number of feature extraction errors, so the tremor scores predicted by the model are only 50-70% correlated with human expert ratings.
[0138] In the field of tremor assessment, this invention is the first to propose a quantitative assessment of tremor symptoms based on the Transformer model. Compared with other techniques or methods, the proposed method achieves the highest classification accuracy. We believe that sequence learning methods can effectively extract kinematic feature information of key points that change spatially over time, which is crucial for tremor quantification. Simply compressing abstract features into single numerical values, such as calculating the average velocity or amplitude of fingertip movements throughout a video, ignores the influence of intentional movements on mild tremors. Extracting these high-level features is a necessary step in traditional machine learning models, which may explain the low classification accuracy in previous works. Furthermore, the method of this invention was tested on a dataset containing 61 ET patients, a larger sample size than most studies, resulting in high confidence in the test results. In contrast, although Buongiorno et al. achieved an accuracy of 95.0% in distinguishing between PD and HC in their experiments, it was trained and evaluated on a dataset with only 30 subjects, resulting in a higher risk of bias.
[0139] Table 8
[0140] Summary of results compared with video-based state-of-the-art methods
[0141]
[0142] Note: HC represents the healthy control group, SCC represents the Spearman correlation coefficient, and ANN represents the artificial neural network.
[0143] This invention proposes a vision-based automated tremor symptom assessment system that utilizes refined whole-body posture estimation. Validation tests of the proposed method demonstrate its effectiveness, showing best classification performance compared to other similar state-of-the-art (SOTA) works. The innovation of this invention lies in automatically scoring the severity of resting tremor and postural tremor in ET patients using only a simple camera (such as a camera module integrated into a smartphone). To obtain accurate scores, the method employs a refined whole-body posture estimation model, HRTNet-Dark, to reduce quantization errors and eliminate keypoint jitter on the video frame timeline. Furthermore, this application proposes a Transformer-based sequence learning model (TDT) capable of calculating the corresponding tremor severity based on extracted kinematic features. Compared to wearable sensors and methods that only capture a single body part, this application captures rich spatial-temporal information from multiple body parts. Therefore, TDT provides clinicians with a non-contact, low-cost, convenient, accurate, and objective tool to assist in tremor diagnosis.
[0144] Compared to wearable sensors, including accelerometers and magnetometers, the most significant advantage of this application is that it can directly obtain kinematic characteristics of key points throughout the body using a remote camera, thus eliminating mechanical contact with the skin. It is well known that attached sensors affect limb movements, which can be significantly different from the characteristics of everyday tremors. No relevant research addresses this issue. Furthermore, these wearable sensors require periodic calibration to avoid accumulated errors, leading to low measurement accuracy and reduced usability.
[0145] On the other hand, some vision-based tremor assessment methods also require placing markers on the hands or other body parts to obtain high-precision detection of key points, which also risks hindering patient movement. Recently, Guo et al. proposed a method for capturing and assessing 3D hand movements in PD patients using a depth camera, achieving a tremor classification accuracy of 81.2% on a dataset. However, this requires a specialized depth camera, and the high cost may limit its application in routine health monitoring. The diagnosis and management of movement disorders, such as ET, require continuous tracking of disease progression, which often increases the burden on clinical units. The method of this invention can use a smartphone-integrated camera for real-time assessment, and this simple and inexpensive approach greatly enriches the application scenarios of telemedicine.
[0146] Finally, compared to key point feature assessment methods that only focus on specific areas, this application allows for joint analysis of multiple body parts, including the hands, limbs, and face. Tremor symptoms are composed of multiple factors and can be classified into pathological and physiological tremors, which respond to different body parts. For example, the nerves in the fingertips are farther from the brain and therefore show more pronounced changes, while arm tremors are less noticeable. Similarly, the nerves in the fingertips are farther from the brain center and therefore show more pronounced changes, while the intensity of arm tremors is easier to control. This application, by jointly analyzing the characteristics of whole-body tremors, can provide some reference for existing assessment systems.
[0147] In summary, this application presents a novel vision-based automated symptom assessment framework for monitoring tremor assessment in non-contact ET patients. Without requiring any wearable sensors or markers, this application achieves an accuracy of 95.6% on a newly created dataset containing 61 ET patients.
[0148] Unless otherwise defined, all technical and / or scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this invention relates. The materials, methods, and embodiments mentioned in this application are illustrative only and not restrictive.
[0149] Although the present invention has been described in conjunction with specific embodiments, those skilled in the art can make appropriate substitutions, modifications and changes within the inventive spirit of this application, and such substitutions, modifications and changes still fall within the protection scope of this application.
Claims
1. A non-contact, video-based, fine-grained whole-body posture estimation method for multi-task assessment of tremor symptoms, comprising: Posture key point estimation module and tremor level estimation module; The pose key point estimation module is used for pose estimation of the subject, extracting a sequence of key point coordinates reflecting the subject's tremor from video frames of the subject performing any of the multi-tasks; Motion feature sequence obtained through key point coordinate sequence; The motion feature sequence is a sequence of amplitude, velocity, acceleration, frequency, or entropy of key points obtained through the key point coordinate sequence; The tremor severity estimation module estimates the subject's tremor severity based on a sequence of motion characteristics; The pose keypoint estimation module is the HRTNet-Dark module; Continuous video frames are processed by the Faster R-CNN human detector to obtain human boundary maps; The human body boundary map is input into the HRTNet-Dark module to obtain a sequence of key point coordinates for the entire body of the subject, which is precisely located. The HRTNet-Dark module includes: the first DARK module, the HRNet model, the Transformer module, and the second DARK module; The first DARK module is used to encode the human body boundary map to obtain a reduced-resolution human body boundary map. The HRNet model is used to obtain heatmaps of key points from a human body boundary map with self-reduced resolution. The Transformer module is used to eliminate jitter in the heatmap of key points and send the processing results to the second DARK module. The second DARK module corresponds to the first DARK module and is used for decoding to reduce quantization errors during encoding and decoding, thereby obtaining a stable sequence of keypoint coordinates.
2. The non-contact, video-based, fine-grained whole-body posture estimation method for multi-task assessment of tremor symptoms according to claim 1, characterized in that: The pose keypoint estimation module eliminates image jitter between consecutive video frames.
3. The non-contact, video-based, fine-grained whole-body posture estimation method for multi-task assessment of tremor symptoms according to claim 1, characterized in that: The tremor level estimation module is the TDT module, which is a module based on the Transformer-based tremor detection algorithm; The motion feature sequence is cut into equal-length segments by a sliding window and then input into the TDT module, which estimates the subject's tremor level.
4. A non-contact, video-based, fine-grained whole-body posture estimation multi-task assessment device for tremor symptoms, comprising a computing unit for running a posture key point estimation module and a tremor level estimation module. The pose key point estimation module is used for pose estimation of the subject, extracting a sequence of key point coordinates reflecting the subject's tremor from video frames of the subject performing any of the multi-tasks; Motion feature sequence obtained through key point coordinate sequence; The motion feature sequence is a sequence of amplitude, velocity, acceleration, frequency, or entropy of key points obtained through the key point coordinate sequence; The tremor severity estimation module calculates the subject's tremor severity based on a sequence of motion characteristics; The pose keypoint estimation module is the HRTNet-Dark module; Continuous video frames are processed by the Faster R-CNN human detector to obtain human boundary maps; The human body boundary map is input into the HRTNet-Dark module to obtain a sequence of key point coordinates for the entire body of the subject, which is precisely located. The HRTNet-Dark module includes: the first DARK module, the HRNet model, the Transformer module, and the second DARK module; The first DARK module is used to encode the human body boundary map to obtain a reduced-resolution human body boundary map. The HRNet model is used to obtain heatmaps of key points from a human body boundary map with self-reduced resolution. The Transformer module is used to eliminate jitter in the heatmap of key points and send the processing results to the second DARK module. The second DARK module corresponds to the first DARK module and is used for decoding to reduce quantization errors during encoding and decoding, thereby obtaining a stable sequence of keypoint coordinates.
5. The contactless, video-based, fine-grained whole-body posture estimation multi-task assessment device for tremor symptoms according to claim 4, characterized in that: The pose keypoint estimation module eliminates image jitter between consecutive video frames.
6. The contactless, video-based, fine-grained whole-body posture estimation multi-task assessment device for tremor symptoms according to claim 4, characterized in that: The tremor level estimation module is the TDT module, which is a module based on the Transformer-based tremor detection algorithm; The motion feature sequence is cut into equal-length segments by a sliding window and then input into the TDT module, which estimates the subject's tremor level.
Citation Information
Patent Citations
Measuring body movement in movement disorder disease
US20190110736A1