Model and device for evaluating severity of tremor based on handwritten image recognition and training method of model
The severity of tremor is evaluated through a deep learning model based on handwritten image recognition, which solves the problem of time-consuming and cost-effective remote tremor assessment, and is suitable for elderly patients with mobility difficulties.
Patent Information
- Application Number
- CN202510354388.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing tremor assessment methods require patients to go to a medical institution for face-to-face assessment, which is time-consuming and costly, and subjective assessment based on doctors is prone to deviations, making it difficult to achieve remote and continuous monitoring.
The tremor severity evaluation model based on handwritten image recognition was used, and deep learning models such as ResNet50, DenseNet, ConvNeXt Tiny, MobileNet-V2 or ETSD-Net were used to evaluate the tremor severity by analyzing Archimedes spiral handwritten images drawn by the patient on paper.
It realizes objective and accurate tremor assessment, supports remote and low-cost assessment, and is suitable for elderly patients with mobility difficulties, improving the convenience and accessibility of assessment.
Smart Images

Figure CN120299639A_ABST
Abstract
Description
Technical Field
[0001] This application relates to tremor assessment technology, and particularly to a tremor severity assessment model and device based on handwritten image recognition, and a training method for the model. Background Art
[0002] Regular assessment of movement disorders helps monitor disease progression, detect disease deterioration or treatment response, and adjust treatment plans accordingly. Currently, due to the lack of specific examinations or biomarkers, it is challenging to assess ET. Usually, doctors need to conduct a face-to-face assessment based on FTM-TRS, asking patients to perform movement tasks such as finger-to-nose, drinking, or drawing to evaluate tremors. Drawing a spiral graph is a commonly used method in clinical practice for assessing tremors. Patients are given a piece of graph paper with a template of an Archimedes spiral guiding line, and two points are marked at the center and outer edge of the spiral. Patients are required to connect these two points with a pen without crossing the guiding line. Doctors evaluate the severity of tremors based on the drawn graph (as shown in Table 1). This method is widely used due to its simplicity and practicality in the clinical environment.
[0003] Table 1 Tremor severity scoring rules for the Archimedes spiral drawing task in FTM-TRS
[0004]
[0005] However, traditional assessment methods require patients to go to a medical institution, make an appointment with a neurologist, and undergo a face-to-face assessment. Many PD and ET patients are elderly people with limited mobility, which makes this process cumbersome and time-consuming for both patients and clinicians. In addition, doctor-based subjective assessments are prone to significant biases, and data is usually difficult to retain, thus unable to continuously monitor patients' conditions. Remote intelligent assessment provides a promising solution for patients by reducing costs and improving convenience and accessibility, enabling patients in remote areas or with limited mobility to receive professional assessments without geographical restrictions.
[0006] In recent decades, the technology for assessing tremors has developed rapidly. Devices such as IMU, EMG, video devices, and electronic tablets have significantly enhanced the objectivity, quantification, and consistency of tremor detection. The application of machine learning and deep learning algorithms in tremor assessment has received increasing attention. Ali et al. (2022) recorded the accelerometer signals of 35 subjects while drawing an Archimedes spiral graph to distinguish ET patients from healthy controls. Sole-Casals et al. (2019) used the information from an electronic tablet as the input for an SVM model to distinguish PD patients from ET patients. Wang et al. (2023) combined a convolutional neural network with an electronic tablet for the diagnosis of ET.
[0007] The inventors conducted a search in the Web of Science database using the search query "Title=essential tremor AND (Topic=drawing OR Topic=writing)". As of January 9, 2025, a total of 151 results were retrieved. After reviewing the titles and abstracts, the inventors excluded 138 studies that focused on genetics, epidemiology, surgery, or drug treatment. Finally, 13 relevant articles were selected for a full-text review, and 9 articles that focused on the assessment of ET severity and potentially included handwriting datasets were analyzed and summarized, as shown in Table 2. Ali et al. (2024) recruited 17 ET patients and 18 healthy controls. The participants performed the Archimedes spiral drawing task while wearing an IMU on their forearm. The SVM model achieved only 68.57% accuracy in estimating the severity of ET. Ma et al. (2023) collected multimodal data from 147 ET patients using a digital writing tablet and pen. By leveraging transfer learning and attention mechanisms, the system achieved an accuracy of 97% in tremor severity classification. Despite the good accuracy, data collection required professional equipment and was supervised by researchers in a laboratory environment, so it was not suitable for remote home assessment.
[0008] Table 2 Overview of handwriting studies for assessing ET
[0009]
[0010] After summarizing these existing technologies, the inventors found that existing studies mainly relied on IMUs or electronic writing tablets. Currently, there is no study that solely focuses on using handwriting images to assess the severity of ET, nor has a dedicated handwriting image dataset been established for ET patients. In addition, existing studies mainly rely on data collected in a laboratory environment and do not address methods for remote assessment. The spiral drawings made with an electronic writing tablet lack template lines, while the Fahn-Tolosa-Marin (FTM-TRS) diagnostic criteria require observing the number of crossings between the handwriting and the template lines, causing difficulties in assessment. In addition, compared with an electronic writing tablet, drawing with a pen and paper is more in line with natural writing habits, making diagnosis and assessment closer to real-life conditions. This method also avoids the costs associated with purchasing and maintaining an electronic writing tablet or IMU device, making the assessment more affordable and convenient and having the potential for remote assessment. Summary of the Invention
[0011] In view of the above problems, the present application aims to propose a model, device, and its training method for assessing the severity of ET based on handwriting images.
[0012] The tremor severity assessment model based on handwritten image recognition of the present application, which is trained to evaluate the ET tremor severity of a subject; the model includes an initial layer and subsequent layers; the initial layer is implemented by one of ResNet50, DenseNet, ConvNeXt Tiny, MobileNet-V2, or ETSD-Net; the subsequent layers are implemented by fully connected layers; the input of the tremor severity assessment model is the handwritten image of the Archimedean spiral drawn by the subject on paper; this image is sent to the initial layer for calculation, and the output of the initial layer is used as the input of the subsequent layer; the output of the subsequent layer corresponds to an evaluation result of one of the five tremor severity levels defined by the Fahn-Tolosa-Marin tremor rating scale.
[0013] Preferably, the image needs to be preprocessed before input, and this preprocessing includes: adjusting the image size to a preset resolution and normalizing the pixel values of the image.
[0014] Preferably, the image is 64×64 pixels after resizing.
[0015] Preferably, when using ETSD-Net as the initial layer, its structure includes: a first convolutional block, a plurality of stacked inverse residual blocks, and a second convolutional block; the first convolutional block includes a convolutional layer, a batch normalization layer, and a ReLU layer, which are used to extract shallow features; this shallow feature is sent to a plurality of inverse residual modules arranged in series, and each inverse residual module integrates a channel-spatial attention module, which is used to extract the semantic features of the handwritten image of the Archimedean spiral; the second convolutional block is used to receive and process the semantic features and output the final features of the initial layer.
[0016] The tremor severity assessment device based on handwritten image recognition of the present application, which includes a calculation unit for running the above-mentioned tremor severity assessment model.
[0017] The training method of the tremor severity assessment model based on handwritten image recognition of the present application, the model is the above-mentioned tremor severity assessment model; the model training is divided into two stages;
[0018] For ResNet50, DenseNet, ConvNeXt Tiny, MobileNet-V2, in the first stage, freeze the initial layer and train the fully connected parameters of the subsequent layers; in the second stage, freeze the parameters of the subsequent layers and fine-tune the parameters of the initial layer.
[0019] For ETSD-Net, use the pre-trained weights in MobileNet-V2 when loading the model. In the first stage, freeze the parameters of the initial layer and fine-tune the parameters of the subsequent layers; in the second stage, freeze the parameters of the subsequent layers and fine-tune the channel-spatial attention module parameters of the initial layer.
[0020] Preferably, ResNet50, DenseNet, ConvNeXt Tiny, and MobileNet-V2 are pre-trained using the ImageNet dataset.
[0021] Preferably, during the fine-tuning process, the fine-tuning process uses mini-batch gradient descent with a batch size of 64 and applies a hierarchical decay learning rate strategy.
[0022] Preferably, each training phase contains 10 training epochs.
[0023] The effective method for remotely evaluating the severity of tremors based on handwritten images in this application provides a practical and feasible method for evaluating the severity of tremors. The model of this application not only achieves objective and accurate evaluation but also supports remote and low-cost evaluation. Description of the Drawings
[0024] Figure 1 It is a flowchart for evaluating the severity of tremors based on handwritten image recognition.
[0025] Figure 2 It is the integration of ETSD-Net in severity prediction.
[0026] Figure 3 It is the performance evaluation of the baseline model and the ETSD network: confusion matrix, ROC curve, and t-SNE visualization of ET severity prediction.
[0027] Figure 4 It is the Grad-CAM heatmap visualization of the baseline model and the ETSD network.
[0028] Figure 5 It is the saliency map visualization of the baseline model and the ETSD network. Detailed Implementation Manner
[0029] Research Subjects
[0030] The Ethics Committee of the Chinese People's Liberation Army General Hospital approved the trial (s 2018 - 021–00 / 01). Multiple neurologists screened patients with typical ET symptoms. Data collection for this experiment began on September 4, 2020. As of March 5, 2024, 315 ET patients at different stages had completed the necessary tests and were confirmed to have no other diseases such as PD, hyperthyroidism, Wilson's disease, or drug-induced tremors and were included in the study. The ages of the participants ranged from 30 to 78 years. To better understand the patient profile, the inventors divided the patients into four groups according to their FTM-TRS writing scores: Group 1 (0 - 7 points), Group 2 (8 - 14 points), Group 3 (15 - 21 points), and Group 4 (22 - 28 points), as shown in Table 3.
[0031] Table 3 Patient Baseline Table (This table includes age, body mass index, gender, disease duration, family history, and FTM-TRS score. FTM-TRS is used to evaluate the severity of tremors, and the higher the score, the more severe it is)
[0032]
[0033]
[0034] 1 Mean ± standard deviation (SD); n(%)
[0035] 2 One-way analysis of variance (ANOVA)
[0036] 3 Pearson chi-square test
[0037] The study of the drawing task found that drawing an Archimedean spiral or a straight line has stronger discriminative ability than writing, probably because it requires continuous movement in multiple planar directions, unlike writing which is mainly vertical movement. Therefore, the inventors chose to have the patients complete FTM-TRS Drawing A (drawing an Archimedean spiral). The patients were randomly divided into two groups at a ratio of 8:2. The medical institution group consisted of 252 patients who completed this task under the guidance of a neurologist, scanned the handwritten images using the scanning function of an HP LaserJet Pro MFP M226dw printer (HP, USA), and saved them as JPEG format images at a preset resolution of 300 dpi. The 63 patients in the remote home group were required to print the PDF template of the drawing task on A4 paper at home, complete the drawing, then take photos with a smartphone or camera at the highest resolution, and upload the images in full-resolution JPEG format.
[0038] A total of 798 high - definition scanned images and 199 photos of drawings made by patients were collected. Each image was independently scored by three neurologists according to the FTM - TRS scale. In case of scoring differences, the neurologists discussed and re - evaluated the images to reach a consistent score. This process produced a high - quality dataset of 997 images, and the FTM - TRS scores were used as labels for training the model. The experimental process is as Figure 1 shown.
[0039] Data pre - processing
[0040] In this study, the inventors utilized transfer learning techniques to classify handwritten spiral images. Transfer learning is a deep - learning method in which a pre - trained model initially developed for a specific task is reused as the starting point for a new - task model. After being trained on a large ImageNet dataset, the pre - trained model is used to enhance performance on a private dataset.
[0041] Pre - processing is a crucial initial step to ensure that the input data is suitable and of high quality for training a deep - learning model. To make our handwritten images compatible with pre - trained deep - learning models, we applied several pre - processing steps, including resizing the images, normalizing pixel values, and data augmentation. Many deep - learning models, especially CNN models like VGG16 and ResNet, are trained with input images of a fixed size, usually 224×224 pixels. To ensure compatibility with these pre - trained models, we resized the images to 224×224 pixels. This size strikes a balance between computational efficiency and image detail. Larger images (e.g., 512×512) increase the computational cost and training time, while smaller images (e.g., 64×64) may lose details and thus reduce performance. Therefore, 224×224 pixels provides the best compromise.
[0042] This resizing ensures that the model can effectively process the images without distortion or loss of key information. The images were normalized using the mean and standard - deviation values from the ImageNet dataset (means are [0.485, 0.456, 0.406] respectively, and standard deviations are [0.229, 0.224, 0.225] respectively). As part of pre - processing, normalization and resizing ensure that the model focuses on learning relevant features from the images rather than being affected by color, brightness, or size variations. This process helps to stabilize the training process and improve the model's convergence.
[0043] In addition to preprocessing, data augmentation plays a crucial role in enhancing the model's capabilities. It artificially expands the training dataset by applying a series of random transformations such as orientation, size, blur effect, contrast change, brightness change, etc., to ensure that the model can effectively learn and extract relevant features. Our dataset is naturally distributed, with more patients having mild tremors, resulting in an imbalance in the number of samples in each group. Before training the model, we also applied data augmentation techniques to balance the groups and ensure a more uniform distribution of the samples used for training. By systematically applying these preprocessing and enhancement techniques, we ensured the reliability of our dataset and were able to train an effective deep learning model for the severity assessment of handwritten spiral images.
[0044] Model Evaluation
[0045] In the task of evaluating the severity of hand-drawn spiral tremors, we adopted four of the most commonly used deep learning models as baseline models.
[0046] 1) ResNet50: The residual structure alleviates the problem of gradient vanishing, making it suitable for extracting complex features. Its mature and stable architecture is a widely used benchmark model in visual tasks.
[0047] 2) DenseNet: The dense connection mechanism supports efficient feature reuse, achieving good performance with fewer parameters. Its characteristics make it particularly suitable for tasks involving small datasets or requiring deep feature fusion.
[0048] 3) ConvNeXt-Tiny: Widely adopting the design principles of modern lightweight convolutional networks, it enables the capture of multi-scale local and global features while taking into account both performance and efficiency.
[0049] 4) MobileNet-V2: By leveraging depthwise separable convolutions and inverted residual structures, it achieves a balance between computational efficiency and prediction performance. The lightweight design is very suitable for deployment on mobile or portable devices.
[0050] Remote diagnosis considering the severity of ET relies on the accuracy and efficiency of the model. We propose an improved model ETSD-Net based on MobileNet-V2. It consists of 2 convolutional blocks, N inverted residual blocks, and 1 fully connected layer, with specific batch normalization layers and activation functions. Specifically, the input of ETSD-Net first passes through a convolutional block containing a convolutional layer, a batch normalization layer, and a ReLU layer to extract shallow features. The shallow features are sent to multiple stacked inverted residual blocks to capture the spatial details and semantic information in the hand-drawn spiral images, which benefits from the introduction of a channel-spatial attention mechanism in the inverted residual modules of the network. Then, the semantic features are refined by another convolutional block with an average pooling layer. Finally, the class-related features are flattened and sent to the fully connected layer to obtain the output.
[0051] Experimental Setup
[0052] To prevent data leakage and ensure an objective evaluation of the baseline model and ETSD-Net proposed in this paper, we adopted a subject-independent data splitting strategy to ensure that data from the same subject only appears in the training, validation, or test sets. Based on this, we divided the collected data and prepared the dataset in a ratio of 6:2:2. To reduce the cost of training the model and accelerate the convergence process, transfer learning was used to train the baseline model and ETSD-Net. The initial layers in the transfer learning model capture generalizable features, while the subsequent layers are more task-specific.
[0053] Therefore, the model undergoes a two-stage training process: the first stage involves adding a new classifier, and the second stage focuses on fine-tuning the model. In the first stage, for ResNet50, DenseNet, ConvNeXt Tiny, and MobileNet-V2, a fully connected layer corresponding to the five severity levels defined by FTM-TRS was added at the end of the four pre-trained models, and the weights of other layers were frozen. For ETSD-Net, since its underlying architecture is MobileNet-V2, we used the pre-trained weights in MobileNet-V2 when loading the model. We froze the weights of the modules that could be matched, and the fine-tuning mainly focused on the spatio-temporal attention module and the final classification layer. In this way, the initial layers of the neural network capture general features such as edges and textures. By keeping these layers unchanged, the model can utilize these learned features without having to retrain.
[0054] During the fine-tuning process, we used a batch size of 64 for iterative training and adopted a learning rate schedule with hierarchical decay. The base learning rate was set to 10^-3. For the feature extraction layer, the learning rate decayed by a factor of 0.5 every three layers, while the classification layer was set to the base learning rate without decay. This is because the feature layers contain pre-trained weights and do not require a large learning rate to find the optimal weights, while the classification layer has not loaded pre-trained weights. The Adam optimizer was used for parameter optimization, and multi-class cross-entropy was used as the loss function for backpropagation. The entire fine-tuning process lasted for 10 epochs. Here, we used a small number of epochs because the model was initialized with pre-trained weights from ImageNet. The experimental results show that this initialization method enables the model to converge within 10 epochs. Therefore, we do not use a large number of training epochs to avoid overfitting on the small-scale dataset. The model was trained on the training set, and the model with the highest accuracy on the validation set was saved as the best model. In this application, the performance metrics used to evaluate the model are accuracy, F1-score, precision, and recall.
[0055] During the fine-tuning process, we used a batch of 64 for iterative training and adopted a learning rate schedule with hierarchical decay. The base learning rate was set to 10 -3 ^-3. For the feature extraction layer, the learning rate decayed by a factor of 0.5 every three layers, while the classification layer was set to the base learning rate without decay. This is because the feature layers contain pre-trained weights and do not require a high learning rate to find the optimal weights, while the classification layer has not loaded pre-trained weights. The Adam optimizer was used for parameter optimization, and multi-class cross-entropy was used as the loss function for backpropagation. The entire fine-tuning process lasted for 10 epochs. The number of epochs was small because the model was initialized with pre-trained weights from ImageNet. The experimental results show that this initialization method enables the model to converge within 10 epochs. Therefore, we do not use more training epochs to avoid overfitting on the small-scale dataset. The model was trained on the training set, and the model with the highest accuracy on the validation set was saved as the best model. The performance metrics used in this paper to evaluate the model are accuracy, F1-score, precision, and recall. This process takes ETSD-Net as an example, as Figure 2 shown.
[0056] Research results
[0057] We compared the performance of four baseline transfer learning models, ResNet50, DenseNet, MobileNet-V2, ConvNeXt-Tiny, and our proposed ETSD-Net model, using four evaluation metrics: accuracy, precision, recall, and F1-score. The results are shown in Table 4.
[0058] Table 4 Comparison of the performance of benchmark models and the ETSD network
[0059] Method Accuracy Precision Recall F1-score 86.43% 88.48% 86.43% 86.73% DenseNet 85.93% 85.92% 85.93% 85.51% ConvNeXt-Tiny 86.93% 88.87% 86.93% 87.21% MobileNet-V2 87.44% 87.63% 87.44% 87.44% ETSD-Net(Ours) 88.44% 88.64% 88.44% 88.45%
[0060] The results show that ETSD-Net outperforms the baseline models on several key metrics. It has the highest accuracy (88.44%), the highest recall (88.44%), and the highest F1-score (88.45%), indicating that the algorithm has good balance and correct classification ability between accuracy and recall. Although ConvNeXt-Tiny shows slightly higher accuracy (88.87%), ETSD-Net achieves good accuracy results (88.64%) while maintaining better overall performance in other metrics.
[0061] The performance of the model was verified through the ROC curve and confusion matrix. The AUC values for each severity level are between 0.98 - 0.99, showing good classification ability. The confusion matrix further highlights the robust performance of the model, with most predictions falling along the diagonal, indicating accurate classification (as Figure 3 shown). The results show that our ETSD-Net has successfully utilized transfer learning techniques and outperforms existing models on multiple evaluation criteria, being more robust and reliable in the task of tremor assessment.
[0062] ETSD-Net is built on top of the MobileNet-V2 architecture, maintaining a lightweight design with comparable computational efficiency to MobileNet-V2. As shown in Table 5, these two models show similar Floating Point Operations per Second (FLOPs) (0.33 GFLOPs for ETSD-Net vs. 0.30 GFLOPs for MobileNet-V2) and parameter sizes (2.26M for ETSD-Net vs. 2.20M for MobileNet-V2). This similarity reflects the architectural choice to retain the computational efficiency of MobileNet-V2 while incorporating improvements to enhance performance.
[0063] However, the inference time of ETSD-Net (20.17 ± 10.56 ms) is significantly higher than that of MobileNet-V2 (7.21 ± 0.93 ms). The increased latency can be attributed to the improvements in feature extraction and overall performance. In contrast, other baseline models such as ResNet50 and ConvNeXt-Tiny achieve faster inference times (e.g., 5.34 ± 1.28 ms for ConvNeXt-Tiny), but at the cost of a significant increase in computational requirements (4.45 GFLOPs and 27.80m parameters for ConvNeXt-Tiny).
[0064] Table 5 Computational Efficiency and Inference Time of the Baseline Model and the ETSD Network
[0065]
[0066]
[0067] In Table 5, FLOPs represent the computational complexity of the model. Params (parameters) affect memory usage and model capacity. Inference time is the time taken for the model to process the input and generate the output, measured in milliseconds (ms).
[0068] Figure 4 and Figure 5 respectively show the Gradient-weighted Class Activation Mapping (Grad-CAM) and Saliency map visualization results of different models in the FTM-TRS input images (0–4). These visualizations together reveal how each model allocates attention to specific regions of the input image during the classification process and provide complementary insights into their focusing mechanisms. The Grad-CAM results highlight a broader attention distribution, while the Saliency map emphasizes sensitivity to task-relevant regions at a finer-grained level.
[0069] ResNet50 and DenseNet show limited and local attention distributions in both Grad-CAM and Saliency map, mainly concentrated on isolated parts of the helix. This incomplete coverage indicates that these models have difficulty capturing the global geometric patterns of the helix, which is crucial for accurate classification. The attention shown by ConvNeXt Tiny is mainly concentrated on small disconnected regions, as shown in both visualizations, indicating a tendency to overfit local details rather than understand the overall structure of the helix. Compared with the baseline models, MobileNet-V2 shows a more balanced attention pattern, with Grad-CAM and Saliency map showing a broader coverage of the helix. However, its focus is still not comprehensive enough, leaving some parts of the structure underrepresented.
[0070] In contrast, ETSD-Net achieved the most comprehensive and consistent attention distribution. As can be seen from the Grad-CAM and Saliency map, the model effectively learned the complete geometric structure of the spiral, and the attention points covered both the central and peripheral regions. This global attention ensured robust feature extraction and better generalization across different input complexities. The combined results of Grad-CAM and Saliency map strongly supported the superior performance of ETSD-Net, whose ability to capture both local and global features exceeded that of the baseline model.
[0071] We collected approximately 1000 high-quality FTM-TRS Archimedean spiral handwritten images from about 400 ET patients to establish a robust dataset with expert scores. Using transfer learning methods, we developed the ETSD-Net model for ET severity assessment. The accuracy of ETSD-Net was 88.44%, which was superior to existing methods. This application made a meaningful contribution to improving the accessibility and reliability of ET assessment, and can be used for tremor severity assessment in remote or resource-limited environments. Elderly patients with limited mobility can also benefit from remote assessment.
Claims
1. A tremor severity assessment model based on handwritten image recognition, which is trained to evaluate the ET tremor severity of a subject; the model includes an initial layer and subsequent layers; the initial layer is implemented by one of ResNet50, DenseNet, ConvNeXtTiny, MobileNet-V2, or ETSD-Net; the subsequent layers are implemented by fully connected layers; the input of the tremor severity assessment model is a handwritten image of an Archimedean spiral drawn by the subject on paper; this image is sent to the initial layer for calculation, and the output of the initial layer is used as the input of the subsequent layer; the output of the subsequent layer corresponds to an evaluation result of one of the five tremor severity levels defined by the Fahn-Tolosa-Marin tremor rating scale.
2. The tremor severity assessment model based on handwritten image recognition according to claim 1, wherein: The image needs to be preprocessed before input, and the preprocessing includes: adjusting the image size to a preset resolution and normalizing the pixel values of the image.
3. The tremor severity assessment model based on handwritten image recognition according to claim 2, wherein: The image is 64×64 pixels after resizing.
4. The tremor severity assessment model based on handwritten image recognition according to claim 1, wherein: When using ETSD-Net as the initial layer, its structure includes: a first convolutional block, multiple stacked inverse residual blocks, and a second convolutional block; the first convolutional block contains a convolutional layer, a batch normalization layer, and a ReLU layer, and is used to extract shallow features; the shallow features are sent to multiple inverse residual modules arranged in series, and each inverse residual module integrates a channel-spatial attention module for extracting semantic features of the handwritten image of the Archimedean spiral; the second convolutional block is used to receive and process the semantic features and output the final features of the initial layer.
5. A tremor severity assessment device based on handwritten image recognition, which includes a computing unit for running the tremor severity assessment model according to any one of claims 1-4.
6. A training method for a tremor severity assessment model based on handwritten image recognition, where the model is the tremor severity assessment model according to any one of claims 1-4; the model training is divided into two stages; For ResNet50, DenseNet, ConvNeXt Tiny, MobileNet-V2, in the first stage, freeze the initial layer and train the fully connected parameters of the subsequent layers; in the second stage, freeze the parameters of the subsequent layers and fine-tune the parameters of the initial layer. For ETSD-Net, use the pre-trained weights in MobileNet-V2 when loading the model. In the first stage, freeze the parameters of the initial layer and fine-tune the parameters of the subsequent layers; in the second stage, freeze the parameters of the subsequent layers and fine-tune the channel-spatial attention module parameters of the initial layer.
7. The training method for a tremor severity assessment model based on handwritten image recognition according to claim 6, wherein: ResNet50, DenseNet, ConvNeXt Tiny, and MobileNet-V2 are pre-trained using the ImageNet dataset.
8. The training method of the tremor severity assessment model based on handwritten image recognition according to claim 7, characterized in that: During the fine-tuning process, the fine-tuning process uses the mini-batch gradient descent method with a batch size of 64 and applies a hierarchical decay learning rate strategy.
9. The training method of the tremor severity assessment model based on handwritten image recognition according to claim 7, characterized in that: Each training phase contains 10 training epochs.