Evaluation method and system, storage medium and terminal for Parkinson's disease standing task

Video data is obtained through self-supervised metric graph convolution network and ordinary cameras, skeleton sequences and joint information are extracted, joint flow networks and skeleton flow networks are trained, and the vertex-specific spatial graph convolution and graph representation supervision losses are combined, which solves the limitations of the existing automatic evaluation method of Parkinson's disease standing task, and realizes efficient and objective automatic evaluation of Parkinson's standing task based on Parkinson's standing task.

CN114140816BActive Publication Date: 2025-06-06SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111106752.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2025-06-06
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

The existing automatic evaluation method for the standing task of Parkinson's disease has problems such as inconvenient data acquisition, lack of objectivity in evaluation results, and inconvenient sensor wearing. Sensor-based methods require calibration and calibration, which affects compliance and initiative in non-clinical environments.

Method used

The self-supervised metric graph convolution network is used to obtain video data through ordinary cameras, extract skeleton sequences and joint information, train joint flow networks and skeleton flow networks, and combine vertex-specific spatial graph convolutions and graphs to represent supervision losses, realizing automatic evaluation of standing actions in the video.

Benefits of technology

Automatic evaluation of Parkinson's standing task based on video is realized, which improves the objectivity and reliability of the evaluation, reduces the cost of evaluation, is suitable for performing in various environments, and does not require patients to wear sensors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140816B_ABST
    Figure CN114140816B_ABST
Patent Text Reader

Abstract

The present invention provides an evaluation method and system for a Parkinson's disease standing up task, a storage medium and a terminal, comprising the following steps: obtaining video information containing a Parkinson's disease patient's standing up action; extracting a skeleton sequence of the Parkinson's disease patient from the video information, and extracting joint information and bone information based on the skeleton sequence; training a self-supervised metric graph convolutional network based on the joint information and the bone information; obtaining the predicted probability of each category of the standing up action of the Parkinson's disease patient to be evaluated based on the trained self-supervised metric graph convolutional network, and selecting the category with the largest predicted probability as the evaluation category of the Parkinson's disease patient to be evaluated. The evaluation method and system, storage medium and terminal for a Parkinson's disease standing up task of the present invention adopt a self-supervised metric graph convolutional network to realize automatic evaluation of a Parkinson's disease standing up task based on a video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular to an evaluation method and system, storage medium and terminal for a Parkinson's disease standing task. Background Art

[0002] Parkinson's disease (PD) can lead to the deterioration of patients' motor function and seriously impair their quality of life. The increasing number of patients has also brought great socioeconomic burden. In clinical practice, neurologists most often use the World Movement Disorders Society-Unified Parkinson's Disease Rating Scale (MDS-UPDRS) to assess the severity and progression of the disease.

[0003] However, this assessment method has the following limitations in clinical trials:

[0004] 1) A skilled neurologist is required to conduct assessments during intermittent clinical visits, which is not conducive to continuous monitoring of symptoms, especially since the COVID-19 pandemic has increased the difficulty of medical consultation and follow-up;

[0005] 2) This method wastes doctors’ valuable time, and the scores vary among different doctors, lacking objectivity;

[0006] 3) When evaluating the motion subtask, multiple motion factors need to be considered simultaneously, such as pauses, speed, etc., which increases the difficulty of evaluating the concise score results.

[0007] Therefore, there is an urgent need for an MDS-UPDRS automatic assessment system to provide objective and reliable assessment results, thereby achieving effective diagnosis of PD. The standing task is an important part of the MDS-UPDRS, and the doctor will score the patient from 0 to 4 based on the patient's completion. Figure 1 As shown in Figure 1, patients are asked to cross their arms over their chest and then stand up. This task helps doctors assess the severity of PD by looking at motor performance.

[0008] The emergence of wearable sensor devices allows accurate motion measurement by analyzing the captured motion signals, so sensor-based methods have gradually emerged to achieve automatic assessment of MDS-UPDRS. The basic idea of ​​this type of method is to collect data using sensors placed at designated locations on PD patients, then extract relevant kinematic features in the time domain or frequency domain, and finally use machine learning algorithms to perform score classification or feature correlation analysis for a task. For example, Martinez-Manzanera et al. analyzed motion data recorded by a nine-degree-of-freedom sensor and used a support vector machine classifier to achieve automatic scoring of three bradykinesia-related items. Jeon et al. measured the tremor signals of patients using a wristwatch-type wearable device, and then used the extracted selected features to automatically assess the severity of tremor through multiple machine learning algorithms. Lin et al. used significant features extracted based on axis-angle representation to train a support vector machine algorithm to objectively assess the severity of bradykinesia.

[0009] All existing automatic evaluation works for PD standing tasks have adopted sensor-based methods. Giuberti et al. used a body sensor network consisting of a single inertial node placed on the patient's chest to collect kinematic variables, and then used k-nearest neighbors (kNN) to construct an automatic UPDRS evaluation system for the standing task with an accuracy of about 65%. Parisi et al. also used chest-mounted sensors to analyze the characteristics of body movements during the task, and finally achieved an accuracy of about 43% using kNN. Although the above works have achieved automatic evaluation of MDS-UPDRS scores for standing tasks, there are still three limitations in data acquisition:

[0010] 1) The sensor is in direct contact with the patient's body, which may increase the patient's discomfort and slightly affect movement;

[0011] 2) Sensors usually require calibration and regular recalibration to maintain good accuracy, thus weakening patient compliance and initiative in non-clinical settings;

[0012] 3) Due to the coordination of movements, patients may need to wear multiple sensors at the same time, which is not convenient and increases costs.

[0013] In addition, in terms of classification algorithms, existing methods have all used machine learning algorithms as classifiers, which require manual development and definition of features from motion trajectories. They lack comprehensiveness and are difficult to automate.

[0014] The latest progress in deep learning-based human posture estimation algorithms has opened up new ideas for the automatic evaluation of MDS-UPDRS. This type of algorithm does not require patients to wear any equipment to collect motion data. It can directly extract joint position coordinates from video data captured by ordinary cameras to characterize motion trajectories. It has low cost and is suitable for evaluation in various environments. Some work has taken advantage of the human posture estimation algorithm, using conventional machine learning algorithms or convolutional neural networks as classification models, and proposed a vision-based system to achieve automatic evaluation of MDS-UPDRS. However, these classification models often require predefined motion features or the order of joint point sequences, and cannot use the natural graph structure of human skeleton data to automatically capture and express the motion dependencies between human joints, which in turn affects the performance of human motion feature modeling and classification.

[0015] To this end, the spatial-temporal graph convolutional network (ST-GCN) proposed by Yan et al. introduced a graph convolutional neural network (GCN) to simulate the spatial structural relationship of the human body and combined it with temporal convolution, so as to directly learn the important spatiotemporal features of the action from the human joint trajectory, without the need to design and calculate the kinematic features from the joint motion trajectory. Since then, the action recognition algorithm based on GCN has been gradually used for the motor function evaluation of PD. For example, Guo et al. embedded the ideas of sparsification and adaptation in GCN, and designed a modeling mechanism of temporal dependency and channel saliency, realizing the effective modeling and reliable evaluation of leg flexibility tasks. Afterwards, ST-GCN combined with attention mechanism and deep supervision strategy was used to develop an automatic evaluation model for PD gait movement disorder. In addition, Hu et al. developed an automatic detection method for Freezing of Gait (FoG) based on GCN and weakly supervised learning strategy, and also expanded FoG detection to a graph sequence modeling task, and then designed a graph sequence recurrent neural network to process spatiotemporal data. However, automatic MDS-UPDRS scoring for the PD standing task using video-based human pose estimation and GCN classification remains to be explored, and there is currently no similar work.

[0016] Self-supervised video representation learning in action recognition aims to design various pre-tasks to train the model to learn good spatiotemporal motion representations from unlabeled datasets, and then directly apply the model to downstream action recognition tasks for feature extraction or fine-tuning.

[0017] For action videos consisting of a series of frames, effective learning of temporal information is very important. Therefore, a series of studies have designed prediction tasks related to temporal dependencies from the temporal domain. For example, Fernando et al. designed a pre-task to predict video clips with incorrect frame order from multiple video clips. Jenni et al. distinguished between original sequences and sequences that have undergone video time transformation as a pre-task. Since action videos also contain important human spatial motion information, there are also studies that design tasks from the overall spatiotemporal domain to learn effective spatiotemporal representations. For example, Wang et al. predicted motion and appearance statistics extracted from the spatial and temporal domains of unlabeled videos in the pre-task. Since image data contains rich feature information, the above-mentioned studies based on RGB videos have achieved good results.

[0018] In addition to the above RGB information, another important form of characterization of action videos is the human skeleton. Since skeleton data only consists of joint position coordinates, the exploration of skeleton sequences in self-supervised video representation learning is more challenging. Si et al. constructed a self-supervised strategy by exploring neighbor relationships, thereby learning discriminative motion features of unlabeled samples with neighborhood consistency as the goal. Lin et al. integrated multiple self-supervised tasks to learn skeleton features through motion prediction of future sequences, jigsaw puzzle recognition of temporal patterns, and contrastive learning of feature spaces.

[0019] Although the above-mentioned work on action recognition has been extensively explored in the field of self-supervised video representation learning, the model based on skeleton data has not been fully studied. In addition, most studies mainly consider the learning of temporal representations, and how to effectively learn spatial representations is still worth considering.

[0020] In recent years, the modeling of human skeleton patterns through spatiotemporal graph convolutional networks has achieved excellent performance and has been widely used in human action recognition. However, previous GCN-based recognition models still have some limitations, such as small receptive field, weak fine-grained information modeling ability, and insufficient flexibility of graph topology. Therefore, many studies have continued to explore and improve various problems. For example, Cheng et al. designed spatial and temporal shift graph convolution operations respectively, and further proposed a shift graph convolution network to reduce computational cost and adaptively adjust the receptive field. Cai et al. used the global motion information of conventional skeleton data and the local motion information captured by joint-aligned optical flow patches to construct a two-stream network to fully characterize human motion. Peng et al. constructed a dynamic graph generation module and introduced a module with high-order connections. Finally, they used neural architecture search to realize the automatic design of GCN, thereby exploring a better GCN architecture. Zhang et al. supplemented the spatial global information in GCN and calculated the context information in different ways, thereby improving the context perception ability of GCN. Shi et al. combined the data-driven idea and designed a graph topology that can be learned in an end-to-end manner along with the input data, thus enhancing the flexibility of the graph structure.

[0021] Although these studies effectively solve many problems of GCN-based models, they often focus on exploring the connection relationship between different joints, while ignoring the specificity of different joints. In addition, many models adopt a two-stream framework, but often do not consider the consistency between the two-stream feature representations in the same task. Summary of the invention

[0022] In view of the shortcomings of the prior art described above, the purpose of the present invention is to provide an evaluation method and system, storage medium and terminal for the Parkinson's disease standing task, which adopts a self-supervised metric graph convolutional network to realize automatic evaluation of the Parkinson's disease standing task based on video.

[0023] To achieve the above-mentioned purpose and other related purposes, the present invention provides an evaluation method for a Parkinson's disease standing-up task, comprising the following steps: obtaining video information containing a standing-up action of a Parkinson's disease patient, wherein the standing-up action is a standard action required for performing the standing-up task evaluation; extracting a skeleton sequence of the Parkinson's disease patient from the video information, and extracting joint information and bone information based on the skeleton sequence; training a self-supervised metric graph convolutional network based on the joint information and the bone information; the self-supervised metric graph convolutional network is used to output the predicted probability of the standing-up action in the video information for each category, and comprises a joint flow network and a bone flow network; the joint flow network and the bone flow network both comprise 9 spatiotemporal units; each spatiotemporal unit is composed of a spatial graph convolution and a temporal convolution; obtaining the predicted probability of the standing-up action of the Parkinson's disease patient to be evaluated for each category based on the trained self-supervised metric graph convolutional network, and selecting the category with the largest predicted probability as the evaluation category of the Parkinson's disease patient to be evaluated.

[0024] In one embodiment of the present invention, the skeleton sequence is extracted from the video information based on the human posture estimation model OpenPose, the joint information is extracted based on the skeleton sequence, and the bone information is extracted based on the coordinate difference between two naturally connected joints.

[0025] In one embodiment of the present invention, training a self-supervised metric graph convolutional network based on the joint information and the skeleton information comprises the following steps:

[0026] Performing self-supervised intra-video quadruple learning on the joint stream network and the skeleton stream network based on the skeleton sequence, and applying constraints of graph representation supervision loss on output features of the joint stream network and the skeleton stream network;

[0027] The parameters obtained by the self-supervised video quadruple learning are used as the initial parameters of the joint flow network and the skeleton flow network, the joint flow network and the skeleton flow network are trained based on the joint information and the skeleton information, and the outputs of the joint flow network and the skeleton flow network are spliced ​​and input into the fully connected layer and softmax.

[0028] In one embodiment of the present invention, the self-supervised intra-video quadruple learning of the skeleton sequence comprises the following steps:

[0029] A data enhancement algorithm based on similarity transformation is used to generate positive samples that are similar to the global context of the skeleton sequence in time and space, and the skeleton sequence is used as an anchor sample;

[0030] Based on a spatial global perturbation algorithm, generating spatial negative samples having the same temporal context as the skeleton sequence and a different spatial context;

[0031] Based on a temporal global perturbation algorithm, generating temporal negative samples that are temporally chaotic with the skeleton sequence and have normal spatial context;

[0032] The anchor samples, the positive samples, the spatial negative samples and the temporal negative samples are all input into the joint flow network and the skeleton flow network, so that the joint flow network and the skeleton flow network meet the preset loss function target.

[0033] In one embodiment of the present invention, the loss function objective is to maximize the relative distance between the anchor sample, positive sample pairs and the anchor sample, temporal negative sample pairs, to maximize the relative distance between the anchor sample, positive sample pairs and the temporal negative sample, spatial negative sample pairs, and to make the minimum distance between the temporal negative sample and the spatial negative sample pairs greater than the maximum distance between the anchor sample and the positive sample pairs.

[0034] In one embodiment of the present invention, when the joint flow network and the skeleton flow network are trained based on the joint information and the skeleton information, the cost function is composed of a cross entropy term and a graph representation supervision loss term.

[0035] In one embodiment of the present invention, in the spatiotemporal unit, vertex-specific spatial graph convolution is used in the spatiotemporal unit with the same number of input and output channels.

[0036] The present invention provides an evaluation system for a Parkinson's disease standing task, comprising an acquisition module, an extraction module, a training module and an evaluation module;

[0037] The acquisition module is used to acquire video information containing standing-up movements of Parkinson's disease patients, where the standing-up movements are standard movements required for standing-up task assessment;

[0038] The extraction module is used to extract a skeleton sequence of a Parkinson's patient from the video information, and extract joint information and bone information based on the skeleton sequence;

[0039] The training module is used to train a self-supervised metric graph convolution network based on the joint information and the skeleton information; the self-supervised metric graph convolution network is used to output the prediction probability of the standing up action in the video information for each category, including a joint flow network and a skeleton flow network; the joint flow network and the skeleton flow network each include 9 spatiotemporal units; each spatiotemporal unit is composed of a spatial graph convolution and a temporal convolution;

[0040] The evaluation module is used to obtain the predicted probability of the standing-up action of the Parkinson's patient to be evaluated for each category based on the trained self-supervised metric graph convolutional network, and select the category with the largest predicted probability as the evaluation category of the Parkinson's patient to be evaluated.

[0041] The present invention provides a storage medium on which a computer program is stored. When the program is executed by a processor, the evaluation method for the Parkinson's disease standing task is realized.

[0042] The present invention provides an evaluation terminal for a Parkinson's disease standing task, comprising: a processor and a memory;

[0043] The memory is used to store computer programs;

[0044] The processor is used to execute the computer program stored in the memory, so that the evaluation terminal for the Parkinson's disease standing up task executes the above-mentioned evaluation method for the Parkinson's disease standing up task.

[0045] As described above, the Parkinson's disease standing task evaluation method and system, storage medium and terminal of the present invention have the following features:

[0046] Beneficial effects:

[0047] (1) Only ordinary cameras can be used to collect patient movement data. This model can be easily migrated to all scenarios, further improving the feasibility of remote assessment;

[0048] (2) The dataset comes from patient evaluation videos collected in clinical practice. The dataset is large and has excellent evaluation performance;

[0049] (3) The self-supervised deep learning scheme used significantly enhances the model's ability to automatically extract spatiotemporal features from standing-up videos, and combines the vertex-specific spatial graph convolution operation (VSGCO) and graph representation supervision loss (GRSL), with an accuracy rate that is much better than traditional feature engineering methods.

[0050] (4) The spatial and temporal knowledge of the action is used to construct positive and negative sample pairs, and the spatial and temporal sensitivity of the model learning is simultaneously improved through metric learning of quadruple. By assigning specific adaptive weight parameters to each joint and different features of each joint, joint-specific spatial feature aggregation is achieved. By maximizing the consistency between the two-stream feature representations, the similarity and discriminative learning of the two-stream features in the classification task is promoted. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A schematic diagram showing a Parkinson's patient completing a standing-up task in one embodiment;

[0052] Figure 2 Shown is a flow chart of an embodiment of a method for evaluating a Parkinson's disease standing task according to the present invention;

[0053] Figure 3 It is a schematic diagram showing a framework of the SSM-GCN model of the present invention in one embodiment;

[0054] Figure 4 Shown is a schematic diagram of the structure of positive and negative samples of the present invention in one embodiment;

[0055] Figure 5 The figure shows a flow chart of calculation of the supervision loss in one embodiment of the present invention;

[0056] FIG6( a ) is a schematic diagram showing the ROC curve of the classification results of various categories using the data set of the present invention;

[0057] FIG6( b ) is a schematic diagram showing a confusion matrix of classification results of various categories using a data set of the present invention;

[0058] Figure 7 Shown is a schematic diagram of the distribution of results of ten repeated experiments of the evaluation method for the Parkinson's disease standing task of the present invention;

[0059] Figure 8 Shown is a schematic structural diagram of an evaluation system for Parkinson's disease standing task in one embodiment of the present invention;

[0060] Fig. 9 Shown is a schematic structural diagram of an evaluation terminal for the Parkinson's disease standing up task in one embodiment of the present invention. DETAILED DESCRIPTION

[0061] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0062] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0063] The Parkinson's disease standing task evaluation method and system, storage medium and terminal of the present invention adopt a self-supervised metric learning scheme with a graph convolutional network (SSM-GCN) to realize automatic evaluation of the Parkinson's disease standing task based on video. Among them, a human posture estimation algorithm is used to extract joint data from the video, and the skeleton data is calculated based on it, thus constructing a two-stream framework; a self-supervised in-video quadruple learning strategy is designed, and the global spatial and temporal relationship of the skeleton sequence is regarded as a self-supervisory signal for video representation learning. A positive sample (data enhancement) and two negative samples (spatial / temporal global perturbations) are constructed for the skeleton sequence of each video. Through metric learning, the positive samples are brought closer and the negative samples are pushed further away, while the spatial negative samples and the temporal negative samples are distinguished, greatly improving the representation ability of the overall spatiotemporal features of the video; a vertex-specific spatial graph convolution operation is proposed, which assigns specific trainable weight parameters to each vertex and each attribute, and realizes effective spatial feature aggregation by learning the importance of different vertices and different attribute features of the same vertex; a graph representation supervision loss is proposed, which constructs a graph representation relationship between the joint graph and the skeleton graph, and then maximizes the consistency between the high-level features of the joints and bones, thereby learning the similar and complementary information of the joint and skeleton data to the greatest extent.

[0064] like Figure 2 As shown, in one embodiment, the Parkinson's disease standing task assessment method of the present invention comprises the following steps:

[0065] Step S1, obtaining video information containing a standing-up action of a Parkinson's disease patient, wherein the standing-up action is a standard action required for performing a standing-up task assessment.

[0066] Specifically, only an ordinary camera can be used to collect the standard actions required for the evaluation of the standing-up task for Parkinson's patients, and sent to the evaluation terminal of the Parkinson's disease standing-up task of the present invention by wired or wireless means. Among them, the device used for recording can be a multimedia device such as a smart phone, a tablet, a camera, and the frame resolution of the recorded video is 720×1280 or 1080×1920, and the frame rate is 30 frames per second. During the recording process, the multimedia device with a camera needs to be placed in front of the Parkinson's patient and remain fixed. Let the Parkinson's patient sit on a straight-backed chair with armrests, put his feet on the ground and sit back, cross his arms on his chest and stand up. If the Parkinson's patient fails, repeat it up to two more times. If it still fails, ask the Parkinson's patient to sit forward on the chair, then cross his arms on his chest and stand up, and try again. If it still fails, the Parkinson's patient can be allowed to stand up with his hands on the armrests, and this action can be repeated up to three times. If it still fails, the Parkinson's patient needs to be assisted to stand up.

[0067] Step S2: extracting a skeleton sequence of a Parkinson's patient from the video information, and extracting joint information and bone information based on the skeleton sequence.

[0068] Specifically, the human pose estimation model OpenPose is used as the human pose estimator to extract all video information frame by frame to obtain the corresponding skeleton sequence. The position coordinates of 25 joints of the human body are extracted for each frame, all expressed in the form of (x, y). In order to unify the input sequence length of the model, the first 50 frames of each skeleton sequence are set as the input video length. If the video length is less than 50 frames, it is padded with 0. After that, z-score normalization is performed for each video to unify the measurement. Finally, based on the 2D joint information of the skeleton sequence extracted from the video, the bone information is calculated by the coordinate difference between the naturally connected joints.

[0069] Step S3, training a self-supervised metric graph convolutional network based on the joint information and the skeleton information; the self-supervised metric graph convolutional network is used to output the predicted probability of the standing up action in the video information for each category, including a joint flow network and a skeleton flow network; the joint flow network and the skeleton flow network each include 9 spatiotemporal units; each spatiotemporal unit is composed of a spatial graph convolution and a temporal convolution.

[0070] like Figure 3As shown, the SSM-GCN model of the present invention integrates a joint stream network and a skeleton stream network. The joint stream network and the skeleton stream network both contain 9 spatiotemporal units. Each spatiotemporal unit is composed of a spatial graph convolution and a temporal convolution (classic τ×1 convolution). Vertex-specific spatial graph convolution operations are embedded in the spatiotemporal units with the same number of input and output channels to adaptively mine and express the importance of different vertices in the current task. Other spatiotemporal units use ordinary spatial graph convolutions. During the pre-training process, based on the metric learning model, the proposed self-supervised video intra-quadruple learning is used to enhance the spatiotemporal sensitivity of the model, while the constraints of the graph representation supervision loss are imposed on the features of the dual-stream output to enhance the similarity of the dual-stream features. In the fine-tuning stage, a well-optimized pre-trained model is used as initialization, and a fully connected layer and softmax are added at the end of the model. The classification error is minimized by the cross entropy loss, and the graph representation supervision loss is also embedded.

[0071] In one embodiment of the present invention, training a self-supervised metric graph convolutional network based on the joint information and the skeleton information mainly consists of two stages, namely a pre-training stage and a fine-tuning stage.

[0072] 31) In the pre-training stage, i.e., the self-supervised pre-task training stage, the main goal is to improve the model's ability to extract spatial and temporal fine-grained features. Specifically, the joint stream network and the skeleton stream network are self-supervised for in-video quadruple learning based on the skeleton sequence, and the constraints of the graph representation supervision loss are imposed on the output features of the joint stream network and the skeleton stream network.

[0073] Among them, Self-supervised Intra-video Quadruplet Learning (SSIQL) uses spatial and temporal prior knowledge to design appropriate pre-tasks without using category labels, and pre-trains the model through self-supervised learning, thereby improving the model's ability to represent the space and time of skeleton sequences. Therefore, based on the metric learning model, construct positive samples similar to the original samples ( Figure 3 ), and two negative samples that are significantly different from the original samples in the spatial and temporal dimensions, respectively, and then the spatiotemporal discrimination of the model is trained by designing a similarity measurement target.

[0074] In the prior art, graph convolution operations usually take the following forms:

[0075]

[0076] in, There are V nodes and C in The input feature matrix of channels, is the output feature matrix. is the adjacency matrix of the graph with dimension V×V, where A represents the connection between different nodes and I represents the self-connection of the node. for The degree matrix of . is a trainable filter parameter matrix shared by all vertices and used to transform C for each vertex in the graph in the channel dimension. in The dimension feature map is C out That is, the parameter matrix of each vertex is Therefore, the features of each vertex are based on the same trainable filter parameters Therefore, It is not vertex-specific, that is, it is not possible to adaptively use different parameters for feature transformation for different vertices.

[0077] Therefore, in order to solve this limitation, the present invention proposes a vertex-specific spatial graph convolution operation, which can be formulated as:

[0078]

[0079] where ⊙ is the element-wise product operation. is a trainable weight parameter matrix with the same dimensions as the input feature matrix X. Therefore, the process A specific weight can be assigned to each channel of each vertex and adaptively updated as the model is trained, thereby adaptively distinguishing the importance of each vertex in the graph while maintaining channel specificity.

[0080] According to the natural connection of the human skeleton, if joints i and j are connected, then the element A in the adjacency matrix ij =1, otherwise A ij = 0. Therefore, The natural representation of the spatial structure of the human body is achieved. However, when a person performs an action, there are potential connections between some unconnected joints in the human skeleton, and these connections are also very important for the representation and recognition of the action. Therefore, the present invention introduces the mask It is different for each graph convolution layer of the network and is a trainable model parameter, so that the adjacency matrix of the graph continues to mine and add joint connections related to the action task based on the natural connection of the human body. In this case, the conventional graph convolution operation in formula (8) can be expressed as:

[0081]

[0082] Furthermore, the vertex-specific spatial graph convolution operation proposed in formula (9) can be expressed as:

[0083]

[0084] Therefore, the vertex-specific spatial graph convolution operation shown in formula (11) can not only mine the spatial structural relationship related to the current task as the model is trained, but also adaptively allocate vertex-specific and channel-specific trainable weight parameters, thereby effectively aggregating the spatial features of the human skeleton sequence.

[0085] In one embodiment of the present invention, Figure 4 As shown, the self-supervised intra-video quadruple learning of the skeleton sequence includes the following steps:

[0086] a) A data augmentation algorithm based on similarity transformation (DABST) generates positive samples that are similar to the global context of the skeleton sequence in time and space, and uses the skeleton sequence as an anchor sample.

[0087] Specifically, the basic idea of ​​constructing positive samples is to keep the motion semantics contained in the original video basically unchanged through data enhancement. The random similarity transformation in the image is applied to each frame in the video to complete the video data enhancement. Assume that the frame length of the skeleton sequence extracted from each video is T, and randomly sample a small rotation, x-direction translation, y-direction translation and scaling candidate factor for the first frame and the Tth frame respectively, expressed as {a 1 ,a T},{t_x 1 ,t_x T},{t_y 1 ,t_y T} and {s 1 ,s T}. Since they are all 0, they are the same as the original video, so these factors are restricted to not all be 0. In order to simulate the smoothness of motion changes in the time dimension of the video, according to the length of the frame sequence, all intermediate frames between the 1st frame and the Tth frame are sampled at equal intervals to create a candidate factor sequence. Therefore, the rotation, x-direction translation, y-direction translation and scaling candidate factors of the tth frame can be expressed as:

[0088]

[0089] Finally, for the position coordinates {x, y} of each joint in the tth frame, there are four degrees of freedom a as above t ,t_x t ,t_y t and t , the similarity transformation result {x',y'} can be obtained by the following formula:

[0090]

[0091] By performing the above coordinate system rotation, translation and scaling transformation operations on each frame, we can obtain the similarity transformation result of the video and complete the data enhancement. In this way, the global context of space and time is maintained.

[0092] b) Based on a spatial global perturbation algorithm, a spatial negative sample having the same temporal context as the skeleton sequence and a different spatial context is generated.

[0093] Specifically, the spatial global disturbance algorithm (SGD) is used to generate spatial negative samples with the same temporal context but different spatial context. Assume that the human skeleton contains S joints, and the joint point set can be expressed as V = {v 1 ,v 2 ,…,v S Taking each video as a unit, the index of the joint point sequence {1,2,…,S} is randomly shuffled, and then the joints of the current video are reordered according to the new index, thereby constructing a skeleton sequence randomly ordered in the spatial dimension.

[0094] c) Based on a temporal global perturbation algorithm, temporal negative samples are generated that are temporally chaotic with the skeleton sequence and have normal spatial contexts.

[0095] Specifically, the temporal global disturbance algorithm (TGD) is used to generate temporal negative samples that are temporally chaotic but have normal spatial context. Assume that the frame set of the current video is F = {f 1 ,f 2 ,…,f T}, randomly sort the indexes of the frame sequence {1, 2, …, T}, and then construct a randomly chaotic frame sequence in the time dimension.

[0096] d) inputting the anchor sample, the positive sample, the spatial negative sample and the temporal negative sample into the joint flow network and the skeleton flow network, so that the joint flow network and the skeleton flow network meet a preset loss function target.

[0097] Specifically, a four-tuple loss function based on metric learning is used as the objective function to optimize the self-supervised learning process. The original skeleton sequence is regarded as an anchor sample, and the positive sample is obtained after data enhancement. After the global perturbation of time and space, the temporal negative sample and the spatial negative sample are obtained respectively. These four samples are passed through the two-stream network respectively, and then the two-stream representation z' from the same sample is obtained. Joint and z' BoneAfter concatenating z' in the channel dimension, the final feature representation z can be obtained by normalizing based on the Euclidean norm using the following formula:

[0098]

[0099] Among them, z' and z represent the feature representation before and after normalization, respectively, and C is the number of channels (i.e., features). Finally, the calculation process of the quadruple loss can be expressed as follows:

[0100]

[0101] where N is the batch size, d(·,·) represents the distance between the representations of two samples (i.e., a similarity measure), and d(x,y) = ‖xy‖ 2 Calculate. a, p, nt, ns represent anchor, positive, temporal negative and spatial negative samples respectively. margin1 and margin2 are two interval parameters. The goal of this loss function is to maximize the relative distance between anchor samples, positive sample pairs and anchor, temporal negative sample pairs to enhance temporal sensitivity, and at the same time maximize the relative distance between anchor samples, positive sample pairs and temporal, spatial negative sample pairs, so that the minimum distance between temporal negative samples and spatial negative sample pairs needs to be greater than the maximum distance between anchor and positive sample pairs, so as to force the distinction between temporal negative samples and spatial negative samples, and enhance temporal and spatial sensitivity at the same time.

[0102] Since the model consists of two streams, joint and skeleton, they ultimately serve the same goal. Therefore, the present invention also designs and embeds the supervisory signal between the joint graph representation and the skeleton graph representation. Figure 5 As shown in FIG. 1 , the present invention uses the graph representation supervision loss to constrain the high-level features extracted from the two streams and model the potential consistency relationship between them. Specifically, for sample i, the graph representation z' of the output of the two streams is first solved. Joint and z' Bone The consistency probability between (note that in the pre-training of self-supervised learning, only anchor samples are considered):

[0103] P i =δ([z' Joint(i) ,z' Bone(i) ]w) (12)

[0104] Among them, z' Joint and z' Bone are the feature representations of the dual-stream framework output. [·,·] is a concatenation operation used to concatenate the two feature matrices along the feature dimension. w is a trainable weight parameter used to fuse the dual-stream graph representation. δ(·) is a sigmoid function used to map the fused dual-stream graph representation to a probability value P between [0,1].i , which can represent the strength of consistency between the joint graph and the skeleton graph representation.

[0105] Ideally, P i The larger the value of , that is, the closer it is to 1, the stronger the consistency between the two-stream graph representations, and vice versa. Therefore, the graph representation supervision loss can be expressed as:

[0106]

[0107] Where, N is the batch size, It is the consistency probability matrix between the two-flow graph representations calculated according to formula (12), and each sample produces a final probability value. is an all-one matrix, representing the consistency between the two-stream graph representations in an ideal situation. Dist is the result of element-by-element subtraction of J and P, which measures the distance between the two. By minimizing formula (13), the distance between the current consistency probability matrix and the ideal perfect consistency matrix can be reduced, thereby maximizing the consistency relationship between the joint graph representation and the bone graph representation extracted by the network, promoting similarity learning under two-stream features.

[0108] The self-supervised pre-task training phase is a self-supervised video quadruple learning process, the main goal of which is to improve the model's ability to extract spatial and temporal fine-grained features. The cost function of the model is:

[0109]

[0110] Among them, λ is a hyperparameter that balances the graph representation supervision loss term.

[0111] 32) In the fine-tuning stage, i.e., the supervised classification task training stage, the parameters obtained by the self-supervised video quadruple learning are used as the initial parameters of the joint stream network and the skeleton stream network, and the joint stream network and the skeleton stream network are trained based on the joint information and the skeleton information, and the outputs of the joint stream network and the skeleton stream network are spliced ​​and input into the fully connected layer and softmax. In this stage, the cost function is composed of the cross entropy term and the graph representation supervision loss term, which is expressed as:

[0112]

[0113] Among them, y i and are the true label and predicted label of the i-th sample respectively, and <·,·> is the dot product operation.

[0114] Step S4: obtaining the predicted probability of the standing-up action of the Parkinson's patient to be evaluated for each category based on the trained self-supervised metric graph convolutional network, and selecting the category with the largest predicted probability as the evaluation category of the Parkinson's patient to be evaluated.

[0115] Specifically, for the standing-up action of the Parkinson's patient to be evaluated, the trained self-supervised metric graph convolutional network is input to obtain its prediction probability for each category, and the category with the largest prediction probability is selected as the evaluation category of the Parkinson's patient to be evaluated.

[0116] For the evaluation results, the accuracy of each category and the overall rate (Acc, the number of samples predicted correctly / the total number of samples), the acceptable accuracy (Acceptable Acc, the number of samples with a prediction error of no more than 1 / the total number of samples), the precision of each category and the average of the five categories, the recall rate, the F1 score and the receiver operating characteristic curve (ROC) and the area under the curve (AUC) can be reported. Among them, the acceptable accuracy reflects the actual clinical situation, that is, due to the subjective differences in the scores of different neurologists and the natural properties of the MDS-UPDRS scale [, the evaluation results predicted to be adjacent scores are often considered acceptable.

[0117] The following specific examples are used to verify the evaluation method of the Parkinson's disease standing task of the present invention.

[0118] Table 1 shows the evaluation indicators of the Parkinson's disease standing task evaluation method of the present invention for each category and overall prediction performance. Each category achieved good classification performance, and the accuracy rate was higher than 50%. Figure 6 (a) shows the ROC curve of the results shown in Table 2. The curves of each category are relatively close to the upper left corner, and the area under the curve (i.e., AUC) is greater than 0.80, which has a high prediction accuracy. It can be seen that the evaluation method of the Parkinson's disease standing task of the present invention achieves an overall accuracy of 70.60% and an overall acceptable accuracy of 98.65%. As shown in the confusion matrix of Figure 6 (b), most of the samples with differences in the predicted results and the true labels are divided into adjacent categories, which is consistent with the properties of MDS-UPDRS and the subjective scoring differences of different doctors.

[0119] Table 1. Performance parameters of the present invention

[0120]

[0121] In order to illustrate the contribution of each component of the SSM-GCN of the present invention, the network after removing SSIQL, VSGCO and GRSL from SSM-GCN is used as the baseline network, and the classification results of adding the three components to the baseline network are reported. The ablation comparison results are shown in Table 2. Obviously, each component has brought about an improvement in accuracy, which shows that these three components have made important contributions to improving the performance of the model. In addition, the present invention repeated the ablation comparison of each group ten times and performed a one-tailed paired sample t-test to verify the significance of the performance improvement brought by each component. The results are shown in Table 2. Figure 7 The experimental results show that the accuracy of the proposed method fluctuates very little and has good stability and reliability. Compared with the baseline network, the improvement of SSM-GCN in accuracy is very significant.

[0122] Table 2. Ablation experimental results of the present invention

[0123]

[0124] The performance of the present invention and the advanced skeleton-based action recognition methods on the standing assessment task are compared, and the classification results are reported in Table 3. Compared with the convolutional neural network (CNN)-based methods, the GCN-based methods have obvious advantages, indicating the adaptability and superiority of GCN in modeling human skeleton sequences. Among the GCN-based comparison methods, the SSM-GCN proposed in the present invention achieves the most advanced performance, outperforming ST-GCN (spatial temporal graph convolutional network), Js-AGCN (joint stream in two-stream adaptive graph convolutional network), Bs-AGCN (bone stream in two-stream adaptive graph convolutional network), 2s-AGCN (two-stream adaptive graph convolutional network) and Motif-STGCN (theme-based spatial temporal graph convolutional network) by 2.13%, 3.87%, 5.80%, 3.10% and 6.00% in accuracy, respectively, and always surpasses other comparison methods in other evaluation indicators.

[0125] Table 3. Comparison results of the present invention with other advanced skeleton-based action recognition models

[0126]

[0127] As shown in Table 4, the present invention is compared with previous work on automatic MDS-UPDRS scoring of Parkinson's patients' standing tasks. According to the investigation, the only two related works currently use sensors to collect motion data, and then use kNN to classify kinematic features, and finally achieve classification accuracy of approximately 65% ​​and 43% respectively, and the acceptable accuracy is also above 90%. However, it can be seen from Table 4 that compared with existing sensor-based methods, the present invention achieves competitive results in both accuracy and acceptable accuracy. Although the data sets are different, the number of samples in the data set used by the present invention is more than ten times higher, so more accurate and reliable performance is obtained on a larger data set.

[0128] Table 4. Comparison results of the present invention and other automatic evaluation studies on the standing-up task

[0129]

[0130] like Figure 8 As shown, in one embodiment, the Parkinson's disease standing task assessment system of the present invention includes an acquisition module 81 , an extraction module 82 , a training module 83 and an assessment module 84 .

[0131] The acquisition module 81 is used to acquire video information containing the standing-up action of a Parkinson's disease patient, where the standing-up action is a standard action required for performing the standing-up task assessment.

[0132] The extraction module 82 is connected to the acquisition module 81 and is used to extract the skeleton sequence of the Parkinson's patient from the video information, and extract joint information and bone information based on the skeleton sequence.

[0133] The training module 83 is connected to the extraction module 82, and is used to train a self-supervised metric graph convolutional network based on the joint information and the skeleton information; the self-supervised metric graph convolutional network is used to output the predicted probability of the standing up action in the video information for each category, including a joint flow network and a skeleton flow network; the joint flow network and the skeleton flow network each include 9 spatiotemporal units; each spatiotemporal unit is composed of spatial graph convolution and temporal convolution.

[0134] The evaluation module 84 is connected to the training module 83, and is used to obtain the predicted probability of the standing-up action of the Parkinson's patient to be evaluated for each category based on the trained self-supervised metric graph convolutional network, and select the category with the largest predicted probability as the evaluation category of the Parkinson's patient to be evaluated.

[0135] Among them, the structures and principles of the acquisition module 81, the extraction module 82, the training module 83 and the evaluation module 84 correspond one to one with the steps in the above-mentioned evaluation method of the Parkinson's disease standing task, so they are not repeated here.

[0136] It should be noted that it should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software called by processing elements, or they can all be implemented in the form of hardware, or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example: the x module can be a separately established processing element, or it can be integrated in a certain chip of the above device. In addition, the x module can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device. The implementation of other modules is similar. These modules can be integrated together in whole or in part, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in the processor element or instructions in the form of software. The above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), one or more microprocessors (DSP), one or more field programmable gate arrays (FPGA), etc. When a module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. These modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0137] The storage medium of the present invention stores a computer program, which, when executed by a processor, implements the above-mentioned Parkinson's disease standing task assessment method. Preferably, the storage medium includes: ROM, RAM, disk, USB flash drive, memory card or optical disk and other media that can store program codes.

[0138] like Fig. 9 As shown, in one embodiment, the evaluation terminal for the Parkinson's disease standing task of the present invention includes: a processor 81 and a memory 92 .

[0139] The memory 92 is used to store computer programs.

[0140] The memory 92 includes: ROM, RAM, disk, USB flash drive, memory card or optical disk and other media that can store program codes.

[0141] The processor 91 is connected to the memory 92 and is used to execute the computer program stored in the memory 92 so that the evaluation terminal for the Parkinson's disease standing up task executes the above-mentioned evaluation method for the Parkinson's disease standing up task.

[0142] Preferably, the processor 91 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0143] In summary, the evaluation method and system, storage medium and terminal of the Parkinson's disease standing task of the present invention can collect the patient's motion data using only ordinary cameras. This mode is easy to migrate to all scenarios, further improving the feasibility of remote evaluation; the data set comes from the evaluation videos of patients collected in clinical practice, the data set is large, and the evaluation performance is excellent; the self-supervised deep learning scheme used significantly enhances the model's ability to automatically extract spatiotemporal features from the standing video, and integrates VSGCO and GRSL, and the accuracy is much better than the traditional feature engineering method; the spatial and temporal knowledge of the action is used to construct positive and negative sample pairs, and the sensitivity of the model learning to space and time is simultaneously improved through the metric learning of the quadruple; by assigning specific adaptive weight parameters to each joint and each joint's different features, joint-specific spatial feature aggregation is achieved; by maximizing the consistency between the dual-stream feature representations, the similarity and discriminative learning of the dual-stream features in the classification task is promoted. Therefore, the present invention effectively overcomes the various shortcomings of the prior art and has a high industrial utilization value.

[0144] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A method for assessing the standing-up task in Parkinson's disease. Features: The following steps are involved: Acquiring video information containing standing-up movements of a Parkinson's disease patient, wherein the standing-up movements are standard movements required for standing-up task assessment; Extracting a skeleton sequence of a Parkinson's patient from the video information, and extracting joint information and bone information based on the skeleton sequence; A self-supervised metric graph convolution network is trained based on the joint information and the skeleton information; the self-supervised metric graph convolution network is used to output the prediction probability of the standing up action in the video information for each category, including a joint flow network and a skeleton flow network; the joint flow network and the skeleton flow network each include 9 spatiotemporal units; each spatiotemporal unit is composed of a spatial graph convolution and a temporal convolution; Based on the trained self-supervised metric graph convolutional network, the predicted probability of the standing-up action of the Parkinson's patient to be evaluated for each category is obtained, and the category with the largest predicted probability is selected as the evaluation category of the Parkinson's patient to be evaluated; Training a self-supervised metric graph convolutional network based on the joint information and the skeleton information comprises the following steps: Performing self-supervised intra-video quadruple learning on the joint stream network and the skeleton stream network based on the skeleton sequence, and applying constraints of graph representation supervision loss on output features of the joint stream network and the skeleton stream network; Using the parameters obtained by the self-supervised video quadruple learning as the initial parameters of the joint stream network and the skeleton stream network, training the joint stream network and the skeleton stream network based on the joint information and the skeleton information, and splicing the outputs of the joint stream network and the skeleton stream network and inputting them into a fully connected layer and a softmax; Self-supervised in-video quadruple learning of the skeleton sequence comprises the following steps: A data enhancement algorithm based on similarity transformation is used to generate positive samples that are similar to the global context of the skeleton sequence in time and space, and the skeleton sequence is used as an anchor sample; Based on a spatial global perturbation algorithm, generating spatial negative samples having the same temporal context as the skeleton sequence and a different spatial context; Based on a temporal global perturbation algorithm, generating temporal negative samples that are temporally chaotic with the skeleton sequence and have normal spatial context; The anchor samples, the positive samples, the spatial negative samples and the temporal negative samples are all input into the joint flow network and the skeleton flow network, so that the joint flow network and the skeleton flow network meet the preset loss function target.

2. The method for evaluating the Parkinson's disease standing task according to claim 1, Features: The skeleton sequence is extracted from the video information based on the human posture estimation model OpenPose, the joint information is extracted based on the skeleton sequence, and the bone information is extracted based on the coordinate difference between two naturally connected joints.

3. The method for evaluating the Parkinson's disease standing task according to claim 1, Features: The loss function aims to maximize the relative distance between anchor samples, positive sample pairs and anchor samples, temporal negative sample pairs, maximize the relative distance between anchor samples, positive sample pairs and temporal negative samples, spatial negative sample pairs, and make the minimum distance between temporal negative samples and spatial negative sample pairs greater than the maximum distance between anchor samples and positive sample pairs.

4. The method for evaluating the Parkinson's disease standing task according to claim 1, Features: When the joint flow network and the skeleton flow network are trained based on the joint information and the skeleton information, the cost function is composed of a cross entropy term and a graph representation supervision loss term.

5. The method for evaluating the Parkinson's disease standing task according to claim 1, Features: In the space-time unit, vertex-specific spatial graph convolution is adopted in the space-time unit with the same number of input and output channels.

6. An evaluation system for the Parkinson's disease standing task, Features: It includes acquisition module, extraction module, training module and evaluation module; The acquisition module is used to acquire video information containing standing-up movements of Parkinson's disease patients, where the standing-up movements are standard movements required for standing-up task assessment; The extraction module is used to extract a skeleton sequence of a Parkinson's patient from the video information, and extract joint information and bone information based on the skeleton sequence; The training module is used to train a self-supervised metric graph convolutional network based on the joint information and the skeleton information; the self-supervised metric graph convolutional network is used to output the prediction probability of the standing up action in the video information for each category, including a joint flow network and a skeleton flow network; the joint flow network and the skeleton flow network each include 9 spatiotemporal units; Each spatiotemporal unit consists of spatial graph convolution and temporal convolution; The evaluation module is used to obtain the predicted probability of the standing-up action of the Parkinson's patient to be evaluated for each category based on the trained self-supervised metric graph convolutional network, and select the category with the largest predicted probability as the evaluation category of the Parkinson's patient to be evaluated; Training a self-supervised metric graph convolutional network based on the joint information and the skeleton information comprises the following steps: Performing self-supervised intra-video quadruple learning on the joint stream network and the skeleton stream network based on the skeleton sequence, and applying constraints of graph representation supervision loss on output features of the joint stream network and the skeleton stream network; Using the parameters obtained by the self-supervised video quadruple learning as the initial parameters of the joint stream network and the skeleton stream network, training the joint stream network and the skeleton stream network based on the joint information and the skeleton information, and splicing the outputs of the joint stream network and the skeleton stream network and inputting them into a fully connected layer and a softmax; Self-supervised in-video quadruple learning of the skeleton sequence comprises the following steps: A data enhancement algorithm based on similarity transformation is used to generate positive samples that are similar to the global context of the skeleton sequence in time and space, and the skeleton sequence is used as an anchor sample; Based on a spatial global perturbation algorithm, generating spatial negative samples having the same temporal context as the skeleton sequence and a different spatial context; Based on a temporal global perturbation algorithm, generating temporal negative samples that are temporally chaotic with the skeleton sequence and have normal spatial context; The anchor samples, the positive samples, the spatial negative samples and the temporal negative samples are all input into the joint flow network and the skeleton flow network, so that the joint flow network and the skeleton flow network meet the preset loss function target.

7. A storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method for evaluating the Parkinson's disease standing task according to any one of claims 1 to 5 is implemented.

8. An evaluation terminal for Parkinson's disease standing task, It is characterized in that include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory, so that the evaluation terminal for the Parkinson's disease standing up task performs the evaluation method for the Parkinson's disease standing up task according to any one of claims 1 to 5.