Limb rehabilitation action recognition and evaluation method based on deep learning

Through deep learning technology combined with multimodal data, a hybrid model and SGP-Net network are constructed, which solves the shortcomings in the recognition and evaluation of limb rehabilitation movements in the existing technology, and achieves higher precision and fine-grained evaluation effects.

CN120108045AActive Publication Date: 2025-06-06XUZHOU MEDICAL UNIVERSITY +1

Patent Information

Application Number
CN202510587646.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The prior art has shortcomings in the recognition and evaluation of rehabilitation movements of patients with limb dysfunction, especially the IMU accuracy is difficult to guarantee, and the existing methods lack the ability to evaluate complex movements in a detailed manner.

Method used

Using a deep learning-based method, combining the multimodal information of RGB visual and skeleton model data, the features of video and skeleton data are extracted and action quality evaluation is performed by building a hybrid model and deep learning network SGP-Net.

Benefits of technology

It significantly improves the recognition accuracy and evaluation accuracy of limb rehabilitation movements, and can provide a finer-grained evaluation of subtle differences in rehabilitation movements, improving the accuracy and scientificity of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108045A_ABST
    Figure CN120108045A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer application, and relates to a limb rehabilitation action recognition and evaluation method based on deep learning, and the method comprises the steps: collecting video data and skeleton data of limb function rehabilitation actions, carrying out the preprocessing, and constructing a training set and a verification set of the video data and the skeleton data; a hybrid model of SE-ResBlock3D and LSTM is constructed, and recognition and classification of limb function rehabilitation actions are achieved on input video data through an execution convolution layer, four residual block sequences and an LSTM module in sequence by the hybrid model; and a deep learning network SGP-Net is constructed, and the deep learning network SGP-Net extracts and evaluates space and time sequence characteristics of the limb function rehabilitation actions by executing an STGCN module, a GAT module and a PMTC module on the input preprocessed skeleton data in sequence, so that quality evaluation of the limb function rehabilitation actions is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and specifically relates to a limb rehabilitation movement recognition and evaluation method based on deep learning. Background Art

[0002] At present, there are three main deficiencies in the known methods and devices for motion recognition and motion quality assessment: (1) Most of them are applied to healthy people, and there are few methods and devices for motion recognition and motion quality assessment for patients with limb dysfunction; (2) The application scenarios mainly include ordinary life scenarios (such as standing, walking, squatting, sitting, etc.) and sports scenarios (such as diving competitions, dancing, etc.), and lack application scenarios for rehabilitation training for patients with limb dysfunction; (3) Most existing methods are based on single-modal video data, while this method combines the multi-modal information of RGB vision and skeleton model data, significantly improving the algorithm performance and accuracy. In contrast, this method can provide a more fine-grained evaluation effect for rehabilitation movements with subtle differences.

[0003] paper( Mourchid, Youssef, and Rim Slama. "D-STGCNT: A Dense Spatio- Temporal Graph Conv-GRU Network based on transformer for assessment of patient physical rehabilitation." Computers in Biology and Medicine 165 (2023): 107420. ) proposed a new deep learning model— Dense Spatio-Temporal Graph Conv-GRU Network with Transformer (D-STGCNT ), for automatically assessing the quality of unsupervised physical rehabilitation exercises for patients. The model combines a modified version of STGCN and Transformer Architecture,efficiently handles spatiotemporal data, especially human skeleton data. D-STGCNT Leveraging dense connections and GRU Mechanism to quickly handle large 3D Skeleton input,effectively modeling temporal dynamics. Transformer The encoder's attention mechanism focuses on relevant parts of the input sequence, which is suitable for evaluating rehabilitation exercises. KIMORE and UI-PRMD Evaluation on the dataset shows that this method is superior to existing methods in terms of accuracy and computational time, achieving faster and more accurate learning and evaluation. The dataset used has few rehabilitation movements and cannot meet the needs of actual limb motor function rehabilitation, which makes it difficult for the model to cover all types of patients with limb dysfunction.

[0004] A rehabilitation action recognition method and system based on multimodal information fusion ( CN 117137435 A ) When the user is doing rehabilitation training, use IMU The collection device continuously collects the position data of 17 key nodes of the user's body and uses RGB-DThe camera continuously collects multiple action image data of the user; the multimodal data alignment algorithm is used to align the data in time and space to obtain multimodal data after time and space alignment; based on the multimodal data after time and space alignment, the action recognition algorithm of the lightweight modal screening decision network is used to identify the user's rehabilitation actions. IMU The collected data and RGB-D The data collected by the camera are aligned in time and space, providing reliable standards and basic data for subsequent rehabilitation movement recognition and quality assessment of rehabilitation movements.

[0005] The biggest disadvantage of this method is IMU The accuracy is difficult to guarantee. IMU Sensors can be affected by integration errors when operating for long periods of time, causing position data to drift. This drift can accumulate over time, affecting long-term tracking and assessment of motion.

[0006] A kind of Transformer Rehabilitation movement assessment method ( CN 117671787 A ) proposed a Transformer-based rehabilitation action quality assessment method, which effectively solved the problem of difficulty in learning the temporal relationship between long feature sequences. Transformer predictive models and use public datasets KIMORE The dataset is then used to train the model. KinectV2 The camera collects skeleton node data, and the skeleton node sequence of the rehabilitation action to be evaluated is input into the trained prediction model for prediction, and the final result is output. The disadvantages of this method are: ① Using the relatively old Kinect V2 Depth camera, not the latest model Kinect Azure DK The depth camera is insufficient in terms of resolution and number of skeleton points; ② The motion relationship between skeleton nodes is not fully considered and no ST-GCN or related algorithms to fully extract the motion relationship between skeleton nodes, but only use Transformer The algorithm is used to evaluate the movement quality; ③ Only single-modal data is used instead of multi-modal data, which makes it difficult to deeply obtain the movement characteristics behind the limb movement, limiting the accuracy of the limb movement quality evaluation.

[0007] Fitness action recognition method, device and related equipment based on human posture estimation ( CN202211531164.8 ) proposed a fitness action recognition method based on human posture estimation and its device, computer equipment and storage medium. The method mainly includes the following steps: ① Video stream processing and human key point extraction: extract human key point data from the video stream, through the target detection algorithm and human posture tracking algorithm (using the improved BlazeDark② Action recognition: The extracted human key point data is input into the trained fitness action classification network (multi-layer neural network) to detect and estimate the human body in the video frame, obtain the coordinates of the human body key points in the preset number and area, and perform serialization and centering processing. MLP Long Short-Term Memory Neural Network LSTM ③Similarity calculation: The similarity between the recognized action data and the standard action data of the corresponding category in the preset standard fitness action dictionary is calculated to obtain a similarity value. ④Action quality evaluation: Fitness action quality evaluation is performed based on the similarity value to obtain a quality evaluation result. The evaluation process may include normalization of the similarity value. The patent uses single-modal data to ensure classification and evaluation accuracy, and lacks the ability to finely evaluate complex actions. For example, the patent uses similarity calculation for action quality evaluation, which makes it difficult to fully capture the subtle differences in actions, resulting in insufficiently refined evaluation results. In addition, the adaptability to dynamic changes is poor, and the user's actions will have dynamic changes and irregularities, which may cause the action recognition model and similarity calculation method to perform poorly when facing non-standard actions, affecting the accuracy of fitness action quality evaluation.

[0008] Human action recognition technology HAR ) has important application value in the fields of rehabilitation, sports analysis and security monitoring. HAR Methods mainly rely on handcrafted features and classical machine learning techniques, which usually require rich domain knowledge and have limited generalization capabilities in different scenarios. With the development of deep learning technology, the research focus has shifted to using neural networks to automatically learn temporal and spatial features from raw videos to improve the accuracy and robustness of action recognition.

[0009] However, existing deep learning methods still have some problems when processing complex video data, such as insufficient feature extraction and insufficient capture of temporal dependencies. In addition, existing methods usually only focus on the classification of actions, while ignoring the evaluation of action quality. This is particularly disadvantageous for application scenarios such as medical rehabilitation that require accurate evaluation of action quality. Therefore, there is an urgent need for a new method that can both efficiently classify actions and accurately evaluate action quality. Summary of the invention

[0010] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method for limb rehabilitation movement recognition and evaluation based on deep learning.

[0011] In order to achieve the purpose of the present invention, the present invention will be implemented by adopting the following technical solutions.

[0012] A method for limb rehabilitation movement recognition and evaluation based on deep learning, comprising the following steps: S1. Preprocess the collected video data and skeleton data of limb function rehabilitation movements, and construct a training set and a validation set of the video data and skeleton data; S2, training and verifying the constructed hybrid model using the training set and validation set of the video data, and saving the optimized hybrid model for video feature extraction and classification of limb function rehabilitation movements to obtain correct video features of the movements; S3, use the training set and validation set of skeleton data to train and validate the constructed deep learning network SGP-Net, and save the optimized deep learning network SGP-Net to extract the spatial and temporal features of the evaluation of limb function rehabilitation movements; S4. Input the correct action video features obtained in step S2 and the spatial and temporal features extracted in step S3 into the action quality assessment module for quality assessment, and obtain personalized guidance and suggestions through the quality assessment report.

[0013] As a preferred solution of the present invention, the deep learning network SGP-Net is composed of a STGCN module including a GAT module and a PMTC module arranged in sequence; wherein: The STGCN module including the GAT module is used to capture temporal features while retaining the structural characteristics of spatial features. It dynamically calculates the attention weights between nodes and adaptively adjusts the connection weights between nodes. The STGCN module including the GAT module includes a jump positioning mechanism of preliminary convolution and two-level graph convolution, which performs feature extraction operations of the first jump and the second jump respectively. Among them: Preliminary convolution, using (9, 1) convolution kernel and ReLU activation function to process input features, capture basic spatiotemporal features, and merge them with the original input through concatenation operation; The jump positioning mechanism first performs a (1, 1) convolution operation on the features after the preliminary convolution, and then uses ConvLSTM2D The convolutional layer captures the temporal features and then passes GAT The module dynamically calculates the attention weights between nodes and adaptively adjusts the connection weights between nodes.

[0014] PMTC module, including dynamic convolution kernel generator, multi-scale convolution operation, spatiotemporal joint attention mechanism, time series modeling module and temporal convolution module, where: Dynamic convolution kernel generator, which generates convolution kernels with different odd dilation rates according to the temporal characteristics of the STGCN module input; Multi-scale convolution operations, in parallel with convolution kernels with different odd dilation rates, perform multi-level temporal feature extraction on the input features to form different feature maps; The spatiotemporal joint attention mechanism is used to perform weighted processing on the high-dimensional feature map concatenated from different feature maps to form multi-scale features; The time series modeling module is used to simultaneously capture the temporal and spatial features of the multi-scale feature vectors after integration and compression by the fully connected layer, and to perform bidirectional processing on the time series information of the multi-scale feature vectors; The temporal convolution module captures the changing trends in action sequences at different time scales through convolution kernels with different even dilation rates.

[0015] As a preferred embodiment of the present invention, the limb function rehabilitation movements are Brunnstrom graded movements, including 6 upper limb movements and 7 lower limb movements, wherein: the 6 upper limb movements include basket-carrying posture, touching the lumbar spine, arm flexion 90°, arm abduction 90°, arm flexion 180° and finger-nose test; the 7 lower limb movements include bilateral knee flexion, bilateral ankle dorsiflexion, bilateral heel sliding, standing knee flexion, bilateral ankle plantar flexion, hip abduction under knee extension and bilateral foot lateral sliding.

[0016] As a preferred solution of the present invention, the preprocessing of the video data comprises the following steps: S31, unify the video frame rate to 30 frames per second; S32, scaling the pixels of the original video data to 270x240; S33, performing frame extraction processing on RGB and Depth videos; S34, performing time alignment on each frame and the joint point data to maintain the temporal and spatial consistency of the data; S35, checking and recording the start and end time of each action, and extracting a frame sequence image containing the complete action; S36, using a Kalman filter to correct the noise of the data to improve the data quality; S37, downsampling the extracted image sequence, that is, retaining one frame of image every 10 frames, which can fully retain the complete semantic information of the action; S38, using horizontal flipping and noise reduction processing to perform data enhancement processing; S39, the resolution of the video frame is randomly cropped to 224x224, and all pixel values ​​are normalized to the range of 0-1; The input data format of the hybrid model of S310, SE-ResBlock3D and LSTM is unified to (3, 24, 224,224).

[0017] As a preferred embodiment of the present invention, the skeleton data is the three-dimensional space coordinate data of 32 joints of the human body, and the preprocessing of the three-dimensional space coordinate data of the 32 joints of the human body includes the following steps: S41, after the 3D spatial coordinate data of the 32 joint points are time-aligned with the video frames, every 100 frames are taken as an input; S42, when a motion data has multiple segments input, if the last segment is less than 100 frames, the first frame of the motion data is filled up to 100 frames; S43, according to the child-parent node correspondence table of 32 joints of the human body provided by the official website of Azure Kinect, the serial numbers of the 32 joints are linked to form a graph for use in the graph convolutional network; The input data format of S44 and SGP-Net models is unified to (100, 32, 3).

[0018] As a preferred solution of the present invention, the 3D convolutional neural network is composed of a 3D convolutional layer, a batch normalization layer and a ReLU activation function; wherein: the 3D convolutional layer has 32 convolution kernels, the size of the convolution kernel is (1,7,7), the step size is (1,2,2), and padding is performed in the height and width directions to maintain the size of the feature map.

[0019] As a preferred solution of the present invention, two adjacent residual sequences in the at least two consecutively arranged residual sequences are connected through a downsampling convolution layer, each residual sequence comprises at least three residual blocks, each residual block is composed of two consecutively arranged 3D convolution layers, and an SE module is arranged after each residual block; As a preferred solution of the present invention, the LSTM is used to capture long-term dependencies.

[0020] As a preferred solution of the present invention, the different odd expansion rates are d=3, d=5, d=7, d=9, d=11, and the number is 16.

[0021] As a preferred embodiment of the present invention, the different even-numbered expansion rates are d=2, d=4, d=8, and the number is 32.

[0022] Beneficial effects: This invention aims to achieve a more comprehensive and accurate evaluation of the motor rehabilitation of hemiplegic patients through deep learning and multimodal fusion analysis. We focus on two core issues: first, how to effectively fuse the RGB video and skeleton data obtained by the depth camera to improve the information richness and accuracy of rehabilitation action discrimination and evaluation; second, we focus on using advanced action quality evaluation methods to analyze the quality of rehabilitation patients in exercise execution, especially focusing on the differences between individuals in the rehabilitation stage. Ultimately, solving the above two problems will help improve the accuracy and scientificity of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1It is a structural diagram of the SE-R3D-LSTM module of the present invention; Figure 2 It is a structural diagram of the SGP-Net module of the present invention; Figure 3 It is a structural diagram of the PMTC module of the present invention; Figure 4 is an accuracy graph of the hybrid model of the present invention; Figure 5 is a loss graph of the hybrid model of the present invention; Figure 6 A confusion matrix diagram of the hybrid model of the present invention; Figure 7 The present invention STGCN Model prediction results graph; Figure 8 The present invention SGP-Net The network’s prediction results graph; Fig. 9 is a flow chart of the method of the present invention; Fig.10 A block diagram of processing raw video data in the method of the present invention; Fig.11 This is a skeleton data processing block diagram in the method described in the present invention. DETAILED DESCRIPTION

[0024] In order to facilitate the understanding of those skilled in the art, the present invention is further described below in conjunction with embodiments and drawings. The contents mentioned in the implementation modes are not intended to limit the present invention.

[0025] As an embodiment of the present invention, Figures 1 to 11 As shown, a method for limb rehabilitation movement recognition and evaluation based on deep learning includes the following steps: 1. Dataset description and preprocessing: We collected 13 different actions from 48 volunteers and selected Brunnstrom The staged movements are shown in Table 1, which include 6 upper limb and 7 lower limb movements. The age range of the volunteers was mainly 30 to 70 years old, with a balanced gender ratio.

[0026] Table 1

[0027] In the present invention, the action execution process of all volunteers was accurately recorded, including the start and end time of the action, and the frame sequence image containing the complete action was extracted. In order to increase the inter-frame distinction, the present invention downsampled the extracted image sequence, that is, one frame of image was retained every 10 frames. When the device frame rate is 30 frames per second ( FPS), the downsampling operation can fully retain the complete semantic information of the action. Finally, 575 action sequence images were sorted out, and the sequence length of each action ranged from 9 to 30 frames. The dataset is divided into a training set and a validation set, of which 85% of the data is used for model training and the remaining 15% is used to verify the model performance to ensure the generalization ability of the model.

[0028] In the early data processing stage, the present invention first scales the pixels of the original video data to 270x240, and performs a variety of data enhancement processes, such as horizontal flipping and noise reduction, to generate more robust training and test data. In order to further optimize the matching of data and neural network input, the resolution of the video frame is randomly cropped to 224x224, and all pixel values ​​are normalized to the range of 0-1. It is worth noting that the number of frames in the video sequence is uniformly adjusted to 24 frames after interval selection. For sequences with less than 24 frames, the last frame is copied to make it full to 24 frames to ensure that all input data has a consistent shape and size.

[0029] Finally, the input data format of the model was unified into (3, 24, 224, 224), which not only ensured the consistency and standardization of data input, but also improved the generalization ability of the model in different environments. These data processing steps laid a solid foundation for the training and verification of subsequent models, enabling the present invention to make full use of deep learning technology to automatically learn effective features from video data, thereby significantly improving the accuracy and robustness of human action recognition.

[0030] 2. Limb Function Rehabilitation Action Recognition and Classification Module The present invention proposes a method that combines SE-ResBlock3D and LSTM A new method for human action recognition and classification based on a hybrid model. The overall architecture of the model is as follows Figure 1 This model combines 3D Convolutional Neural Networks ( Convolutional Neural Network, CNN ), residual block ( Residual Block )、 SE ( Squeeze-and-Excitation ) modules and LSTM To achieve efficient video feature extraction and classification.

[0031] First, the video data is input into a 3D A convolutional layer with 32 filters, kernel size (1, 7, 7), stride (1, 2, 2), and padding in height and width to maintain the size of the feature map. The convolutional layer is followed by a batch normalization layer and ReLU The data then passes through the first residual block sequence, which consists of three Residual Block, each Residual Block It consists of two 3D convolutional layers and an SE module. The residual block is designed so that the input can be skipped to alleviate the gradient vanishing problem. Similarly, the feature map passes through the second and third residual block sequences in sequence, each of which contains multiple Residual Block , the number of channels is 64 and 128 respectively. And after each residual block sequence, the feature map passes through a downsampling convolution layer to further reduce the spatial size and increase the number of channels. Finally, the feature map passes through the fourth residual block sequence, which contains three Residual Block , the number of channels is 256. Then, the feature map undergoes global average pooling to reduce the spatial dimension to 1.

[0032] At the same time, the features are compressed to 256 dimensions through a 1D convolution layer, and the feature tensor is transposed so that the frame number dimension is at the front to meet LSTM The input requirements are LSTM It is used to capture long-term dependencies in videos. The features are input into a three-layer LSTM In the network, take the hidden state of the last time step of its output, through ReLU After the activation function, it enters the fully connected layer for classification to obtain the final probability distribution of each action category.

[0033] SE-ResBlock3D It is our main module for extracting spatial features, which combines 3D Convolution, residual connection and SE , through the synergy of these components to enhance the feature extraction capability. In specific implementation, a SE Module, recalculate channel weights and weight feature maps. This method not only enhances the model's focus on key features, but also demonstrates excellent computational efficiency and robustness, and can provide significant performance improvements in computer vision tasks such as video classification and action recognition, with wide compatibility and application potential.

[0034] 3. Spatial and temporal feature extraction and quality assessment module of limb function rehabilitation movements The action quality assessment module of the present invention is based on STGCN, GAT and PMTC Built an innovative deep learning model architecture SGP-Net , which is used to efficiently extract and evaluate the spatial and temporal features of human motion. By deeply integrating these advanced neural network modules, the model achieves accurate quality assessment of complex human motions and has broad application potential, especially in the fields of rehabilitation medicine and motion analysis.

[0035] SGP-Net The network structure, such as Figure 2 As shown, in the implementation process, first STGCN The module extracts features in spatial and temporal dimensions from the input action video data. STGCN The module uses a jump positioning mechanism of preliminary convolution and two-level graph convolution to perform the first jump and second jump feature extraction operations respectively. In the preliminary convolution, the model uses a (9, 1) convolution kernel with a number of 64 convolution kernels. ReLU The activation function processes the input features, captures the basic spatiotemporal features, and merges them with the original input through a concatenation operation, providing richer feature input for subsequent jump connections.

[0036] In each level of jump positioning, the model first performs a (1, 1) convolution operation on the input features, with 64 convolution kernels, which mainly serves to fuse and compress feature channels. ConvLSTM2D The convolution layer further captures the temporal features, and the number of convolution kernels is 32. ConvLSTM2D Combining the convolution operation with LSTM The temporal dependency modeling capability of this layer can maintain the structure of features in the spatial dimension while capturing dynamic changes in the temporal dimension. The output of this layer is GAT The module further processes GAT The module dynamically calculates attention weights between nodes through Leaky ReLU Activation function and Softmax Normalization processing adaptively adjusts the connection weights between nodes, so that the feature propagation process focuses more on important nodes and suppresses irrelevant noise.

[0037] Specifically, GAT The module performs a linear transformation on each node feature, calculates the similarity between nodes through a dot product operation, and normalizes the attention coefficient to dynamically weight feature propagation. This mechanism ensures that the model can adaptively select and aggregate features in complex node relationships, significantly enhancing the efficiency and accuracy of feature extraction in different scenarios. The output features of the two-level skip connection are merged through a splicing operation to form a rich feature set, which is used as a subsequent PMTC Input to the module.

[0038] PMTC The module is designed to solve the problem of extracting and analyzing temporal dynamic features in complex action sequences. Figure 3 As shown, PMTC The module greatly improves the model's ability to capture features at different time scales through the combination of dynamic convolution kernel generation, multi-scale convolution operations and spatiotemporal joint attention mechanism, and is particularly suitable for modeling complex actions over long time series.

[0039] first, PMTC Dynamic convolution kernel generator in the module ( Dynamic Kernel Generator)Generate convolution kernels with different expansion rates according to the input time features. The expansion rates are d=3, d=5, d=7, d=9, and d=11 respectively, to ensure that the features can be effectively captured at multiple time scales. The size of the convolution kernel is (3,1). Its original design intention is to ensure that the receptive field of the time dimension is wide enough and reduce the interference of the spatial dimension on time modeling. By dynamically generating convolution kernels, PMTC The module can adaptively adjust the perception range of the convolution kernel according to the input data, ensuring that both short-term and long-term temporal features can be effectively captured.

[0040] Next, the system extracts multi-level temporal features from the input data through parallel multi-scale convolution operations. Each convolution layer uses an independent dilation rate to perform convolution operations, ensuring that the model can handle both short-term and long-term temporal changes at multiple time scales. The convolutional features are then extracted through ReLU The activation function performs nonlinear mapping and passes Dropout Prevent model overfitting. The feature maps output by each convolutional layer are concatenated ( Concatenate ) to form a high-dimensional feature representation, ensuring that multi-scale features are preserved and providing input for the subsequent spatiotemporal joint attention mechanism.

[0041] After completing the multi-scale convolution, the spatiotemporal joint attention mechanism is introduced ( Spatio-Temporal Joint Attention ), this mechanism weights the concatenated feature maps to ensure that the model can focus on the most important features at different times and scales. The attention weights are calculated through spatiotemporal joint calculations, dynamically adjusting the weight relationship between features to ensure that the model focuses on the most representative temporal dynamic features. This adaptive weighting method can not only highlight the key changes in the action, but also effectively reduce the interference of irrelevant features, improving the robustness and accuracy of the model.

[0042] Next, the multi-scale features weighted by the spatiotemporal joint attention mechanism enter the fully connected layer ( Fully Connected, FC ) for further integration and compression. The fully connected layer compresses the multi-dimensional features into a more compact vector representation through nonlinear mapping, ensuring that the features can better represent the key action changes in the time series and provide more effective input for the subsequent time series modeling module.

[0043] Then, the feature vector processed by the fully connected layer is input into ConvLSTM2D and bidirectional LSTM ( Bi- LSTM ) module. ConvLSTM2DThe module further enhances the model's ability to model spatiotemporal dependencies by simultaneously capturing temporal and spatial features, and can accurately extract the dynamic changes of complex actions. Bi-LSTM The module enhances the model's ability to capture long-term dependencies by processing time series information bidirectionally, ensuring that key changes in action sequences can be fully identified and extracted. The module's bidirectional architecture allows the model to process both forward and backward information in action sequences simultaneously, ensuring high accuracy and robustness in action recognition and evaluation.

[0044] After the time series modeling is completed, the features are input into the temporal convolution module TCN middle. TCN The module uses multiple convolution kernels with different expansion rates (d=2, d=4, d=8) to capture the changing trends in action sequences at different time scales. Through multi-scale time modeling, TCN The module can effectively extract long-term dependencies, thus ensuring that the model has strong temporal dynamic analysis capabilities in the evaluation of long-sequence action quality. TCN The output features of the module provide highly precise temporal dynamic analysis for the comprehensive evaluation of motion quality, ensuring the accuracy and reliability of the motion quality evaluation results.

[0045] PMTC After the module completes feature extraction, the next step of the model processing will be PMTC The output multi-scale features are input to the (1,1) convolution layer for further feature compression and adjustment. At this point, the model integrates the high-dimensional features through the small convolution kernel and compresses them into a more compact feature vector, ensuring the computational efficiency of the subsequent stages and retaining key information. Next, these processed feature vectors are passed to the last module of the model - LSTM module. LSTM Used to capture long-term dependencies in video sequences, LSTM The model can effectively identify the long-term characteristics and complex temporal dynamic changes of actions in video sequences. This is crucial for the quality assessment of human actions, because the smoothness and coordination of actions often require comprehensive analysis through long-term sequence information. LSTM The module adopts a stacked layer design to ensure that the timing information of complex action sequences can be extracted and integrated layer by layer. LSTM The output of generates a high-dimensional feature vector containing the dynamic features of the complete action sequence. This feature vector provides detailed data support for subsequent evaluation tasks, ensuring that key indicators such as the smoothness, accuracy and consistency of the action can be fully reflected when evaluating the action quality.

[0046] The technical route of the present invention is through STGCN The module's powerful spatiotemporal feature extraction, GAT Adaptive attention weighting of modules, and PMTC The multi-scale dynamic modeling of the module enables a comprehensive evaluation of human motion. LSTM The module's time series processing capabilities enable the model to efficiently handle complex backgrounds and diverse action scenes. It is widely used in rehabilitation therapy, motion analysis and other fields, providing strong support for the industry's technological development.

[0047] More accurate limb function rehabilitation movement recognition and movement quality assessment algorithms have significant advantages in limb movement rehabilitation and home rehabilitation. It can help patients improve the effectiveness of rehabilitation training through personalized training plan optimization, real-time feedback and correction, avoid secondary injuries caused by improper movements, and quantify rehabilitation progress, so that rehabilitation doctors can evaluate the treatment effect. In terms of home rehabilitation, accurate recognition algorithms ensure that patients can receive high-quality rehabilitation guidance at home, reducing treatment costs. At the same time, through remote monitoring and support, it improves patient compliance and rehabilitation experience, making rehabilitation more reliable and efficient.

[0048] During the training and validation process, the proposed model verifies the accuracy and loss graphs on the validation set, as shown in Figure 4 and Figure 5 The verification accuracy of the proposed method is shown in epoch The initial validation loss of 0.3121 drops to a loss of 0.0001 at the end of the iteration. The results show that the proposed method is very effective, with extremely low training and validation losses and high accuracy during training and validation.

[0049] According to the experimental results, all categories perform well. For further analysis, we calculate the recall, accuracy, precision and F1 score of the proposed model on the test set, as shown in Table 2.

[0050] Table 2

[0051] The confusion matrix shows how the predictions of the hybrid model compare to the actual results. We can clearly see which categories the model performs well on and which categories it misclassifies. Figure 6 As shown, the confusion matrix shows accurate predictions along the diagonal, darker colors indicate that the proposed model has higher classification accuracy for the relevant class, while lighter colors indicate that there are misclassified samples. When tested on 13 human action classifications, the model achieved excellent results on all performance evaluation metrics.

[0052] Experimental results show that our method achieves high accuracy and robustness in 13 action classification tasks, verifying its effectiveness in complex video data processing. The evaluation results of the confusion matrix further prove the excellent performance of the model in various categories, especially in distinguishing similar actions. At the same time, the accuracy of distinguishing the first and fourth category actions needs to be further improved. In addition, our dataset covers different age groups and genders, ensuring the wide applicability and generalization ability of the model, showing great potential in practical applications. In general, the model proposed in this invention is not only significantly innovative in theory, but also shows great potential in practical applications.

[0053] IV. Evaluation of Model Prediction Performance This paper introduces a graph attention network ( GAT ) and spatial multi-scale temporal convolution ( PMTC ) modules, fused into a new network structure SGP-Net , compared with the traditional baseline model ( STGCN ), the performance on the test set has been significantly improved. In this invention, we use the mean absolute error ( MAD ), Root Mean Square Error ( RMS ), mean absolute percentage error ( MAPE ) and mean square error ( MSE ) These four evaluation indicators comprehensively evaluate the prediction performance of the model from different angles.

[0054] ; MAD It measures the average level of absolute error between predicted values ​​and actual values, reflecting the average size of the prediction error. MAD What is calculated is the average of the absolute values ​​of the differences between the predicted values ​​and the true values, so it can directly reflect the accuracy of the prediction, and the unit is the same as the original data.

[0055] ; RMS It is used to evaluate the difference between the predicted value and the actual value, and is obtained by taking the square root of the mean of the sum of squares of the errors. RMS It can amplify larger errors. The smaller its value is, the better the prediction effect of the model is.

[0056] ; MAPE It is the percentage of the estimated prediction error relative to the actual value, reflecting the proportion of the predicted error to the actual value. MAPE It is particularly suitable for comparing data of different magnitudes. By standardizing the error into a percentage, it facilitates error comparison between different data sets. MAPEThe smaller it is, the lower the relative error of the model is and the better the prediction effect is.

[0057] ; MSE Measures the average size of the squared forecast errors by calculating the square of the differences between the predicted and actual values ​​and taking the average. MSE It can amplify large prediction errors, is sensitive to outliers, and is suitable for model optimization and improvement. MSE The smaller the value, the higher the prediction accuracy of the model.

[0058] These evaluation indicators can comprehensively measure the accuracy, robustness and consistency of model predictions, and provide a multi-dimensional reference for model performance evaluation.

[0059] Table 3 shows the test set results. GAT and PMTC After the module, MAD From 0.03540 of the baseline model to 0.02727, RMS A significant decrease from 0.09582 to 0.07257, MAPE It dropped significantly from 0.40777 to 0.31336, MSE The overall optimization of these indicators shows that the present invention significantly improves the evaluation accuracy and generalization ability of the model for action quality.

[0060] MAD and RMS The significant decrease in RMS The decrease in indicates that the overall distribution of forecast errors is more concentrated, reducing the impact of large deviations. MSE The significant decrease in further confirms the effectiveness of the model in reducing large errors, and reflects the powerful ability of the model in processing complex action sequences after the introduction of the new module. MAPE The decrease in indicates that the improved model has significantly improved its ability to handle proportional errors, which means that the model can more accurately identify and evaluate the quality of actions and perform better in dealing with diverse action patterns.

[0061] Introduction GAT After the module, the model improves the focus on key features by adaptively adjusting the weights between nodes, thereby maintaining a high recognition accuracy in complex scenarios. GAT The dynamic weighting mechanism effectively suppresses the propagation of noise features, allowing the model to maintain robust performance on unseen test data. PMTCThe module further enhances the model's sensitivity to time series changes by accurately modeling temporal dynamic features, allowing the model to not only capture static feature information, but also track dynamic changes during the execution of actions. This dynamic feature extraction capability enables the model to accurately evaluate the quality of actions in complex time series data, especially in tasks that require distinguishing between action details and fluency.

[0062] Overall, the present invention introduces GAT and PMTC Module, integrated with the baseline model into a new network architecture SGP- Net After that, the significant performance improvement on the test set fully demonstrated the innovation and practicality of this technical route. This improvement not only significantly improved the evaluation accuracy of the model, but also enhanced the robustness and generalization ability of the model, enabling it to cope with complex and changing movement patterns in practical applications, and providing valuable technical support for fields such as rehabilitation medicine and movement analysis. Figure 7 and Figure 8 shown.

[0063] Table 3

[0064] The present invention provides a system and method based on multimodal deep learning for human motion classification and quality assessment, which is particularly suitable for rehabilitation therapy and motion analysis. The system achieves efficient and accurate classification and quality assessment of human motion through the collection and fusion of multimodal data and advanced deep learning algorithms. The specific implementation process is as follows: Fig. 9 As shown, the video data processing module is as follows Fig.10 As shown, the skeleton data processing module is as follows Fig.11 shown.

Claims

1. A method for limb rehabilitation movement recognition and evaluation based on deep learning, characterized by: The steps include: S1. Preprocess the collected video data and skeleton data of limb function rehabilitation movements, and construct a training set and a verification set of the video data and skeleton data; S2, training and verifying the constructed hybrid model using the training set and validation set of the video data, and saving the optimized hybrid model for video feature extraction and classification of limb function rehabilitation movements; S3, use the training set and validation set of skeleton data to train and validate the constructed deep learning network SGP-Net, and save the optimized deep learning network SGP-Net to extract the spatial and temporal features of the evaluation of limb function rehabilitation movements; S4. Input the video features obtained in step S2 and the spatial and temporal features extracted in step S3 into the motion quality assessment module for quality assessment, and obtain personalized guidance and suggestions through the quality assessment report.

2. The method for recognizing and evaluating limb rehabilitation movements based on deep learning according to claim 1, characterized in that: The deep learning network SGP-Net is composed of a STGCN module including a GAT module and a PMTC module arranged in sequence, wherein the STGCN module including the GAT module includes a jump positioning mechanism of preliminary convolution and two-stage graph convolution, and performs feature extraction operations of the first jump and the second jump respectively; wherein: Preliminary convolution, using (9, 1) convolution kernel and ReLU activation function to process input features, capture basic spatiotemporal features, and merge them with the original input through concatenation operation; The jump positioning mechanism first performs a (1, 1) convolution operation on the features after the preliminary convolution, and then uses ConvLSTM2D The convolutional layer captures the temporal features and then passes GAT The module dynamically calculates the attention weights between nodes and adaptively adjusts the connection weights between nodes.

3. The method for recognizing and evaluating limb rehabilitation movements based on deep learning according to claim 2, characterized in that: The PMTC module includes a dynamic convolution kernel generator, a multi-scale convolution operation, a spatiotemporal joint attention mechanism, a temporal modeling module, and a temporal convolution module, where: Dynamic convolution kernel generator, which generates convolution kernels with different odd dilation rates according to the temporal characteristics of the STGCN module input; Multi-scale convolution operation, in parallel with convolution kernels with different odd dilation rates, performs multi-level temporal feature extraction on the input features to form different feature maps; wherein: the different odd dilation rates are d=3, d=5, d=7, d=9, d=11, and the number is 16; The spatiotemporal joint attention mechanism is used to perform weighted processing on the high-dimensional feature map concatenated from different feature maps to form multi-scale features; The time series modeling module is used to simultaneously capture the temporal and spatial features of the multi-scale feature vectors after integration and compression by the fully connected layer, and to perform bidirectional processing on the time series information of the multi-scale feature vectors; The temporal convolution module captures the changing trends in the action sequence at different time scales through convolution kernels with different even dilation rates; wherein: the different even dilation rates are d=2, d=4, d=8, and the number is 32.

4. The method for identifying and evaluating limb rehabilitation movements based on deep learning according to claim 1, characterized in that: The limb function rehabilitation movements are Brunnstrom graded movements, including 6 upper limb movements and 7 lower limb movements, among which: the 6 upper limb movements include basket-carrying posture, touching the lumbar spine, arm flexion 90°, arm abduction 90°, arm flexion 180° and finger-nose test; the 7 lower limb movements include bilateral knee flexion, bilateral ankle dorsiflexion, bilateral heel sliding, standing knee flexion, bilateral ankle plantar flexion, hip abduction under knee extension and bilateral foot lateral sliding.

5. The method for identifying and evaluating limb rehabilitation movements based on deep learning according to claim 1, characterized in that: The preprocessing of the video data comprises the following steps: S31, unify the video frame rate to 30 frames per second; S32, scaling the pixels of the original video data to 270x240; S33, performing frame extraction processing on RGB and Depth videos; S34, performing time alignment on each frame and the joint point data to maintain the temporal and spatial consistency of the data; S35, checking and recording the start and end time of each action, and extracting a frame sequence image containing the complete action; S36, using a Kalman filter to correct the noise of the data to improve the data quality; S37, downsampling the extracted image sequence, that is, retaining one frame of image every 10 frames, which can fully retain the complete semantic information of the action; S38, using horizontal flipping and noise reduction processing to perform data enhancement processing; S39, the resolution of the video frame is randomly cropped to 224x224, and all pixel values ​​are normalized to the range of 0-1; The input data format of the hybrid model of S310, SE-ResBlock3D and LSTM is unified to (3, 24, 224, 224).

6. The method for identifying and evaluating limb rehabilitation movements based on deep learning according to claim 1, characterized in that: The skeleton data is the three-dimensional space coordinate data of 32 joints of the human body. The preprocessing of the three-dimensional space coordinate data of the 32 joints of the human body includes the following steps: S41, after the 3D spatial coordinate data of the 32 joint points are time-aligned with the video frames, every 100 frames are taken as an input; S42, when a motion data has multiple segments input, if the last segment is less than 100 frames, the first frame of the motion data is filled up to 100 frames; S43, according to the child-parent node correspondence table of 32 joints of the human body provided by the official website of Azure Kinect, the serial numbers of the 32 joints are linked to form a graph for use in the graph convolutional network; The input data format of S44 and SGP-Net models is unified to (100, 32, 3).

7. The method for identifying and evaluating limb rehabilitation movements based on deep learning according to claim 1, characterized in that: The hybrid model is composed of 3D Convolutional neural network, at least two consecutive residual sequences and LSTM , To achieve video feature extraction and classification.

8. The method for identifying and evaluating limb rehabilitation movements based on deep learning according to claim 7, characterized in that: The 3D convolutional neural network is composed of a 3D convolutional layer, a batch normalization layer and a ReLU activation function; wherein: the 3D convolutional layer has 32 convolution kernels, the size of the convolution kernel is (1,7,7), the step size is (1,2,2), and padding is performed in the height and width directions to maintain the size of the feature map.

9. The method for identifying and evaluating limb rehabilitation movements based on deep learning according to claim 7, characterized in that: Two adjacent residual sequences in the at least two consecutively arranged residual sequences are connected through a downsampling convolution layer, each residual sequence includes at least three residual blocks, each residual block is composed of two consecutively arranged 3D convolution layers, and an SE module is arranged after each residual block.

10. The method for identifying and evaluating limb rehabilitation movements based on deep learning according to claim 7, characterized in that: The LSTM is used to capture long-term dependencies.

Citation Information

Patent Citations

  • Fitness action recognition method, device and related equipment based on human posture estimation

    CN115880774B

  • Rehabilitation action recognition method and system based on multi-modal information fusion

    CN117137435A

  • Transform-based rehabilitation action evaluation method

    CN117671787A

  • Upper limb rehabilitation robot rehabilitation training motion function assessment method

    CN104997523A

  • Motion recognition method and system based on fusion graph convolutional network and Transform network

    CN115100574A

Cited By

  • Neural rehabilitation action detection method based on domain generalization neural network

    CN120804582A

  • A neural rehabilitation action detection method based on domain generalization neural network

    CN120804582B

  • Exercise rehabilitation data prediction system based on deep learning

    CN121483489A

  • Intelligent rehabilitation training evaluation method based on multi-scale space-time diagram convolutional network

    CN121709139A