Intelligent Rehabilitation Assistance Training System for Spinal Degenerative Diseases Based on Deep Learning

Through an intelligent rehabilitation assisted training system based on deep learning, combined with OpenPose and sparse optical flow tracking technology, the patients' traditional Chinese medicine guided surgery skeleton sequence is obtained and evaluated in real time, solving the problem of difficult to achieve automated, precise and intelligent rehabilitation training evaluation in the existing technology, and achieving efficient action recognition and correction.

CN114092854BActive Publication Date: 2025-06-03FUDAN UNIVERSITY

Patent Information

Application Number
CN202111295019.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-06-03
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

It is difficult for the prior art to achieve automated, precise and intelligent evaluation and correction of traditional Chinese medicine guidance rehabilitation training in patients with degenerative spinal lesions.

Method used

The intelligent rehabilitation assisted training system based on deep learning is adopted, combined with OpenPose two-dimensional pose estimation and sparse optical flow tracking, and the skeleton sequence is obtained in real time, and video behavior recognition and evaluation is carried out through the hierarchical limb attention LSTM network to realize action segmentation, error correction and scoring.

Benefits of technology

It has realized the automation, precision and intelligent evaluation and correction of traditional Chinese medicine guidance training for patients with degenerative spinal lesions, which has reduced the burden on medical staff and improved the patient's practice effect and flexibility of rehabilitation training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092854B_ABST
    Figure CN114092854B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent rehabilitation assistance training system for spinal degenerative diseases based on deep learning. The system of the present invention includes a real-time classification module for traditional Chinese medicine guiding techniques videos based on deep learning and a video sequence division and evaluation module based on human skeleton representation; the former obtains two-dimensional human skeleton data as the training data of the learning model, conducts deep learning training to obtain a generalized deep learning model, and finally obtains the real-time frame classification result; the latter, according to the frame classification result, segments and corrects the skeleton sequences of the same category in real time, and compares and scores the segmented sequence segments with the skeleton sequence segments of the expert group videos of the corresponding category. The system of the present invention does not require the guidance and intervention of medical staff, enables patients to perform traditional Chinese medicine guiding techniques training by themselves at any time, is applicable to families and primary medical and health institutions, can relieve the pressure of medical staff, and improve the flexibility and accuracy of patients' rehabilitation training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision video understanding, and specifically relates to a video sequence recognition and evaluation system based on deep learning. Background Art

[0002] Traditional Chinese Medicine (TCM) guiding techniques are a health preservation and treatment method with distinct Chinese characteristics. By guiding patients to perform body movement regulation on the basis of mental regulation and breathing regulation, that is, the main activities of the limbs, it promotes the recovery of limb movement. With the continuous development of computer hardware and artificial intelligence technologies, the research and development of intelligent assisted rehabilitation training systems has become a hot topic in related fields at home and abroad. In computer vision tasks, video behavior recognition is a highly challenging field. RGB frame input methods often fail to meet the real-time processing performance requirements, and skeleton sequence-based methods have lower time complexity but rely on the extraction of skeleton information during the inference process. The present invention uses a recurrent neural network based on skeleton sequence input as the basic network architecture for video behavior recognition, and combines two-dimensional pose estimation with sparse optical flow tracking through OpenPose [1] to obtain the skeleton sequence in real time, and finally performs skeleton sequence segmentation, patient exercise score evaluation, and movement correction reminder according to the classification results. The present invention combines the fields of computer vision and TCM guiding technique rehabilitation training, and develops an automatic evaluation and assisted training system for rehabilitation training actions for spinal degenerative diseases, realizing automated, precise, and intelligent rehabilitation training. Summary of the Invention

[0003] In order to reduce the burden on TCM rehabilitation medical staff when patients with spinal degenerative diseases perform rehabilitation training through TCM guiding techniques and improve the exercise effect of patients, the present invention provides an intelligent rehabilitation assistance training system for spinal degenerative diseases based on deep learning, which combines artificial intelligence technology to automatically evaluate and correct the TCM guiding technique exercises of patients.

[0004] The intelligent rehabilitation assistance training system for spinal degenerative diseases based on deep learning provided by the present invention is composed of a real-time classification module for TCM guiding technique videos based on deep learning and a video sequence partitioning and evaluation module based on human skeleton representation.

[0005] The real-time classification module for TCM guiding technique videos based on deep learning obtains two-dimensional human skeleton data through Openpose [1] as the training data of the deep learning model, performs supervised deep learning training to obtain a generalized deep learning model; then, through OpenPose frame-by-frame detection combined with sparse optical flow tracking, the input video is preprocessed in real time, and the result is fed as input into the pre-trained classification model to obtain the real-time frame classification result.

[0006] The human skeleton representation and video sequence segmentation and evaluation module segments and corrects the skeleton sequences of the same category in real time according to the frame classification results of the real-time classification module, and compares and scores the segmented sequence segments with the video skeleton sequence segments of the corresponding category expert group.

[0007] In the intelligent rehabilitation assistance training system for spinal degenerative diseases based on deep learning proposed by the present invention:

[0008] Corresponding to the real-time classification module of traditional Chinese medicine guiding exercise videos based on deep learning, its working content is as follows;

[0009] (1) Obtain training data for deep learning;

[0010] (2) Train a deep learning model;

[0011] (3) Use OpenPose to detect every other frame and combine sparse optical flow tracking to preprocess the video in real time and classify it;

[0012] Corresponding to the human skeleton representation and video sequence segmentation and evaluation module, its working content is as follows:

[0013] (4) Segment and correct the skeleton sequence in real time based on the classification results;

[0014] (5) Compare and score the sequence segments after segmentation;

[0015] The specific operation process for obtaining the deep learning training data described in content (1) is as follows:

[0016] (11) Process the video data. According to the design of the rehabilitation training exercise movements, all video data are clipped to obtain short video data with a length of no more than 1000 frames;

[0017] (12) Flip all the short video data horizontally left and right as data augmentation;

[0018] (13) Use the Openpose model pre-trained on the BODY_25 dataset to extract 2D poses from the processed video data samples in process (11), and obtain a skeleton data sequence represented by the 2D spatial coordinates of 25 key points. Among them, the 25 key points are Nose, Neck, RShoulder, RElbow, RWrist, LShoulder, LElbow, LWrist, MidHip, RHip, RKnee, Rankle, LHip, LKnee, LAnkle, REye, LEye, Rear, LEar, LBigToe, LSmallToe, LHeel, RBigToe, RSmallToe, Rheel, Background in the index order of 0-24;

[0019] (14) Perform data preprocessing on all skeleton data sequences. First, perform a translation operation to subtract the mid-hip coordinates of the first valid frame from the coordinates of all frames of the skeleton sequence; then perform a normalization operation to calculate the average shoulder-to-shoulder distance d of the skeleton sequence s , and then scale the coordinates of all frames of the skeleton sequence. The scaling factor is Finally, perform a zero-padding operation to fill all zeros at the end of the skeleton sequence and fix the length of the skeleton sequence to 1000 frames.

[0020] The specific process of training the deep learning model described in content (2) is as follows:

[0021] (21) The deep learning model uses a classification network model, specifically a hierarchical limb attention LSTM network, which includes a perspective transformation module, a limb attention module, a limb-level classification module, and a body-level classification module; the hierarchical limb attention LSTM takes the skeleton data as input. First, the perspective transformation module performs a 2D translation transformation on the input sequence coordinates. Then, each frame coordinate of the skeleton sequence is divided into 8 limb parts and respectively passes through the LSTM layer and the Dropout layer of the limb-level classification module; then the feature vectors of the 8 limb parts are concatenated and passed through the limb attention module to calculate the spatio-temporal attention weights of each limb; finally, the feature vectors corresponding to the 8 limbs are concatenated and weighted according to the spatio-temporal attention weights, and pass through the LSTM layer, the Dropout layer, and the softmax layer of the body-level classification module to finally obtain the classification scores of the skeleton sequence at each frame.

[0022] (22) Set the model hyperparameters;

[0023] The main hyperparameters in the model are: training conditions, batch size, learning rate, dropout rate, LSTM orthogonal initialization multiplication factor, maximum number of iterations;

[0024] (23) Start training. Based on the validation loss value of the model during training, when the validation loss value of the model no longer decreases and continues for 30 iterations, it indicates that the network has converged, and training ends;

[0025] (24) Adjust the hyperparameters multiple times to obtain the model with the best generalization performance;

[0026] Among them, the operation process of the perspective transformation module in the hierarchical limb attention LSTM network in step (21) is as follows:

[0027] To reduce the impact of changes in the shooting angle on the classification performance of the model, the input skeleton sequence is subjected to adaptive two-dimensional translation and rotation operations through the perspective transformation module, and the coordinates of each frame of the input skeleton sequence are adjusted. Among them, the specific calculation process is defined as:

[0028] S′ t,j =[x′ t,j ,y′ t,j ′=R t (S t,j -d t ), (1)

[0029] Among them, S t,j =[x t,j ,y t,j ′ represents the two-dimensional coordinates of the j-th key point of the t-th frame of the input skeleton sequence; represents the translation vector corresponding to the t-th frame, and R t represents the two-dimensional rotation matrix corresponding to the t-th frame, which is specifically expressed as:

[0030]

[0031] Among them, respectively represent the translation amounts along the horizontal and vertical axes of all coordinates of the t-th frame of the skeleton sequence, and α t represents the radian of counterclockwise rotation of all coordinates of the t-th frame of the skeleton sequence.

[0032] In process (21), the skeleton information passing through the perspective transformation module is divided into 8 limbs with a certain overlap according to the human body distribution, namely the head, left arm, right arm, torso, left leg, right leg, left foot, and right foot. Subsequently, the information of different limbs is calculated through separate LSTM layers and Dropout layers and then cascaded again into the overall skeleton information.

[0033] In process (21), the operation process of the limb attention module is as follows:

[0034] H t =LSTM(concat(H t,1 ,...,H t,L)) , (3)

[0035] a t = W 1 tanh(W 2 H t + b 2 ) + b 1

[0036]

[0037] Wherein, H t,i represents the i-th limb information of the t-th frame skeleton sequence, 1 ≤ i ≤ L, (L = 8), H t represents the feature information obtained after the 8 limb feature information is cascaded and passed through the LSTM layer and the Dropout layer. W 1 , W 2 is a learnable parameter matrix, b 1 , b 2 is a bias vector, and then the weight vector α of each limb is calculated through the softmax activation t,l , and finally the feature vector of each weighted limb is obtained, and all the weighted limb feature vectors are cascaded as the input of the subsequent module:

[0038] H' t,l = α t,l · H t,l , (5)

[0039] H' t = concat(H' t,1 ,..., H' t,L ), (6)

[0040] The cascaded weighted limb feature vectors pass through two LSTM layers, Dropout, and one fully connected layer, and then the classification scores are obtained through the softmax activation.

[0041] Content (3) uses the trained model to classify the behavior of the video to be classified. The specific operation process is as follows:

[0042] (31) Obtain and process the skeleton sequence from the input video signal source, and classify it. Use the method of combining frame-by-frame detection with sparse optical flow tracking by Openpose to obtain the skeleton sequence information in real time. Perform OpenPose pose estimation every 5 frames to obtain the coordinates of 25 human body key points. Among them, the 25 key points are respectively Nose, Neck, RShoulder, RElbow, RWrist, LShoulder, LElbow, LWrist, MidHip, RHip, RKnee, Rankle, LHip, LKnee, LAnkle, REye, LEye, Rear, LEar, LBigToe, LSmallToe, LHeel, RBigToe, RSmallToe, Rheel, Background in the index order of 0-24. Use the Lucas-Kanade method to track the two-dimensional coordinate information of 25 human body key points in each of the subsequent 4 frames. The specific process is as follows:

[0043] S t+1 = calOpticalFlowPyrLK(I t ,I t+1 ,S t ), (7)

[0044] Among them, S t ,I t respectively represent the grayscale image and skeleton information of the t-th frame. calOpticalFlowPyrLK is an implementation of the Lucas-Kanade method in the OpenCV open source library.

[0045] (32) The classification of the skeleton sequence frames is slightly different in the inference process and the training process. In the inference process, after obtaining 10 frames of skeleton information each time, preprocess and classify the skeleton information. The preprocessing process includes translation, normalization, and zero padding. Among them, the translation process no longer depends on the coordinates of the mid-hip (MidHip) of the first frame, but maintains a mid-hip coordinate as the origin during the operation of the system, and this coordinate is updated every 10 seconds; the normalization and zero padding processes are the same as in step (14).

[0046] (33) In the inference process, since classification is performed every 10 frames instead of for the entire video, during the operation of the system, it is necessary to make all LSTMs in the model maintain their parameter information after each classification, that is, turn on the LSTM stateful mode. In this mode, the hierarchical limb attention LSTM network will retain the states of all LSTM layers after each inference and use them as the initial state for the next inference.

[0047] In content (4), the skeleton sequence is segmented and corrected in real time based on the classification results, and the process is as follows:

[0048] (41) Skeleton sequence segmentation based on classification results:

[0049] During the operation of the system, the exercises performed by the patient generally include multiple different actions. In order to accurately compare the skeleton sequence information of the expert group of the corresponding category in the subsequent scoring process, real-time sequence segmentation needs to be performed after sequence frame classification. Specifically: at the first frame of the sequence or the end frame of the previous sequence segment is determined, the category of the current sequence segment is judged according to the classification result of the current frame, and the basis for determining the end of a sequence segment is either that the frame category is continuously inconsistent with the current sequence segment category and reaches the maximum error tolerance length (100 frames) or the input sequence ends.

[0050] (42) Real-time error correction based on classification results, and the process is as follows:

[0051] During the operation of the system, according to the skeleton information and classification result of each frame, different parameters are calculated and error correction text information is generated. The relationship between the action category and the calculation parameters is shown in Table 1.

[0052] Table 1, Reference parameter names for generating error correction text for different action categories

[0053] Action category Parameter name 0 Relative upward movement amplitude of both hands, angle between upper arm and lower arm 1 Relative amplitude of both hands spreading left and right, head rotation amplitude left and right 2 Relative vertical height of both elbows, relative horizontal width of both axes 3 Relative vertical height of both elbows, relative horizontal width of both axes 4 Relative downward extension amplitude of both hands 5 Ratio of relative downward extension amplitude of both hands to average calf length 6 Longitudinal distance between hip and neck 7 Ratio of difference in longitudinal distance between both feet to longitudinal distance from hip to neck

[0054] In content (5), the segmented sequence segments are compared and scored

[0055] During the operation of the system, whenever a new sequence segment is determined, sequence comparison and scoring are performed according to the category of the sequence segment. First, the system maintains a piece of skeleton sequence information Qn of the expert group for each action category. For a sequence segment C m , the similarity between the two sequences is calculated by the dynamic time warping (DTW) algorithm. The DTW algorithm, through the idea of dynamic programming, finds the best alignment path between the two sequences and calculates the sum of the Euclidean distances of the sequence frames on the path:

[0056]

[0057] Cost(i, j) = D(i, j) + min[Cost(i - 1, j), Cost(i, j - 1), Cost(i - 1, j - 1)], (9)

[0058] Among them, x, y represent the two-dimensional coordinates in the sequence segment, and D(i, j) represents the i-th frame of sequence Q n and the j-th frame of sequence C mThe Euclidean distance of the j-th frame, Cost(i, j) is the cumulative Euclidean distance sum at the (i, j) position of the optimal alignment path of the two sequences, and Cost(n, m) is the similarity of the two sequences. According to the similarity of the sequences, calculate the alignment score (in percentage system) between the current sequence segment and the expert group skeleton sequence. To make the score distribution more uniform, the scores are divided into 4 grades according to the similarity of the sequences, namely 90 - 100 (corresponding to similarity 0 - 2), 75 - 90 (corresponding to similarity 2 - 4), 60 - 75 (corresponding to similarity 4 - 6), 40 - 60 (corresponding to similarity 6 - ∞). For the similarity of each grade, the score calculation process is as follows:

[0059] score = low + (len - 10 * 2 cost-highCost ), (10)

[0060] where low and highCost represent the lowest score and the highest similarity of the current grade, len represents the score range length of the current grade, and cost represents the similarity of the sequence.

[0061] In summary, the present invention innovatively combines artificial intelligence technology with traditional Chinese medicine rehabilitation medicine to realize an intelligent rehabilitation assistance training system for patients with spinal degenerative diseases. It includes: (1) A real-time classification model for traditional Chinese medicine guiding exercise videos based on deep learning: The two-dimensional human skeleton data obtained by Openpose is used as the training data of the deep learning model for supervised deep learning training to obtain a generalized deep learning model. Then, the video is preprocessed in real time through Openpose frame-by-frame detection combined with sparse optical flow tracking, and the result is fed into the pre-trained deep learning model as input to obtain the real-time classification result; (2) Video sequence division and evaluation based on human skeleton representation: Based on the above real-time classification result, the skeleton sequences of the same category are segmented and corrected, and the segmented subsequences are scored. Among them, in the scoring process, the dynamic time warping algorithm is used to compare the subsequences with the expert group video sequences, calculate the video similarity and convert it into a percentage system score; in the correction process, based on the two-dimensional skeleton data and the classification result, reminder information is predefined according to the characteristics of the two-dimensional skeleton data for each action, and the patient is reminded in the form of voice broadcast when the predefined feature conditions are met. Experiments show that the real-time classification model proposed in the present invention achieves a good balance in terms of accuracy and speed performance, and the intelligent rehabilitation assistance training system for spinal degenerative diseases based on this classification model has high application value.

[0062] The present invention applies artificial intelligence technology to the field of rehabilitation training of traditional Chinese medicine guiding techniques, and promotes the recovery of limb motor function by automatically guiding the limb activities of patients on the basis of regulating the mind and breathing. Through technologies such as computer vision and deep learning, the present invention can, during the process of patients' training of traditional Chinese medicine guiding techniques, identify and evaluate the accuracy of patients' movement postures in real time, give corrections in a timely manner in the form of voice reminders, and score the patients' training according to the results of the comparison between the training process and the movements of the expert group after the patients' training ends. The intelligent rehabilitation assistance training system designed by the present invention for patients in the remission stage of spinal degenerative diseases can enable patients to perform traditional Chinese medicine guiding technique training by themselves at any time without the guidance and intervention of medical staff, is applicable to families and primary medical and health institutions, can greatly reduce the pressure on medical staff, and improve the flexibility and accuracy of patients' rehabilitation training. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is the overall flowchart of the present invention.

[0064] Figure 2 On the left is a schematic diagram of the joint points of the BODY_25 dataset in OpenPose, and on the right is a schematic diagram of dividing 25 key points into 8 parts of the limbs in the classification method of the present invention.

[0065] Figure 3 It is the architecture diagram of the hierarchical limb attention LSTM network of the present invention, which is divided into four modules from left to right: perspective transformation module, body level module, limb attention module, and limb level module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] In the present invention, the structure of the real-time classification model of traditional Chinese medicine guiding technique movements (hierarchical limb attention LSTM network) is as Figure 3 shown, where after the input skeleton is preprocessed, it only contains the two-dimensional coordinate information of 25 key points, as Figure 2 shown, where each skeleton sequence can be expressed as where T = 1000 and J = 25, representing the time length of the skeleton sequence and the key point data of the skeleton respectively.

[0067] After the original skeleton sequence data undergoes a preprocessing process including translation, normalization, and zero padding, it is input into the hierarchical limb attention LSTM for classification, and finally the classification scores of each frame of the skeleton sequence are output.

[0068] The implementation of the present invention includes two parts: the training of the classification model (hierarchical limb attention LSTM network) and the implementation of the intelligent rehabilitation assistance training system, and the specific implementation process is as follows:

[0069] (1) Preparation of the dataset

[0070] Due to the unique applicability of the present invention, we conducted experiments using our self - collected rehabilitation exercise dataset for spinal cord (RDSD: rehabilitation movements for spinal degenerative diseases). RDSD consists of 1012 videos in 9 action categories (8 exercise actions and a standing action), with the average length of each video being 18.4 seconds (30 frames per second). All videos were flipped horizontally and vertically as data augmentation, and finally divided into a training set and a test set in a ratio of 75:25.

[0071] (2) Data pre - processing

[0072] For the RDSD dataset proposed in the present invention, two - dimensional pose estimation was carried out through the open - source multi - person pose skeleton code library OpenPose [1] to extract the skeleton sequence, where each frame of the skeleton sequence consists of the two - dimensional spatial coordinates of 25 key points.

[0073] Based on each skeleton sequence, the present invention pre - processes the dataset, including three steps: translation, normalization, and zero padding. First, perform the translation operation, subtracting the hip - middle coordinates of the first valid frame from all coordinates of all frames of the skeleton sequence. Then, perform the normalization operation, calculate the average shoulder - to - shoulder distance d s of all frames of the skeleton sequence, and then scale the coordinates of all frames of the skeleton sequence with a scaling factor of Finally, perform the zero - padding operation, padding all zeros after the skeleton sequence with less than 1000 frames, and truncating the skeleton sequence with more than 1000 frames to ensure that the time length of all skeleton sequences is 1000 frames.

[0074] (3) Model training

[0075] The main hyperparameters in the model are: training conditions, batch size, learning rate, dropout rate, LSTM orthogonal initialization multiplication factor, and the maximum number of iterations.

[0076] In the present invention, the settings of the model's hyperparameters are as follows: training conditions: a single - block GTX1070 GPU; batch size: set to 64; learning rate: the initial learning rate is set to 0.005, and ReduceLROnPlateau is used as the learning rate adjustment strategy, that is, if the validation loss value of the model does not decrease every 10 iterations, the learning rate is reduced by 10 times; dropout rate: set to 0.1; LSTM orthogonal initialization multiplication factor: set to 0.001; the maximum number of iterations: 300. When the validation loss value of the model no longer decreases and continues for 30 iterations, the training is terminated early, and the total number of training times is generally more than 100 times.

[0077] (4) Experimental results

[0078] To study the effectiveness of each module in the classification network (Hierarchical Limb Attention LSTM network), the present invention conducts comparative experiments on adding / removing each model and comparing it with the baseline network. The baseline network is RNNs, which consists of three layers of LSTM + Dropout; HRNNs indicates that the network consists of one layer of Limb-level LSTM + Dropout and two layers of Body-level LSTM + Dropout; VT-RNNs adds a View-transformation Module as shown in Figure 3 to RNNs; VT-HRNNs adds a View-transformation Module as shown in Figure 3 to HRNNs; VT-HRNNs-ATT adds a Limb-Attention module as shown in Figure 3 to VT-HRNNs, which is the Hierarchical Limb Attention LSTM network of the present invention. As can be seen from Table 2, using a hierarchical structure (adding Limb-level) improves the test accuracy by 1.22% compared to using only Body-level LSTM. Adding the View-transformation Module improves the test accuracy by 1.04%, and adding the Limb-Attention module improves the test accuracy by 2.41%. This demonstrates the effectiveness of each module in the Hierarchical Limb Attention LSTM network of the present invention.

[0079] Table 2, Ablation experiments of each module of the Hierarchical Limb Attention LSTM network of the present invention on the RDSD dataset

[0080] Network structure Training accuracy Testing accuracy RNNs(Baseline) 95.30% 92.01% HRNNs 96.30% 93.23% VT-RNNs 95.43% 93.05% VT-HRNNs 96.06% 92.93% VT-HRNNs-ATT 96.64% 95.34%

[0081] (5) Implementation of the intelligent rehabilitation assistance training system

[0082] The intelligent rehabilitation assistance training system in the present invention includes the following four parts of functions: real-time acquisition and classification of skeleton sequences based on frame-by-frame detection of 0penPose[1] combined with sparse optical flow tracking, segmentation and real-time error correction of skeleton sequences, and scoring of skeleton sequence segments, as shown in Figure 1 .

[0083] The system receives the image frame signal source from an ordinary two-dimensional RGB camera and obtains the skeleton sequence in real time. Specifically, it runs OpenPose once every 5 frames for [1] two-dimensional pose estimation, and based on the obtained two-dimensional coordinates of 25 human body key points, obtains the two-dimensional coordinates of the human body key points in the next 4 frames through the Lucas-Kanade sparse optical flow tracking method.

[0084] Every time the latest 10-frame human skeleton sequence is obtained, the classification network (Hierarchical Limb Attention LSTM network) in the present invention is run once to obtain the action category labels of these 10 frames of sequences. Specifically, after preprocessing the 10-frame human skeleton sequence including translation, normalization, and zero-padding, it is fed into the pre-trained Hierarchical Limb Attention LSTM network to obtain the real-time classification result. Note that the Hierarchical Limb Attention LSTM network maintains the memory state of all LSTM units from the start of the system operation and clears the memory state every 30 s.

[0085] For each frame of the skeleton sequence, according to its key point two-dimensional coordinate information and the action category of this frame, the corresponding text error correction information is generated to prompt the patient with the key points of the current exercise method. Specifically, in the present invention, different reference parameters are defined for 8 exercise method actions as shown in Table 1, which are used to dynamically calculate and generate text information and remind the user by voice broadcast.

[0086] The system dynamically manages all classified frames that have not been segmented into action segments, that is, action segment segmentation. Specifically, the category of the action segment is determined according to the action categories of several latest unclassified frames, and the inconsistency of the action category is tolerated for up to 100 frames.

[0087] For the segmented action segments, the system selects the corresponding pre-loaded expert group video skeleton sequence information according to their action categories, calculates the similarity distance between the current action segment and the expert group action segment through the dynamic time warping algorithm, and then calculates the percentage score of the current action segment according to the predefined distance-score conversion rule.

[0088] The system is implemented based on the Python3 language, mainly using public code libraries such as TensorFlow, OpenCV, and multi-process processing, and runs in real time on a portable min-PC in a BS manner, which has high intelligent auxiliary application value for the rehabilitation training of patients with spinal degenerative diseases.

[0089] References

[0090] [1]Cao Z, Hidalgo G, Simon T, et al. OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018.

[0091] [2]Donald J Berndt and James Clifford.1994.Using Dynamic Time Wrapingto Find Patterns in Time Series.Proceedings of the AAAI Conference onArtificial Intelligence(AAAI).359- 370pages。

Claims

1. An intelligent rehabilitation assistance training system for spinal degenerative diseases based on deep learning, characterized in that, it consists of a real-time classification module for traditional Chinese medicine guiding technique videos based on deep learning and a video sequence division and evaluation module based on human skeleton representation; The real-time classification module for traditional Chinese medicine guiding technique videos based on deep learning obtains two-dimensional human skeleton data through Openpose as the training data of the deep learning model, conducts supervised deep learning training, and obtains a generalized deep learning model; Then, through OpenPose frame-by-frame detection combined with sparse optical flow tracking, the input video is preprocessed in real time, and the result is fed as input into the pre-trained classification model to obtain the real-time frame classification result; The video sequence division and evaluation module based on human skeleton representation and segmentation, according to the frame classification result of the real-time classification module, segments and corrects the skeleton sequences of the same category in real time, and compares and scores the segmented sequence segments with the skeleton sequence segments of the expert group videos of the corresponding category; among them, for the skeleton sequence of each frame, according to its key point two-dimensional coordinate information and the action category of this frame, the corresponding text error correction information is generated to prompt the patient of the key points of the current exercise.

2. The intelligent rehabilitation assistance training system for spinal degenerative diseases based on deep learning according to claim 1, characterized in that: Corresponding to the real-time classification module for traditional Chinese medicine guiding technique videos based on deep learning, its working content is as follows; (1) Obtain the training data of deep learning; (2) Train the deep learning model; (3) OpenPose frame-by-frame detection combined with sparse optical flow tracking preprocesses the video in real time and classifies it; Corresponding to the video sequence division and evaluation module based on human skeleton representation, its working content is as follows: (4) Segment and correct the skeleton sequence in real time based on the classification result; (5) Compare and score the segmented sequence segments; In content (1), the specific operation process of obtaining the deep learning training data is as follows: (11) Process the video data, and according to the design of the rehabilitation training exercise actions, clip all the video data to obtain short video data with a length of no more than 1000 frames; (12) Flip all the short video data horizontally and vertically as data augmentation; (13) Use the Openpose model pre-trained on the BODY_25 dataset to extract the two-dimensional pose of the video data samples processed in process (11), and obtain a skeleton data sequence represented by the two-dimensional spatial coordinates of 25 key points. Among them, the 25 key points are indexed in the order of 0-24 as Nose, Neck, RShoulder, RElbow, RWrist, LShoulder, LElbow, LWrist, MidHip, RHip, RKnee, Rankle, LHip, LKnee, LAnkle, REye, LEye, Rear, LEar, LBigToe, LSmallToe, LHeel, RBigToe, RSmallToe, Rheel, Background; (14)Perform data preprocessing on all skeleton data sequences. First, perform a translation operation by subtracting the hip center coordinates of the first valid frame from the coordinates of all frames in the skeleton sequence. Then, perform a normalization operation by calculating the average shoulder-to-shoulder distance d of the skeleton sequence s , and then scale the coordinates of all frames in the skeleton sequence. The scaling factor is Finally, perform a zero-padding operation by filling all zeros at the end of the skeleton sequence to fix the length of the skeleton sequence to 1000 frames. In content (2), the training of the deep learning model specifically includes: (21) The deep learning model uses a classification network model, specifically a hierarchical limb attention LSTM network, which includes a perspective transformation module, a limb attention module, a limb-level classification module, and a body-level classification module; The hierarchical limb attention LSTM takes skeleton data as input. First, the perspective transformation module performs a two-dimensional translation transformation on the input sequence coordinates. Then, each frame coordinate of the skeleton sequence is divided into 8 limb parts, and each passes through the LSTM layer and Dropout layer of the limb-level classification module; Then, the feature vectors of the 8 limb parts are concatenated and passed through the limb attention module to calculate the spatio-temporal attention weights of each limb; Finally, the feature vectors corresponding to the 8 limbs are concatenated and weighted according to the spatio-temporal attention weights, and pass through the LSTM layer, Dropout layer, and softmax layer of the body-level classification module to finally obtain the classification scores of the skeleton sequence at each frame. (22) Set the model hyperparameters; The main hyperparameters in the model are: training conditions, batch size, learning rate, dropout rate, LSTM orthogonal initialization multiplication factor, maximum number of iterations; (23) Start training. Based on the validation loss value of the model during training, when the validation loss value of the model no longer decreases and continues for 30 iterations, it means that the network has converged and the training ends; (24) Adjust the hyperparameters multiple times to obtain the model with the best generalization performance; In content (3), use the trained model to classify the behavior of the video to be classified. The specific operation process is as follows: (31) Obtain and process the skeleton sequence and classification from the input video signal source. Use the method of combining Openpose frame-by-frame detection with sparse optical flow tracking to obtain the skeleton sequence information in real time. Perform OpenPose pose estimation every 5 frames to obtain the coordinates of the 25 human body key points; Use the Lucas-Kanade method to track the two-dimensional coordinate information of the 25 human body key points in each of the subsequent 4 frames. The specific process is as follows: S t+1 = calOpticalFlowPyrLK(I t , I t+1 , S t ), Among them, S t , I t respectively represent the grayscale image and the skeleton information of the t-th frame, and calOpticalFlowPyrLK is an implementation of the Lucas-Kanade method in the OpenCV open source library; (32) The classification of the skeleton sequence frames is slightly different during the inference process and the training process. During the inference process, after every 10 frames of skeleton information are obtained, a preprocessing and classification of the skeleton information are performed. The preprocessing process includes translation, normalization, and zero padding; (33) During the inference process, make all LSTM in the model retain their parameter information after each classification, that is, turn on the LSTM stateful mode. In this mode, the hierarchical limb attention LSTM network retains the states of all LSTM layers after each inference and uses them as the initial state for the next inference; In content (4), based on the classification results, the skeleton sequence is segmented and corrected in real time. The process is as follows: (41) Skeleton sequence segmentation based on classification results: Perform real-time sequence segmentation after sequence frame classification, specifically: when it is the first frame of the sequence or the end frame of the previous sequence segment is determined, judge the category of the current sequence segment according to the classification result of the current frame. The basis for determining the end of a sequence segment is either that the frame category is continuously inconsistent with the category of the current sequence segment and reaches the maximum error tolerance length or the input sequence ends; (42) Real-time error correction based on the classification result, and its process is as follows: According to the skeleton information and classification result of each frame, calculate different parameters and generate error correction text information; In content (5), compare and score the sequence segments that have ended segmentation, including: Whenever a new sequence segment is determined, sequence alignment scoring is performed according to the category of the sequence segment; first, the system maintains a piece of expert group skeleton sequence information Q for each action analogy n , for a sequence segment C m , the similarity of the two ends of the sequence is calculated by the dynamic time warping (DTW) algorithm. The DTW algorithm, through the idea of dynamic programming, finds the best alignment path of the two sequences and calculates the sum of the Euclidean distances of the sequence frames on the path: Cost(i, j) = D(i, j) + min[Cost(i - 1, j), Cost(i, j - 1), Cost(i - 1, j - 1)], Among them, x and y represent two-dimensional coordinates in the sequence segment, and D(i, j) represents the Euclidean distance between the i-th frame of sequence Q n and the j-th frame of sequence C m The cumulative Euclidean distance sum at the (i, j) position of the optimal alignment path of the two sequences is Cost(i, j), and the similarity of the two sequences is Cost(n, m); according to the similarity of the sequences, calculate the alignment score between the current sequence segment and the expert group skeleton sequence, using a percentage system; in order to make the score distribution more uniform, the scores are divided into 4 levels according to the similarity of the sequences: namely 90 - 100, corresponding to a similarity of 0 - 2; 75 - 90, corresponding to a similarity of 2 - 4; 60 - 75, corresponding to a similarity of 4 - 6; 40 - 60, corresponding to a similarity of 6 - ∞. For each level of similarity, the score calculation process is as follows: score = low+(len - 10 * 2 cost-highCost ) where low and highCost represent the lowest score and the highest similarity of the current file, len represents the length of the score range of the current file, and cost represents the similarity of the sequence.

3. The intelligent rehabilitation assistance training system for spinal degenerative diseases based on deep learning according to claim 2, characterized in that the operation process of the perspective conversion module in the hierarchical limb attention LSTM network in step (21) is as follows: In order to reduce the influence of the change of the shooting angle on the classification performance of the model, the input skeleton sequence is subjected to adaptive two-dimensional translation and rotation operations through the perspective conversion module, and the coordinates of each frame of the input skeleton sequence are adjusted. Among them, the specific calculation process is defined as: S′ t,j = [X′ t,j , y′ t,j ′ = R t (S t,j - d t ), Among them, S t,j = [x t,j , y t,j ′ represents the two-dimensional coordinates of the j-th key point in the t-th frame of the input skeleton sequence; represents the translation vector corresponding to the t-th frame; R t represents the two-dimensional rotation matrix corresponding to the t-th frame, specifically expressed as: Among them, respectively represent the translation amounts of all coordinates of the t-th frame of the skeleton sequence along the horizontal axis and the vertical axis, and α t represents the radian of counterclockwise rotation of all coordinates of the t-th frame of the skeleton sequence; The skeleton information passing through the perspective conversion module is divided into 8 limbs with a certain overlap according to the human body distribution, namely the head, left arm, right arm, torso, left leg, right leg, left foot, and right foot. Subsequently, the information of different limbs is calculated through separate LSTM layers and Dropout layers and then cascaded again into the overall skeleton information.

4. The intelligent rehabilitation assistance training system for spinal degenerative diseases based on deep learning according to claim 3, characterized in that the operation process of the limb attention module in the hierarchical limb attention LSTM network in step (21) is as follows: Among them, H t,i represents the i-th limb information of the t-th frame skeleton sequence, 1 ≤ i ≤ L, (L = 8), H t represents the feature information obtained after the 8 limb feature information is cascaded and passed through the LSTM layer and the Dropout layer. W 1 , W 2 is a learnable parameter matrix, b 1 , b 2 is a bias vector, and then the weight vector α of each limb is calculated through the softmax activation t,l , and finally the feature vector of each weighted limb is obtained, and all the weighted limb feature vectors are cascaded as the input of the subsequent module: H′ t,l = α t,l · H t,l , H′ t = ConCat(H′ t,1 , …, H′ t,L ), The cascaded weighted limb feature vectors pass through two LSTM layers, Dropout, and one fully connected layer, and then obtain the classification score through softmax activation.

5. The intelligent rehabilitation assistance training system for spinal degenerative diseases based on deep learning according to claim 3, characterized in that in process (42), according to the skeleton information and classification result of each frame, calculate different parameters, and the relationship between the action category and the calculated parameters is as follows: The action categories are: 0, 1, 2, 3, 4, 5, 6, 7; the corresponding parameter names are: the relative upward lifting amplitude of both hands, the angle between the upper arm and the lower arm; the relative amplitude of the left and right expansion of both hands, the left and right rotation amplitude of the head; the relative longitudinal height of both elbows, the relative transverse width of both axes; the relative longitudinal height of both elbows, the relative transverse width of both axes; the relative downward extension amplitude of both hands; the ratio of the relative downward extension amplitude of both hands to the average length of the lower legs; the longitudinal distance between the buttocks and the neck; the ratio of the longitudinal distance difference between both feet to the longitudinal distance between the buttocks and the neck.

Citation Information

Patent Citations

  • Video object tracking method based on feature optical flow and online ensemble learning

    CN102903122A

  • Behavior classification method based on skeleton and video feature fusion

    CN112560618A

Cited By

  • Spinal assist intervention device and method based on detection of knee joint load distribution

    CN122701316A