Exercise rehabilitation evaluation method and system based on limb posture and emotion recognition
Through a deep learning model that monitors the changes in limb posture and mood in real time, the accuracy and convenience of rehabilitation assessment in the existing technology are solved, and a more comprehensive rehabilitation effect assessment and treatment quality improvement are achieved.
Patent Information
- Application Number
- CN202510215961.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, when evaluating the rehabilitation effect of patients with chronic diseases and limb dysfunction caused by accidental injuries, there are problems such as limited medical resources, low patient compliance, and lack of real-time feedback, resulting in poor rehabilitation results.
Through computer vision technology, the patient's limb posture and mood changes are monitored in real time, and the rehabilitation training video is analyzed using deep learning models, the key points and emotional characteristics of the limb bones are extracted, and the characteristics are fusion are carried out to form quantitative rehabilitation evaluation results.
It provides a more comprehensive rehabilitation assessment, which improves the accuracy and convenience of the assessment. Patients can conduct professional evaluations at home, enhances the accessibility of rehabilitation services, and supports doctors to make scientific decisions and improves the quality of rehabilitation treatment.
Smart Images

Figure CN120340110A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a motion rehabilitation evaluation method and system based on limb posture and emotion recognition, and belongs to the technical fields of video analysis and pattern recognition. Background Art
[0002] In modern society, chronic diseases such as cardiovascular and cerebrovascular diseases and stroke have become the main causes of death globally. Especially in northern Chinese cities, the incidence of stroke is relatively high, resulting in a large number of patients with limb motor dysfunction. These chronic disease patients and those with limb dysfunction caused by accidental injuries need long-term physical function rehabilitation training to restore muscle strength, balance, and mobility. However, problems such as limited medical resources, low patient compliance, and lack of real-time feedback have seriously hindered the improvement of rehabilitation effects.
[0003] In recent years, the emergence of rehabilitation robots, motion sensors, and somatosensory game technologies has provided new ways for physical rehabilitation evaluation and guidance. For example, the patent with the application number 202110326924.0 discloses a limb rehabilitation training and evaluation system based on visual tracking control; the method described in the above patent relies on specific devices. At the same time, the similarity algorithm it uses may not accurately identify certain abnormalities in complex gait patterns. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a motion rehabilitation evaluation method and system based on limb posture and emotion recognition, which can monitor and evaluate the rehabilitation training effect of patients in real time through computer vision technology. The system provides accurate and quantitative data for medical professionals by analyzing the limb postures and emotional changes of patients during the movement process.
[0005] To achieve the above purpose, the present invention is implemented by the following technical solutions:
[0006] In the first aspect, the present invention provides a motion rehabilitation evaluation method based on limb posture and emotion recognition, including the following steps:
[0007] Collect the action videos of patients during rehabilitation training, record the age and gender information of the patients, and construct a self-made dataset;
[0008] Obtain and process the video stream, define the candidate regions containing the rehabilitation actions of the patients according to the self-made dataset, and construct a basic action posture dataset;
[0009] According to the age and gender information of the patients, set personalized key point extraction parameters and weights corresponding to different age stages and genders, and extract personalized limb bone key points based on the basic action posture dataset and the pose estimation algorithm.
[0010] Based on the emotion recognition model, according to the basic action gesture dataset, emotion features are obtained through facial key point detection and facial action unit analysis;
[0011] Based on the spatio-temporal graph convolutional network model, according to the basic action gesture dataset and the personalized limb bone key points, the movement trajectory of the knee joint is analyzed, and limb pose information is extracted;
[0012] The limb pose information and emotion features are fused to form a quantitative rehabilitation assessment result.
[0013] Furthermore, the obtaining and processing of the video stream to define a candidate region containing the patient's rehabilitation actions includes:
[0014] Based on the target detection network model Faster R-CNN, according to the video stream, each frame of image is obtained in real time as an input image , and the input image is input into the feature extraction network. In the feature extraction network, the input image is processed layer by layer using multiple convolutional layers , and a deep feature map containing the patient's limb contour, pose, and action trajectory information is output;
[0015] The deep feature map is input into the region proposal network RPN. After obtaining local features through convolutional layers, the region proposal network RPN outputs candidate boxes, and the human body region in the image is framed by the candidate boxes, thereby defining the candidate region containing the patient's rehabilitation actions.
[0016] Furthermore, the deep feature map output by the feature extraction network is , where k is the number of candidate regions;
[0017] The feature extraction process is expressed as: , where represents the feature extraction function of the feature extraction network, represents the obtained deep feature map, and the deep feature map is used to capture the basic contour information of the human body.
[0018] Furthermore, it also includes: embedding a channel and spatial attention module CBAM in the region proposal network RPN to filter out irrelevant background information:
[0019] Channel attention module: generating channel weights , as follows:
[0020] ;
[0021] where is an activation function that enhances the attention to the patient area by integrating the information of global pooling;
[0022] Spatial attention module: Apply spatial weights to the feature map after channel attention processing as follows: as follows:
[0023] ;
[0024] The feature map processed by the channel and spatial attention module CBAM is defined as:
[0025] .
[0026] Furthermore, it also includes:
[0027] Add a balance mechanism to the region proposal network RPN to construct an improved region proposal network RPN;
[0028] Input the feature map into the improved region proposal network RPN. After obtaining local features through the convolutional layer, use RPN to output candidate boxes to frame the human body region in the image.
[0029] Furthermore, according to the age and gender information of the patient, corresponding to different age stages and genders, set personalized key point extraction parameters and weights, and based on the basic action pose dataset, extract personalized limb bone key points based on the pose estimation algorithm, including:
[0030] Perform real-time processing on each frame in the video stream through the Transformer-based pose estimation algorithm:
[0031] Use the Transformer-based pose estimation algorithm to detect the human key points in each frame, and output the key point coordinates of each frame as where represents the three-dimensional coordinates of the human key points in the th frame, is the number of key points of the human body;
[0032] The key point data is represented as: where, is the three-dimensional coordinates of the th key point;
[0033] For each key point, set an initial weight; use the Xavier initialization method to randomly initialize the weights of all key points;
[0034] During the training process, the attention mechanism is used to dynamically adjust the weights of key points; according to age and gender information, attention weights are assigned to each key point, as shown in the following formula:
[0035] ;
[0036] Where: is the attention score of the -th key point, and the input is the key point position , the age range of the patient and gender ; is the adaptive weight of the -th key point.
[0037] Furthermore, based on the basic action posture dataset, through facial key point detection and facial action unit analysis, features related to emotions are extracted to obtain emotion features, including:
[0038] Based on facial key point recognition and quantification of facial action units, emotion features are output through the emotion recognition model OpenFace ;
[0039] is a multi-dimensional vector representing information including facial action unit AU and head pose, as shown in the following formula:
[0040] ;
[0041] Among them, AU is the facial action unit, is the number of facial action units.
[0042] Furthermore, based on the basic action posture dataset, the movement trajectory of the knee joint is analyzed based on the personalized limb bone key points, and limb posture information is extracted, including:
[0043] The joint position data in consecutive frames is processed by the spatio-temporal graph convolutional network model ST-GCN to extract the motion features of the joints;
[0044] Among them, the rate of change of the joint position is the speed , and the calculation formula is as follows:
[0045] ;
[0046] Where, is the position of the joint at the -th frame moment; is the time interval; is the motion speed at the -th frame moment;
[0047] Extract the relative motion relationship of each joint through the spatio-temporal graph convolutional network model ST-GCN, and calculate the flexion and extension angles of each joint , and the calculation formula is as follows:
[0048] ;
[0049] Wherein, , , and are the three-dimensional coordinates of joints , , and ; is the angle formed by joints , and ;
[0050] Generate the limb posture information according to the position change rate of the joint and the flexion and extension angle of the joint
[0051] Furthermore, the feature fusion of the limb posture information and the emotion feature includes:
[0052] Map the motion features extracted by the spatio-temporal graph convolutional network model ST-GCN to the scoring criteria of the FMA scale, and output the FMA score result as ;
[0053] Input the emotion feature obtained by the emotion recognition model OpenFace into the BiGRU time series network for processing, as shown in the following formula:
[0054] ;
[0055] Fuse the output feature and the output feature of the FMA scoring processing network through weighted average:
[0056] ;
[0057] Wherein, and are weighted coefficients used to adjust the influence of the FMA score and the emotion feature on the final evaluation result, is the fused feature vector
[0058] In a second aspect, the present invention also provides a motion rehabilitation evaluation system based on limb posture and emotion recognition, which is used to implement any one of the above-mentioned motion rehabilitation evaluation methods based on limb posture and emotion recognition, including:
[0059] A video acquisition module, which is used to collect action videos of a patient performing rehabilitation training from multiple perspectives;
[0060] A data preprocessing module, which records the age and gender information of the patient, combines the collected action videos to construct a self-made dataset; according to the self-made dataset, demarcates candidate regions containing the patient's rehabilitation actions, and generates a basic action posture dataset;
[0061] A posture estimation module, which sets personalized key point extraction parameters and weights corresponding to different age stages and genders according to the age and gender information of the patient, and extracts personalized limb bone key points based on the posture estimation algorithm according to the basic action posture dataset;
[0062] An emotion detection module, which is used to obtain emotion features according to the basic action posture dataset through facial key point detection and facial action unit analysis;
[0063] A rehabilitation effect evaluation module, which analyzes the movement trajectory of the knee joint according to the basic action posture dataset and the personalized limb bone key points, and extracts limb posture information; fuses the limb posture information and the emotion features to form a quantitative rehabilitation evaluation result.
[0064] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0065] The present invention provides a motion rehabilitation evaluation method and system based on limb posture and emotion recognition, which can simultaneously evaluate the limb posture and emotion state of a patient, provide a more comprehensive rehabilitation evaluation perspective, so as to more accurately reflect the overall rehabilitation situation of the patient. Only by using a portable camera in the rehabilitation scenario, the system can use the deployed deep learning model to evaluate the rehabilitation state of the patient in real time.
[0066] The present invention does not require any complex hardware devices, such as sensors, etc., and realizes non-contact evaluation.
[0067] The rehabilitation plan provided by the present invention is natural and easy to operate, enabling patients to perform professional rehabilitation evaluations at home and improving the accessibility of rehabilitation services.
[0068] The data support provided by the system helps doctors make more scientific and accurate medical decisions and improve the quality of rehabilitation treatment. Description of the Drawings
[0069] Figure 1 is a flowchart of a motion rehabilitation evaluation method based on limb posture and emotion recognition in an embodiment of the present invention;
[0070] Figure 2It is a structural block diagram of a motion rehabilitation evaluation system based on limb posture and emotion recognition in an embodiment of the present invention;
[0071] Figure 3 It is a schematic diagram of camera placement of a motion rehabilitation evaluation method based on limb posture and emotion recognition in an embodiment of the present invention;
[0072] Figure 4 It is a sub - flowchart of step S400 of a motion rehabilitation evaluation method based on limb posture and emotion recognition in an embodiment of the present invention;
[0073] Figure 5 It is a spatio - temporal graph convolutional network model diagram of a motion rehabilitation evaluation method based on limb posture and emotion recognition in an embodiment of the present invention; Specific implementation mode
[0074] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0075] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "plurality" is two or more.
[0076] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0077] Embodiment 1
[0078] Please refer to Figure 1, this embodiment provides a motion rehabilitation assessment method based on limb posture and emotion recognition, including the following steps:
[0079] Step S100, collect the action videos of patients during rehabilitation training from multiple perspectives to generate a self-made dataset.
[0080] Specifically, it includes the following steps:
[0081] Use a fixed-position camera to collect training videos of a sufficient number of patients during limb rehabilitation exercises, and sequentially extract the video segments containing specific rehabilitation actions during the training. Perform category annotation according to the rehabilitation actions in the video segments to construct a rehabilitation action posture dataset for training an action recognition model, that is, the spatio-temporal graph convolutional network model ST-GCN. To ensure capturing the complete limb movements of patients during rehabilitation training, the system is designed with at least four fixed-position cameras. As Figure 3 shown, they are respectively set at the front, side, back, and top view positions. The specific positions, angles, and heights of each camera are precisely planned to maximize the capture of various actions of patients during rehabilitation.
[0082] Exemplarily, the front-view camera is installed in front of the patient, about 2 to 3 meters away from the patient, and at the same height as the patient's waist, forming a top-down angle of about 15 to 20 degrees with the ground, capturing the actions of the patient's upper and lower bodies from the front, especially suitable for capturing the activities of the patient's arms, chest, and torso; the side-view camera is fixed on the side of the patient, about 2 meters away from the patient, and the camera height is set at the position parallel to the patient's shoulder. The angle of the camera should be horizontal or slightly downward looking to ensure complete capture of the actions of the patient's waist, legs, and lower limbs, especially the dynamic changes during gait training or joint activities; the back-view camera is fixedly installed behind the patient, about 1 to 2 meters away from the patient, and parallel or slightly higher than the patient's shoulder. The camera angle should ensure comprehensive capture of the action changes of the patient's back, spine, and shoulder, especially suitable for those rehabilitation trainings that require observing the back posture and arm movements; the top-view camera is installed on the ceiling directly above the patient, looking directly down at the patient's body. This camera provides a top-down view of the whole body, accurately capturing the overall movement trajectory of the patient during rehabilitation training, especially coordinated movements and complex limb interaction movements, such as standing, squatting, rotating, etc. It should be noted that all cameras ensure the precise synchronization of the video recording time through a synchronous trigger device or timestamp marking function, guaranteeing the accuracy of the post-processing of video segments and the training of the action recognition model.
[0083] Making a self-made dataset also includes: obtaining pictures of various emotional expressions from a public database, classifying the obtained pictures of emotional expressions, ensuring that each picture has a correct emotional label, such as happy, sad, angry, etc. Constructing the labeled pictures of emotional expressions into an emotional dataset, which is used to train and validate an emotion recognition model. In this embodiment, the emotion recognition model uses the OpenFace model.
[0084] In addition, making a self-made dataset also includes: recording the identity information of the patient including age and gender information. At the same time, face information for identifying the patient's identity information to call age and gender information can also be collected.
[0085] Step S200, obtaining and processing the video stream to be detected, accurately defining the candidate region containing the patient's rehabilitation actions, and ensuring subsequent tracking of the posture characteristics and emotional characteristics of the patient's limbs.
[0086] More specifically, for the detection of the candidate region, the target detection network model Faster R-CNN is used, which is an optimized target detection network model introducing a self-attention mechanism and a Region Proposal Network (RPN). For the patient video stream in the rehabilitation scenario, each frame of the image is obtained in real time as the input image. . The input image is set to a size of 224×224×3 to standardize the input size for subsequent feature extraction processing. The input image passes through a feature extraction network (ResNet-50) to extract multi-level feature maps. In the feature extraction network, multiple convolutional layers are used to process the image layer by layer, and deep feature maps containing information such as the patient's limb contours, postures, and movement trajectories are gradually extracted. The feature maps output by the feature extraction network are , where k is the number of candidate regions. The feature extraction process can be described as: , where represents the feature extraction function of the feature extraction network (ResNet-50), and represents the obtained deep feature map, which is used to capture the basic contour information of the human body and provide a basis for subsequent region recommendation. To ensure that the model focuses on the human body region of the patient, the present invention embeds a Channel and Spatial Attention Module (CBAM) in the region recommendation network to filter out irrelevant background information in the image.
[0087] Among them, the channel attention module generates channel weights to improve the response of the feature channels related to the human body, as follows:
[0088] ,
[0089] Among them, is an activation function that enhances the attention to the patient area by integrating the information of global pooling.
[0090] The spatial attention module focuses on the spatial area where the human body is located and reduces external background interference. It applies spatial weights to the feature map after channel attention processing as follows: as follows:
[0091] ,
[0092] Furthermore, a balance mechanism is added to the RPN to increase the attention to the foreground (i.e., the human body area) features while reducing the over-response to the background area. By setting different loss weights, the RPN is more inclined to identify the human body area, further reducing background interference. The feature map is input into the improved Region Proposal Network (RPN). After obtaining local features through a 3×3 convolutional layer, the RPN outputs candidate bounding boxes , which can accurately frame the human body area in the image.
[0093] Step S300: Using deep learning methods, a face recognition model is established to identify the patient's identity information from the video stream, and information vectors corresponding to the age and gender of the patient are extracted from the self-made dataset.
[0094] More specifically, MobileNetV2 is used as a feature extractor to extract the key features of the input facial image; the age and gender information of the patient are extracted.
[0095] According to the age and gender information of the patient, the data is divided. The age is classified according to different age groups: teenagers (13 - 18 years old) are labeled 0; young people (19 - 35 years old) are labeled 1; middle-aged people (36 - 55 years old) are labeled 2; the elderly (56 years old and above) are labeled 3; the gender is represented by binary classification: 0 represents female, and 1 represents male; the age classification label and the gender label are combined to form a two-dimensional vector. In this embodiment, the patients are divided into eight categories according to age and gender for targeted rehabilitation assessment.
[0096] Step S400: According to the age and gender information of the patient, corresponding to different age stages and genders, personalized key point extraction parameters and weights are set, and based on the basic action posture dataset and the pose estimation algorithm, personalized limb bone key points are extracted. According to the personalized limb bone key points, the spatio-temporal graph convolutional model is used to analyze the patient's rehabilitation information and evaluate the actions against the template video of the rehabilitation doctor to achieve a normative evaluation.
[0097] Among them, please refer to Figure 4 , step S400 further includes the following sub-steps:
[0098] Step S401, embed age and gender information into the pose estimation algorithm to form key point extraction based on age and gender information.
[0099] Exemplarily, each frame in the video stream is processed in real time through a Transformer-based pose estimation algorithm. This algorithm is used to detect human key points (such as joints, limb parts, etc.) in each frame, and the key point coordinates of each frame are output as , where represents the three-dimensional coordinates of the human key points in the th frame, and is the number of human key points. The key point data can be expressed as: , where, is the three-dimensional coordinates of the th key point. The Transformer-based pose estimation algorithm can capture the long-term dependence relationship of the poses between frames by modeling the self-attention mechanism of the video frames, so as to accurately extract the human key point information in each frame.
[0100] During the training process, the model automatically learns to assign appropriate weights to patients of different ages and genders. The following are the steps and methods to achieve this goal:
[0101] Take the vector representations of age and gender together with the key point data of pose estimation as input and pass it into the model;
[0102] For each key point, set an initial weight. Use the Xavier initialization method to randomly initialize the weights of all key points;
[0103] During the training process, use the attention mechanism to dynamically adjust the weights of the key points. Through the neural network, assign attention weights to each key point according to patient characteristics such as age and gender:
[0104] ;
[0105] Where: is the attention score of the th key point, and the input is the key point position , the age of the patient, and the gender . is the Adaptive weights for key points. This approach can ensure that the model focuses on important key points in patients of different ages and genders in a targeted manner.
[0106] Furthermore, through action coherence and importance assessment, frames containing key rehabilitation actions are extracted from the video stream. The coherence of an action can be judged by calculating the change amount of poses between adjacent frames. By calculating the Euclidean distance of pose data between two adjacent frames, as shown in the following formula:
[0107] ,
[0108] If is less than a threshold, it indicates that the pose change between adjacent frames is small and does not contain key action information. In this way, frames with significant changes (such as the start, end, or important moments of rehabilitation actions) are extracted.
[0109] Step S402, as Figure 5 shown, analyzes the pose information in the key frames through a spatio-temporal graph convolutional network model (ST-GCN).
[0110] More specifically, exemplarily, the spatio-temporal graph convolutional network model ST-GCN can model the relationships between various key points spatially and the dynamic changes between key points temporally. The human pose in each frame can be represented as a graph structure, with nodes being the human key points and edges being the spatial connection relationships between key points. Specifically, the input of the spatio-temporal graph convolutional network model ST-GCN is a spatio-temporal graph , where represents the set of key points (nodes), and each node corresponds to the position of a key point, represents the connection relationships between key points (edges), usually using known skeleton connection relationships, represents the time step, connecting graphs at multiple moments. The convolutional operation of the spatio-temporal graph convolutional network model ST-GCN can be represented by the following formula:
[0111] where, represents the feature representation of all nodes at time step t, is the element of the adjacency matrix between node and node , represents the feature representation of node i at time step t, ∑ i∈N(t) represents the sum over all adjacent nodes i of node t, represents node The set of adjacent nodes, and σ is a non-linear activation function. The spatio-temporal graph convolutional network model ST-GCN generates spatio-temporal feature representations of each key point of the patient through multiple layers of convolution, and evaluates the accuracy of the rehabilitation actions based on these features.
[0112] Step 403, perform action matching between the standard template action video of the rehabilitation doctor and the rehabilitation video of the patient.
[0113] More specifically, exemplarily, extract the key frame data of both through a pose estimation algorithm, and calculate the similarity between each frame:
[0114] ;
[0115] Among them, and are the key frames in the patient's and the rehabilitation doctor's template videos respectively, represents the similarity between two frames. Frames with higher similarity indicate that the patient's actions are more consistent with the standard actions. By calculating the similarity of the entire action sequence, the normality and execution of the patient's actions are evaluated.
[0116] Step S500, by fusing the speed and frequency features, emotional features and other diversified information of the patient's actions, evaluate the patient's rehabilitation status using the Fugl-Meyer Assessment (FMA) scale, and generate the patient's own rehabilitation trend report.
[0117] More specifically, use the OpenFace model to obtain emotional features. The OpenFace model extracts emotion-related features through facial key point detection and Facial Action Unit (AU) analysis. Extract 68 facial key points, and calculate changes in their displacements, angles, etc. as emotional features; further, through facial key point recognition and quantification of facial action units (such as represents the "frowning" action of the face, represents a smile), these facial action units are usually related to specific emotional states (such as anger, happiness, etc.). The emotional features output by the OpenFace model are usually a multi-dimensional vector, representing information such as facial action units (AUs) and head poses. Suppose we extract facial action units , then the facial emotion feature can be represented as a vector:
[0118] ,
[0119] Among them, is the facial action unit, is the number of facial action units.
[0120] Furthermore, the movement speed and frequency of the patient reflect their movement fluency and rehabilitation progress. The Spatio-Temporal Graph Convolutional Network (ST-GCN) can process joint position data in consecutive frames to extract the movement features of the joints. The rate of change of the joint position is the speed. The calculation formula:
[0121] ;
[0122] where, is the position of the joint at the frame moment; is the time interval; is the movement speed at the frame moment. Through the time convolutional layer of the Spatio-Temporal Graph Convolutional Network (ST-GCN), the model can automatically learn the movement speed of each joint and extract the dynamic pattern of the speed changing over time from the graph structure.
[0123] Furthermore, the movement amplitude can be represented by the angle of the joint, usually the included angle formed by two joints and the bone connecting them. For example, the flexion and extension angle of the elbow is the angle formed by the humerus, forearm, and upper arm. The Spatio-Temporal Graph Convolutional Network model (ST-GCN) can calculate the flexion and extension angles of each joint by extracting the relative movement of each joint.
[0124] Calculation formula: Assume that joints , and form an angle (such as the angle of the elbow joint):
[0125] ;
[0126] where, , , and are the three-dimensional coordinates of joints , , and ; is the angle formed by joints , and . The graph convolutional layer of the Spatio-Temporal Graph Convolutional Network model (ST-GCN) can capture the relative angle changes between joints, and then learn the flexion and extension amplitude and posture changes of the joints.
[0127] Preferably, map the movement features extracted by the Spatio-Temporal Graph Convolutional Network model (ST-GCN) to the scoring criteria of the FMA scale.
[0128] Exemplarily, taking the evaluation of the upper limb motor function (UE Motor Function) part of a certain patient, especially the finger motor function, as an example, the specific steps are as follows:
[0129] 0 points: The patient is unable to perform any finger movements; 1 point: The patient can make partial finger movements, but there are obvious incoordination and movement limitations; 2 points: The patient can complete all basic finger movements, and the movements are smooth and coordinated. By analyzing the finger movement trajectories (speed, frequency, etc.) of the patient during rehabilitation training, map the features extracted by the spatio-temporal graph convolutional network model ST-GCN to the FMA score: If the finger movement speed of the patient is very low, and the movement is not smooth or incomplete, give 0 or 1 point. If the finger movement frequency of the patient is normal, the movement is coordinated and the amplitude is large, give 2 points.
[0130] The formulaic representation of the above steps is:
[0131] ,
[0132] where, is the movement speed of the patient's finger; is the amplitude and coordination of the finger movement.
[0133] Specifically, The calculation of contains the following two parts:
[0134] Movement amplitude , this parameter measures the displacement or movement amplitude of the finger, and usually calculates the displacement of the finger within the time period . For example, if represents the position of the patient's finger at time , then the amplitude is calculated as:
[0135] ,
[0136] If it is necessary to measure the maximum displacement from the starting point to the current moment, the following can be used:
[0137] ,
[0138] Here is the initial position, representing the change of the finger from the initial state to the maximum displacement.
[0139] Movement coordination , coordination reflects the smoothness and precision of finger movements. For example, the coordination can be calculated by the smoothness of the speed or acceleration curve of finger movement. One calculation method of coordination is to calculate the ratio of the variance to the mean according to the speed sequence of the finger:
[0140] ;
[0141] This formula represents the volatility of speed. Smaller fluctuations usually indicate better coordination and smoother movements. Additionally, when conducting two-handed coordination training, the synchrony of the movements of two fingers or hands can also be considered, typically measured by calculating the angular difference between the movement trajectories:
[0142] ;
[0143] where is the angular difference between the movement trajectories of two fingers or hands.
[0144] The final combines amplitude and coordination and can be integrated through weighted summation:
[0145] ;
[0146] where and are parameters that control the weights of amplitude and coordination in the comprehensive evaluation. According to specific rehabilitation goals or the needs of the patient, the values of these two factors can be adjusted to better reflect the characteristics of finger movements.
[0147] By calculating these parameters, scores are given for each evaluation item, and thus a score is generated for each dimension of the FMA scale. Similarly, the same method can be used to evaluate the lower limb motor function. Taking the evaluation of the flexion and extension function of a patient's knee as an example: 0 points: unable to bend the knee; 1 point: able to bend the knee, but with a small and unsmooth angle; 2 points: able to fully bend the knee with a smooth movement and good amplitude. By analyzing the movement trajectory of the knee joint (such as the movement speed, frequency, and amplitude of the knee joint), corresponding movement characteristics can be generated and mapped to the FMA score, and finally the FMA score result is .
[0148] Furthermore, an late fusion strategy is adopted to fuse the FMA score result with emotional characteristics:
[0149] The emotional characteristics obtained by OpenFace are input into a BiGRU time series network for processing, and this sub-network can learn the time series changes of the patient's emotional state:
[0150] ,
[0151] The characteristics output by the emotion network and the output of the FMA score processing network are fused through weighted averaging:
[0152] ,
[0153] Among them, and are weighting coefficients used to adjust the influence of the FMA score and emotional characteristics on the final evaluation result. is the fused feature vector.
[0154] Furthermore, the fused feature vector is input into the final regression model for prediction to generate a personalized patient rehabilitation progress report.
[0155] More specifically, in view of the significant differences in the physical functions and recovery abilities of patients in different age groups and genders, in order to more accurately promote the rehabilitation process, the regression model corrects the rehabilitation progress prediction by introducing adjustment coefficients for gender and age. Let the rehabilitation progress output by the regression model be , and the output can be adjusted according to gender and age through the following adjustment function:
[0156] ;
[0157] Where: is the gender label (0 represents female, 1 represents male). is the age label (0 represents juvenile, 1 represents young, 2 represents middle-aged, 3 represents elderly). and are the adjustment coefficients for gender and age respectively. Based on the settings of rehabilitation doctors, specifically, elderly patients may have a longer rehabilitation time during the recovery process. Therefore, has a larger value; while men may recover strength faster, so the gender adjustment coefficient will enhance the influence on the recovery of strength training.
[0158] In this embodiment, the patient rehabilitation progress report includes:
[0159] 1. Motor function assessment: Combining the patient's emotional fluctuations, analyzing its impact on motor recovery, and generating personalized suggestions.
[0160] 2. Comprehensive assessment: Based on the overall analysis, predicting the patient's rehabilitation progress and expected rehabilitation time, and providing suggestions on the impact of emotions on motor function recovery.
[0161] Embodiment 2
[0162] Please refer to Figure 2, this embodiment provides a motion rehabilitation evaluation system based on limb posture and emotion recognition, which is applicable to the above-mentioned motion rehabilitation evaluation method based on limb posture and emotion recognition, and includes a video acquisition module, a data preprocessing module, a posture estimation module, an emotion detection module, and a rehabilitation effect evaluation module. The modules referred to in the present invention refer to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory.
[0163] The video acquisition module is used to collect action videos of patients during rehabilitation training from multiple perspectives.
[0164] The data preprocessing module is used to record the age and gender information of the patient, combine the collected action videos to construct a self-made dataset; according to the self-made dataset, define the candidate areas containing the patient's rehabilitation actions, and generate a basic action posture dataset.
[0165] The posture estimation module sets personalized key point extraction parameters and weights corresponding to different age stages and genders according to the age and gender information of the patient, and extracts personalized limb bone key points based on the posture estimation algorithm according to the basic action posture dataset.
[0166] The emotion detection module is used to obtain emotion features according to the basic action posture dataset through facial key point detection and facial action unit analysis.
[0167] The rehabilitation effect evaluation module analyzes the movement trajectory of the knee joint according to the basic action posture dataset and the personalized limb bone key points, and extracts limb posture information; fuses the limb posture information with the emotion features, and evaluates the patient's rehabilitation situation with the Fugl-Meyer Assessment (FMA) scale to generate the patient's own rehabilitation trend report and form a quantitative rehabilitation evaluation result.
[0168] Among them, obtaining the patient's remote rehabilitation exercise video segment specifically includes:
[0169] Obtain the video stream of the patient during rehabilitation training through the camera. The video stream contains the content of the patient's rehabilitation training actions, and the rehabilitation training actions include upper limb actions, lower limb actions, standing actions, sitting actions, hand actions, and facial expressions. The video can be obtained through a fixed or mobile camera to ensure that the patient's limb movements can be comprehensively captured.
[0170] The working steps of the data preprocessing module specifically include: using intelligent algorithms to preliminarily analyze the scenes captured by the camera to identify and mark possible interference sources, such as unrelated personnel and non-rehabilitation-related objects. Through background segmentation and human detection algorithms, accurately define the candidate regions containing the patient's rehabilitation movements. Record the patient's age and gender information, and combine the collected action videos to construct a self-made dataset to provide basic data for subsequent processing.
[0171] The working steps of the pose estimation module specifically include: according to the patient's age and gender information, corresponding to different age stages and genders, set personalized key point extraction parameters and weights. Based on the basic data, perform real-time or frame-by-frame analysis on each frame in the video stream, and use pose estimation algorithms to track the movement trajectories and pose changes of the patient's limbs. The pose estimation algorithm based on Transformer performs targeted real-time detection of human key points. Based on the evaluation of action coherence and importance, extract the frames containing key rehabilitation action information from the dynamic video. Input the extracted key frames into the evaluation model, and use the spatio-temporal graph convolutional network (ST-GCN) deep learning model to automatically evaluate the patient's rehabilitation progress and effect according to the pose information in the key frames. This model conducts rehabilitation evaluation by analyzing the time series of pose key points related to the rehabilitation part, providing a basis for the patient's rehabilitation status. At the same time, the system matches the standard template action video taken by the rehabilitation doctor with the patient's rehabilitation training video, and through comparing the similarity of the actions, realizes accurate action recognition and normative evaluation, ensuring that the actions performed by the patient meet the rehabilitation requirements and tracking their rehabilitation progress.
[0172] The emotion detection module specifically obtains the patient's emotion features based on the data preprocessing module through the OpenFace facial expression analysis tool.
[0173] The working steps of the rehabilitation effect evaluation module specifically include: adopting a weighted fusion strategy to comprehensively evaluate diversified information such as the patient's action speed and frequency characteristics, emotion characteristics, and limb abnormal trends. Adopt a late fusion strategy to fuse the scoring results of the Fugl-Meyer Assessment of Motor Function (FMA) scale with the emotion characteristics. Specifically, input the emotion characteristics obtained by OpenFace into the BiGRU time series network for processing, and this network learns the time series changes of the patient's emotional state. Then, fuse the features output by the emotion network and the output of the FMA scoring processing network through weighted averaging, where the weighting coefficients are used to adjust the influence of the FMA score and emotion characteristics on the final evaluation result.
[0174] Further, the fused feature vectors are input into a regression model for prediction, and an adjustment function is used to adjust the output of the regression model according to gender and age to generate a personalized patient rehabilitation progress report. The report includes a motor function assessment and a comprehensive assessment, predicts the patient's rehabilitation progress and estimated rehabilitation time based on an overall analysis, and provides suggestions on the impact of emotions on motor function recovery. Through this method of weighted fusion of features, the present invention can more comprehensively evaluate the patient's rehabilitation status and provide more accurate rehabilitation guidance and decision support for medical professionals.
[0175] For the specific function implementation of each of the above modules, reference may be made to the relevant content in the method of Embodiment 1, and details are not elaborated herein.
[0176] Embodiment 3
[0177] The embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of a method for motion rehabilitation assessment based on limb posture and emotion recognition in Embodiment 1 are implemented.
[0178] Embodiment 4
[0179] This embodiment provides a computer device, including: a memory and a processor. The memory is used to store computer programs / instructions. The processor is used to execute the computer programs / instructions to implement the steps of a method for motion rehabilitation assessment based on limb posture and emotion recognition in Embodiment 1.
[0180] Embodiment 5
[0181] This embodiment provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of a method for motion rehabilitation assessment based on limb posture and emotion recognition in Embodiment 1 are implemented.
[0182] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
[0183] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as methods, systems, or computer program products. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0184] This disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks or multiple blocks.
[0185] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks or multiple blocks.
[0186] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks or multiple blocks.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure rather than to limit the scope of its protection. Although the present disclosure has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: after reading the present disclosure, those skilled in the art can still make various changes, modifications, or equivalent substitutions to the specific embodiments of the invention. However, these changes, modifications, or equivalent substitutions are all within the scope of the protection of the pending claims of the disclosure.
Claims
1. A motion rehabilitation assessment method based on limb posture and emotion recognition, characterized in that, It includes the following steps: Collect the action videos of the patient during rehabilitation training, record the age and gender information of the patient, and construct a self-made dataset; Obtain and process the video stream, and define the candidate regions containing the patient's rehabilitation actions according to the self-made dataset, and construct a basic action posture dataset; According to the age and gender information of the patient, corresponding to different age stages and genders, set personalized key point extraction parameters and weights, and based on the basic action posture dataset, extract personalized limb bone key points based on the pose estimation algorithm; Based on the emotion recognition model, according to the basic action posture dataset, obtain emotion features through facial key point detection and facial action unit analysis; Based on the spatio-temporal graph convolutional network model, according to the basic action posture dataset and the personalized limb bone key points, analyze the movement trajectory of the knee joint and extract limb posture information; Fuse the limb posture information and emotion features to form a quantitative rehabilitation evaluation result.
2. The motion rehabilitation assessment method based on limb posture and emotion recognition according to claim 1, wherein The obtaining and processing the video stream and defining the candidate regions containing the patient's rehabilitation actions include: Based on the object detection network model Faster R-CNN, according to the video stream, each frame of image is obtained in real time as the input image , the input image is input into the feature extraction network, and in the feature extraction network, the input image is processed layer by layer using a multi-level convolutional layer , and a deep feature map containing information on the patient's limb contours, postures, and movement trajectories is output; Input the deep feature map into the Region Proposal Network (RPN), obtain local features through the convolutional layer, then use the RPN to output candidate boxes, and frame the human body region in the image through the candidate boxes, so as to define the candidate regions containing the patient's rehabilitation actions.
3. The motion rehabilitation evaluation method based on limb posture and emotion recognition according to claim 2, characterized in that The deep feature map output by the feature extraction network is , where k is the number of candidate regions; The feature extraction process is expressed as: , where represents the feature extraction function of the feature extraction network, represents the obtained deep feature map, and the deep feature map is used to capture the basic contour information of the human body.
4. The method for evaluating motion rehabilitation based on limb posture and emotion recognition according to claim 3, wherein It also includes: Embed the Channel and Spatial Attention Module (CBAM) in the RPN to filter out irrelevant background information: Channel attention module: generate channel weights , as follows: ; Among them, is an activation function that enhances the attention to the patient area by improving the information fused with global pooling; Spatial attention module: Apply spatial weights to the feature map after channel attention processing as follows: as shown in the following formula: ; The feature map processed by the CBAM is defined as: 。 5. The motion rehabilitation assessment method based on limb posture and emotion recognition according to claim 4, wherein, It also includes: Add a balance mechanism to the RPN to construct an improved RPN; Input the feature map into the improved Region Proposal Network (RPN). After obtaining local features through the convolutional layer, use the RPN to output candidate bounding boxes to frame the human body region in the image.
6. The method for evaluating motion rehabilitation based on limb posture and emotion recognition according to claim 5, wherein The training of personalized key point extraction parameters and weights according to the age and gender information of the patient, corresponding to different age stages and genders, and the extraction of personalized limb bone key points based on the basic action posture dataset and the pose estimation algorithm include: Perform real-time processing on each frame in the video stream through the Transformer-based pose estimation algorithm: Use the Transformer-based pose estimation algorithm to detect human key points in each frame, and output the key point coordinates of each frame as , where represents the three-dimensional coordinates of the human key points in the th frame, and is the number of key points of the human body; The key point data is represented as: , where is the three-dimensional coordinate of the th key point; For each key point, set an initial weight; use the Xavier initialization method to randomly initialize the weights of all key points; During the training process, use the attention mechanism to dynamically adjust the weights of the key points; update the attention weights for each key point according to the age and gender information, as shown in the following formula: ; Wherein: is the attention score of the th key point, and the input is the position of the key point , the age range of the patient and gender ; is the adaptive weight of the th key point.
7. The method for evaluating motion rehabilitation based on limb posture and emotion recognition according to claim 6, wherein, The extraction of emotion-related features and obtaining emotion features according to the basic action posture dataset through facial key point detection and facial action unit analysis include: Based on facial key point recognition and quantification of facial action units, emotional features are output through the emotion recognition model OpenFace ; is a multi-dimensional vector, representing information including facial action unit AU and head pose, as shown in the following formula: ; where AU is the Action Unit, and is the number of Action Units.
8. The motion rehabilitation evaluation method based on limb posture and emotion recognition according to claim 7, wherein, The analysis of the movement trajectory of the knee joint and extraction of limb posture information based on the personalized limb bone key points according to the basic action posture dataset include: Process the joint position data in consecutive frames through the spatio-temporal graph convolutional network model ST-GCN to extract the motion features of the joints; Among them, the rate of change of the joint position is the speed , and the calculation formula is as follows: ; Among them, is the position of the joint at the frame moment; is the time interval; is the motion speed at the frame moment; Extract the relative motion relationships of each joint through the spatio-temporal graph convolutional network model ST-GCN, and calculate the flexion and extension angles of each joint , and the calculation formula is as follows: ; wherein, , , and are joints , , and 's three-dimensional coordinates are the angles formed by joints , and ; According to the rate of change of the position of the joint and the flexion / extension angle of the joint generate the limb posture information.
9. The method for evaluating sports rehabilitation based on limb posture and emotion recognition according to claim 8, wherein, The feature fusion of the limb posture information and emotion features includes: Map the motion features extracted by the spatio-temporal graph convolutional network model ST-GCN to the scoring criteria of the FMA scale, and output the FMA score result as ; The emotion features obtained by the emotion recognition model OpenFace are input into the BiGRU time series network for processing, as shown in the following formula: ; The output features and the output features of the FMA scoring processing network perform feature fusion through weighted averaging: ; Among them, and are weighting coefficients used to adjust the influence of the FMA score and emotional features on the final evaluation result, is the fused feature vector.
10. A motion rehabilitation evaluation system based on limb posture and emotion recognition, which is used to implement the motion rehabilitation evaluation method based on limb posture and emotion recognition according to any one of claims 1 to 9 above, including: Video acquisition module, which is used to collect action videos of patients with multiple perspectives during rehabilitation training; Data preprocessing module, which records the age and gender information of patients, and combines the collected action videos to construct a self-made dataset; Based on the self-made dataset, candidate regions containing the patients' rehabilitation actions are defined, and a basic action posture dataset is generated; Posture estimation module, according to the age and gender information of the patients, corresponding to different age groups and genders, sets personalized key point extraction parameters and weights, and based on the basic action posture dataset and the posture estimation algorithm, extracts personalized limb bone key points; Emotion detection module, which is used to obtain emotion features according to the basic action posture dataset through facial key point detection and facial action unit analysis; Rehabilitation effect evaluation module, according to the basic action posture dataset and the personalized limb bone key points, analyzes the movement trajectory of the knee joint and extracts limb posture information; fuses the limb posture information and the emotion features to form a quantitative rehabilitation evaluation result.
Citation Information
Patent Citations
Limb rehabilitation training and evaluation system based on visual tracking control
CN113100755A
Cited By
Limb conflict behavior identification method and device and storage medium
CN120766366A
A method and device for recognizing limb conflict behavior and a storage medium
CN120766366B
Rehabilitation evaluation system and method based on data feedback
CN120938361A
Human body posture recognition method and system based on double-attention structured position coding
CN121600558A
Multi-source data driven limb movement rehabilitation quantitative evaluation method and device
CN122436126A