Deep learning system for joint motion range evaluation
Through a deep learning system, combined with image acquisition and data processing modules, the deep learning model is used to calculate joint mobility, which solves the subjectivity and error problems of joint mobility measurement in the existing technology, and achieves high accuracy and convenient evaluation.
Patent Information
- Application Number
- CN202510067580.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-27
AI Technical Summary
The current joint mobility measurement methods have problems such as strong subjectivity, low efficiency and large measurement errors, making it difficult to achieve high accuracy and convenient evaluation.
The deep learning system is adopted, including an image acquisition module, a data processing module and a joint mobility recognition module, and the relationship between the human joint activity trajectory and angle is learned through the deep learning model, the joint mobility is calculated, and visual output is provided.
It improves the objectivity and accuracy of joint mobility measurement, reduces the dependence on professional skills, is suitable for ordinary users and recovered patients, and achieves convenient joint health assessment.
Smart Images

Figure CN120047969A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and specifically refers to a deep learning system for joint range of motion evaluation. Background Art
[0002] Range of Motion (ROM) of joints is a key indicator for evaluating joint health status and is widely used in sports rehabilitation, medical diagnosis, and physical therapy.
[0003] However, current joint ROM measurements mainly rely on manual evaluation tools (such as goniometers), which have problems of strong subjectivity, low efficiency, and large measurement errors. With the development of computer vision and deep learning technologies, automated joint ROM evaluation systems based on images or videos are expected to improve evaluation efficiency and accuracy. Summary of the Invention
[0004] In view of the above situation, to overcome the defects of the prior art, the present invention provides a deep learning system for joint range of motion evaluation to solve the problems of measurement accuracy and convenience in joint range of motion evaluation.
[0005] The technical solution adopted by this application is as follows:
[0006] This solution provides a deep learning system for joint range of motion evaluation, including:
[0007] An image acquisition module: used to acquire joint movement trajectory images or videos at different angles and static pose data;
[0008] A data processing module: used to preprocess the acquired videos or images, including denoising, background segmentation, and extraction of human key points;
[0009] A joint range of motion recognition module: based on a deep learning model, used to construct a human joint point recognition model, learn the relationship between human joint movement trajectories and angles, analyze and identify the key point data using the trained human joint point recognition model, and calculate the joint range of motion;
[0010] A result output module: used to visually output the measurement results of the joint range of motion (including real-time movement trajectories and angle curve graphs, etc.), and can provide rehabilitation suggestions.
[0011] Preferably, the image acquisition module includes an RGB camera or a depth camera. The RGB camera or depth camera is used to acquire human joint movement videos or pictures. The RGB camera is used to acquire RGB images or videos for visual feature extraction to obtain two-dimensional position information, and the depth camera is used to obtain depth maps for capturing the three-dimensional spatial position information of human key points.
[0012] Preferably, the data processing module first predicts the two-dimensional key point coordinates of the human joints in the obtained image, and then for each located joint pixel point, uses the binocular positioning principle to measure its three-dimensional key point coordinates. After removing outliers through a filtering algorithm, the accurate spatial information of the human joints is obtained. By extracting the two-dimensional / three-dimensional key point sequence, the movement trajectory of the human body is reflected, thereby obtaining the skeleton mode.
[0013] Preferably, the deep learning model includes an LSTM network for time series analysis, a CNN network for predicting the degree of activity, and a Transformer network for fusing multi-modal input data.
[0014] Preferably, the joint activity recognition module is composed of:
[0015] Model data input layer: used to input key point data, and also includes time series and derivative features;
[0016] The key point data includes the (x, y) coordinates of the two-dimensional key point data and the (x, y, z) coordinates of the three-dimensional key point data, which is used to increase the depth information;
[0017] The time series is the movement trajectory information of the key points, usually a sliding window of several past frames; the derivative features are the joint angles (such as the angles of the knee joint and elbow joint), angular velocity, and angular acceleration calculated based on the key points.
[0018] Feature extraction module: includes time series feature extraction and spatial feature extraction;
[0019] Among them, the time series feature extraction is used to capture the dynamic changes of the joint activity data and is implemented based on the LSTM network / Transformer network; the spatial feature extraction is used to obtain the activity degree information of the spatial structure of the joint key points and is implemented based on the CNN network / GCN network. The convolutional neural network CNN is used to extract the human body shape features in the image, and the graph convolutional network GCN is used to extract the bone structure features.
[0020] Specifically, it is manifested as: using the CNN / GCN network to extract spatial features from the data, and then inputting the feature sequence extracted by the CNN / GCN into the LSTM / Transformer network, and using the LSTM / Transformer network to model the sequence data to learn the time series features of the joint activity.
[0021] Multi-feature fusion module: combines time series features and spatial features to generate the final activity degree prediction;
[0022] Integrate the time features generated by the LSTM / Transformer and the spatial features extracted by the GCN / CNN through weighted fusion.
[0023] It also includes a multi-task learning function, which predicts the range of motion of different joints simultaneously by setting up multiple branch networks, such as the knee joint, elbow joint, shoulder joint, etc.
[0024] Output layer: It includes the prediction of the range of joint motion and auxiliary output;
[0025] Predict the angular values of the range of motion of each joint and output the predicted values with the dimension of the number of joints. The auxiliary output is to predict the health status (including normal activities, limited activities, abnormal activities) and provide the activity trajectory curve for dynamic analysis.
[0026] The beneficial effects achieved by the present invention using the above solution are as follows:
[0027] 1. Using deep learning algorithms combined with human key points for pose estimation and angle prediction, overcoming the subjectivity and limitations of traditional measurement tools, and improving the objectivity and accuracy of joint range of motion measurement.
[0028] 2. Combining RGB images and depth images to achieve the fusion of multi-modal data, and the evaluation model is more accurate.
[0029] 3. The system can be integrated into intelligent devices, supporting convenient joint health assessment effects, reducing the dependence on professional skills for assessment, and being suitable for ordinary users and rehabilitation patients. Brief Description of the Drawings
[0030] Figure 1 It is a schematic diagram of the composition of a deep learning system for joint range of motion evaluation provided by this solution. Detailed Embodiments
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention.
[0032] Embodiment 1
[0033] Please refer to Figure 1 As shown, this embodiment provides a deep learning system for joint range of motion evaluation, including:
[0034] (1) Image acquisition module: It is used to acquire joint motion trajectory images or videos and static pose data at different angles.
[0035] The image acquisition module includes an RGB camera or a depth camera. The RGB camera or the depth camera is used to acquire human joint activity videos or pictures. The RGB camera is used to acquire RGB images or videos for visual feature extraction to obtain two-dimensional position information, and the depth camera is used to obtain depth maps for capturing the three-dimensional spatial position information of human key points.
[0036] (2) Data processing module: used to preprocess the collected videos or images, including denoising, background segmentation, and extraction of human key points.
[0037] The data processing module is the core preprocessing link of the entire system, extracting relevant information about joints from the original image or video data, and providing accurate and reliable input data for the subsequent joint range of motion recognition module. The data processing module first predicts the two-dimensional key point coordinates of the human joints in the obtained image, and then for each located joint pixel point, uses the binocular positioning principle to measure its three-dimensional key point coordinates. After removing outliers through a filtering algorithm, the precise spatial information of the human joints is obtained. Through the extracted two-dimensional / three-dimensional key point sequence, the movement trajectory of the human body is reflected, thereby obtaining the skeleton mode.
[0038] Furthermore, the specific process of the data processing module includes the following steps:
[0039] Step 1: Data acquisition and preprocessing:
[0040] 1.1 Data acquisition
[0041] Multi-source data acquisition is performed through RGB cameras and depth cameras to obtain two-dimensional key point data and three-dimensional key point data. Among them, the RGB camera is used to collect RGB images or videos for visual feature extraction to obtain two-dimensional position information. The depth camera is used to obtain depth maps to capture the three-dimensional spatial position information of human key points.
[0042] As a further embodiment, data acquisition also includes wearable devices such as IMU sensors, which can obtain precise attitude angle information such as angular velocity and linear acceleration.
[0043] 1.2 Data preprocessing
[0044] 1.2.1 Denoising: Apply Gaussian filtering or bilateral filtering to the image to reduce noise interference;
[0045] 1.2.2 Resolution adjustment: Adjust the input image to a unified resolution (such as 256×256) to adapt to the input requirements of the deep learning model;
[0046] 1.2.3 Frame sampling: Uniformly sample video frames according to the set frequency (such as 30FPS) to reduce data redundancy and reduce the computational pressure through segmented processing.
[0047] Step 2: Human key point detection and extraction:
[0048] 2.1 Key point detection
[0049] Extract the key point positions of the human body in each frame of the image through a human pose estimation algorithm, and the algorithms include but are not limited to:
[0050] OpenPose: Supports 2D / 3D human pose estimation and locates 33 key points (such as head, shoulders, elbows, knees, etc.).
[0051] HRNet: A high-resolution pose estimation network with high positioning accuracy.
[0052] MediaPipe Pose: A lightweight key point detection model suitable for real-time applications.
[0053] In this embodiment, the OpenPose algorithm is preferably used for human key point detection, and the joint angles are calculated based on the detection results.
[0054] 2.2 Key point coordinate normalization
[0055] 2.2.1 Normalize the detected key point coordinates to the range of [0, 1] to eliminate the influence of different device resolutions on the data.
[0056] 2.2.2 For 3D key point data, normalize it to [0, 1, Z] according to the depth information.
[0057] 2.3 Joint angle calculation
[0058] Calculate the angles related to joint mobility based on the detected key points.
[0059] Taking the calculation of the knee joint angle as an example:
[0060] Knee joint angle
[0061] Among them, is the vector from the hip joint to the knee joint, is the vector from the knee joint to the ankle joint.
[0062] Step 3: Data cleaning and filtering:
[0063] 3.1 Key point data denoising
[0064] 3.1.1 Use time series smoothing methods (such as moving average, Kalman filter) to remove the jitter and detection noise of key points.
[0065] 3.1.2 If a key point is missing (such as detection failure due to occlusion), interpolation methods (such as linear interpolation or spline interpolation) can be used to complete the data.
[0066] 3.2 Data validity check
[0067] 3.2.1 Check whether the key points conform to the human anatomical structure. For example, the head cannot be lower than the knees, and the shoulders cannot cross to the opposite side.
[0068] 3.2.2 Mark and eliminate abnormal data (such as key point coordinates with large offsets).
[0069] 3.3 Data Alignment and Normalization
[0070] 3.3.1 Ensure that all key point sequences in each frame of the image are aligned to avoid loss or disorder.
[0071] 3.3.2 For multi-modal data, ensure consistent sampling frequencies and reconstruct the time series through interpolation or sampling.
[0072] Step 4 Feature Extraction and Construction
[0073] 4.1 Spatial Features
[0074] 4.1.1 Key point position relationships: Extract the relative positions (Euclidean distance, direction angle) between key points in each frame of the image.
[0075] 4.1.2 Skeletal connection features: Construct a relationship graph between joints according to the human skeletal structure (such as the connection strength from the shoulder to the elbow).
[0076] 4.2 Temporal Features
[0077] 4.2.1 Speed and acceleration: Calculate the angular velocity and angular acceleration of joints through the first and second derivatives of key point coordinates.
[0078] 4.2.2 Time series trajectory: Use Fourier transform to encode the motion trajectory of key points.
[0079] 4.3 High-Order Features
[0080] 4.3.1 Range of motion: Statistically calculate the range of motion (such as the maximum angle of the knee joint) through the key point trajectory over a period of time.
[0081] 4.3.2 Dynamic feature encoding: Encode the dynamic change information in the time series into a feature vector for model input.
[0082] Step 5 Deep Learning Input Preprocessing
[0083] 5.1 Data Formatting
[0084] Convert the cleaned and extracted feature sequences into a tensor format suitable for deep learning models.
[0085] 5.2 Feature Normalization
[0086] Normalize the input features (such as standardize to a mean of 0 and a variance of 1) to improve the convergence of the model.
[0087] The final output of the data processing module is the cleaned and optimized feature data, which can be directly input into the deep learning model, including:
[0088] Key point coordinates: two-dimensional / three-dimensional time series;
[0089] Derived features: joint angles, velocities, accelerations, etc.
[0090] (3) Joint range of motion recognition module: Based on the deep learning model, it is used to construct a human joint point recognition model, learn the relationship between the human joint movement trajectory and the angle, and use the trained human joint point recognition model to analyze and recognize the key point data to calculate the joint range of motion.
[0091] The composition of the joint range of motion recognition module includes:
[0092] Model data input layer: Used to input key point data, and also includes time series and derived features;
[0093] The key point data includes the (x, y) coordinates of two-dimensional key point data and the (x, y, z) coordinates of three-dimensional key point data, which is used to increase the depth information;
[0094] The time series is the motion trajectory information of the key points, usually a sliding window of several past frames; the derived features are the joint angles (such as the angles of the knee joint and elbow joint), angular velocity, and angular acceleration calculated based on the key points.
[0095] Feature extraction module: Includes time series feature extraction and spatial feature extraction;
[0096] Among them, the time series feature extraction is used to capture the dynamic changes of joint movement data and is implemented based on the LSTM network / Transformer network; the spatial feature extraction is used to obtain the range of motion information of the spatial structure of joint key points and is implemented based on the CNN network / GCN network. The convolutional neural network CNN is used to extract the human body shape features in the image, and the graph convolutional network GCN is used to extract the bone structure features.
[0097] Specifically, it is manifested as: using the CNN / GCN network to extract spatial features from the data, and then inputting the feature sequence extracted by the CNN / GCN into the LSTM / Transformer network, and using the LSTM / Transformer network to model the sequence data to learn the time series features of joint movement.
[0098] Multi-feature fusion module: Combines time series features and spatial features to generate the final range of motion prediction;
[0099] Integrate the time features generated by the LSTM / Transformer and the spatial features extracted by the GCN / CNN through weighted fusion.
[0100] It also includes a multi-task learning function, which predicts the range of motion of different joints simultaneously by setting up multiple branch networks, such as the knee joint, elbow joint, shoulder joint, etc.
[0101] Output layer: It includes the prediction of joint range of motion and auxiliary output;
[0102] Predict the angular values of the range of motion of each joint and output the predicted values with the dimension of the number of joints. The auxiliary output is to predict the health status (including normal activities, limited activities, abnormal activities) and provide the activity trajectory curve for dynamic analysis.
[0103] Preferably, the training and optimization of the human joint point recognition model include:
[0104] Data preprocessing and augmentation:
[0105] Use human pose datasets (such as MPII, COCO) and combine them with the labeled range of motion information, and perform data augmentation work by adding random noise, angle rotation, time series perturbation, etc.
[0106] In addition, data samples of different ages, genders, and body types can be added to reduce model bias.
[0107] Loss function:
[0108] Mean squared error (MSE): It is used to evaluate the error between the predicted range of motion and the true value, where:
[0109]
[0110] Optimizer:
[0111] Use Adam or SGD to optimize the model.
[0112] (4) Result output module: It is used to visually output the measurement results of joint range of motion (including real-time activity trajectory, angle curve graph, etc.) and can provide rehabilitation suggestions.
[0113] In addition, the result output module includes the display function of the dynamic angle change curve and the joint health analysis report. It can provide personalized rehabilitation plans and dynamic progress monitoring based on the historical data of the user.
[0114] Embodiment 2
[0115] As further elaborated, this embodiment provides a deep learning method for joint range of motion evaluation based on the above system, including the following steps:
[0116] S1 The user stands in front of the camera and performs joint activities, such as knee bending or shoulder rotation;
[0117] S2 acquires the original images at different angles during joint movement and at rest, processes the original images, and the system collects the user's activity data through a camera and calculates the coordinates of joint key points in real time;
[0118] S3 The human joint point recognition model predicts the joint range of motion according to the input data and generates a health assessment report.
[0119] Embodiment 3
[0120] For the convenience of expanding and integrating this system into other systems for use, an API interface is developed. Through the API interface, this system can be integrated into intelligent devices to support convenient joint health assessment effects. Other systems such as:
[0121] Integration with rehabilitation equipment: Link with physical therapy equipment (such as robotic arms, treadmills) to provide real-time feedback;
[0122] Integration with sports fitness APP: Generate activity reports to help users monitor their sports performance;
[0123] Integration with medical diagnosis systems: Link with the electronic health record (EHR) system to provide auxiliary diagnosis data for doctors.
[0124] Embodiment 4
[0125] The proposed actual application scenarios of this embodiment include, but are not limited to, fields such as medical treatment and rehabilitation, fitness and sports, and health monitoring of the elderly, with a wide application range and meeting diverse needs.
[0126] Medical treatment and rehabilitation:
[0127] Used to evaluate the joint recovery of patients and help doctors formulate rehabilitation plans. It can also be used to monitor the postoperative recovery effect and detect abnormal activities.
[0128] Fitness and sports:
[0129] Provide suggestions for correcting sports postures and optimizing the range of motion for fitness enthusiasts. It can also evaluate functional indicators such as the flexibility of athletes.
[0130] Health monitoring of the elderly:
[0131] Used for early detection of joint degeneration or limited mobility and provide relevant health maintenance suggestions.
[0132] It should be noted that although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, without departing from the spirit of the present invention, without creative design, structures and embodiments similar to this technical solution should all fall within the protection scope of the present invention.
Claims
1. A deep learning system for joint range of motion evaluation, characterized in that: include: Image acquisition module: used to collect joint activity trajectory images or videos at different angles and static posture data; Data processing module: used to pre-process the collected videos or images; Joint activity recognition module: Based on the deep learning model, it is used to build a human joint point recognition model, learn the relationship between the human joint activity trajectory and angle, use the trained human joint point recognition model to analyze and identify key point data, and calculate the joint activity; Result output module: used to visualize the measurement results of joint range of motion and provide rehabilitation suggestions.
2. A deep learning system for joint range of motion evaluation according to claim 1, characterized in that: The image acquisition module includes an RGB camera and a depth camera for acquiring videos or pictures of human joint activities. The RGB camera is used to acquire RGB images or videos for visual feature extraction to obtain two-dimensional position information, and the depth camera is used to obtain a depth map to capture the three-dimensional spatial position information of key points of the human body.
3. A deep learning system for joint range of motion evaluation according to claim 2, characterized in that: The preprocessing includes denoising, background segmentation, and extraction and detection of human key points.
4. A deep learning system for joint range of motion evaluation according to claim 3, characterized in that: The data processing module first predicts the two-dimensional key point coordinates of the human joints in the obtained image, and then measures the three-dimensional key point coordinates of each located joint pixel point using the binocular positioning principle. After removing outliers through a filtering algorithm, the precise spatial information of the human joints is obtained. The extracted two-dimensional / three-dimensional key point sequence is used to reflect the movement trajectory of the human body, thereby obtaining the skeleton mode.
5. A deep learning system for joint range of motion evaluation according to claim 1, characterized in that: The deep learning model includes an LSTM network for time series analysis, a CNN network for predicting activity, and a Transformer network for fusing multimodal input data.
6. A deep learning system for joint range of motion evaluation according to claim 5, characterized in that: The joint range of motion identification module comprises: Model data input layer: used to input key point data, including time series and derived features; The key point data includes the (x, y) coordinates of the two-dimensional key point data and the (x, y, z) coordinates of the three-dimensional key point data, which are used to increase the depth information; the time series is the motion trajectory information of the key points; the derived features are the joint angle, angular velocity and angular acceleration calculated based on the key points; Feature extraction module: including temporal feature extraction and spatial feature extraction; Temporal feature extraction is used to capture the dynamic changes of joint activity data, which is implemented based on LSTM network / Transformer network; spatial feature extraction is used to obtain the activity information of the spatial structure of joint key points, which is implemented based on CNN network / GCN network. Convolutional neural network CNN is used to extract human morphological features in the image, and graph convolutional network GCN is used to extract bone structure features; Multi-feature fusion module: combines temporal features and spatial features to generate the final activity prediction; Integrate the temporal features generated by LSTM / Transformer and the spatial features extracted by GCN / CNN through weighted fusion; Output layer: includes joint range of motion prediction and auxiliary output; Predict the angle value of each joint range of motion and output the predicted value with the dimension of the number of joints; the auxiliary output is the predicted health status and provides the activity trajectory curve for dynamic analysis.
7. A deep learning system for joint range of motion evaluation according to claim 6, characterized in that: The feature extraction module uses the CNN / GCN network to extract spatial features of the data, and then inputs the feature sequence extracted by the CNN / GCN into the LSTM / Transformer network, uses the LSTM / Transformer network to model the sequence data, and learns the temporal characteristics of joint activities.
8. A deep learning system for joint range of motion evaluation according to claim 3, characterized in that: The OpenPose algorithm is used to detect key points of the human body, and the joint angles are calculated based on the detection results.
Citation Information
Cited By
Artificial intelligence-based shoulder joint motion range automatic identification and analysis system, equipment and medium
CN120726700A
Joint motion range evaluation method and device based on 3D human body measurement and video guidance
CN121549805A