Joint angle prediction method based on depth camera and AI algorithm

Through the joint angle prediction method based on depth camera and AI algorithm, combined with human body markers and inverse kinematics, the problems of large error and poor real-time performance in traditional methods are solved, and real-time and accurate prediction of joint angles is achieved.

CN120766313APending Publication Date: 2025-10-10SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510862841.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Traditional joint angle acquisition methods have problems such as large errors and poor real-time performance.

Method used

A method based on depth camera and AI algorithm is adopted to obtain RGBD images of the human body, use the trained human landmark prediction model and joint angle prediction model, and combine it with inverse kinematics to achieve real-time prediction of joint angles.

Benefits of technology

It reduces environmental interference and manual measurement errors, realizes real-time and accurate prediction of joint angles, and improves the real-time and accuracy of acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766313A_ABST
    Figure CN120766313A_ABST
Patent Text Reader

Abstract

The invention discloses a joint angle prediction method based on a depth camera and an AI algorithm, and relates to the field of joint angle prediction. The invention aims to solve the problems of large error and poor real-time performance of the existing joint angle acquisition method. The method comprises the steps that human body RGBD images are collected, and a coordinate set of human body mark points of a subject corresponding to each human body RGBD image is extracted; the space coordinate set of the preprocessed human body mark points and the RGBD image form an own data set, the 2D public data set, the 3D public data set and the own data set form a training set, and the training set is utilized to train a human body mark point prediction model; obtaining a to-be-detected human body RGBD image, inputting the to-be-detected human body RGBD image into the trained human body mark point prediction model, and obtaining human body mark point coordinates in the to-be-detected human body image; and inputting human body mark point coordinates in a human body image to be detected into the trained joint angle prediction model to obtain a predicted human body joint angle. The method is used for predicting the joint angle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of joint angle prediction, and in particular to a joint angle prediction method based on a depth camera and an AI algorithm. Background Art

[0002] Changes in joint angles during human movement are crucial indicators for assessing motor function, objectively reflecting flexibility, coordination, and motor control. Accurately measuring parameters such as joint range of motion, trajectory, and rate of angle change can provide crucial data support for sports injury rehabilitation, motor skill optimization, fall prevention in the elderly, and neurological disease assessment.

[0003] Traditional joint angle acquisition methods rely primarily on manual measurement or motion capture systems to determine the angle between adjacent skeletal segments. However, these methods are susceptible to environmental interference, leading to large errors in joint angle measurement. Furthermore, these methods cannot generate results in real time, resulting in poor real-time performance. Summary of the Invention

[0004] The purpose of this invention is to solve the problems of large errors and poor real-time performance in existing joint angle acquisition methods, and to propose a joint angle prediction method based on a depth camera and AI algorithm.

[0005] A joint angle prediction method based on a depth camera and AI algorithm, specifically:

[0006] Step 1: Obtain an RGBD image of the human body to be tested, input the RGBD image of the human body to be tested into the trained human body landmark point prediction model, and obtain the coordinates of the human body landmark points in the RGBD image of the human body to be tested;

[0007] The trained human landmark prediction model is obtained by:

[0008] A1. Collect human RGBD images and extract the coordinate set of the subject's body landmark points corresponding to each human RGBD image;

[0009] A2. Preprocess the RGBD image and the coordinates of the human body marker points. Combine the preprocessed coordinate set of the human body marker points and the preprocessed RGBD image into a proprietary dataset. Combine the 2D public dataset, the 3D public dataset, and the proprietary dataset into a training set. Use the training set to train a human body marker point prediction model to obtain a trained human body marker point prediction model.

[0010] Step 2: Input the coordinates of the human body marker points in the RGBD image of the human body to be tested into the trained joint angle prediction model to obtain the predicted human joint angles.

[0011] Furthermore, the positions of the human body marker points include: sacrum, twelfth thoracic vertebra, fourth cervical vertebra, midpoint of the forehead-occipital center line, left acromion, left olecranon, midpoint of left wrist, midpoint of left third metacarpal bone, right acromion, right olecranon, midpoint of right wrist, midpoint of right third metacarpal bone, left femoral head, midpoint of left lateral-medial femoral condyle, midpoint of left lateral-medial malleolus, center point of left tarsal bone, right femoral head, midpoint of right lateral-medial femoral condyle, midpoint of right lateral-medial malleolus, center point of right tarsal bone, seventh cervical vertebra, end of left middle finger, end of left thumb, end of right middle finger, end of right thumb, left calcaneus, end of left first metatarsal bone, end of left fifth metatarsal bone, right calcaneus, end of right first metatarsal bone, end of right fifth metatarsal bone, tip of nose, center point of left eye, center point of right eye, center point of left auricle, center point of right auricle.

[0012] Furthermore, the RGBD image and the coordinates of the human body markers are preprocessed as follows:

[0013] Performing data enhancement operations on RGB images; the enhancement operations include: color jittering, rotation, and scaling;

[0014] Add random noise and perform geometric transformation on the depth image.

[0015] Furthermore, the human body landmark prediction model includes: a primary feature extraction module, a feature fusion module, an advanced feature extraction module, and a human body landmark prediction module;

[0016] The primary feature extraction module includes: an RGB primary feature extraction unit and a depth image primary feature extraction unit;

[0017] The RGB primary feature extraction unit is used to extract primary features from the RGB image;

[0018] The depth image primary feature extraction unit is used to extract primary features from the depth image;

[0019] The feature fusion module is used to fuse the primary features in the RGB image and the primary features in the depth image to obtain a feature fusion result;

[0020] The high-level feature extraction module obtains high-level features using feature fusion results;

[0021] The human body landmark prediction module is a fully connected layer, which is used to map high-level features into human body landmark coordinates.

[0022] Furthermore, the RGB primary feature extraction unit is used to extract primary features from the RGB image, specifically:

[0023]

[0024] in, It is the primary feature in the RGB image, is an RGB image, It is the RGB primary feature extraction network;

[0025] The RGB primary feature extraction network includes: a first convolutional layer, a second convolutional layer, a batch normalization layer, and an activation function layer;

[0026] The activation function is a Relu activation function;

[0027] The depth image primary feature extraction unit is used to extract primary features from the depth image, specifically:

[0028]

[0029] in, is the primary feature in the depth image, is the depth image, It is a deep image primary feature extraction network;

[0030] The network structure for extracting primary features of depth images is the same as the network structure for extracting primary features of RGB images.

[0031] Furthermore, the feature fusion module is used to fuse the primary features in the RGB image and the primary features in the depth image to obtain a feature fusion result, specifically:

[0032]

[0033] in, is the result of feature fusion.

[0034] Furthermore, the high-level feature extraction module obtains high-level features using the feature fusion results, specifically:

[0035]

[0036] in, It is the FasterNet-M network, It is an advanced feature.

[0037] Furthermore, the human body landmark prediction module is a fully connected layer, which is used to map high-level features into human body landmark coordinates, specifically:

[0038]

[0039] in, is a fully connected layer, It is the predicted coordinate set of human body landmark points.

[0040] Furthermore, the training of the human body landmark prediction model adopts the following loss function:

[0041]

[0042]

[0043]

[0044] in, is the body marker number, is the total number of body landmarks, is the 3D reprojection error loss function, It is The real coordinates of 3D human body markers, It is 3D human body marker predicted coordinates, is the projection function, is the coordinate of the i-th 2D human body marker point, yes The weight coefficient of yes The weight coefficient of is the 2D reprojection error loss function.

[0045] Furthermore, the trained joint angle prediction model is obtained by:

[0046] B1. Construct a musculoskeletal model. On the musculoskeletal model, construct skeletal markers corresponding to human markers in different motion states, obtain the coordinates of the skeletal markers, and use inverse kinematics to minimize the distance between the skeletal markers and the human markers. Obtain the joint angle corresponding to the shortest distance between the skeletal markers and the human markers.

[0047] B2. Preprocess the coordinates of the human body marker points, and form the preprocessed human body marker point positions and corresponding joint angle sets into a second training set;

[0048] The pre-processing of the position coordinates of the human body markers is specifically as follows: applying a preset rotation to the human body markers;

[0049] B3. Using the second training set to train the joint angle prediction model to obtain a trained joint angle prediction model;

[0050] The joint angle prediction model includes: a BiLSTM module and a Transformer module;

[0051] The BiLSTM module uses the coordinates of human body markers to obtain a feature vector containing global temporal information;

[0052] The Transformer module uses a feature vector containing global temporal information to obtain a set of human joint angles.

[0053] The beneficial effects of the present invention are:

[0054] This paper designs a method for identifying human joint angles based on visual recognition, inverse kinematics, and artificial intelligence. The method comprises four main components: multimodal data acquisition, inverse kinematics modeling, neural network training, and joint angle prediction. The method implements joint angle prediction using skeletal marker data, combined with inverse kinematics and a joint angle prediction model. This method uses a human key point prediction model and a joint angle prediction model to achieve joint angle prediction, reducing environmental interference and errors caused by manual measurement. Furthermore, the method can obtain joint angle prediction results in real time, improving the real-time nature of joint angle acquisition. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Flowchart of the present invention;

[0056] Figure 2 Schematic diagram of human body markers;

[0057] Figure 3 Schematic diagram of bony landmarks in the musculoskeletal model;

[0058] Figure 4 Get a flow chart for joint angles;

[0059] Figure 5 This is a graph showing the knee bending angle during walking;

[0060] Figure 6 A comparison of knee flexion angles in a single gait. DETAILED DESCRIPTION

[0061] Specific implementation method 1: Figure 1 As shown, the specific process of the joint angle prediction method based on a depth camera and an AI algorithm in this embodiment is as follows:

[0062] Step 1: Obtain an RGBD image of the human body to be tested, input the RGBD image of the human body to be tested into the trained human body landmark point prediction model, and obtain the coordinates of the human body landmark points in the RGBD image of the human body to be tested;

[0063] The trained human landmark prediction model is obtained by:

[0064] A1. Use a 3D depth camera to collect RGBD images of the human body, and use a motion capture system to extract the coordinate set of the subject's body marker points corresponding to each RGBD image of the human body. ;

[0065] in, It is a set of 36 body landmarks of any subject;

[0066] In this step, a depth camera array is constructed using the ToF module to form a stereo baseline configuration, supporting a high-definition resolution of 1280×720@60Hz and a depth detection accuracy of ±2mm to identify human landmarks.

[0067] The human body landmarks are distributed in the following locations: head, neck, torso 1-3, left / right shoulder, left / right elbow, left / right wrist, left / right hand, left / right hip, left / right knee, left / right ankle, left / right foot, left / right middle finger, left / right thumb, left / right heel, left / right first metatarsal bone, left / right fifth metatarsal bone, nose, left / right eye, left / right ear. The specific landmark positions and numbers are shown in Table 1 and Figure 2 shown.

[0068] Table 1

[0069]

[0070] In this step, the coordinate axes for obtaining the human body markers are set based on experience, and the relative positions between the coordinate axes obtained subsequently, so the position of the coordinate origin can be set based on experience.

[0071] A2. Preprocess the RGBD image and the coordinates of the human body marker points. Combine the preprocessed coordinate set of the human body marker points and the preprocessed RGBD image into a proprietary dataset. Combine the 2D public dataset, the 3D public dataset, and the proprietary dataset into a training set. Use the training set to train a human body marker point prediction model to obtain a trained human body marker point prediction model.

[0072] The 2D public datasets include: COCO, MPII, etc.

[0073] The 3D public datasets include: Human3.6M, MPI-INF-3DHP, etc.

[0074] The RGBD image is preprocessed, specifically: color jittering, rotation, scaling and other enhancement operations are performed on the RGB image, and random noise addition and geometric transformation are performed on the depth image;

[0075] The preprocessing of the spatial coordinates of the human body marker points specifically includes: randomly perturbing the spatial coordinates of the human body marker points to simulate noise in a real scene, etc.

[0076] The human body landmark prediction model includes: a primary feature extraction module, a feature fusion module, an advanced feature extraction module, and a human body landmark prediction module;

[0077] The primary feature extraction module includes: an RGB primary feature extraction unit and a depth image primary feature extraction unit;

[0078] The RGB primary feature extraction unit is used to extract primary features from the RGB image in RGBD, specifically:

[0079]

[0080] in, It is the primary feature in the RGB image, is an RGB image, It is the RGB primary feature extraction network;

[0081] The RGB primary feature extraction network includes: a first convolutional layer, a second convolutional layer, a batch normalization layer, and an activation function layer;

[0082] The activation function is a Relu activation function;

[0083] The depth image primary feature extraction unit is used to extract primary features from the depth image in RGBD, specifically:

[0084]

[0085] in, is the primary feature in the depth image, is the depth image, It is a deep image primary feature extraction network;

[0086] The network structure for extracting primary features of depth images is the same as the network structure for extracting primary features of RGB images.

[0087] The feature fusion module is used to fuse the primary features in the RGB image and the primary features in the depth image to obtain a feature fusion result, specifically:

[0088]

[0089] in, is the result of feature fusion;

[0090] The high-level feature extraction module uses the feature fusion results to obtain high-level features, specifically:

[0091]

[0092] in, It is the FasterNet-M network, It is an advanced feature;

[0093] The number of input channels of the convolutional layer in the FasterNet-M network is the number of channels;

[0094] The human marker point prediction module is a full connection layer, which is used for mapping high-level features into human marker point coordinates, and specifically comprises:

[0095]

[0096] wherein, is a full connection layer, is a human marker point prediction coordinate set, .

[0097] Step two, input the human marker point coordinates in the to-be-measured human body RGBD image into the trained joint angle prediction model to obtain the predicted human body joint angle;

[0098] The trained joint angle prediction model is obtained by the following way:

[0099] B1, based on OpenSim (or Anybody), a musculoskeletal model is constructed, the bone length in the musculoskeletal model is randomly scaled, the bone marker points corresponding to the human marker points in different motion states are constructed on the musculoskeletal model, the coordinates of the bone marker points are obtained, the bone marker points and the human marker point coordinates are made shortest by the inverse kinematics method, and the corresponding joint angle when the bone marker points and the human marker point coordinates are shortest is obtained, as shown in Figure 4 .

[0100] This step randomly scales the bone length in the musculoskeletal model, specifically: scaling the bone length ± 10%, simulating human bodies of different heights and limb proportions, thereby enhancing the adaptability of the subjects.

[0101] The joint angle includes: hip flexion angle, knee internal rotation angle, etc.

[0102] In this step, when the bone marker points are obtained, the action speed can be randomly changed in the time dimension (such as compressing the gait cycle to 80%-120% length); in the spatial dimension, random motion blur is added to simulate the capture scene of high-speed action.

[0103] B2, the human marker point position coordinates are preprocessed, and the preprocessed human marker point position and the corresponding joint angle set are combined to form a second training set;

[0104] The human marker point position coordinates are preprocessed, specifically: to improve the robustness of the model to measurement noise, a small rotation is applied to the human marker points, simulating the error caused by camera shaking or marker point sliding in actual application.

[0105] B3, the second training set is used to train the joint angle prediction model, and the trained joint angle prediction model is obtained.

[0106] The joint angle prediction model comprises a BiLSTM module and a Transformer module.

[0107] The BiLSTM module obtains a feature vector containing global time sequence information by using human marker point coordinates .

[0108] A bidirectional LSTM network is constructed to analyze the motion sequence from two directions, forward and backward. The LSTM of each layer gradually extracts time features such as rhythm and speed change of human motion. For example, when consecutive leg key points are recognized to move forward, the model can infer the swing phase in the gait cycle. The hidden state of each time step t is calculated as follows:

[0109]

[0110]

[0111] wherein, is the input of 36 human key points of the current frame, represents the bidirectional hidden state output at time step t, which contains sequence features extracted from the current frame and historical and future frames, is the forward hidden state at time step t, is the backward hidden state at time step t, is the forward hidden state at time step t-1, is the backward hidden state at time step t+1.

[0112] The output of the BiLSTM module is a vector set . (dimension: sequence length L x BiLSTM hidden state dimension );

[0113] The Transformer module obtains a set of human joint angles by using the feature vector containing global time sequence information.

[0114] In this step, spatial relationship modeling based on self-attention mechanism is established through multi-head Transformer, which can analyze the cooperative relationship between different joints. For example, when the hand key points are stretched forward, the model will control the linkage change of the shoulder joint and the elbow joint. The establishment method is as follows:

[0115] 1. Self-attention calculation:

[0116]

[0117] 2. Multi-head attention:

[0118]

[0119] in ;

[0120] 3. Encoder layer output:

[0121]

[0122] in, , is the query vector, is the key vector, V is the value vector, is a learnable weight matrix, is the dimension of the key vector, is the normalized exponential function, is the number of attention heads, is the final output weight matrix, It is a feed-forward network. It is the context feature sequence output by the Transformer encoder (dimension is sequence length L x Transformer model dimension ), , j is the attention head label;

[0123] 4. Fully connected output layer: obtain the set of human joint angles;

[0124] This step maps the context feature sequence output by the Transformer module to the final set of joint angles through one or more fully connected layers (MLPs);

[0125] Joint angles: The model outputs single-dimensional angle values ​​for 71 joints. These 71 dimensions directly correspond to the angles on the motion axes of human joints, such as knee flexion angles, hip internal and external rotation angles, etc.

[0126] like Figure 3 As shown, the red balls are human markers set on the human skeleton model, and the blue balls are human markers. This paper uses BiLSTM and Transformer to fuse spatiotemporal features. The BiLSTM captures temporal dynamics and extracts temporal features. The Transformer uses a self-attention mechanism to establish cross-joint associations and global dependencies, enabling the simultaneous capture of temporal dependencies of motion sequences and spatial relationships between joints. Ultimately, these fused spatiotemporal features are mapped to precise 71-dimensional human joint angles corresponding to each time step.

[0127] Specific implementation method 2: The loss function of the human body landmark prediction model trained using the training set is obtained by the following method:

[0128] First, the mean square error between the 3D human body marker coordinate points and the human body marker predicted coordinate points is used as the 3D reprojection error loss function:

[0129]

[0130] in, is the body marker number, is the total number of body landmarks, is the 3D reprojection error loss function, It is The real coordinates of 3D human body markers, It is Predicted coordinates of 3D human body markers;

[0131] In this step, 3D human body markers refer to the human body markers in the 3D dataset and the human body markers in the own dataset;

[0132] Then, the mean square error between the 2D human body marker coordinate points and the human body marker predicted coordinate points is used as the 2D reprojection error loss function, specifically:

[0133]

[0134] in, is the projection function used to map 3D points to the 2D image plane, is the coordinate of the i-th 2D human body marker point;

[0135] Finally, the total loss function is obtained by using the 2D reprojection error loss function and the 3D reprojection error loss function:

[0136]

[0137] in, is the total loss function, yes The weight coefficient of yes The weight coefficient of .

[0138] Example: In order to verify the beneficial effects of the present invention, the present invention conducted the following experiments:

[0139] The present invention is installed and deployed in a device such as an acquisition card, and a model quantization engine combined with a dynamic loading interface is used to implement the deployment of a new joint prediction module. The model quantization engine uses FP32 to INT8 conversion, and the dynamic loading interface supports loading the new joint prediction module at runtime. The following comparison is then performed:

[0140] 1. Robustness comparison:

[0141] Taking the change of knee bending angle during normal walking as an example, the angle obtained by the present invention is basically consistent with the knee joint angle calculated by inverse kinematics. Compared with the joint angle obtained by the traditional point-line method, the joint angle has better robustness. Figure 5 As shown, the knee angle calculated by the dotted line and the angle output by the model are compared with the joint angle of the motion capture system, with an accuracy of 86.3% and 91.4% respectively. Therefore, the joint angle output by the present invention is more accurate.

[0142] 2. Efficiency comparison:

[0143] The present invention can achieve end-to-end calculation based on bone points to joint angles, eliminating the complex parameter adjustment and pairing process of the inverse kinematics model. The model based on AI algorithm and GPU deployment can achieve faster calculation and can achieve real-time calculation compared to the inverse kinematics calculation method. Figure 6 As shown in the figure, using 2000 frames of gait data for calculation, a traditional inverse kinematics algorithm takes approximately 3 seconds to complete the calculation, while the present invention can control the calculation time to less than 75 milliseconds. Furthermore, the present invention eliminates the need for time-consuming inverse kinematics model parameter adjustment, significantly reducing the time required for joint angle calculation, and improving overall computational efficiency by over 40 times.

[0144] This invention simulates the inverse kinematics calculation process through an AI model. Compared to traditional line-to-line angle calculations based on body markers, the inverse kinematics model is more robust and computationally efficient. Therefore, the invention can achieve both higher efficiency and more accurate calculations.

Claims

1. A joint angle prediction method based on a depth camera and AI algorithm, characterized in that The specific process of the method is: Step 1: Obtain an RGBD image of the human body to be tested, input the RGBD image of the human body to be tested into the trained human body landmark point prediction model, and obtain the coordinates of the human body landmark points in the RGBD image of the human body to be tested; The trained human landmark prediction model is obtained by: A1. Collect human RGBD images and extract the coordinate set of the subject's body landmark points corresponding to each human RGBD image; A2. Preprocess the RGBD image and the coordinates of the human body marker points. Combine the preprocessed coordinate set of the human body marker points and the preprocessed RGBD image into a proprietary dataset. Combine the 2D public dataset, the 3D public dataset, and the proprietary dataset into a training set. Use the training set to train a human body marker point prediction model to obtain a trained human body marker point prediction model. Step 2: Input the coordinates of the human body marker points in the RGBD image of the human body to be tested into the trained joint angle prediction model to obtain the predicted human joint angles.

2. The joint angle prediction method based on a depth camera and an AI algorithm according to claim 1, characterized in that: The positions of the human body landmarks include: sacral vertebra, twelfth thoracic vertebra, fourth cervical vertebra, midpoint of the forehead-occipital line, left acromion, left olecranon, midpoint of left wrist, midpoint of left third metacarpal bone, right acromion, right olecranon, midpoint of right wrist, midpoint of right third metacarpal bone, left femoral head, midpoint of left lateral-medial femoral condyle, midpoint of left lateral-medial malleolus, center point of left tarsal bone, right femoral head, midpoint of right lateral-medial femoral condyle, midpoint of right lateral-medial malleolus, center point of right tarsal bone, seventh cervical vertebra, end of left middle finger, end of left thumb, end of right middle finger, end of right thumb, left calcaneus, end of left first metatarsal bone, end of left fifth metatarsal bone, right calcaneus, end of right first metatarsal bone, end of right fifth metatarsal bone, tip of nose, center point of left eye, center point of right eye, center point of left auricle, center point of right auricle.

3. The joint angle prediction method based on a depth camera and an AI algorithm according to claim 2, characterized in that: The RGBD image and the coordinates of the human body markers are preprocessed as follows: Performing data enhancement operations on RGB images; the enhancement operations include: color jittering, rotation, and scaling; Add random noise and perform geometric transformation on the depth image.

4. The joint angle prediction method based on a depth camera and an AI algorithm according to claim 3, characterized in that: The human body landmark prediction model includes: a primary feature extraction module, a feature fusion module, an advanced feature extraction module, and a human body landmark prediction module; The primary feature extraction module includes: an RGB primary feature extraction unit and a depth image primary feature extraction unit; The RGB primary feature extraction unit is used to extract primary features from the RGB image; The depth image primary feature extraction unit is used to extract primary features from the depth image; The feature fusion module is used to fuse the primary features in the RGB image and the primary features in the depth image to obtain a feature fusion result; The high-level feature extraction module obtains high-level features using feature fusion results; The human body landmark prediction module is a fully connected layer, which is used to map high-level features into human body landmark coordinates.

5. The joint angle prediction method based on a depth camera and an AI algorithm according to claim 4, characterized in that: The RGB primary feature extraction unit is used to extract primary features from the RGB image, specifically: in, is the primary feature in the RGB image, is an RGB image, It is the RGB primary feature extraction network; The RGB primary feature extraction network includes: a first convolutional layer, a second convolutional layer, a batch normalization layer, and an activation function layer; The activation function is a Relu activation function; The depth image primary feature extraction unit is used to extract primary features from the depth image, specifically: in, is the primary feature in the depth image, is the depth image, It is a deep image primary feature extraction network; The network structure for extracting primary features of depth images is the same as the network structure for extracting primary features of RGB images.

6. The joint angle prediction method based on a depth camera and an AI algorithm according to claim 5, characterized in that: The feature fusion module is used to fuse the primary features in the RGB image and the primary features in the depth image to obtain a feature fusion result, specifically: in, is the result of feature fusion.

7. The joint angle prediction method based on a depth camera and an AI algorithm according to claim 6, characterized in that: The high-level feature extraction module uses the feature fusion results to obtain high-level features, specifically: in, It is the FasterNet-M network, It is an advanced feature.

8. The joint angle prediction method based on a depth camera and an AI algorithm according to claim 7, characterized in that: The human body landmark prediction module is a fully connected layer that maps high-level features to human body landmark coordinates, specifically: in, is a fully connected layer, It is the predicted coordinate set of human body landmark points.

9. The joint angle prediction method based on a depth camera and an AI algorithm according to claim 8, characterized in that: The training of the human landmark prediction model adopts the following loss function: in, is the body marker number, is the total number of body landmarks, is the 3D reprojection error loss function, It is The real coordinates of 3D human body markers, It is 3D human body marker predicted coordinates, is the projection function, is the coordinate of the i-th 2D human body marker point, yes The weight coefficient of yes The weight coefficient of is the 2D reprojection error loss function.

10. The joint angle prediction method based on a depth camera and an AI algorithm according to claim 9, characterized in that: The trained joint angle prediction model is obtained by: B1. Construct a musculoskeletal model. On the musculoskeletal model, construct skeletal markers corresponding to human markers in different motion states, obtain the coordinates of the skeletal markers, and use inverse kinematics to minimize the distance between the skeletal markers and the human markers. Obtain the joint angle corresponding to the shortest distance between the skeletal markers and the human markers. B2. Preprocess the coordinates of the human body marker points, and form the preprocessed human body marker point positions and corresponding joint angle sets into a second training set; The pre-processing of the position coordinates of the human body markers is specifically as follows: applying a preset rotation to the human body markers; B3. Using the second training set to train the joint angle prediction model to obtain a trained joint angle prediction model; The joint angle prediction model includes: a BiLSTM module and a Transformer module; The BiLSTM module uses the coordinates of human body markers to obtain a feature vector containing global temporal information; The Transformer module uses a feature vector containing global temporal information to obtain a set of human joint angles.

Citation Information

Patent Citations

  • Real-time human body posture inverse kinematics solving method and device

    CN112037312A

  • Human body posture prediction method and system based on attention mechanism fused with multi-stream graph

    CN113642379A

  • Swimming posture evaluation method, system and device and computer readable storage medium

    CN114783001A

  • Skeleton detection model construction method and image data identification method

    CN115620016A

  • Joint angle prediction method based on motion unit and neuromusculoskeletal model

    CN117195024A