Method for hierarchically estimating 3D human body posture from sparse inertial measurement unit (IMU)
Through the hierarchical shared structure of the hierarchical attitude evaluation model and the Mamba module, the conflict problem of sparse IMU's position reconstruction in 3D human pose estimation is solved, and efficient and real-time 3D human pose estimation is achieved to adapt to different environments and user needs.
Patent Information
- Application Number
- CN202510190557.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, when using the sparse inertial measurement unit (IMU) to perform 3D human posture estimation, there is negative motion information transmission and conflict in the posture reconstruction process of different body parts, and the kinematic correlation between different parts is not effectively utilized.
The hierarchical pose evaluation model is used to extract and share the hidden features of body parts hierarchically through the Mamba module, and a hierarchical shared structure based on the selective state space model (S-SSM) is designed, and the motion information of different body parts is gradually learned, and the pose estimation results are reconstructed through linear functions.
Reduce the number of sensors, improve user comfort, reduce costs and energy consumption, improve system adaptability and robustness, enhance algorithm generalization capabilities, support a wide range of application scenarios, adapt to harsh environments, reduce pose estimation complexity, and realize real-time pose estimation.
Smart Images

Figure CN120267274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human pose estimation, and specifically to a method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs). Background Technique
[0002] 3D Human Pose Estimation (HPE) is the process of determining the positions of human joints in a three-dimensional coordinate system through motion signals. This is crucial in many practical applications, including motion sensing games, competitive sports, medical rehabilitation, and emergency rescue. These applications require the analysis of changes in human poses.
[0003] When capturing motion signals, compared with visual sensors and environmental sensors, wearable sensors such as Inertial Measurement Units (IMUs) are hardly affected by environmental factors such as object occlusion, insufficient light, and space limitations. In addition, they also provide privacy protection for users.
[0004] Early work attempted to capture dynamic human motion information through sparse IMU configurations to achieve full-body pose reconstruction, but this task has proven to be challenging. Other complex network models have achieved good performance in full-body pose estimation, and it has been found that the information of the end joints helps to reconstruct human poses. However, the contributions of the end joints of different limbs are still ambiguous.
[0005] The latest research has achieved state-of-the-art results in pose estimation by dividing the end joints into three parts: the torso, lower limbs, and upper limbs. However, studies have shown that their method still has unclear problems, which may lead to negative motion information transmission during the pose reconstruction of different body parts. Specifically, this work uses a shared layer to estimate the poses of the three body parts in parallel and independently. However, different parts have different motion complexities and information, which may lead to conflicts between parts during the reconstruction of human poses. In addition, this learning method treats each task as equally important, resulting in repetitive small pose estimation tasks and not clearly reflecting the kinematic correlations between different parts. Summary of the Invention
[0006] The object of the present invention is to provide a method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), including the following steps:
[0007] 1) Obtain the measurement data of inertial measurement units deployed at key parts of the human body, and perform preprocessing to obtain the input signal
[0008] 2) Construct a hierarchical pose evaluation model based on the Mamba module;
[0009] 3) Process the input signal X using the hierarchical pose evaluation model to obtain the human pose estimation result;
[0010] The human pose estimation result is visualized through the 3D human SMPL model;
[0011] The human pose estimation result includes the body part pose estimation result Y Pose (i.e., the 3D rotation matrix of each joint point of the SMPL model) and the global translation estimation result Y Trans (i.e., the 3D coordinates of the spine root joint point in the SMPL model);
[0012] The body part pose estimation result Y Pose ={Y Torso , Y Lower , Y Uoper}; Y Torso , Y Lower , Y Upper are the torso pose estimation result, the lower limb pose estimation result, and the upper limb pose estimation respectively.
[0013] The global translation estimation result Y Trans includes the 3D coordinates of the spine root joint point in the SMPL model.
[0014] Furthermore, in step 1), the preprocessing steps include: aligning the measurement data of the inertial measurement units deployed at the key human body parts with the spine root joint, and performing normalization and splicing operations;
[0015] The key human body parts include the head, waist, wrist, and knee. Among them, one inertial measurement unit is arranged at the head and waist, and two inertial measurement units are arranged at the wrist and knee.
[0016] Furthermore, the input signal X = {x A , x R}; represents the triaxial acceleration measurement value, represents the joint rotation matrix of the selected joint in 3D space.
[0017] Furthermore, the hierarchical pose evaluation model based on the Mamba module includes the Mamba module and the pose estimation task heads for different body parts;
[0018] The Mamba module processes the input signal X, hierarchically extracts and shares the motion hidden features of different body parts in X;
[0019] The pose estimation task head for different body parts uses a linear function to reconstruct the body part pose estimation result Y from the motion hidden features of different body parts Pose and the global translation estimation result Y Trans .
[0020] Furthermore, the Mamba module is a Mamba module based on the Selective-State Space Models (S-SSM);
[0021] The internal unit structure of the selective state space model S-SSM is as follows:
[0022]
[0023] where h t+1 、h t represent the hidden features of the (t + 1)-th frame and the t-th frame; Δ is the time step; B (Δ) is the input matrix; A (Δ) is the state matrix; C is the output matrix; S t 、S t-1 are the state variables of the t-th frame and the (t - 1)-th frame; Δ, B, and A are all internal calculation results within the Mamba module.
[0024] The state matrix A (Δ) is as follows:
[0025]
[0026] where i and j represent the row and column coordinates of the state matrix inside the Mamba module.
[0027] Furthermore, when processing the input signal X using the pose evaluation model, the time step Δ, the input matrix B (Δ) 、the output matrix C are dynamically updated;
[0028] The dynamic update method of the internal parameters p of the time step Δ, the input matrix B (Δ) 、the output matrix C is as follows:
[0029]
[0030] where h t is the state hidden feature of the Mamba module at the t-th frame; p C (t), p Δ (t) are respectively h tInternal parameters determined after linear mapping by the Linear layer; p are their respective internal parameters, determined by the state selection space structure of Mamba through the Linear mapping of h during model training. t is determined by the Linear mapping of
[0031] Furthermore, the Mamba module is as follows:
[0032]
[0033] where is the input of the current layer; Cap is the motion feature; σ represents the SiLU function.
[0034] is the output of the Mamba module.
[0035] Furthermore, the output Y n (t) of the pose evaluation model is as follows:
[0036]
[0037] where n ∈ {Torso,Lower,Upper,Trans};
[0038] where the motion feature is as follows:
[0039]
[0040] where ⊕ represents concatenation; R init is the human joint rotation matrix; P init is the global position.
[0041] Furthermore, the human joint rotation matrix R unit and the global position P init are as follows:
[0042]
[0043] where represents the frame; T is the time period.
[0044] Furthermore, the hierarchical pose evaluation model based on the Mamba module is trained using a historical dataset; the historical dataset includes measurement data from inertial measurement units deployed at key body parts of the human body, and corresponding human pose estimation results;
[0045] During training, the loss function of the hierarchical pose model is as follows:
[0046]
[0047] Among them, Y n (t) is the output of the hierarchical pose evaluation model, and Y′ n (t) is the actual value.
[0048] It should be noted that the present invention proposes a 3D HPE method (HiPoser) based on hierarchical shared learning. This method sequentially learns the motion information and features of different body parts, and finally performs local body part pose estimation and global translation estimation through different task heads. Specifically, the present invention decomposes the 3D HPE process into four tasks: 1) torso pose estimation, 2) lower limb pose estimation, 3) upper limb pose estimation, and 4) global translation estimation, and designs a hierarchical shared structure with the Mamba module as the backbone to promote the learning and sharing of features between different body parts. As Figure 1 shown, this innovative hierarchical shared structure solves the task conflict problem that appears in the previous 3D HPE methods based on multi-task learning (MTL). In addition, the present invention incorporates the motion state (i.e., joint pose and body position) into the hierarchical shared structure of each action window, enhancing the stability performance of HiPoser during the motion process. Considering the motion coordination between body parts, HiPoser can adjust the task order to achieve different levels of priority performance (such as detail priority and stability priority).
[0049] The technical effects of the present invention are beyond doubt, and the beneficial effects of the present invention are as follows:
[0050] 1. Reduce the number of sensors and improve user comfort: The present invention estimates the 3D human pose through sparsely distributed IMU sensors, which can significantly reduce the number of sensors on the wearable device, reduce the weight and volume of the device, improve the wearing comfort of the user, reduce the interference with daily activities, and is suitable for long-term wearing and sports scenarios.
[0051] 2. Reduce costs and energy consumption: The present invention can reduce the manufacturing cost and maintenance cost of the device, and reduce the energy consumption required for data collection and transmission.
[0052] 3. Improve system adaptability and deployment flexibility: The present invention can be flexibly deployed in different types of users or application scenarios. Even if the number of IMUs is limited, the present invention can still effectively estimate the 3D pose through intelligent algorithms, thus adapting to the needs of different populations.
[0053] 4. Enhance the robustness and generalization ability of the algorithm: The present invention can still maintain high accuracy when the number of sensors is limited. This robustness and generalization ability ensure the effectiveness of the system under different human body sizes, poses or environments, and reduce the dependence on specific training data.
[0054] 5. Support a wide range of application scenarios: This technology can be applied to multiple fields, such as health monitoring, motion analysis, virtual reality, augmented reality, and human motion capture, etc.
[0055] 6. Reduce the computational complexity of pose estimation: Due to the reduction of sensor data input, the present invention can optimize the computational process, reduce the complexity and processing time of pose estimation, making real-time pose estimation possible and meeting the application requirements for low latency and high response speed, especially in real-time feedback or interaction scenarios.
[0056] 7. Adapt to harsh or special environments: The sparse IMU system can be used in special or harsh environments where a large number of sensors cannot be deployed, such as underwater, space, battlefield and other scenarios. Through limited sensor data, the system can still effectively estimate human poses and ensure stable performance in various environments. Brief Description of the Drawings
[0057] Figure 1 is a hierarchical sharing structure;
[0058] Figure 2 is the human pose recognition process;
[0059] Figure 3 is the result visualization comparison. Detailed Embodiments
[0060] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above-mentioned subject scope of the present invention is limited to the following embodiments. Without departing from the above-mentioned technical idea of the present invention, various substitutions and changes made according to ordinary technical knowledge and customary means in the art shall be included within the protection scope of the present invention.
[0061] Embodiment 1:
[0062] See Figures 1 to 3 , a method for hierarchically estimating 3D human poses from sparse inertial measurement units (IMUs), comprising the following steps:
[0063] 1) Obtain the measurement data of inertial measurement units deployed at key parts of the human body, and perform preprocessing to obtain the input signal
[0064] 2) Construct a hierarchical pose evaluation model based on the Mamba module;
[0065] 3) Use the hierarchical pose evaluation model to process the input signal X to obtain the human pose estimation result;
[0066] The human pose estimation result is visualized through a three-dimensional human SMPL model;
[0067] The human body pose estimation result includes the body part pose estimation result Y Pose (i.e., the three-dimensional rotation matrix of each joint point of the SMPL model) and the global translation estimation result Y Trans (i.e., the three-dimensional coordinates of the spine root joint point in the SMPL model);
[0068] The body part pose estimation result Y Pose ={Y Torso ,Y Lower ,Y Upper}; Y Torso ,Y Lower ,Y Upper are the trunk pose estimation result, the lower limb pose estimation result, and the upper limb pose estimation respectively.
[0069] The global translation estimation result Y Trans includes the three-dimensional coordinates of the spine root joint point in the SMPL model.
[0070] In step 1), the steps for preprocessing include: aligning the measurement data of the inertial measurement units deployed at the key parts of the human body with the spine root joint, and performing normalization and splicing operations;
[0071] The key parts of the human body include the head, waist, wrist, and knee. Among them, one inertial measurement unit is arranged on the head and waist, and two inertial measurement units are arranged on the wrist and knee.
[0072] The input signal X = {x A , x R}; represents the three-axis acceleration measurement value, represents the joint rotation matrix of the selected joint in three-dimensional space.
[0073] The pose hierarchical evaluation model based on the Mamba module includes the Mamba module and the pose estimation task heads for different body parts;
[0074] The Mamba module processes the input signal X, and hierarchically extracts and shares the motion hidden features of different body parts in X;
[0075] The pose estimation task heads for different body parts use linear functions to reconstruct the body part pose estimation result Y Pose and the global translation estimation result Y Trans .
[0076] The Mamba module is a Mamba module based on the Selective-State Space Models (S-SSM);
[0077] The internal unit structure based on the selective state space model S-SSM is as follows:
[0078]
[0079] Among them, h t+1 , h t represent the hidden features of the (t + 1)-th frame and the t-th frame; Δ is the time step; B (Δ) is the input matrix; A (Δ) is the state matrix; C is the output matrix; S t , S t-1 are the state variables of the t-th frame and the (t - 1)-th frame; Δ, B, and A are all internal calculation results within the Mamba module.
[0080] The state matrix A (Δ) is as follows:
[0081]
[0082] Among them, i and j represent the horizontal and vertical coordinates of the state matrix inside the Mamba module.
[0083] When using the pose evaluation model to process the input signal X, the time step Δ, the input matrix B (Δ) , and the output matrix C are dynamically updated;
[0084] The dynamic update methods of the internal parameters p of the time step Δ, the input matrix B (Δ) , and the output matrix C are as follows:
[0085]
[0086] Among them, h t is the state hidden feature of the Mamba module at the t-th frame; p c (t), p Δ (t) are respectively the internal parameters determined after linear mapping of h t through the Linear layer; p is their respective internal parameter. It is determined by the Linear linear mapping of h t in the state selection space structure of Mamba during model training.
[0087] The Mamba module is as follows:
[0088]
[0089] Among them, is the input of the current layer; Cap is the motion feature; σ represents the SiLU function. Is the output of the Mamba module.
[0090] The output Y of the pose evaluation model n (t) is as follows:
[0091]
[0092] Where n ∈ {Torso, Lower, Upper, Trans};
[0093] Among them, the motion feature is as follows:
[0094]
[0095] Among them, ⊕ represents concatenation; R init is the human joint rotation matrix; P init is the global position.
[0096] The human joint rotation matrix R init and the global position P init are as follows:
[0097]
[0098] Among them, represents the frame; T is the time period.
[0099] The hierarchical pose evaluation model based on the Mamba module is trained with a historical dataset; the historical dataset includes measurement data of inertial measurement units deployed on key body parts of the human body and corresponding human pose estimation results;
[0100] During the training process, the loss function of the hierarchical pose model is as follows:
[0101]
[0102] Among them, Y n (t) is the output of the hierarchical pose evaluation model, and Y' n (t) is the actual value.
[0103] Example 2:
[0104] A method for hierarchically estimating 3D human poses from sparse inertial measurement units (IMUs) includes the following steps:
[0105] 1) Obtain the measurement data of inertial measurement units deployed on key parts of the human body and perform preprocessing to obtain the input signal
[0106] 2) Construct a hierarchical pose evaluation model based on the Mamba module;
[0107] 3) Process the input signal X using the hierarchical pose evaluation model to obtain the human pose estimation result;
[0108] The human pose estimation result is visualized through a three-dimensional human SMPL model;
[0109] The human pose estimation result includes the body part pose estimation result Y Pose (i.e., the three-dimensional rotation matrix of each joint point of the SMPL model) and the global translation estimation result Y Trans (i.e., the three-dimensional coordinates of the spine root joint point in the SMPL model);
[0110] The body part pose estimation result Y Pose ={Y Torso , Y Lower , Y Upper}; Y Torso , Y Lower , Y Upper are respectively the trunk pose estimation result, the lower limb pose estimation result, and the upper limb pose estimation.
[0111] The global translation estimation result Y Trans includes the three-dimensional coordinates of the spine root joint point in the SMPL model.
[0112] Example 3:
[0113] A method for hierarchically estimating 3D human poses from sparse inertial measurement units (IMUs), the technical content is the same as that in Example 2. Further, in step 1), the steps for preprocessing include: aligning the measurement data of the inertial measurement units deployed at the key human body parts with the spine root joint, and performing normalization and splicing operations;
[0114] The key human body parts include the head, waist, wrist, and knee. Among them, one inertial measurement unit is arranged at the head and waist, and two inertial measurement units are arranged at the wrist and knee.
[0115] Example 4:
[0116] A method for hierarchically estimating 3D human poses from sparse inertial measurement units (IMUs), the technical content is the same as any one of Examples 2-3. Further, the input signal X = {x A , x R}; represents the three-axis acceleration measurement value, represents the joint rotation matrix of the selected joint in three-dimensional space.
[0117] Example 5:
[0118] A method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), the technical content is the same as any one of Examples 2-4. Further, the pose hierarchical evaluation model based on the Mamba module includes the Mamba module and pose estimation task heads for different body parts;
[0119] The Mamba module processes the input signal X, hierarchically extracts and shares the motion hidden features of different body parts in X;
[0120] The pose estimation task heads for different body parts use linear functions to reconstruct the body part pose estimation result Y from the motion hidden features of different body parts Pose and the global translation estimation result Y Trans .
[0121] Example 6:
[0122] A method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), the technical content is the same as any one of Examples 2-5. Further, the Mamba module is a Mamba module based on the Selective-State Space Models (S-SSM);
[0123] The internal unit structure of the Selective State Space Model S-SSM is as follows:
[0124]
[0125] where h t+1 、h t represent the hidden features of the (t + 1)-th frame and the t-th frame; Δ is the time step; B (Δ) is the input matrix; A (Δ) is the state matrix; C is the output matrix; S t 、S t-1 are the state variables of the t-th frame and the (t - 1)-th frame; Δ, B, and A are all internal calculation results within the Mamba module.
[0126] The state matrix A (Δ) is as follows:
[0127]
[0128] where i and j represent the horizontal and vertical coordinates of the state matrix inside the Mamba module.
[0129] Example 7:
[0130] A method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), the technical content is the same as any one of Embodiments 2-6. Further, when using the pose evaluation model to process the input signal X, the time step Δ, the input matrix B (Δ) , and the output matrix C are dynamically updated;
[0131] The dynamic update methods of the time step Δ, the input matrix B (Δ) , and the internal parameter p of the output matrix C are as follows:
[0132]
[0133] where h t is the state hidden feature of the Mamba module at the t-th frame; p C (t) and p Δ (t) are respectively the internal parameters determined after the linear mapping of h t through the Linear layer; p is their respective internal parameter. The internal parameter is determined by the Linear linear mapping of h t in the state selection space structure of Mamba during model training.
[0134] Embodiment 8:
[0135] A method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), the technical content is the same as any one of Embodiments 2-7. Further, the Mamba module is as follows:
[0136]
[0137] where is the input of the current level; Cap is the motion feature; σ represents the SiLU function.
[0138] Embodiment 9:
[0139] A method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), the technical content is the same as any one of Embodiments 2-8. Further, the output of the pose evaluation model is as follows:
[0140]
[0141] In the formula, n ∈ {Torso, Lower, Upper, Trans};
[0142] where the motion feature is as follows:
[0143]
[0144] Among them, ⊕ represents concatenation; R init is the human joint rotation matrix; P init is the global position.
[0145] Example 10:
[0146] A method for hierarchically estimating 3D human poses from sparse inertial measurement units (IMUs), the technical content is the same as any one of Examples 2-9. Further, the human joint rotation matrix R init , the global position P init is as follows:
[0147]
[0148] Among them, represents a frame; T is the time period.
[0149] Example 11:
[0150] A method for hierarchically estimating 3D human poses from sparse inertial measurement units (IMUs), the technical content is the same as any one of Examples 2-10. Further, the hierarchical pose evaluation model based on the Mamba module is trained through a historical dataset; the historical dataset includes the measurement data of inertial measurement units deployed on key body parts of the human body and the corresponding human pose estimation results;
[0151] During the training process, the loss function of the hierarchical pose model is as follows:
[0152]
[0153] Among them, Y n (t) is the output of the hierarchical pose evaluation model, and Y′ n (t) is the actual value.
[0154] Example 12:
[0155] A method for hierarchically estimating 3D human poses from sparse inertial measurement units (IMUs) is as follows:
[0156] The SMPL model can use its linear blend skinning technology to present a realistic human shape and effectively represent dynamic human poses through joint rotation matrices. The IMU configuration of the present invention refers to related work and is consistent with the joint positions of the SMPL model, so as to efficiently capture the poses of body parts during movement.
[0157] Given a sample set containing action events This sample set includes samples, through An IMU records the motion changes of the human body posture. Each sample is a sequence of T frames collected at a unified time scale. The present invention aligns the IMU measurement values with the spinal root joints, and performs normalization and splicing operations to obtain the input signal where represents the triaxial acceleration measurement value, represents the joint rotation matrix of the selected joint in three-dimensional space. The output is Y = {Y Pose , Y Trans}, where Y Pose ∈R 24×3×3 is the rotation matrix of 24 joints of the SMPL model, represents the position in the global triaxial coordinate system. The model of the present invention obtains the output Y(t) of the t-th frame by estimating the input X(t) at time t, where t = {1,…,T}.
[0158] The overall architecture of HiPoser is shown in the figure. The present invention designs a hierarchical sharing structure, enabling the motion features of the torso, lower limbs, and upper limbs to be shared at different levels. As the core of the hierarchical sharing structure, the Mamba module can infer the correct posture based on the current motion information and motion state, which is crucial for reducing the posture accumulation error in IMU-based three-dimensional human posture estimation. Finally, a simple linear function serves as the task head to efficiently reconstruct the posture of each body part and the translation of the entire body. The hierarchical sharing structure of HiPoser allows for flexible setting of the task order, realizing three-dimensional human posture estimation with different priorities.
[0159] The Mamba module in action processing
[0160] Previous work has chosen RNN, Bi-RNN, or Transformer to capture motion-related features. However, these methods may not be effective solutions for IMU-based three-dimensional human posture estimation in terms of resource consumption and historical inference. To meet the requirements of efficient inference of motion postures and capturing all necessary information in the context, the present invention adopts the Mamba module based on the selective state space model (S-SSM) as the backbone of the hierarchical sharing structure. Specifically, the S-SSM can selectively process motion features and effectively remember long-term motion information. The present invention sets S t as the state variable of the t-th frame in the shared layer, and the formula is as follows:
[0161] S t = A (Δ) S t-1 + B (Δ) h t ,
[0162] h t+1 = CS t
[0163] where h represents the hidden feature in S-SSM. Δ is the time step and is calculated on the discretized matrices A (Δ) and B (Δ) The input matrix B (Δ) acts directly on h t , and the state matrix A (Δ) stores all the historical motion information. The output matrix C defines the linear mapping relationship from S t to h t+1 , and there is no need to discretize Δ.
[0164] Motion information storage
[0165] Since the action state update depends on the stored historical motion information, A (Δ) plays an important role in action process modeling. To efficiently solve the long-distance dependence problem in action process modeling within limited storage space, A (Δ) uses a high-order polynomial projection operator to compress all the current input information from S t-1 as follows:
[0166]
[0167] where A (Δ) is a coefficient matrix. By calculating A (Δ) S t-1 , all the historical states at time t can be approximated infinitely to handle the long-range dependence relationship.
[0168] Selective mechanism
[0169] During the motion inference process, the model parameters remain unchanged, which may lead to different motion features being calculated with the same state matrix A (Δ) , thus losing the targeted inference ability for the current human pose and historical action information. To solve the problem of the irrelevance between the input motion and the state space, the present invention adopts a simple selection mechanism to dynamically calculate the parameters p of B (Δ) , C, and Δ as follows:
[0170]
[0171] p C (t) = Linear C (h t )
[0172] p Δ (t) = LinearΔ(h t )
[0173] where h tThrough Linear projection, the parameter p C and p Δ are determined during the training process. In fact, the state matrix A (Δ) storing historical motion information is discretized by Δ and combined with B (Δ) to obtain the new state S t , so that A (Δ) realizes data dependence in a parameter - efficient manner, that is, A (Δ) is able to utilize h t to generate the new state S t , enabling the model of the present invention to selectively perform pose inference.
[0174] Mamba module
[0175] To capture as much necessary information in the context as possible, a complete Mamba module (MB) can be represented as follows:
[0176]
[0177] where is the input of the current layer. Cap captures motion features through S - SSM. σ represents the SiLU function. The Linear function increases the dimension of, in order to capture more detailed and complex motion features in a higher - dimensional solution space. The convolutional operation Conv enhances the ability of the Mamba module to capture short - range local motion features, complements the capture of long - term dependencies by S - SSM, and forms a complex representation of motion information.
[0178] Hierarchical shared learning
[0179] In the hierarchical shared structure of the present invention, implementing shallow - layer tasks can enhance the model's utilization of the basic motion information of body parts related to motion, thus being more effective when performing deeper - layer tasks. Compared with directly regressing the full - body pose using IMU measurements or using body - part - level motion information, the present invention solves the problem of motion blur of human body parts and conducts hierarchical shared learning in a more detailed manner, gradually exploring the associations between part - level tasks, which further prevents the problem of conflicting estimation tasks and improves the performance of 3D human pose estimation specifically based on the basic features of different body parts. In addition, the present invention introduces the motion state to enable the model to have a more stable effect in reconstructing the poses of each part during random and complex motions.
[0180] Task determination
[0181] To achieve hierarchical shared learning of part - motion information at different levels, the present invention divides 3D human pose estimation into four tasks: torso (6 joints) pose estimation Lower limb (8 joints) pose estimation Upper limb (10 joints) pose estimation (i.e., Y Pose ={Y Torso , Y Lower , Y Upper}); and three-axis body translation estimation These tasks can sequentially obtain information from the previous task and pass it backward in a hierarchical sharing structure, thus realizing information flow sharing.
[0182] Motion state
[0183] Determining the starting position of the human body and the amplitude of pose change can effectively improve the stability of the model in subsequent motion pose estimation tasks. To obtain more details of the human body in the global coordinate system, the present invention extracts the joint rotation matrix R init and the global position P init , to represent the pose state and position state of the human body, defined as follows:
[0184]
[0185] Where All joints are aligned to the spinal joint, P init and R init are initially set to 0. After that, P init and R init are determined by Y Pose and Y Trans of the last frame of the sequence.
[0186] Hierarchical sharing structure
[0187] Considering that recent work has focused more on pose reconstruction, the present invention temporarily uses the task order Torso→Lower→Upper→Trans as an example to better capture the motion details of body parts. Therefore, the channels of the hierarchical sharing structure of the present invention can be represented as follows:
[0188]
[0189] Where is the shared hierarchical motion feature of task n∈{Torso,Lower,Upper,Trans}, and ⊕ represents concatenation. Since R init represents the local motion change of human joints, and P init represents the relative global position of the human body, the present invention chooses to place R init before each part's pose estimation and place P init before translation estimation, which helps the model of the present invention to maximize the extraction of key information without being affected by local or global factors.
[0190] Loss function
[0191] Finally, the estimate Y′ for each task n n is calculated as follows:
[0192]
[0193] where TaskHead is just a linear layer but can effectively regress the body pose Y from Pose and the global translation Y Trans . Thus, the global loss can be expressed as:
[0194]
[0195] The present invention uses the AMASS, DIP-IMU, and TotalCapture datasets to evaluate the effectiveness of the model. Details of these datasets can be found in the supplementary file.
[0196] The present invention uses DIP, TransPose, TIP, PIP, and DynaIP as baselines to demonstrate the superiority of HiPoser. The present invention evaluates the estimation performance through the following metrics: 1) SIP error (°): the average global rotation error of the upper arm and thigh; 2) Angle (Ang) error (°): the average global rotation error of all human joints; 3) Position (Pos) error (cm): the average Euclidean distance error of all joints; 4) Mesh error (Mesh) (cm): the average Euclidean distance error of all vertices of the body mesh; 5) Distance error (Dist) (cm): the average Euclidean distance error of the global translation of the whole body.
[0197] The present invention fixes the sampling frequency of all datasets at 60 Hz and sets the number of frames T = 300. The present invention uses PyTorch and PyTorch Lightning to build the model and updates the parameters using the Adam optimizer with a learning rate of 1e-3. The number of training epochs is 500, and the batch size is 256. The hidden dimension of MB is set to 128. The training and validation processes are implemented on an NVIDIA GeForce RTX 4090 GPU.
[0198] Due to the relatively small size of the DIP-IMU and TotalCapture datasets, the present invention is only trained on AMASS and validated on DIP-IMU and TotalCapture to evaluate the generalization performance of HiPoser. The present invention compares the average values of the baselines on each metric. As shown in Table 1, HiPoser of the present invention outperforms other methods. The present invention believes that the superiority of HiPoser lies in its hierarchical shared learning structure, which can effectively utilize the basic information of the movements of different body parts and better estimate joint rotations. In addition, the introduction of the motion state helps to maintain the continuity of the motion pattern. Figure 3 The estimation results of HiPoser and other methods are shown in Figure 3 .
[0199] Table 1 Result Comparison
[0200]
Claims
1. A method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), characterized in that, It includes the following steps: 1) Obtain the measurement data of the inertial measurement units deployed at key parts of the human body, and perform preprocessing to obtain the input signal 2) Construct a hierarchical pose evaluation model based on the Mamba module; 3) Use the hierarchical pose evaluation model to process the input signal X to obtain the human pose estimation result; The human pose estimation result is visualized through the 3D human SMPL model; The human body pose estimation result includes the body part pose estimation result Y Pose and the global translation estimation result Y Trans ; The body part pose estimation result Y Pose ={Y Torso , Y Lower , Y Upper}; Y Torso , Y Lower , Y Upper are respectively the torso pose estimation result, the lower limb pose estimation result, and the upper limb pose estimation. Global translation estimation result Y Trans It includes the three-dimensional coordinates of the spine root joint points in the SMPL model.
2. A method for hierarchically estimating 3D human pose from sparse Inertial Measurement Units (IMUs), characterized in that, In step 1), the preprocessing steps include: aligning the measurement data of the inertial measurement units deployed at the key body parts of the human body with the spine root joint, and performing normalization and splicing operations.
3. A method for hierarchically estimating 3D human pose from sparse Inertial Measurement Units (IMUs), characterized in that, The input signal X = {x A , x R}; represents the triaxial acceleration measurement value, represents the joint rotation matrix of the selected joint in three-dimensional space.
4. A method for hierarchically estimating 3D human pose from sparse Inertial Measurement Units (IMUs), according to claim 1, characterized in that, The hierarchical pose evaluation model based on the Mamba module includes the Mamba module and the pose estimation task heads for different body parts; The Mamba module processes the input signal X, hierarchically extracts and shares the motion hidden features of different body parts in X; The pose estimation task heads for different body parts use linear functions to reconstruct the body part pose estimation result Y from the motion hidden features of different body parts Pose and the global translation estimation result Y Trans .
5. A method for hierarchically estimating 3D human postures from sparse inertial measurement units (IMUs), according to claim 1, wherein The Mamba module is a Mamba module based on the selective state space model S-SSM; The internal unit structure based on the selective state space model S-SSM is as follows: where h t+1 and h t represent the hidden features of the (t + 1)-th frame and the t-th frame; Δ is the time step; B (Δ) is the input matrix; A (Δ) is the state matrix; C is the output matrix; S t and S t-1 are the state variables of the t-th frame and the (t - 1)-th frame; State matrix A (Δ) As shown below: where i and j represent the horizontal and vertical coordinates of the state matrix inside the Mamba module.
6. A method for hierarchically estimating 3D human pose from sparse Inertial Measurement Units (IMUs), according to claim 5, characterized in that, When processing the input signal X using the pose evaluation model, the time step Δ and the input matrix B (Δ) and the output matrix C are updated dynamically; Time step Δ, input matrix B (Δ) , and the dynamic update method of the internal parameter p of the output matrix C is as follows: Among them, h t is the state hidden feature of the Mamba module at the t-th frame; p C (t), p Δ (t) are respectively the internal parameters determined after the linear mapping of h t through the Linear layer.
7. A method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), according to claim 4, characterized in that The Mamba module is as follows: Among them, is the input at the current level; Cap is the motion feature; σ represents the SiLU function; is the output of the Mamba module.
8. A method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), characterized in that, The output Y n of the pose evaluation model is as follows: where n ∈ {Torso, Lower, Upper, Trans}; Among them, the motion characteristics are as follows: Among them, ⊕ represents concatenation; R init is the human joint rotation matrix; P init is the global position.
9. A method for hierarchically estimating 3D human poses from sparse Inertial Measurement Units (IMUs), according to claim 8, characterized in that Rotation matrix R of human joints init , global position P init are as follows: Among them, represents a frame; T is a time period.
10. A method for hierarchically estimating 3D human pose from a sparse inertial measurement unit (IMU) according to claim 1, characterized in that, The hierarchical pose evaluation model based on the Mamba module is trained through a historical dataset; the historical dataset includes the measurement data of the inertial measurement units deployed at the key body parts of the human body and the corresponding human pose estimation results; During the training process, the loss function of the hierarchical pose model is as follows: Among them, Y n (t) is the output of the hierarchical attitude evaluation model, and Y' n (t) is the actual value.
Citation Information
Cited By
Human body posture estimation method and device based on visual inertia fusion
CN121170848A