A whole-body posture estimation method and system based on pressure and inertial sensors

By arranging sensors in the insole, wrist and feet, combining neural networks and large language models, the sensor calibration and data fusion problems of human posture reconstruction in the prior art are solved, and precise capture and efficient reconstruction of complex human movements are achieved, which is suitable for applications in multiple fields.

CN119645223BActive Publication Date: 2025-08-08SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411590666.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-08-08
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

The prior art has problems such as inaccurate sensor calibration, external signal interference, difficult to solve the balance between the number of sensors and user comfort, low efficiency of multi-sensor data fusion, high computing cost, and difficulty in real-time operation in human posture reconstruction, resulting in poor reconstruction accuracy and user experience.

Method used

Using a combination of pressure sensors and inertial sensors, the sensors are arranged in the insoles, wrists and feet, combined with ground reaction forces and pressure centers to estimate body posture, use neural networks and loss functions of physical constraints to perform posture reconstruction, and expand the data set through virtual IMU data, and generate detailed motion descriptions in combination with large language models to improve the accuracy of posture estimation.

Benefits of technology

It realizes accurate capture of complex human movements, improves the accuracy and robustness of posture reconstruction, and is suitable for research and application in multiple fields, especially in complex motion scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645223B_ABST
    Figure CN119645223B_ABST
Patent Text Reader

Abstract

The present invention discloses a whole-body posture estimation method and system based on pressure and inertial sensors, which relates to human body posture reconstruction technology. This solution is proposed to address the problems of low technical integration in the existing technology. Three inertial sensors are used to generate real IMU data; the human body posture reconstruction algorithm fuses the real IMU data and virtual IMU data through the following steps to perform whole-body posture estimation: S1. Biological estimation based on pressure: Obtain pressure distribution through pressure sensor to estimate ground reaction force GRF and pressure center CoP; S2. Estimate body posture by combining ground reaction force GRF and pressure center CoP. The advantage is that only three inertial sensors are used in combination with plantar pressure data to achieve accurate capture of complex human body movements. By combining virtual IMU data with actually collected IMU data, a larger-scale data set is constructed, which can be applied to multiple human body research fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a human body posture reconstruction technology, and in particular to a whole body posture estimation method and system based on pressure and inertial sensors. Background Art

[0002] In the digital health and interactive entertainment sectors, particularly in rapidly developing areas such as rehabilitation, fitness applications, and virtual reality (VR) gaming, accurate reconstruction of human pose in three-dimensional (3D) space has become crucial for improving perception accuracy and user experience. Providing detailed, real-time feedback on the user's movements and posture not only significantly enhances the effectiveness of physical training and rehabilitation programs but also makes interactions in digital environments more natural and intuitive. Furthermore, in VR gaming, accurate pose reconstruction can provide a more immersive and interactive experience, bridging the gap between virtual and real life. With the increasing popularity of these applications, the demand for advanced 3D human pose reconstruction techniques is increasing. This field holds promising research prospects and has profound implications for the health, fitness, and entertainment industries. To meet this demand, innovative technologies are urgently needed to effectively capture and analyze human motion, improving reconstruction accuracy while minimizing impact on user comfort. Significant progress has been made in the field of human pose reconstruction in recent years, particularly with the emergence of deep learning-based machine vision solutions. However, these approaches still face numerous challenges, such as occlusion, viewing angle limitations, and sensitivity to lighting conditions.

[0003] Inertial measurement units (IMUs), which typically include accelerometers, gyroscopes, and sometimes magnetometers, play a key role in capturing human motion and reconstructing posture. By providing acceleration and angular velocity data, IMU sensors are crucial in estimating joint positions and the overall orientation of the body. To offset sensor drift and optimize multi-sensor data fusion, Kalman filtering and sensor fusion algorithms have traditionally been used, which have significantly improved the accuracy of posture reconstruction. In recent years, the introduction of deep learning models has provided powerful tools for interpreting IMU data and facilitated the processing of nonlinear relationships between sensor outputs and actual human motion. The development of deep learning has enabled effective learning from large amounts of IMU data, thereby achieving accurate reconstruction of complex human motion.

[0004] However, IMU-based posture reconstruction faces many challenges in practical applications, such as inaccurate sensor calibration and external signal interference, which may affect the stability and accuracy of the system. At the same time, when designing wearable devices, how to find a balance between the number of sensors and user comfort remains a key issue that needs to be addressed. Some current IMU-based wearable sports solutions use a small number of sensors (such as six IMUs) to estimate whole-body posture. Although effective, they usually need to be installed in parts such as the knees or thighs, which may not be convenient for athletes or rehabilitation patients. In addition, these solutions are generally not suitable for complex sports scenarios, so further exploration and verification are still needed in fields such as sports science.

[0005] Pressure sensors, on the other hand, can capture subtle changes in pressure with their high sensitivity and precision, providing higher data granularity for posture reconstruction. Integrating pressure sensors into wearable devices and environmental settings (such as mats and chairs) has proven to be extremely valuable for human posture reconstruction. These sensors detect the distribution of forces exerted by the body on different surfaces, identifying weight distribution and contact points of various body parts. However, pressure-based posture reconstruction systems also have limitations, particularly when detecting postures that do not directly contact the sensor surface, such as mid-air movements. During running and walking, the tactile interaction between the foot and the ground can vividly depict 2D pressure maps during different activities, providing unique insights into ground dynamics that are unavailable with inertial units and visual sensors. By recording the pressure distribution as a person moves on a surface, these sensors not only reveal details of movement patterns but also enhance our understanding of how humans interact with their environment.

[0006] Human posture reconstruction technology is currently divided into three main categories, each of which has inherent defects: (1) Vision-based posture reconstruction: low recognition accuracy in occluded or low-light environments, which may lead to privacy issues, poor model interpretability, and strict requirements on the location of the visual device, making it unsuitable for large-scale daily applications. (2) Pressure-based posture reconstruction: generally low recognition accuracy and more limitations. (3) Wearable device-based posture reconstruction: high data collection cost and complex labeling, and the device may affect the user experience (such as poor comfort).

[0007] Furthermore, most existing technologies fail to effectively integrate multi-sensor data, resulting in inefficient information utilization and the unrealized potential of multimodal data fusion. Some methods employ complex machine learning or deep learning algorithms, which can improve pose reconstruction accuracy but are computationally expensive and difficult to run in real time on mobile devices. These factors limit their widespread application in daily life. Summary of the Invention

[0008] The present invention aims to provide a whole-body posture estimation method and system based on pressure and inertial sensors to solve the problems existing in the above-mentioned prior art.

[0009] The whole-body posture estimation system based on pressure and inertial sensors described in the present invention includes: a pressure sensor, a posture reconstruction model, and three inertial sensors;

[0010] The pressure sensor is arranged on the insole to generate plantar pressure data; the three inertial sensors are used to generate real IMU data, one of which is arranged on the wrist, and the other two inertial sensors are arranged on both feet in a one-to-one correspondence;

[0011] The human pose reconstruction algorithm fuses real IMU data and virtual IMU data to estimate the whole body pose through the following steps:

[0012] S1. Biological pressure estimation: Obtain pressure distribution through pressure sensors to estimate ground reaction force (GRF) and center of pressure (CoP).

[0013] S2. Estimate body posture by combining ground reaction force GRF and center of pressure CoP.

[0014] The whole-body posture estimation method based on pressure and inertial sensors described in the present invention utilizes the system to perform whole-body posture estimation.

[0015] The whole-body posture estimation method and system based on pressure and inertial sensors described in the present invention have the advantage of using only three inertial sensors in combination with plantar pressure data to accurately capture complex human body movements. In view of the difficulty in collecting human body posture data in wearable devices, a virtual IMU data generation method is also designed. By combining virtual IMU data with the actual collected IMU data, a larger-scale data set is constructed. This data set is used in conjunction with plantar pressure data to reconstruct human motion posture. It can be applied to research in multiple fields such as sports science, gait analysis, virtual reality interaction, and ergonomics to further understand the human movement mechanism. Its effectiveness has been verified in multiple daily action scenes (such as running, walking, basketball, badminton, boxing, and dancing). BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic diagram of the principle of the system described in the present invention.

[0017] Figure 2 It is a schematic diagram of the principle of human body posture reconstruction described in the present invention.

[0018] Figure 3 It is a bar chart of the verification results of the system described in the present invention among different individuals. DETAILED DESCRIPTION

[0019] like Figure 1 As shown in the figure, the whole-body posture estimation system based on pressure and inertial sensors described in the present invention includes: a pressure sensor, a posture reconstruction model, and three inertial sensors. The pressure sensor is placed on the insole to generate plantar pressure data. The three inertial sensors are used to generate real IMU data, with one inertial sensor placed on the wrist and the other two inertial sensors placed on the feet in a one-to-one correspondence.

[0020] The human pose reconstruction algorithm fuses real IMU data and virtual IMU data to estimate the whole body pose through the following steps:

[0021] Pressure-based biological estimation: The pressure distribution is obtained by the plantar pressure sensor to estimate the ground reaction force GRF and center of pressure CoP.

[0022] First, the spatial distribution of foot pressure is collected and represented by a heat map H(x,y), where x and y represent the spatial coordinates within the insole surface area. Then, the GRF is estimated by integrating the pressure distribution over the entire insole area, using formula (1):

[0023] GRF=∫∫H(x,y)d x d y #(1);

[0024] Then, the center of pressure (CoP) is estimated using formulas (2) and (3). The information about the center of pressure reveals how the body adjusts its posture and redistributes its weight during dynamic motion to maintain balance and stability. This data is of great significance in improving the accuracy of posture estimation during dynamic motion of the human body.

[0025]

[0026] The body posture is estimated by combining the ground reaction force GRF and the center of pressure CoP. The human body is considered as a multi-body system consisting of multiple joints, each of which corresponds to the movement of a different part of the body. The dynamics of human motion posture can be modeled using the Newton-Euler equation or Lagrangian mechanics. Based on the data of GRF and CoP, the human motion posture reconstruction formula (4) is as follows:

[0027]

[0028] Where M(q) represents the human body mass matrix associated with the joint structure q, Represents the acceleration vector of the joint, which represents the inertial force acting on the system; is the Coriolis and centrifugal force matrix, which is used to describe the joint rotation offset caused by inertia, is the velocity vector of the joint, which describes the Coriolis force and centrifugal force generated by the relative motion of the various parts of the system; G(q) is the gravitational torque vector, J T (q) is the transposed vector of the Jacobian matrix of the system, which maps the force in the task space back to the joint torque. F represents the external force vector such as GRF, which represents the external force influence acting on the system.

[0029] Jacobian matrix J T (q) connects joint motion with CoP and GRF, thus showing how joint motion affects the change of CoP, as shown in formula (5):

[0030]

[0031] In activities such as walking, optimization algorithms are used to predict and enhance body movements, focusing on the stability of body posture and gait. In order to accurately reconstruct posture from sensor data, physical constraints are integrated into the learning process of the neural network. The input of the model includes kinematic data of body joints, plantar GRF and CoP data. The joint position loss function improves accuracy by reducing the difference between the predicted joint position and the actual joint position. Its mean square error formula (6) is as follows, where p i and Represent the actual position and predicted position of the i-th joint respectively, and N is the total number of joints;

[0032]

[0033] At the same time, the center of mass CoM of the body can be calculated based on the position and mass distribution of each joint of the body. The center of mass position loss formula (7) is added to ensure that the posture meets the physical constraints. h represents the estimated height of CoM from the support base, and g represents the acceleration due to gravity:

[0034]

[0035] In addition, a dynamic consistency loss is added to verify that the dynamics of the estimated position conforms to Newton's law, especially the effect of GRF on the body acceleration. Formula (8) uses the dynamic equation to calculate the loss component, where M is set to the mass matrix of the body, is the predicted acceleration of the joint:

[0036]

[0037] By integrating the above loss functions, the entire neural network loss function for estimating body posture is a combination of joint position loss, center of mass loss, and dynamic consistency loss, formula (9):

[0038]

[0039] The model first uses a fully connected layer to process and integrate the input kinematic data. It then introduces a Transformer module to capture the complex temporal dependencies and interactions between different joints. The self-attention mechanism within the Transformer is particularly well suited for modeling the diverse spatial relationships of body joints and effectively identifying the dynamic connections between them. Finally, another set of fully connected layers maps the Transformer outputs to predicted joint positions, resulting in accurate pose reconstruction.

[0040] In order to learn 3D human motion based on dynamic chains of global features, a scheme combining local and global modeling is designed. Due to the weak correlation between body parts, directly using the data of three IMU sensors as model input may lead to motion blur, where the data includes real IMU and virtual IMU data. To solve this problem, a local area modeling strategy is adopted to divide the human body into two main dynamic chains, such as Figure 2 As shown, this improves the accuracy of motion capture and analysis.

[0041] Lower Limb and Trunk Chain: For the combined kinetic chain of the lower limbs and trunk, insole pressure sensors are used to calculate the ground reaction force (FRF) and center of pressure (CoP), providing essential data for force distribution and management in the lower body and core. Furthermore, an inertial measurement unit (IMU) on the foot further enriches the dataset, providing orientation and acceleration information, thus supporting a comprehensive dynamic analysis of the lower body and trunk.

[0042] Upper Limb Chain: The second dynamic chain consists of the two arms, with wrist-worn IMUs serving as the primary sensors. Inertial information is crucial for capturing the complex movements of the upper limbs, particularly during tasks like reaching, lifting, or throwing.

[0043] However, the main challenge facing local region-based methods is how to maintain global consistency across body regions. To this end, global features are incorporated into the body pose network of the dynamic chain. The global features are denoted as G(IMU1, IMU2, IMU3), where G(·) represents the global feature aggregation function extracted by the fully connected layer, which roughly integrates the whole-body motion information from the three IMUs. The global features G are then combined with the input of the Physical-aware Body PoseNetwork, which specifically models the two local dynamic chains. By fusing global and local dynamic information, the network is able to maintain coherence when reconstructing the whole-body pose, making the local motion consistent with the overall body dynamics.

[0044] The obtained model is evaluated using three indicators: mean joint error MPJPE, body jitter, and translation. The performance of the proposed body posture reconstruction system is comprehensively evaluated:

[0045] MPJPE: Measures the spatial accuracy of pose reconstruction by calculating the average Euclidean distance between the predicted joint positions and the true positions Groundtruth, which is the main indicator for evaluating model accuracy.

[0046] Jitter: Measures the temporal variation between consecutive frames to evaluate the smoothness of motion prediction. Lower jitter values indicate more stable and consistent motion estimation, which is an important metric for dynamic posture analysis.

[0047] Translation: This evaluates the model's accuracy in predicting the overall body joint positions in space. Even if individual joint positions are accurately predicted, if the overall skeleton position deviates significantly from the true position, the model's practical applicability and accuracy will be affected.

[0048] Finally, the body pose estimation network was trained end-to-end using the PyTorch and PyTorch Lightning frameworks. Weights were updated using the Adam optimizer with a batch size of 512 and a learning rate of 3e-4. The loss function was designed based on the combined loss from step 3 to ensure accurate prediction of joint positions and overall skeleton consistency.

[0049] In some practical applications, the amount of data available for deep learning training may be insufficient. Therefore, the present invention also provides a method for expanding data, which can perform deep learning training even when there is only a small amount of real collected data, thereby better estimating the human body motion skeletal morphology.

[0050] A cross-modal transfer learning network is built to expand the training dataset by integrating data from various sources, such as images, videos, and sensors. This not only enriches the data pool but also enhances the model's robustness by simulating various real-world scenarios. Furthermore, the algorithm utilizes a large language model (LLM) to automatically generate motion descriptions, combined with motion synthesis techniques, to improve the accuracy and realism of the pose reconstruction model. The specific implementation process is as follows:

[0051] (1) Text-to-Action Generation

[0052] The accuracy of IMU data generated from videos, particularly those in social media, can be limited by factors such as video resolution, lighting, and occlusion. To address these limitations, natural language text provides clear action descriptions, such as "running" and "jumping," which are directly used to generate virtual IMU data. This overcomes the limitations imposed by varying video quality and effectively replaces visual information.

[0053] (2) Text description generation for posture reconstruction

[0054] LLMs are applied to pose reconstruction to generate detailed descriptions of activities, better modeling the variability of human behavior. By generating context-specific text descriptions, such as "A robust athlete runs steadily on the track," virtual IMU data accurately reflects real human motion, enhancing model training effectiveness. These descriptions focus on describing the details of human motion, eliminating environmental or emotional factors, to ensure accurate pose reconstruction.

[0055] An advanced Large Language Model (LLM) is used to generate precise text descriptions to inform the synthesis of virtual IMU data. These descriptions, such as "a robust athlete runs steadily on the track" or "a muscular hiker carefully adjusts their pace on uneven terrain," are used to describe body movements in detail while ignoring irrelevant environmental or emotional factors, ensuring that these descriptions are directly applicable to the task of pose reconstruction. The prompt generation process is carefully designed to produce detailed and specific text guides that clarify the actions involved in various activities. For example, the descriptions include comprehensive details of the upper and lower body movements, supplemented by descriptive adjectives such as "robust" and "smooth." By strictly focusing on the necessary action details, maintaining a specific length and expressiveness, these texts are tailored to support the system in accurately estimating poses without distracting irrelevant information.

[0056] (3) Motion generation model

[0057] 3D human motion is generated from text descriptions through a diffusion model, and a probabilistic mapping from text to motion is established using a denoising step, formula 10:

[0058]

[0059] x0 represents the initial skeleton joint data of human motion, q(x0) represents the probability distribution of real motion skeleton joint data, and the sequence x0~x T Represents the intermediate latent features. The core process of the model is diffusion and backpropagation. The diffusion process runs through the Markov chain mechanism, gradually adding Gaussian noise to the data until it approaches the underlying Gaussian distribution N(0,1). This process is carefully controlled by a variance, expressed as formula (11) and formula (12):

[0060]

[0061] In contrast, the reverse process aims to reconstruct the original data by a single denoising step, which uses the T )=N(x T; 0, I) to describe the inverse Markov chain, expressed as formula (13) and formula (14):

[0062]

[0063] These operations enable the model to effectively generate high-fidelity output from an initial noisy data distribution. MotionGPT further transforms text descriptions into specific motion sequences. For example, "A person stands up from a seated position and takes a few steps forward." The natural language model converts the text into motion symbols, which are then reconstructed into a continuous motion sequence by the decoder, faithfully reproducing the described action.

[0064] (4) IMU data integration

[0065] To accurately estimate human motion, inverse kinematics (IK) is used to determine the rotational motion of each joint relative to its parent joint and the translational motion of the root joint, usually the pelvis. Inverse kinematics adjusts the rotation of the skeleton joints so that end effectors such as hands and feet can accurately reach preset target positions in three-dimensional space. Specifically, the IK process first initializes the animation skeleton in a baseline pose, sets a target position for each joint, and then uses a robust inverse kinematics algorithm such as the Jacobi method or cyclic coordinate descent to iteratively optimize the joint angles to achieve the specified target position. This process relies on predefined joint positions and bone hierarchies to ensure the alignment of joints with target trajectories, thereby improving the realism and accuracy of joint motion modeling.

[0066] Next, the IMUsim tool was used to calculate joint accelerations and angular velocities based on the local joint rotations and root translations derived from inverse kinematics. This generated data from 22 strategically placed virtual IMU sensors. Furthermore, IMUsim introduced noise to simulate the typical noise characteristics found in real IMU data. This addition of noise made the virtual IMU data more closely aligned with actual sensor readings, thereby improving the accuracy and realism of the human motion dynamics modeling.

[0067] Finally, to ensure consistency between the virtual IMU data and the real-world data, the generated virtual data was further calibrated to match the distribution of the IMU data collected by Xsens. Combining the calibrated virtual data with the real-world data expanded the initial training dataset, significantly improving the training results of the pose estimation system despite limited data.

[0068] To demonstrate the robustness and versatility of the proposed full-body pose estimation system in a variety of dynamic environments, experiments were conducted across a range of real-world activity scenarios. This evaluation measured the accuracy of the system's reconstruction of human pose during various types of motion, which is crucial for high-fidelity motion tracking applications under diverse conditions. A total of 25 participants were recruited, 10 of whom completed all six designated activities, as listed in Table 1.

[0069] Table 1 Reconstruction accuracy of different activity postures

[0070] Activity running walk basketball badminton boxing dance MPJPE(cm) 4.91 5.22 7.29 11.54 10.81 7.68 <![CDATA[Jitter(10 2 m / s 3 )]]> 2.35 1.98 3.93 4.37 3.42 3.69 Translation(cm) 3.18 2.71 3.35 5.23 3.14 2.61

[0071] In each activity, participants performed a 5-minute routine within a designated 4m x 5m area. The routines were as follows, reflecting the typical dynamic patterns of each sport or activity:

[0072] (1) Basketball: including dribbling in place, stride dribbling, three-step layup and shooting;

[0073] (2) Badminton: including forehand and backhand receiving in the front court and forehand and backhand hitting in the back court;

[0074] (3) Boxing: including straight punch, swing punch, hook punch and forward punch;

[0075] (4) Dance: including conventional dance movements.

[0076] In each activity, the accuracy of the pose reconstruction results generated by the system is evaluated by measuring the mean joint position error (MPJPE), translation, and jitter metrics. As shown in Table 1, the full-body pose estimation system exhibits high reconstruction accuracy and consistency.

[0077] Based on the analysis of the results, we concluded that while running and walking are cyclical movements with relatively stable patterns, badminton, boxing, and other sports are highly dynamic and volatile, characterized by rapid speed changes, accelerations and decelerations, and rapid and large transitions between movements. Therefore, the accuracy of posture reconstruction in these highly volatile sports is lower than that in running and walking.

[0078] Without using virtual IMU data, higher MPJPE values were observed in sports such as basketball with an MPJPE of 10.5cm, badminton with an MPJPE of 14.3cm, and boxing with an MPJPE of 14.8cm. This indicates that the accuracy of posture estimation is significantly reduced in these dynamic and fast-paced sports. In contrast, more predictable and repetitive sports such as running with an MPJPE of 6.3cm and walking with an MPJPE of 6.4cm showed lower errors. This shows that the model performs well in regular and repetitive movement patterns, but performs poorly in situations with high unpredictability and fast movements such as basketball, badminton, and boxing. In order to further verify the performance variations of the system between different individuals during specific operations, the leave-one-out method was used for analysis, and the results are shown in the figure. Figure 3 shown.

[0079] While the model demonstrates good generalization across users, we also observed that deep learning-based models can face challenges when tasks become more complex or involve untrained motions. As shown in Table 2, the model's performance varies significantly across activities, reflecting the complexity and variability of human movement patterns. This finding emphasizes the importance of fully considering motion diversity when designing and training pose estimation systems to improve their adaptability to unknown motion patterns.

[0080] To this end, it was necessary to validate the solution after incorporating virtual IMU data. The results are shown in Table 2. The introduction of virtual IMU data significantly increased the diversity and quantity of training examples. Notably, MPJPE values decreased across all activities, with a particularly significant improvement in high-dynamic sports. For example, the MPJPE value in basketball decreased from 10.5 cm to 9.2 cm, while the MPJPE value in boxing decreased from 14.8 cm to 12.5 cm. This result demonstrates that the introduction of virtual data can effectively improve the model's generalization ability across a wider range of human activities, thereby enhancing its performance in different sports scenarios.

[0081] Table 2 Comparison of leave-one-out validation performance with and without virtual IMU data

[0082]

[0083]

[0084] Accurately reconstructing human posture is crucial in sports science, especially in outdoor environments and during dynamic movements. However, current posture reconstruction methods suffer from various issues. In contrast, our system enables detailed 3D posture reconstruction, which is essential for advanced sports training and performance evaluation. This reconstruction system not only provides coaches with real-time feedback on athletes' posture and technique in various sports, thereby improving training effectiveness, but also plays a role in equipment design. A deeper understanding of the interaction between athletes and their equipment can foster innovation, improve performance, and reduce the risk of injury.

[0085] The method for estimating whole-body posture based on pressure and inertial sensors described in the present invention utilizes the system to perform whole-body posture estimation.

[0086] Compared with the prior art, the present invention has at least the following advantages:

[0087] (1) It can efficiently and accurately reconstruct complex body movements.

[0088] (2) By leveraging the text understanding and representation capabilities of large language models and combining them with motion generation technology, the accuracy and authenticity of generated virtual IMUs are improved.

[0089] (3) We utilize a variety of human motion postures, focusing on daily movements (such as running, walking, basketball, boxing, etc.), and at the same time, by reconstructing specific skeleton morphologies, verifying the effectiveness and robust applicability of posture estimation in accurately capturing and analyzing complex movements in various dynamic activities.

[0090] Those skilled in the art can make various other corresponding changes and deformations based on the technical solutions and concepts described above, and all of these changes and deformations should fall within the scope of protection of the claims of the present invention.

Claims

1. A whole-body posture estimation system based on pressure and inertial sensors, characterized in that: include: pressure sensor, posture reconstruction model, and three inertial sensors; The pressure sensor is arranged on the insole to generate plantar pressure data; The three inertial sensors are used to generate real IMU data, wherein one inertial sensor is arranged on the wrist, and the other two inertial sensors are arranged on the feet in a one-to-one correspondence; The human pose reconstruction algorithm fuses real IMU data and virtual IMU data to perform full-body pose estimation through the following steps: S1. Biological pressure estimation: Obtain pressure distribution through pressure sensors to estimate ground reaction force (GRF) and center of pressure (CoP). S2. Estimate body posture by combining ground reaction force GRF and center of pressure CoP; The step S1 is specifically as follows: Collect the spatial distribution of foot pressure and represent it as a heat map H(x,y), where x and y represent the spatial coordinates within the insole surface area; By integrating the pressure distribution over the entire insole area, the ground reaction force GRF = ∫∫H(x,y)d x d y ; Estimate the center of pressure CoP based on the ground reaction force: The step S2 is specifically as follows: Using the formula Reconstruct human body motion posture; Where M(q) represents the human body mass matrix associated with the joint structure q, represents the velocity of the joint, represents the acceleration of the joint, is the Coriolis force and centrifugal force matrix used to describe the joint rotation offset caused by inertia, G(q) is the gravitational torque vector, J T (q) is the Jacobian matrix of the body joints and CoP, and F represents the external force; Jacobian matrix J T (q) is to connect joint motion with CoP and GRF: In step S2, the joint position loss function improves the accuracy by reducing the difference between the predicted joint position and the actual joint position. Its mean square error formula is: Where p i represents the actual position of the i-th joint, They represent the predicted position of the i-th joint, and N is the total number of joints; The center of mass CoM of the body is calculated based on the position and mass distribution of each joint of the body, and the center of mass position loss L is added CoM Make sure the pose complies with the physical constraints; h is the estimated height of the CoM from the support base, and g is the acceleration due to gravity: Add dynamic consistency loss to verify that the dynamics of the estimated position conforms to Newton's law, and calculate the loss component using the dynamic equation where M is set to be the mass matrix of the body, is the predicted acceleration of the joint; Integrating each loss function to obtain the neural network loss function for estimating body posture: L total =λ1L pose +λ2L CoM +λ3L dyn ; The pose reconstruction model uses a fully connected layer to process and integrate the input kinematic data; then introduces the Transformer to capture the temporal dependencies and interactions between different joints; finally, another set of fully connected layers is used to map the Transformer output to the predicted joint positions to generate the pose reconstruction results.

2. The whole body posture estimation system based on pressure and inertial sensors according to claim 1, characterized in that: The learning modeling of the posture reconstruction model is specifically as follows: Incorporating global features into a body posture network of dynamic chains; the dynamic chains include lower limb and trunk chains, and upper limb chains; The global feature is denoted as G(IMU1, IMU2, IMU3), where G(·) represents the global feature aggregation function extracted by the fully connected layer, which preliminarily integrates the whole-body motion information from the three inertial sensors; Subsequently, the global feature G is combined with the input of the physical perception body posture network to maintain coherence when reconstructing the whole body posture by fusing global and local dynamic information, so that the local motion is consistent with the overall body dynamics; Finally, the body pose estimation network is trained end-to-end using PyTorch and PyTorch Lightning frameworks; the training adopts the Adam optimizer with a batch size of 512 and a learning rate of 3e-4 for weight updates.

3. The whole body posture estimation system based on pressure and inertial sensors according to claim 2, characterized in that: The virtual IMU data is generated based on the large language model, as follows: Text-to-action generation: Provide action descriptions through natural language text and directly generate virtual IMU data; Textual description generation for pose reconstruction: Apply LLMs to generate detailed descriptions of activities in pose reconstruction. This ensures that virtual IMU data accurately reflects real human motion by generating context-specific textual descriptions, enhancing model training effectiveness. Motion Generation Model: Generate 3D human motion from text descriptions through a diffusion model, and use a denoising step to build a probabilistic mapping from text to motion: p θ (x0)∶=∫q(x0∶T)d x1 ∶T; x0 represents the initial skeleton joint data of human motion, q(x0) represents the probability distribution of real motion skeleton joint data; the sequence x0~x T Represents the intermediate latent features, and the model is completed through diffusion and backpropagation; The diffusion process operates through a Markov chain mechanism, gradually adding Gaussian noise to the data until it approaches the underlying Gaussian distribution N(0,1). This process is controlled by the variance formula: On the contrary, the reverse process reconstructs the original data by a single denoising, which uses the T )=N(x T ; 0, I) to describe the inverse Markov chain, expressed as: p θ (x t-1 |x t )∶=N(x t-1 ;μ θ (x t ,t),Σ θ (x t ,t)); It enables the model to generate high-fidelity output from an initial noisy data distribution. MotionGPT further converts text descriptions into specific motion sequences, converting the text into motion symbols through a natural language model, and then reconstructing it into a continuous motion sequence through a decoder, truly reproducing the actions in the description. IMU data synthesis: Inverse kinematics is used to determine the rotational motion of each joint relative to its parent joint and the translational motion of the root joint. Based on the local joint rotation and root translation data obtained by inverse kinematics, the joint's motion acceleration and angular velocity are calculated to generate data from multiple strategically placed virtual IMU sensors. Noise is also introduced to simulate the typical noise characteristics in real IMU data. Calibrate and adjust the generated virtual data to match the distribution of the collected IMU data; The calibrated virtual data is combined with the real collected data to expand the initial training dataset.

4. A whole-body posture estimation method based on pressure and inertial sensors, characterized in that: Full body posture estimation is performed using the system described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Multi-person attitude estimation method based on inertial sensor and multifunctional camera

    CN114627490A

  • A method, a system and a computer program product for estimating positions of a subject's feet and centre of mass relative to each other

    WO2021010832A1