Human pose estimation apparatus based on global pressure and method thereof

By using a global pressure-based human pose estimation device, which employs a spatial feature encoder and a long short-term attention module, the error problem of monocular image methods in scenarios with high requirements for lighting and privacy is solved. This enables unified estimation and physical constraints of whole-body contact movements with the ground, thereby improving the accuracy and application range of pose estimation.

CN119992645BActive Publication Date: 2025-11-04NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510020551.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-11-04
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

Existing methods for estimating human pose based on monocular images have errors in scenarios with uncertain lighting conditions, high privacy requirements, visual occlusion, or self-occlusion. Furthermore, they cannot uniformly estimate actions that involve the whole body in contact with the ground or actions involving only the feet in contact with the ground, lacking physical constraints and characteristics.

Method used

A human posture estimation device based on global pressure is adopted. The spatial features of global pressure of the human body are extracted by a spatial feature encoder. Combined with a long and short-term temporal attention module and an action regressor, nonlinear regression calculation of the spatiotemporal features of pressure is realized to obtain the posture and displacement parameters of the human body parameterized model.

Benefits of technology

It achieves non-invasive human pose estimation in special scenarios, introduces physical constraints and characteristics, breaks through the limitations of action types, and broadens the application prospects and space of human pose estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992645B_ABST
    Figure CN119992645B_ABST
Patent Text Reader

Abstract

The application provides a human posture estimation device based on global pressure and a method thereof. The device comprises: a spatial feature encoder configured to extract human global pressure spatial features from a global pressure frame sequence of human motion; a long short-term time sequence attention module configured to extract human global pressure time sequence features from the human global pressure spatial features and fuse the human global pressure time sequence features with the human global pressure spatial features to obtain pressure space-time features; and a motion regressor configured to perform nonlinear regression calculation on the pressure space-time features to obtain posture and displacement parameters of a human parameterized model. The application replaces visual information with pressure information which is independent of lighting conditions, does not infringe privacy and is not affected by visual occlusion, thereby achieving non-invasive expansion of human posture estimation in special scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human pose estimation in the field of deep learning technology, and particularly to a human pose estimation device and method based on global pressure. Background Technology

[0002] Human pose estimation is a method that transforms human movements into quantifiable position and angle parameters. With the rapid advancement of deep learning technology, this technique has demonstrated enormous potential in accurately generating human movements, especially in the fields of virtual characters and humanoid robots, where its application value and research prospects have attracted considerable attention. Currently, monocular image-based human pose estimation methods have become mainstream. This method uses monocular human images as input and leverages deep learning technology to accurately fit the parameters of 3D joint points or human skin models, thereby achieving a detailed representation of human movements.

[0003] However, existing monocular image-based human pose estimation methods suffer from several limitations. Firstly, they are highly dependent on lighting conditions (unable to shoot in low light or dim light); secondly, they significantly infringe on the privacy of the subject (unacceptable in places with high privacy requirements, such as hospitals and homes); thirdly, they are highly dependent on the subject's movements (significant errors occur when the subject is obscured by objects or body parts); and fourthly, they focus only on the subject and not on its environment (often resulting in unreasonable interactions with the ground). These limitations restrict their application in scenarios with uncertain lighting conditions, high privacy requirements, frequent visual or self-occlusion, and a focus on human-ground interactions. Furthermore, these image-based motion estimation methods often fail to consider physical constraints or characteristics, potentially leading to physically inaccurate phenomena when integrated into realistic environments with physics engines, such as human models floating in mid-air or exhibiting unnatural excessive forward or backward tilts.

[0004] Current methods for estimating human posture using pressure often only focus on the posture of the human body in a lying position, or only on the pressure exerted by the feet on the ground. They cannot provide a unified estimation method for movements that involve the entire body in contact with the ground, or movements where only the feet are in contact with the ground. Therefore, providing a method for estimating the whole-body posture that does not rely on a monocular camera, satisfies real physical laws, and is unrestricted by the type of human movement is a problem that urgently needs to be solved. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a human posture estimation device and method based on global pressure.

[0006] The technical solution adopted in this invention is as follows:

[0007] A human posture estimation device based on global pressure includes: a spatial feature encoder for extracting global pressure spatial features from a global pressure frame sequence of human movements; a long short-term attention module for extracting global pressure temporal features from the global pressure spatial features and fusing them with the global pressure spatial features to obtain pressure spatiotemporal features; and a motion regressor for performing nonlinear regression calculations on the pressure spatiotemporal features to obtain the posture and displacement parameters of a human parameterized model.

[0008] The present invention also provides a method using the above-described human posture estimation device based on global pressure, the method comprising the following steps:

[0009] Step S1: Construct a network model for the human pose estimation device and train the network model;

[0010] Step S2: Acquire a global pressure frame sequence of human body movements; input the global pressure frame sequence into a spatial feature encoder to obtain the global pressure spatial features F of the human body. s ;

[0011] Step S3: Calculate the global pressure spatial characteristics F of the human body. s Inputting long and short-term time series self-attention modules yields the spatiotemporal stress features F. st ;

[0012] Step S4: Transfer the spatiotemporal characteristics of pressure F st Input the motion regressor to obtain the final human posture and displacement parameters.

[0013] The beneficial effects of this invention are as follows: First, it addresses the problem of erroneous estimation in vision-based human posture estimation methods under uncertain lighting conditions, high privacy requirements, and frequent visual occlusion or self-occlusion scenarios. By replacing visual information with pressure information that is independent of lighting conditions, does not infringe on privacy, and is unaffected by visual occlusion, it achieves a non-invasive extension of human posture estimation in special scenarios. Second, through high-dimensional feature extraction of pressure information, it achieves in-depth extraction and utilization of human posture and real physical information, introducing physical constraints and physical characteristics into the human posture estimation method. Finally, this invention breaks through the limitations of action types in pressure-based human posture estimation, realizing human posture estimation for all types of actions, from full-body contact with the ground to non-limited body parts contacting the ground, thus broadening the application prospects and space of the estimation method. Attached Figure Description

[0014] Figure 1 This is a structural diagram of the device of the present invention;

[0015] Figure 2 This is a structural diagram of the feature encoder in an embodiment of the present invention;

[0016] Figure 3 This is a comparison diagram of the human posture and displacement estimation effects in an embodiment of the present invention. Detailed Implementation

[0017] Embodiments of the present invention will now be described in detail with reference to the accompanying drawings, examples of which are illustrated in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0018] like Figure 1 As shown, this embodiment provides a human posture estimation device based on global pressure, including: a spatial feature encoder, a long short-term temporal attention module, and an action regressor. The spatial feature encoder is used to extract spatial features of global human pressure information from a sequence of global pressure frames of human movements. The long short-term temporal attention module is used to extract temporal features of global human pressure from continuous spatial features of global human pressure and fuse them with the spatial features of global human pressure to obtain spatiotemporal features of pressure. The action regressor is used to perform nonlinear regression calculations on the spatiotemporal features of human pressure to obtain the posture and displacement parameters of the human parameterized model, ultimately driving the human parameterized model.

[0019] The specific implementation device includes a memory and one or more processors. The memory stores the code and executable files of the spatial feature encoder, the long short-term attention module, and the action regressor. When the processor executes the executable files of the spatial feature encoder, the long short-term attention module, and the action regressor, it is used to implement the human posture estimation method based on global pressure of the present invention.

[0020] The spatial feature encoder includes a spatial feature extraction network, such as... Figure 2 As shown, it includes an initial convolutional layer, a max pooling layer, a four-stage convolutional layer, an average pooling layer, and a fully connected layer. Each of the four-stage convolutional layers consists of two residual blocks, and each residual block consists of two convolutional layers.

[0021] The long short-term attention module includes a long short-term attention network, which consists of a gated recurrent structure and a self-attention module. The gated recurrent structure is composed of two gating mechanisms: an update gate and a reset gate. The self-attention module includes a fully connected layer, a self-attention mechanism, a normalized exponential function, a random deactivation layer, and a layer normalization function.

[0022] The motion regressor includes a motion regression network, which consists of fully connected layers and random deactivation layers. It performs nonlinear regression on the fused features to output human pose parameters and displacement parameters.

[0023] This embodiment also provides a human posture estimation method based on global pressure, including the following steps:

[0024] Step S1: Construct the network model of the above-mentioned human pose estimation device and train the network model;

[0025] Step S2: Acquire a global stress frame sequence of human body movements; input the global stress frame sequence into a spatial feature encoder to obtain the global stress spatial features F of the human body. s The global pressure frame sequence includes the pressure generated by various parts of the human body that can contact the ground, such as the feet, knees, hips, torso, elbows, hands, and head, as well as the relative positional relationships between the parts of the human body that are in contact with the ground.

[0026] The spatial characteristics specifically include the magnitude of the pressure exerted by the human body on each pressure measurement unit of the pressure pad, the two-dimensional spatial position of each pressure measurement unit relative to the pressure pad's coordinate system, and spatial information such as the distance, angle, and area between each pressure measurement unit. For example, when a person performs a plank exercise, the wrists, elbows, and feet apply pressure to the ground. The pressure pad is composed of multiple small pressure measurement units arranged together, which can record the pressure values ​​exerted by the wrists, elbows, and feet on the ground. The distribution of pressure values ​​from multiple pressure measurement units not only reflects the magnitude of the pressure and the size of the contact area between each part of the body and the ground, but also characterizes positional information such as the distance and offset angle between the pressure areas of the wrists, elbows, and feet.

[0027] Spatial feature encoders, such as Figure 2 As shown, the architecture includes an initial convolutional layer with a kernel size of 3*3 and 64 channels, a 2*2 max pooling layer, a four-stage convolutional layer, an average pooling layer, and a fully connected layer. Each of the four-stage convolutional layers consists of two residual blocks, and each residual block is composed of two 3*3 convolutional layers. As the number of stages increases, the number of channels in the convolutional kernel increases from 64 to 128, 256, and 512, thus achieving high-dimensional feature extraction. Since the stress frame sequence is a single-channel image, and stress information accounts for a relatively small proportion during actions such as standing and squatting, the classic ResNet18 feature extractor is modified by changing the input channel to 1, reducing the initial convolutional kernel size and pooling layer size to obtain more refined stress spatial features.

[0028] Step S3: Calculate the global pressure spatial characteristics F of the human body. s Inputting into the long short-term self-attention module, since the single-frame pressure information representing action has significant ambiguity, the global pressure spatial features are processed in the long short-term self-attention module by paying attention to local neighboring frames and global long-term frames to obtain pressure spatiotemporal features F that better characterize human action. st .

[0029] Among them, long and short time series self-attention modules are such as Figure 1 As shown, it includes a gated loop structure and a self-attention module. The pressure spatial features F obtained from the spatial feature encoder in step S1 are... s First, the pressure features between adjacent frames are integrated using a gated loop structure, and the pressure spatial features F are then combined. s Introducing local time characteristics and combining them with the initial pressure space characteristics F s The summation yields the local temporal spatial features F. slocal =GRU(F s )+F s , where GRU(F s F represents the pressure space characteristic. s The output of the gated loop structure implicitly represents the changes in distance and angle of pressure measurement units in the pressure feature representation of nearby time frames. Local temporal spatial features of pressure are learned from these changes in distance and angle of pressure measurement units in the pressure feature representation. Then, the local temporal spatial features F... slocal After passing through the self-attention module, the global temporal information within the stress feature frame is integrated using the global cross-weighting mechanism of the self-attention module, and then combined with the initial local temporal spatial feature F. slocal The summation yields the pressure-space-time characteristic F. st =Attention(F slocal )+F slocal , where Attention(F slocal The output of the self-attention module represents the local temporal spatial features, which include representations of changes in pressure measurement distance and angle between time frames. It also learns pressure spatiotemporal features that better represent important information about human posture. Residual connections allow the output features to retain some of the original features, better representing pressure features at both the temporal and spatial levels. Pressure spatiotemporal feature F st It includes the spatial representation of the interaction between the human body and the ground in terms of pressure, as well as the spatial positional relationship between the human body parts in contact with the ground. It also includes the mutual perception between adjacent frames of human body movement and frames with a large global distance, and information on the physical dynamics mechanism related to contact and ground interaction.

[0030] Step S4: Transfer the spatiotemporal characteristics of pressure F st The model is fed into the SMPL (A Skinned Multi-Person Linear Model) motion regressor to obtain the final human posture θ and displacement parameters T.

[0031] When training the parameters of the spatial feature encoder, long short-term attention module, and action regressor in step S1, the training sample data is input into the network model, and the training loss function is calculated as: L = L pose +L 3d +L trans +L contact The network model was trained iteratively, and the AdamW optimizer was used to optimize the network parameters. The learning rate was set to 5e-4, and the total number of iterations was 1000. The network model parameters were determined based on the training loss value.

[0032] Among them, the human pose loss function θ represents the predicted human pose parameters. This serves as the baseline true value for human posture parameters;

[0033] Human body 3D joint loss function J(θ,T) represents the coordinates of the three-dimensional joints of the SMPL human three-dimensional parametric model under the control of the human posture parameter θ and the human displacement parameter T. True reference values ​​for human posture parameters and the true reference value of human body displacement parameters 3D joint coordinates of a SMPL human body 3D parametric model under control;

[0034] The human body displacement parameter loss function is T represents the predicted human body displacement parameter. This serves as the baseline true value for human body displacement parameters;

[0035] The whole-body contact loss function is J c (θ,T) represents the coordinates of the three-dimensional joints of the SMPL human three-dimensional parametric model in contact with the ground, controlled by the human posture parameter θ and the human displacement parameter T. True values ​​of human posture parameters and the true reference value of human body displacement parameters The coordinates of the three-dimensional joints of the SMPL human body three-dimensional parametric model in contact with the ground under control.

[0036] Among them, the three-dimensional joint J in contact with the ground c (θ,T) refers to the projection point obtained by projecting the three-dimensional joint point onto the ground. The sum of pressure values ​​within a certain neighborhood The height of the Z-axis of the three-dimensional joint is greater than the threshold τ1. Joints less than the threshold τ2. The formula is expressed as: Typically, the neighborhood is 25 square centimeters, τ1 is 5 centimeters, and τ2 is 5 centimeters. J(θ,T) represents the coordinates of the three-dimensional joints of the SMPL human 3D parametric model under the control of the human posture parameter θ and the human displacement parameter T.

[0037] Compared to existing technologies, firstly, this invention uses pressure information to determine whether a 3D joint is in contact with the ground. A 3D joint is considered to be in contact with the ground only if it has a pressure value on its projection onto the ground and its height is close to the ground height. Other methods often only determine the height in the Z-axis direction. This greatly improves the accuracy of determining whether a body part is in contact with the ground and indirectly improves the accuracy of human posture estimation methods. Secondly, the whole-body contact loss function of this invention can not only handle the case of foot contact with the ground, but also realize the contact determination of joints that may be in contact with the ground through two-dimensional feature determination. This expands the method's posture estimation range for a wider range of movements and can handle human posture estimation for different types of ground contact movements such as standing, handstand, plank, sitting, and kneeling.

[0038] Example:

[0039] A pressure mat, approximately 2 meters long and 1.5 meters wide or larger, is laid out in the human activity area. Participants perform a series of daily movements on the mat, including standing, lying down, sitting, and planking, without any part of their body contacting the ground outside the mat. This method allows for the precise acquisition of pressure frame sequence data generated by the entire body during various movements. The overall pressure frame data is divided into a T=20-frame sequence. This global pressure frame sequence is then input into a spatial feature encoder composed of convolutional layers, pooling layers, and fully connected layers to obtain the spatial features F of global ground pressure, which includes the parts and locations of the human body that generate pressure on the ground, as well as the relative spatial relationships between the pressure-generating parts of the body and the ground. s In this embodiment, the pressure pad consists of 19,200 pressure measurement units arranged in a matrix of 120 rows and 160 columns. When a person moves on the pressure pad, pressure is applied to some of the pressure measurement units, causing them to generate pressure values. Due to their two-dimensional matrix arrangement, these pressure measurement units with pressure values ​​have two-dimensional spatial positions relative to the pressure pad's coordinate system. After passing through a spatial feature encoder, spatial information such as the relative distances, relative angles, and combined areas between the pressure measurement units with pressure values ​​is obtained.

[0040] The stress frame sequence first passes through a 2D convolutional layer with 1 input channel, 64 output channels, a 3x3 kernel size, and a stride of 2. Next, it passes through a 2D max-pooling layer with a 2x2 kernel size and a stride of 1. Then, it goes through four stages of convolutional layers, each consisting of two residual blocks, and each residual block is composed of two 3x3 convolutional layers. As the number of stages increases, the number of channels in the convolutional kernel increases from 64 to 128, 256, and 512, achieving high-dimensional feature extraction. Finally, it passes through an average pooling layer and a fully connected layer with 512 input channels and 1024 output channels to further enhance the dimensionality, increasing the extraction of high-dimensional information about the global stress spatial features of the human body.

[0041] The pressure spatial feature F passed through the spatial feature encoder s A gated recurrent structure is used to integrate the pressure features between adjacent frames. This gated recurrent structure consists of two layers of bidirectional gated recurrent units, with an input dimension of 1024 and a hidden layer dimension of 1024. The pressure spatial features F are then integrated. s After introducing a gated loop structure, a pressure characteristic GRU1(F) with local temporal properties is obtained. s ), and with the initial pressure space characteristics F s Perform residual connections to obtain combined features of local temporal characteristics and original pressure space characteristics. Among them GRU(F) s F represents the pressure space characteristics. s The output of the gated loop structure. To fully integrate local temporal features, the local temporal spatial features at this point... It needs to go through a gated loop structure again and be combined with the original local temporal spatial features. Residual connections are performed to obtain the final local temporal spatial features. in Indicating pressure space characteristics The output of the gated loop structure is then processed. Next, the local temporal spatial features F are... slocal A self-attention module is incorporated, which is a multi-head self-attention module with 4 heads and an embedding dimension of 1024. Through the global cross-weighting mechanism of the self-attention module, the intra-frame global temporal information of the stress feature is integrated and then combined with the initial local temporal spatial feature F. slocal The summation yields the pressure-space-time characteristic F. st =Attention(F slocal )+F slocal , where Attention(F slocal ) represents the local temporal spatial feature F slocal The specific calculation formula for the output of the self-attention module is as follows: Where SoftMax is the normalized exponential function. For video key vector F slocal The dimension size is typically 1024. For local temporal spatial features F slocal The transpose of .

[0042] The spatiotemporal characteristics of pressure F st The parameters are fed into an SMPL model parameter regressor consisting of multiple fully connected layers to obtain the final human pose parameters θ and displacement parameters T. These estimated human pose and displacement parameters then drive the SMPL 3D parametric human model.

[0043] To verify the superiority of this invention in human pose and displacement estimation, it was compared with other human pose estimation methods. By deploying these models on a unified dataset, the performance differences between this invention and other methods can be clearly revealed. Method 1 uses a convolutional neural network for human pose estimation; Method 2 is trained on a larger dataset based on Method 1. This invention measures the accuracy of human pose and displacement estimation for each model using the whole-body average joint position error, lower-body average joint position error, whole-body average PK alignment joint position error, lower-body average PK alignment joint position error, and global average joint position error. The whole-body average joint position error, lower-body average joint position error, whole-body average PK alignment joint position error, and lower-body average PK alignment joint position error measures the root mean square error between the estimated whole-body joints and lower-body joints and the true values. The smaller the error of these two indicators, the more subtle the difference between the estimated value and the true value, indicating a more accurate estimated human pose. The global average joint position error measures the error between the estimated and actual global coordinates of all joints in the body. It measures the accuracy of the global position estimation for each joint; a smaller index indicates a more accurate global position estimation for each joint. Table 1 compares the present invention with other methods on five indices: global average joint position error, lower body average joint position error, global average Pseudo-alignment joint position error, lower body average Pseudo-alignment joint position error, and global average joint position error.

[0044] Table 1 compares the human posture and displacement estimation indices of existing commonly used methods with those of the present invention.

[0045]

[0046] As shown in Table 1, compared with existing human posture estimation methods, the average position error per joint of the whole body decreased by 42.4-208.7 mm, the average position error per joint of the lower body decreased by 36.4-181.2 mm, the average position error per joint of the whole body with standard alignment decreased by 30.6-134.9 mm, and the global average position error per joint decreased by 15.7-295.2 mm. Since the position error of the lower body joints is lower than that of the whole body, this indicates that pressure has a stronger ability to characterize the stability of the lower body and a stronger ability to estimate the posture and displacement of body parts in contact with the ground.

[0047] To demonstrate the effectiveness of this invention in estimating human posture and displacement, a visual comparison with other methods was performed on the same data, such as... Figure 3 As shown. In the three striding and standing movements, the present invention surpasses Method 2 in both human posture and global displacement. This demonstrates that the present invention has the ability to accurately estimate human posture and global displacement when using only sparse pressure information for human posture estimation.

Claims

1. A human posture estimation device based on global pressure, characterized in that, The device includes: A spatial feature encoder is used to extract global stress spatial features of the human body from a global stress frame sequence of human movements. The long-short-term temporal attention module includes a gated loop structure and a self-attention module. The gated loop structure consists of two gate mechanisms: an update gate and a reset gate. This gated loop structure is used to extract local human pressure spatiotemporal features from the global human pressure spatial features. The self-attention module includes a fully connected layer, a self-attention mechanism, a normalized exponential function, a random deactivation layer, and a layer normalization function. This self-attention module is used to integrate the local human pressure spatiotemporal features through a global cross-weighting mechanism, and then add the integrated features to the local human pressure spatiotemporal features to obtain the pressure spatiotemporal features. An action regressor is used to perform nonlinear regression calculations on the spatiotemporal characteristics of the pressure to obtain the posture and displacement parameters of the human body parameterized model.

2. The human posture estimation device based on global pressure according to claim 1, characterized in that, The global pressure frame sequence includes the pressure exerted on the ground by body parts in contact with the ground, as well as the relative positional relationship between the body parts in contact with the ground.

3. The human posture estimation device based on global pressure according to claim 1, characterized in that, The global pressure spatial characteristics of the human body include the magnitude of the pressure value generated by the human body to each pressure measurement unit of the pressure pad, the two-dimensional spatial position of each pressure measurement unit relative to the coordinate system of the pressure pad, and the spatial information between each pressure measurement unit.

4. The human posture estimation device based on global pressure according to claim 1, characterized in that, The spatial feature encoder includes an initial convolutional layer, a max pooling layer, a four-stage convolutional layer, an average pooling layer, and a fully connected layer. Each of the four-stage convolutional layers consists of two residual blocks, and each residual block consists of two convolutional layers.

5. The human posture estimation device based on global pressure according to claim 1, characterized in that, The motion regressor includes a fully connected layer and a randomly deactivated layer. The motion regressor performs nonlinear regression on the fused features to output human posture parameters and displacement parameters.

6. The method using the human posture estimation device based on global pressure as described in claim 1, characterized in that, The method includes the following steps: Step S1: Construct a network model for the human pose estimation device and train the network model; Step S2: Acquire a global pressure frame sequence of human body movements; input the global pressure frame sequence into a spatial feature encoder to obtain the global pressure spatial features F of the human body. s ; Step S3: Calculate the global pressure spatial characteristics F of the human body. s Inputting long and short-term time series self-attention modules yields the spatiotemporal stress features F. st ; Step S4: Transfer the spatiotemporal characteristics of pressure F st Input the motion regressor to obtain the final human posture and displacement parameters.

7. The method according to claim 6, characterized in that, In step S1, the training loss function is: L = L pose +L 3d +L trans +L contact , where L pose Let L be the human pose loss function. 3d L is the loss function for the three-dimensional joints of the human body. trans Let L be the loss function for human body displacement parameters. contact This is the whole-body contact loss function.

8. The method according to claim 7, characterized in that, The whole-body contact loss function is J c (θ,T) represents the coordinates of the three-dimensional joints of the SMPL human three-dimensional parametric model in contact with the ground, controlled by the human posture parameter θ and the human displacement parameter T. True values ​​of human posture parameters and the true reference value of human body displacement parameters The coordinates of the three-dimensional joints of the SMPL human body three-dimensional parametric model in contact with the ground under control.

9. The method according to claim 8, characterized in that, 3D joint coordinates Where J(θ,T) represents the coordinates of the three-dimensional joints of the SMPL human three-dimensional parametric model under the control of the human posture parameter θ and the human displacement parameter T. The projection points obtained by projecting 3D joints onto the ground The sum of pressure values ​​within a certain neighborhood. τ1 represents the Z-axis height of the 3D joint, and τ2 and τ1 represent the threshold values.

Citation Information

Patent Citations

  • Feature interaction fusion method and system for 3D human body posture estimation

    CN117115915A

  • Human body posture estimation method and device fusing vision and pressure and medium

    CN117593762A