Human body posture estimation device and method based on whole body pressure and monocular video

By integrating whole-body pressure and monocular video features in human posture estimation, the problem of failure to fully consider the interaction between the human body and the ground in the prior art is solved, and a more accurate and reasonable human posture estimation is achieved, which is suitable for fields such as animation and virtual reality.

CN119992646AActive Publication Date: 2025-05-13NANJING UNIV

Patent Information

Application Number
CN202510020594.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-13
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The existing human posture estimation method based on monocular images fails to fully consider the interaction between the human body and the ground environment, resulting in jitter, drift, penetration and sliding phenomena in applications in animation, virtual reality and other fields, and lacks physical constraints, resulting in pose abnormalities that do not conform to physical laws in the real environment.

Method used

A human posture estimation device based on whole-body pressure and monocular video is adopted. The device includes a video feature encoder, a pressure feature encoder, a feature fusion device and an action regressor. By fusing the video features and the whole-body pressure features, non-linear regression calculation is performed to obtain the posture and displacement parameters of the human parametric model.

Benefits of technology

The multi-dimensional representation of human body movements is realized, the accuracy of posture estimation and the rationality of interaction with the ground are enhanced, the unreal problem in the real physics engine environment is solved, the scale limitation is exceeded, and more accurate human body posture and displacement estimation is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992646A_ABST
    Figure CN119992646A_ABST
Patent Text Reader

Abstract

The invention provides a human body posture estimation device and method based on whole body pressure and a monocular video. The device comprises a video feature encoder used for extracting video features related to human body actions from a monocular front view angle video frame sequence of the human body actions; the pressure feature encoder is used for extracting whole body pressure features related to human body actions from the whole body ground pressure frame sequence; the feature fusion device is used for fusing the whole body pressure features and the video features to obtain fused features; and the action regression device is used for performing nonlinear regression calculation according to the fusion features to obtain posture and displacement parameters of the human body parameterized model. According to the method, the interaction relation between the human body and the ground is concerned, the pressure information and the monocular video information are fused, multi-dimensional representation of human body actions is achieved, the defect that a traditional monocular video human body posture estimation method neglects environment information is overcome, and the accuracy of human body posture estimation and the reasonability of interaction between the human body posture estimation method and the ground are remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human body posture estimation in the field of deep learning technology, and in particular to a human body posture estimation device and method based on whole body pressure and monocular video. Background Art

[0002] Human pose estimation is a method that converts human motion into parameters that can represent position, angle, etc. With the rapid development of deep learning technology, accurate and reasonable human motion generation technology and its application in virtual humans or humanoid robots have shown broad application value and research prospects. At present, the mainstream human pose estimation method is the human pose estimation method based on monocular images. This method takes a monocular human image as input, and fits the parameters of the human three-dimensional joint points or human skin model through deep learning technology to obtain the corresponding human motion representation.

[0003] However, the existing human posture estimation methods based on monocular images mainly focus on the human body itself in the image, and fail to fully consider the interaction with environmental factors such as the ground. Therefore, when these actions are mapped to three-dimensional space or real environment, unreasonable interactions with the ground often occur, such as jitter and drift in time, and penetration and sliding in space. These phenomena will limit the further application of human motion representation in animation, virtual reality, embodied intelligence and other fields. At the same time, since these human body actions estimated based on images lack physical constraints or physical information, when driven in a real environment with a physical engine, phenomena that do not conform to physical laws often occur, such as floating in the air or unreasonable forward and backward posture anomalies. There are also some existing methods that consider the interaction between the human body and the ground, but they often only consider the interaction information between the feet and the ground, and do not consider the interaction information between other parts of the human body and the ground, such as knees, hips, torso, elbows, hands, head and other parts; and these explorations only stay on the contact information between body parts and the environment, without considering the deeper interaction information based on physical dynamics mechanisms and the relative position relationship between various parts of the human body. Therefore, how to provide a whole-body human posture estimation method and system that can meet the real physical laws is a problem that needs to be solved urgently. Summary of the invention

[0004] In view of the deficiencies in the prior art, the object of the present invention is to provide a human body posture estimation device and method based on whole body pressure and monocular video.

[0005] The technical solution adopted by the device of the present invention is:

[0006] A human posture estimation device based on whole-body pressure and monocular video comprises: a video feature encoder for extracting video features related to human body movements from a monocular frontal perspective video frame sequence of human body movements; a pressure feature encoder for extracting whole-body pressure features related to human body movements from a whole-body ground pressure frame sequence; a feature fusion device for fusing the whole-body pressure features with the video features to obtain fusion features; and a motion regressor for performing nonlinear regression calculations based on the fusion features to obtain posture and displacement parameters of a human body parameterized model.

[0007] The present invention also provides a method for using the above-mentioned human posture estimation device based on whole body pressure and monocular video, the method comprising the following steps:

[0008] Step S1, constructing a network model of a human posture estimation device and training the network model;

[0009] Step S2, collecting a monocular frontal perspective video frame sequence of human body movements; inputting the video frame sequence into a video feature encoder to obtain a monocular video feature of the human body movements;

[0010] Step S3, collecting a whole body ground pressure frame sequence of human body movements; inputting the whole body ground pressure frame sequence into a pressure feature encoder to obtain a whole body pressure feature of the human body movements;

[0011] Step S4, inputting the whole body pressure feature and the monocular video feature of the human body motion into a feature fusion device, fusing the whole body pressure feature and the video feature, and outputting a fusion feature containing the whole body pressure information and the video information of the human body motion;

[0012] Step S5: input the fused features into the action regressor to obtain the final human body posture and displacement parameters.

[0013] Compared with the prior art, the present invention has the following beneficial effects:

[0014] (1) The present invention focuses on the interaction between the human body and the ground. By fusing pressure information with monocular video information, it realizes the multi-dimensional representation of human body movements, making up for the deficiency of traditional monocular video human body posture estimation methods that ignore environmental information, and significantly enhancing the accuracy of human body posture estimation and the rationality of its interaction with the ground.

[0015] (2) By extracting high-dimensional features from pressure information, we can deeply extract and utilize real physical information, solving the problem of unreality of human posture estimation when driven by an environment with a real physical engine.

[0016] (3) The present invention breaks through the scale limitation of posture estimation and utilizes the large-scale detection capability of the ground pressure pad to comprehensively capture and analyze the pressure information of various parts of the body in contact with the ground and their spatial relationship, thereby achieving more accurate posture and displacement estimation of various parts of the human body. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a structural diagram of the device of the present invention;

[0018] Figure 2 This is a structural diagram of a feature fusion device in an embodiment of the present invention;

[0019] Figure 3 A comparison diagram of human body posture estimation effects in an embodiment of the present invention;

[0020] Figure 4 4 is a comparison diagram of the human body displacement estimation effect in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] Embodiments of the present invention are described in detail below with reference to the accompanying drawings, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.

[0022] like Figure 1 As shown, this embodiment provides a human posture estimation device based on whole-body pressure and monocular video, including: a video feature encoder, used to extract video features related to human body movements from a monocular frontal perspective video frame sequence of human body movements; a pressure feature encoder, used to extract whole-body pressure features related to human body movements from a whole-body ground pressure frame sequence; a feature fusion device, used to fuse the whole-body pressure features and the video features to obtain fusion features; and a motion regressor, used to perform nonlinear regression calculation based on the fusion features to obtain the posture and displacement parameters of the human body parameterized model, and finally drive the human body parameterized model.

[0023] The specific implementation device includes a memory and one or more processors, wherein the memory stores codes and executable files of a video feature encoder, a pressure feature encoder, a feature fusion device, and a motion regressor. When the processor executes the executable files of the video feature encoder, the pressure feature encoder, the feature fusion device, and the motion regressor, it is used to implement a human body posture estimation method based on whole-body pressure and monocular video of the present invention.

[0024] The video feature encoder includes a video feature extraction network, which consists of four-stage convolutional layers. Each stage of the convolutional layer is composed of a two-dimensional convolution function, a batch normalization function, and a linear rectification function. Each stage downsamples the feature map of the previous stage and fuses the downsampled feature map with the feature map of each previous stage to obtain the final video features that fuse all scales.

[0025] The pressure feature encoder includes a pressure feature extraction network, which includes an initial convolution layer, a maximum pooling layer, a four-stage convolution layer, an average pooling layer and a fully connected layer. Each of the four-stage convolution layers consists of two residual blocks, and each residual block consists of two convolution layers. As the number of stages increases, the spatial resolution of the features will decrease and the number of channels will increase, thereby realizing high-dimensional feature extraction.

[0026] The feature fuser includes a feature fusion network, which includes a fully connected layer, a cross attention mechanism, a normalized exponential function, and a position-by-position feedforward network, such as Figure 2 As shown in the figure, the feature fusion network adopts a multi-head attention mechanism, passes the pressure feature through a fully connected layer to obtain a pressure query vector, passes the video feature through a fully connected layer to obtain a video key vector and a video numerical vector, calculates the correlation matrix of the pressure query vector and the video key vector, passes the correlation matrix through a normalized exponential function to obtain the attention weight, and then uses the obtained attention weight to perform matrix product calculation on the video numerical vector to obtain the cross-attention fusion feature, and the cross-attention fusion feature is residually connected with the initial pressure feature and the initial video feature to obtain the attention fusion feature with residual information, and the attention fusion feature with residual information is input into a position-by-position element-by-element feedforward network composed of two fully connected layers, a random dropout layer and a layer normalization structure to obtain the attention fusion feature with residual information and local features, and finally the attention fusion feature with residual information and local features is concatenated with the original pressure feature to obtain the final fusion feature.

[0027] The action regressor includes an action regression network, which consists of a fully connected layer and a random dropout layer. The fused features are subjected to nonlinear regression to output human posture parameters and displacement parameters.

[0028] This embodiment also provides a method for estimating human posture based on whole body pressure and monocular video, comprising the following steps:

[0029] Step S1, constructing a network model of the above-mentioned human posture estimation device and training the network model;

[0030] Step S2, collecting a monocular frontal perspective video frame sequence of human body movements; inputting the video frame sequence into a video feature encoder to obtain a monocular video feature of the human body movements.

[0031] Step S3, collecting the whole body ground pressure frame sequence of human body movements; inputting the whole body ground pressure frame sequence into the pressure feature encoder to obtain the whole body pressure feature of the human body movements. The whole body ground pressure frame sequence includes the pressure generated by the contact between the various parts of the human body that can contact the ground and the ground, such as the downward pressure generated by the feet, knees, buttocks, trunk, elbows, hands, head, etc. on the ground in different movements.

[0032] Step S4, input the whole body pressure feature and monocular video feature of human body motion into the feature fusion device, fuse the whole body pressure feature and video feature, and output the fusion feature containing the whole body pressure information and video information of human body motion. The whole body pressure feature of human body motion includes the pressure exerted on the ground by the body parts in contact with the ground, and also includes the relative position relationship between the parts in contact with the ground. The relative position relationship between the parts in contact with the ground refers to the position relationship such as the distance and angle between the parts in contact with the ground. For example, when the human body performs a plank support, the wrist, elbow and foot will exert pressure on the ground. The pressure pad is composed of a plurality of small pressure units arranged in an array, and these small units can record the pressure values ​​exerted on the ground by the wrist, elbow and foot. The pressure value distribution of multiple pressure units not only reflects the pressure and contact area of ​​each part in contact with the ground, but also can characterize the position information such as the distance and offset angle of the pressure area between the wrist, elbow and foot.

[0033] The overall structure of the feature fusion unit is as follows: Figure 2 As shown, the whole body pressure feature F can be obtained from step S1 and step S2. p and video feature F v , the whole body pressure characteristic F p As the pressure query vector Q p , the video feature F v As the video value vector V v and video key vector K v The multimodal cross-fusion attention mechanism of pressure video is used to capture the long-range dependencies in the sequence, thereby obtaining the cross-attention fusion feature that combines the whole body pressure feature and the video feature. Among them, SoftMax is a normalized exponential function, d k is the video key vector K v The dimension size of is the video key vector K vThe transposed vector of . Through the cross-attention mechanism, the relative relationship between the dynamic mechanism represented by the pressure feature and the texture information represented by the video feature can be well obtained. The pressure size, distance and angle information between pressure units represented in the pressure feature are integrated with the human body texture contour information represented in the video feature to obtain the human body motion feature representation from the two-dimensional angle of the camera direction and the vertical direction of the ground. It can better integrate the visual information represented by the video feature with the physical dynamic mechanism represented by the pressure feature. p ,K v ,V v ) and initial pressure characteristics F p and video feature F v Perform matrix addition to obtain the attention fusion feature F with residual information ca =CrossAttention(Q p ,K v ,V v )+F p +F v This will make the fusion feature have a certain degree of original features, better characterize the pressure size of the pressure feature and the distance and angle between the pressure units, as well as the human body motion texture and contour of the visual feature. ca Put it into the position element-by-element feedforward network, and after learning the local features, we get the attention fusion feature F′ with residual information and local features ca , and the original pressure characteristic F p After splicing, the final fusion feature F is obtained fusion . The purpose of splicing with the original pressure feature is to enable the fusion feature to contain more information about the pressure magnitude and distance and angle between pressure units represented in the pressure feature, so as to achieve implicit representation of the physical dynamic mechanism. The pressure feature contains more information related to contact and physical dynamic mechanism, so after splicing, it can represent more relative relationships between the human body and the ground, while maintaining the whole body and local posture of the video feature. At the same time, it can also represent the relative position relationship between the parts of the human body that are in contact with the ground, and supplement the video features that are missing due to self-occlusion.

[0034] Step S5: Fusion feature F fusion Put it into the SMPL (A Skinned Multi-Person Linear Model) model action regressor to get the final human posture θ and displacement parameter T.

[0035] When the parameters of the video feature encoder, pressure feature encoder, feature fusion device and motion regressor are trained in step S1, the training sample data is input into the network model, and the training loss function L = L is calculated.pose +L 3d +L 2d +L trans +L contact , the network model is iteratively trained, the AdamW optimizer is used to optimize the network parameters, the learning rate is set to 1e-4, the total number of iterations is 1500, and the network model parameters are determined based on the training loss value.

[0036] The human posture loss function θ is the predicted human posture parameter, is the benchmark truth value of human posture parameters.

[0037] Human body 3D joint loss function J(θ,T) is the 3D joint coordinates of the SMPL human body 3D parametric model controlled by the human body posture parameter θ and the human body displacement parameter T. is the true value of human body posture parameter and the true value of human displacement parameters The 3D joint point coordinates of the SMPL human 3D parametric model under control.

[0038] Human body 2D joint loss function The error calculation is performed after the three-dimensional joint points of the three-dimensional human body template are orthogonally projected O(·) in the Z-axis direction. The calculation formula of the orthogonal projection is (x, y) = O(x, y, z). The difference between the two-dimensional joint point loss function of the human body and the commonly used two-dimensional joint point loss function of the human body is that the commonly used two-dimensional joint point loss function of the human body adopts a weak perspective projection model to perform two-dimensional projection on the three-dimensional joint points of the human body, and the present invention adopts an orthogonal projection camera model to perform two-dimensional projection on the three-dimensional joint points of the human body; different from the traditional weak perspective projection camera model, the orthogonal projection camera model does not depend on the distance between the human body and the camera in the Z-axis direction, and only constrains the two-dimensional joint points on the XY plane of the human body, realizing the decoupling of the human body posture parameters and the human body displacement parameters, making the loss function calculation of the two-dimensional joint points more simplified and accurate, avoiding the complex calculation in the depth direction, reducing the mutual interference between the estimation of the human body posture parameters and the human body displacement parameters, and thus realizing a more accurate estimation of the human body posture parameters and the human body displacement parameters.

[0039] The human body displacement parameter loss function is: T is the predicted human body displacement parameter, It is the reference truth value of human body displacement parameters.

[0040] The human body contact loss function is: J c (θ,T) is the coordinate of the three-dimensional joint point of the SMPL human body three-dimensional parametric model in contact with the ground under the control of the human body posture parameter θ and the human body displacement parameter T. is the true value of human body posture parameter and the true value of human displacement parameters The coordinates of the three-dimensional joint points of the SMPL human body three-dimensional parametric model under control in contact with the ground.

[0041] The three-dimensional joint point J in contact with the ground c (θ,T) refers to the projection point obtained by projecting the three-dimensional joint point onto the ground The sum of the pressure values ​​within a certain neighborhood Greater than the threshold τ1, and the Z-axis height of the three-dimensional joint point The joint points that are less than the threshold τ2. The formula is expressed as: Typically, the neighborhood range is 25 square centimeters, τ1 is 5, and τ2 is 5 centimeters. Where J(θ,T) represents the three-dimensional joint point coordinates of the SMPL human body three-dimensional parametric model under the control of the human body posture parameter θ and the human body displacement parameter T.

[0042] Compared with the prior art, firstly, the present invention uses pressure information to determine whether a three-dimensional joint is in contact with the ground. A three-dimensional joint is determined to be in contact with the ground only if there is a pressure value on the ground projection and the height is close to the ground height. Other methods often only have height determination in the Z-axis direction, which greatly improves the accuracy of determining whether a body part is in contact with the ground, and indirectly improves the accuracy of the human posture estimation method; secondly, the whole-body contact loss function of the present invention can not only handle the situation where the feet are in contact with the ground, but also realize the contact determination of the joints of the whole body that may be in contact with the ground through feature determination in two dimensions, which expands the method's posture estimation range for a wider range of actions, and can handle human posture estimation for actions with different types of contact with the ground, such as standing, handstand, plank support, sitting down, kneeling, etc.

[0043] Example:

[0044] An RGB camera is placed in front of the human body to collect a frontal perspective video frame sequence of human motion. The object detection method YOLO is used to extract the bounding box of the human body in the entire video, and the human body is cut out from the entire image containing the human body and the environment. The human body video cut frame sequence is input into the video feature encoding network HRNet to obtain the monocular video features of human motion.

[0045] A pressure pad is placed on the ground where the human body moves. Different human movements (including common daily movements such as standing, lying down, sitting, and plank support) are performed on the pressure pad, and no part of the body is in contact with the ground outside the pressure pad. The size of the pressure pad in this embodiment is about two meters long and one and a half meters wide or more. The pressure pad data recording program is turned on to collect the whole body pressure frame sequence of the human body movement, and the whole body pressure frame sequence is directly input into the pressure feature encoding network Resnet to obtain the whole body pressure feature of the human body movement.

[0046] In order to represent the temporal information in the attention mechanism, the whole body pressure feature F after temporal alignment is first p and video feature F v Position encoding is performed, and the general position encoding based on sine and cosine functions is respectively combined with the whole body pressure feature F p and video feature F v Add together so that the whole body pressure feature and video feature have position information. The dimensions of the pressure feature and video feature are set to 2048. The pressure query vector Q p is the whole body pressure characteristic F p The linear transformation is obtained, Q p =F p W Q , where W Q is the trainable query vector parameter matrix, the video value vector V v and video key vector K v is the video feature F v The linear transformation is obtained, V v =F v W V ,K v =F v W K , where W V is a trainable numerical vector parameter matrix, W K is the trainable key vector parameter matrix.

[0047] Adopt the multi-modal cross-fusion attention mechanism of pressure video, and use the multi-head attention mechanism to calculate the cross-attention fusion feature Among them, SoftMax is a normalized exponential function, d k is the video key vector K v The dimension size is usually 2048. is the video key vector K v The transposed vector of , the number of attention heads is 4. The common residual structure is used to fuse the initial feature information to obtain the attention fusion feature F with residual information ca , the attention fusion feature F with residual information caPut it into the position element-by-element feedforward network to obtain the attention fusion feature F with residual information and local features c ′ a , where the position element-by-element feedforward network consists of two fully connected layers, a random dropout layer, and a layer normalization structure, with an input dimension of 2048, a hidden layer dimension of 2048, and a random dropout probability set to 0.1. Finally, the original whole body pressure feature F p After splicing, the final fusion feature F is obtained fusion .

[0048] The fusion feature F fusion The model is put into the SMPL model parameter regression network composed of multiple fully connected layers, and the final human body posture and displacement parameters are output. The estimated human body posture and displacement parameters can drive the SMPL human body 3D parametric model.

[0049] In order to prove the advantages of the present invention over other human posture estimation methods, common methods are used to estimate posture on the same batch of data and compared. Method 1 estimates human posture by convolutional neural network and gated recurrent unit; Method 2 estimates human posture by convolutional neural network and multi-layer perceptron; Method 3 estimates human posture by building a general model through visual converter; Method 4 estimates image optical flow information by convolutional neural network for human posture estimation; Method 5 estimates human posture by convolutional neural network, recurrent neural network and multi-layer perceptron. The present invention measures the human posture estimation accuracy of each model by averaging per-joint position error, per-model vertex position error and acceleration error. The average per-joint position error and per-model vertex position error indicators measure the root mean square error between the estimated human joints and human skin vertices and the true value. The smaller the error of the two indicators, the smaller the difference between the estimated value and the true value, that is, the more accurate the estimated human posture. The acceleration error indicator measures the jitter of the joint point in the timing. The smaller the acceleration error indicator, the smoother the posture estimation. Table 1 compares the present invention with other methods in terms of the three indicators of average per-joint position error, per-model vertex position error and acceleration error:

[0050] Table 1 Comparison of human posture estimation indicators between existing common methods and the method of the present invention

[0051]

[0052]

[0053] As shown in Table 1, compared with the existing human posture estimation method, the average position error of each joint of the present invention is reduced by 9.8-19.6mm, the position error of each model vertex is reduced by 10.0-24.3mm, and the acceleration is reduced by 11.6-434.3m / s2 .

[0054] In order to prove the ability of the present invention to estimate human body displacement parameters relative to other methods, this embodiment uses commonly used models for comparison on the same batch of data. The present invention measures the human body displacement estimation effect between various models through global displacement error, global average per-joint position error, and global joint point jitter error. The global displacement error refers to the error between the estimated value and the true value of the human body's global displacement; the global average per-joint error refers to the root mean square error between the estimated value and the true value of the human body's joints; these two indicators measure the effect of the global displacement estimation of the human body, and the smaller the indicator, the more accurate the human body displacement estimation. The global joint point jitter index refers to the jitter error of adjacent frames of the global joint points of the human body. The smaller the indicator, the higher the smoothness of the human body displacement. Table 2 is a comparison of the present invention and other methods in terms of three indicators: global displacement error, global average per-joint position error, and global joint point jitter error (method one, method two, and method three do not have the ability to estimate global displacement):

[0055] Table 2 Comparison of human body displacement estimation indicators between existing common methods and the method of the present invention

[0056] method Global displacement error Global average per-joint position error Global joint jitter error Method 4 1193(1151.4↓) 141.2(80.4↓) 686(626↓) Method 5 1023(981.4↓) 75.6(14.8↓) 92(32↓) The present invention 41.6 60.8 60

[0057] As shown in Table 2, compared with the existing human body displacement estimation method, the global displacement error of the present invention is reduced by 981.4-1151.4mm, the global average position error of each joint is reduced by 14.8-80.4mm, and the global joint point jitter error is reduced by 32-626m / s. 2 .

[0058] In order to demonstrate the human body posture estimation effect of the present invention, this embodiment is compared with other methods on the same data. Figure 3 As shown. In the plank action, the human body's limbs estimated by the present invention are all in contact with the ground, while the feet of the human body estimated by method three are not in contact with the ground. There is an estimation error in the left foot of method two, and there are certain estimation errors in the limbs of method three. In the stride action, the human body's feet estimated by the present invention are all in contact with the ground, while methods one, two and three all have certain degrees of foot contact estimation errors, and even method three has an estimation error of floating in the air.

[0059] In order to demonstrate the human body displacement estimation effect of the present invention, this embodiment is compared with other methods on the same data. Figure 4 In the circle movement, the human body displacement trajectory estimated by the present invention has a small error with the true value, the human body displacement estimated by method five has the problem of wrong direction, and method four does not estimate the real human body displacement.

Claims

1. A human posture estimation device based on whole body pressure and monocular video, characterized in that: The device includes: A video feature encoder is used to extract video features related to human actions from a monocular frontal view video frame sequence of human actions; A pressure feature encoder, used for extracting whole-body pressure features related to human body movements from a whole-body ground pressure frame sequence; A feature fusion device, used for fusing the whole body pressure feature and the video feature to obtain a fusion feature; The action regressor is used to perform nonlinear regression calculation according to the fusion features to obtain the posture and displacement parameters of the human body parameterized model.

2. The human body posture estimation device based on whole body pressure and monocular video according to claim 1, characterized in that: The whole-body pressure characteristics include the pressure exerted on the ground by the parts of the human body in contact with the ground, and also include the relative positional relationship between the parts of the human body in contact with the ground.

3. The human body posture estimation device based on whole body pressure and monocular video according to claim 1, characterized in that: The video feature encoder includes four-stage convolutional layers, each of which is composed of a two-dimensional convolution function, a batch normalization function and a linear rectification function. Each stage downsamples the feature map of the previous stage, and fuses the downsampled feature map with the feature map of each previous stage to obtain the final video features that fuse various scales.

4. The human body posture estimation device based on whole body pressure and monocular video according to claim 1, characterized in that: The pressure feature encoder includes an initial convolution layer, a maximum pooling layer, a four-stage convolution layer, an average pooling layer and a fully connected layer, wherein each of the four-stage convolution layers is composed of two residual blocks, and each residual block is composed of two convolution layers. As the number of stages increases, the spatial resolution of the features will decrease and the number of channels will increase, thereby realizing high-dimensional feature extraction.

5. The human body posture estimation device based on whole body pressure and monocular video according to claim 1, characterized in that: The feature fusion device includes a fully connected layer, a cross-attention mechanism, a normalized exponential function and a position-by-element feedforward network; wherein, the whole-body pressure feature is passed through a fully connected layer to obtain a pressure query vector, the video feature is passed through a fully connected layer to obtain a video key vector and a video numerical vector, the correlation matrix of the pressure query vector and the video key vector is calculated, the correlation matrix is ​​passed through a normalized exponential function to obtain an attention weight, and then the obtained attention weight is used to perform matrix product calculation on the video numerical vector to obtain a cross-attention fusion feature, the cross-attention fusion feature is residually connected with the whole-body pressure feature and the video feature to obtain an attention fusion feature with residual information, the attention fusion feature with residual information is input into a position-by-element feedforward network composed of two layers of fully connected layers, a random inactivation layer and a layer normalization function to obtain an attention fusion feature with residual information and local features, and finally the feature fusion device splices the attention fusion feature with residual information and local features with the whole-body pressure feature to obtain a final fusion feature.

6. The human body posture estimation device based on whole body pressure and monocular video according to claim 1, characterized in that: The action regressor includes a fully connected layer and a random dropout layer. The action regressor performs nonlinear regression on the fused features to output human body posture parameters and displacement parameters.

7. A method for using a human posture estimation device based on whole body pressure and monocular video as claimed in claim 1, characterized in that: The method comprises the following steps: Step S1, constructing a network model of a human posture estimation device and training the network model; Step S2, collecting a monocular frontal perspective video frame sequence of human body movements; inputting the video frame sequence into a video feature encoder to obtain a monocular video feature of the human body movements; Step S3, collecting a whole body ground pressure frame sequence of human body movements; inputting the whole body ground pressure frame sequence into a pressure feature encoder to obtain a whole body pressure feature of the human body movements; Step S4, inputting the whole body pressure feature and the monocular video feature of the human body motion into a feature fusion device, fusing the whole body pressure feature and the video feature, and outputting a fusion feature containing the whole body pressure information and the video information of the human body motion; Step S5: input the fused features into the action regressor to obtain the final human body posture and displacement parameters.

8. The method according to claim 7, characterized in that In step S1, the training loss function is: L = L pose +L 3d +L 2d +L trans +L contact, Among them, L pose is the human posture loss function, L 3d is the loss function of the 3D joint points of the human body, L 2d is the loss function of the two-dimensional joint points of the human body, L trans is the human body displacement parameter loss function, L contact is the human body contact loss function.

9. The method according to claim 8, characterized in that Human body 2D joint loss function Where J(θ, T) is the three-dimensional joint coordinates of the SMPL human body three-dimensional parametric model controlled by the human body posture parameter θ and the human body displacement parameter T. is the true value of human body posture parameter benchmark and the true value of human displacement parameters The 3D joint coordinates of the SMPL human 3D parametric model under control, and O(·) represents the orthogonal projection.

10. The method according to claim 8, characterized in that The human body contact loss function is: J c (θ, T) is the coordinate of the three-dimensional joint point of the SMPL human body three-dimensional parametric model in contact with the ground under the control of the human body posture parameter θ and the human body displacement parameter T, is the true value of human body posture parameter benchmark and the true value of human displacement parameters The coordinates of the three-dimensional joint points of the SMPL human body three-dimensional parametric model under control in contact with the ground.

Citation Information

Patent Citations

  • Behavioral disorder detection method considering vision and plantar pressure multimode perception

    CN115601840A

  • Monocular video-based multi-stage human motion capture method and device, and medium

    CN116386141A

  • Human body posture estimation method and device fusing vision and pressure and medium

    CN117593762A

  • Pose fusion estimation

    US20230316734A1

Cited By

  • Monocular video-based humanoid robot whole body motion generation method and system

    CN122200206A