Human pose estimation apparatus based on whole body pressure and monocular video and method thereof

By combining whole-body pressure and monocular video to create a human pose estimation device, the problem of unreasonable human-ground interaction in monocular image methods is solved, achieving more accurate human pose and displacement estimation and meeting the application requirements of real physical laws.

CN119992646BActive Publication Date: 2025-11-11NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510020594.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-11-11
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

Existing monocular image-based human pose estimation methods fail to adequately consider the interaction between the human body and environmental factors such as the ground, resulting in unreasonable motion representations in three-dimensional space, such as temporal jitter and drift, and spatial penetration and sliding phenomena. Furthermore, the lack of physical constraints leads to abnormal poses that do not conform to physical laws in real-world environments.

Method used

A human posture estimation device based on whole-body pressure and monocular video is adopted. Video and whole-body pressure features are extracted by video feature encoder and pressure feature encoder, and then fused by feature fusion. A motion regressor performs nonlinear regression calculation to obtain the posture and displacement parameters of the human parametric model.

Benefits of technology

It achieves multi-dimensional representation of human movements, enhances the accuracy of posture estimation and the rationality of interaction with the ground, breaks through the limitations of scale, and can more accurately capture the pressure information of various parts of the body in contact with the ground and their spatial relationships, meeting the needs of real physical information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992646B_ABST
    Figure CN119992646B_ABST
Patent Text Reader

Abstract

This invention provides a human pose estimation device and method based on whole-body pressure and monocular video. The device includes: a video feature encoder for extracting video features related to human movement from a monocular frontal view video frame sequence; a pressure feature encoder for extracting whole-body pressure features related to human movement from a whole-body ground pressure frame sequence; a feature fusion unit for fusing the whole-body pressure features and video features to obtain fused features; and a motion regressor for performing nonlinear regression calculations based on the fused features to obtain the pose and displacement parameters of a parametric human body model. This invention focuses on the interaction between the human body and the ground. By fusing pressure information and monocular video information, it achieves a multi-dimensional representation of human movement, overcoming the shortcomings of traditional monocular video human pose estimation methods that neglect environmental information, and significantly enhancing the accuracy of human pose estimation and the rationality of its interaction with the ground.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human pose estimation in the field of deep learning technology, and particularly to a human pose estimation device and method based on whole-body pressure and monocular video. Background Technology

[0002] Human pose estimation is a method that converts human actions into representative parameters such as position and angle. With the rapid development of deep learning technology, accurate and reasonable human motion generation techniques, and their applications in virtual humans or humanoid robots, have demonstrated broad application value and research prospects. Currently, the mainstream human pose estimation method is based on monocular images. This method uses a monocular human image as input and fits the parameters of the human body's 3D joints or skinned model using deep learning technology to obtain the corresponding human motion representation.

[0003] However, existing monocular image-based human pose estimation methods primarily focus on the human body itself within the image, failing to adequately consider interactions with environmental factors such as the ground. Therefore, when these movements are mapped to 3D space or a real-world environment, unreasonable interactions with the ground often occur, such as temporal jitter and drift, and spatial penetration and sliding phenomena. These phenomena limit the further application of human motion representation in animation, virtual reality, and embodied intelligence. Furthermore, because these image-based human movements lack physical constraints or information, when driven in a real-world environment with a physics engine, they often exhibit physical inconsistencies, such as floating in the air or unreasonable forward or backward tilts. Some existing methods consider human-ground interaction, but often only consider the interaction between the feet and the ground, neglecting the interaction between other body parts and the ground, such as the knees, hips, torso, elbows, hands, and head. Moreover, these explorations only address contact information between body parts and the environment, without considering deeper-level physical dynamics-based interactions and the relative positional relationships between different body parts. Therefore, how to provide a method and system for estimating the whole-body human posture that can meet the real physical laws is an urgent problem to be solved. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, the purpose of this invention is to propose a human posture estimation device and method based on whole-body pressure and monocular video.

[0005] The technical solution adopted by the device of the present invention is as follows:

[0006] A human pose estimation device based on whole-body pressure and monocular video includes: a video feature encoder for extracting video features related to human movements from a monocular frontal view video frame sequence; a pressure feature encoder for extracting whole-body pressure features related to human movements from a whole-body ground pressure frame sequence; a feature fusion unit for fusing the whole-body pressure features and the video features to obtain fused features; and a motion regressor for performing nonlinear regression calculations based on the fused features to obtain the pose and displacement parameters of a human parametric model.

[0007] This invention also provides a method for estimating human pose using the above-described human posture estimation device based on whole-body pressure and monocular video, the method comprising the following steps:

[0008] Step S1: Construct a network model for the human pose estimation device and train the network model;

[0009] Step S2: Collect a monocular frontal view video frame sequence of human movements; input the video frame sequence into a video feature encoder to obtain monocular video features of human movements;

[0010] Step S3: Collect the whole-body ground pressure frame sequence of human movement; input the whole-body ground pressure frame sequence into the pressure feature encoder to obtain the whole-body pressure features of human movement;

[0011] Step S4: Input the whole-body pressure features of human movement and the monocular video features into the feature fusion unit to fuse the whole-body pressure features and video features, and output the fused features containing the whole-body pressure information of human movement and video information.

[0012] Step S5: Input the fused features into the motion regressor to obtain the final human posture and displacement parameters.

[0013] Compared with the prior art, the present invention has the following beneficial effects:

[0014] (1) This invention focuses on the interaction between the human body and the ground. By integrating pressure information and monocular video information, it realizes the multi-dimensional representation of human movements, which makes up for the shortcomings of traditional monocular video human posture estimation methods that ignore environmental information, and significantly enhances the accuracy of human posture estimation and the rationality of its interaction with the ground.

[0015] (2) By extracting high-dimensional features from pressure information, we can achieve in-depth extraction and utilization of real physical information, thus solving the problem of unrealistic human posture estimation in environments driven by real physical engines.

[0016] (3) This invention breaks through the scale limitation of attitude estimation. By utilizing the large-scale detection capability of the ground pressure pad, it comprehensively captures and analyzes the pressure information of each part of the body in contact with the ground and their spatial relationship, so as to achieve more accurate attitude and displacement estimation of each part of the human body. Attached Figure Description

[0017] Figure 1 This is a structural diagram of the device of the present invention;

[0018] Figure 2 This is a structural diagram of the feature fusion unit in an embodiment of the present invention;

[0019] Figure 3 This is a comparison chart of human pose estimation results in embodiments of the present invention;

[0020] Figure 4 This is a comparison chart of the human body displacement estimation effects in embodiments of the present invention. Detailed Implementation

[0021] Embodiments of the present invention will now be described in detail with reference to the accompanying drawings, examples of which are illustrated in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0022] like Figure 1 As shown, this embodiment provides a human pose estimation device based on whole-body pressure and monocular video, including: a video feature encoder for extracting video features related to human movements from a monocular frontal view video frame sequence; a pressure feature encoder for extracting whole-body pressure features related to human movements from a whole-body ground pressure frame sequence; a feature fusion unit for fusing the whole-body pressure features and video features to obtain fused features; and a motion regressor for performing nonlinear regression calculations based on the fused features to obtain the pose and displacement parameters of the human parameterized model, ultimately driving the human parameterized model.

[0023] The specific implementation device includes a memory and one or more processors. The memory stores the code and executable files of a video feature encoder, a pressure feature encoder, a feature fusion unit, and a motion regressor. When the processor executes the executable files of the video feature encoder, the pressure feature encoder, the feature fusion unit, and the motion regressor, it is used to implement the human posture estimation method based on whole-body pressure and monocular video of the present invention.

[0024] The video feature encoder includes a video feature extraction network, which consists of four convolutional layers. Each convolutional layer is composed of a two-dimensional convolutional function, a batch normalization function, and a linear rectified function. Each stage downsamples the feature map of the previous stage and fuses the downsampled feature map with the feature map of each previous stage to obtain the final video features fused from various scales.

[0025] The stress feature encoder includes a stress feature extraction network, which consists of an initial convolutional layer, a max pooling layer, a four-stage convolutional layer, an average pooling layer, and a fully connected layer. Each of the four-stage convolutional layers consists of two residual blocks, and each residual block consists of two convolutional layers. As the number of stages increases, the spatial resolution of the features decreases while the number of channels increases, thus achieving high-dimensional feature extraction.

[0026] The feature fusion processor includes a feature fusion network, which comprises fully connected layers, a cross-attention mechanism, a normalized exponential function, and a position-by-position element-wise feedforward network, such as... Figure 2 As shown, the feature fusion network employs a multi-head attention mechanism. The stress features are passed through a fully connected layer to obtain a stress query vector, and the video features are passed through a fully connected layer to obtain video key vectors and video numerical vectors. The correlation matrix between the stress query vector and the video key vectors is calculated. The correlation matrix is ​​then processed through a normalized exponential function to obtain attention weights. The obtained attention weights are then used to perform matrix multiplication on the video numerical vectors to obtain cross-attention fusion features. These cross-attention fusion features are residually connected to the initial stress features and initial video features to obtain attention fusion features with residual information. These residual attention fusion features are then input into a position-by-element feedforward network consisting of two fully connected layers, a random deactivation layer, and a layer normalization structure to obtain attention fusion features with residual information and local features. Finally, these attention fusion features with residual information and local features are concatenated with the original stress features to obtain the final fused feature.

[0027] The motion regressor includes a motion regression network, which consists of fully connected layers and random deactivation layers. It performs nonlinear regression on the fused features to output human pose parameters and displacement parameters.

[0028] This embodiment also provides a human pose estimation method based on whole-body pressure and monocular video, including the following steps:

[0029] Step S1: Construct the network model of the above-mentioned human pose estimation device and train the network model;

[0030] Step S2: Collect a monocular frontal view video frame sequence of human body movements; input the video frame sequence into the video feature encoder to obtain the monocular video features of human body movements.

[0031] Step S3: Acquire a full-body ground pressure frame sequence of human movements; input the full-body ground pressure frame sequence into the pressure feature encoder to obtain the full-body pressure features of human movements. The full-body ground pressure frame sequence includes the pressure generated by various parts of the human body that can contact the ground, such as the feet, knees, hips, torso, elbows, hands, and head, which exert downward pressure on the ground during different movements.

[0032] Step S4: Input the full-body pressure features of human movement and the monocular video features into the feature fusion unit to fuse the full-body pressure features and video features, outputting fused features containing full-body pressure information and video information. The full-body pressure features of human movement include the pressure exerted on the ground by the body parts in contact with the ground, as well as the relative positional relationships between these parts. The relative positional relationships between the body parts in contact with the ground refer to the distance, angle, and other positional relationships between the parts that exert pressure on the ground. For example, when a person performs a plank exercise, the wrists, elbows, and feet apply pressure to the ground. The pressure pad is composed of multiple small pressure units arranged in a row. These small units can record the pressure values ​​exerted on the ground by the wrists, elbows, and feet. The distribution of pressure values ​​from multiple pressure units not only reflects the magnitude of the pressure and the size of the contact area between each part and the ground, but also characterizes the positional information such as the distance and offset angle between the pressure areas of the wrists, elbows, and feet.

[0033] The overall structure of the feature fusion unit is as follows: Figure 2 As shown, the whole-body pressure characteristics F can be obtained from steps S1 and S2. p and video features F v The whole-body pressure characteristics F p As the pressure query vector Q p , to use video features F v As the video numerical vector V v and video key vector K v A stress video multimodal cross-fusion attention mechanism is employed to capture long-range dependencies in the sequence, thereby obtaining cross-attention fusion features that integrate whole-body stress features and video features. Where SoftMax is the normalized exponential function, d k For video key vector K v Dimension size, For video key vector K vThe transpose of the vector. Through the cross-attention mechanism, the relative relationship between the dynamic mechanism represented by pressure features and the texture information represented by video features can be effectively obtained. The pressure magnitude, distance between pressure units, and angle information represented in the pressure features are fused with the human body texture contour information represented in the video features, resulting in a two-dimensional representation of human motion features from the camera direction and the vertical direction of the ground. This better integrates the visual information represented by video features with the physical dynamic mechanism represented by pressure features. The cross-attention feature is fused using CrossAttention(Q...). p ,K v V v ) and initial pressure characteristics F p and video features F v By performing matrix addition, we obtain the attention fusion feature F containing residual information. ca =CrossAttention(Q) p ,K v V v )+F p +F v This allows the fused features to retain some of the original features, better representing the pressure magnitude and the distance and angle between pressure units in the pressure features, as well as the human motion texture and contour in the visual features. Attention fusion features F, which contain residual information, are then used. ca By feeding the data into a position-wise element-wise feedforward network, an attention fusion feature F′ containing residual information and local features is obtained after learning local features. ca , with the original pressure characteristic F p The final fused feature F is obtained after splicing. fusion The purpose of stitching the features together with the original pressure features is to fuse the features so that they contain more information about the pressure magnitude, distance, and angle between pressure units, thus achieving an implicit representation of the physical dynamics mechanism. Pressure features contain more information related to contact and physical dynamics, so the stitched feature can represent more of the relative relationship between the human body and the ground, while maintaining the full-body local pose of the video features. It can also represent the relative positional relationship between the parts of the human body in contact with the ground, supplementing some video features missing due to occlusion.

[0034] Step S5: Fuse the feature F fusion The model is fed into the SMPL (A Skinned Multi-Person Linear Model) motion regressor to obtain the final human posture θ and displacement parameters T.

[0035] During parameter training of the video feature encoder, pressure feature encoder, feature fusion unit, and action regressor in step S1, training sample data is input into the network model, and the training loss function L = L is calculated.pose +L 3d +L 2d +L trans +L contact The network model was trained iteratively, and the AdamW optimizer was used to optimize the network parameters. The learning rate was set to 1e-4, and the total number of iterations was 1500. The network model parameters were determined based on the training loss value.

[0036] Among them, the human pose loss function θ represents the predicted human pose parameters. This represents the baseline true value of human posture parameters.

[0037] Human body 3D joint loss function J(θ,T) represents the coordinates of the three-dimensional joints of the SMPL human three-dimensional parametric model under the control of the human posture parameter θ and the human displacement parameter T. True reference values ​​for human posture parameters and the true reference value of human body displacement parameters The coordinates of the three-dimensional joints of the SMPL human body three-dimensional parametric model under control.

[0038] Human body two-dimensional joint loss function This method calculates the error after orthogonally projecting the 3D joints of a 3D human body template onto the Z-axis direction using the formula O(·). The formula for orthogonal projection is (x,y)=O(x,y,z). The difference between this 2D human body joint loss function and commonly used 2D human body joint loss functions lies in the following: Commonly used 2D human body joint loss functions use a weak perspective projection model to project the 3D human body joints into 2D form, while this invention uses an orthogonal projection camera model. Unlike traditional weak perspective projection camera models, the orthogonal projection camera model does not depend on the distance between the human body and the camera in the Z-axis direction; it only constrains the 2D joints on the XY plane of the human body. This decouples the human body posture parameters from the human body displacement parameters, simplifying and improving the accuracy of the 2D joint loss function calculation. It avoids complex calculations in the depth direction, reduces mutual interference between the estimation of human body posture parameters and human body displacement parameters, and thus achieves more accurate estimation of human body posture parameters and human body displacement parameters.

[0039] The human body displacement parameter loss function is T represents the predicted human body displacement parameter. This serves as the baseline true value for human body displacement parameters.

[0040] The whole-body contact loss function is J c (θ,T) represents the coordinates of the three-dimensional joints of the SMPL human three-dimensional parametric model in contact with the ground, controlled by the human posture parameter θ and the human displacement parameter T. True reference values ​​for human posture parameters and the true reference value of human body displacement parameters The coordinates of the three-dimensional joints of the SMPL human body three-dimensional parametric model in contact with the ground under control.

[0041] Among them, the three-dimensional joint J in contact with the ground c (θ,T) refers to the projection point obtained by projecting the three-dimensional joint point onto the ground. The sum of pressure values ​​within a certain neighborhood The height of the Z-axis of the three-dimensional joint is greater than the threshold τ1. Joints less than the threshold τ2. The formula is expressed as: Typically, the neighborhood is 25 square centimeters, τ1 is 5 centimeters, and τ2 is 5 centimeters. J(θ,T) represents the coordinates of the three-dimensional joints of the SMPL human 3D parametric model under the control of the human posture parameter θ and the human displacement parameter T.

[0042] Compared to existing technologies, firstly, this invention uses pressure information to determine whether a 3D joint is in contact with the ground. A 3D joint is considered to be in contact with the ground only if it has a pressure value on its projection onto the ground and its height is close to the ground height. Other methods often only determine the height in the Z-axis direction. This greatly improves the accuracy of determining whether a body part is in contact with the ground and indirectly improves the accuracy of human posture estimation methods. Secondly, the whole-body contact loss function of this invention can not only handle the case of foot contact with the ground, but also realize the contact determination of joints that may be in contact with the ground through two-dimensional feature determination. This expands the method's posture estimation range for a wider range of movements and can handle human posture estimation for different types of ground contact movements such as standing, handstand, plank, sitting, and kneeling.

[0043] Example:

[0044] An RGB camera is placed directly in front of the human body to capture a sequence of video frames from the frontal view of the human body's movements. The bounding box of the human body in the entire video is extracted using the object detection method YOLO. The human body is then extracted from the entire image containing the human body and the environment. The extracted human body video frame sequence is input into the video feature coding network HRNet to obtain the monocular video features of the human body's movements.

[0045] A pressure mat is placed on a surface suitable for human movement. Various movements (including standing, lying down, sitting, planking, and other common daily actions) are performed on the pressure mat, with no body part outside the mat in contact with the ground. In this embodiment, the pressure mat is approximately two meters long and one and a half meters wide or larger. A pressure mat data recording program is activated to collect a full-body pressure frame sequence of human movement movements. This sequence is then directly input into a ResNet pressure feature encoding network to obtain the full-body pressure features of the human movement.

[0046] To characterize temporal information in the attention mechanism, the temporally aligned whole-body pressure feature F is first... p and video features F v Position encoding is performed, and the general position encoding based on sine and cosine functions is combined with the whole-body pressure feature F. p and video features F v The summation ensures that both the full-body pressure features and video features possess location information. The dimensions of both the pressure features and video features are set to 2048. The pressure query vector is Q. p The whole body pressure characteristics F p Q is obtained by performing a linear transformation. p =F p W Q W Q The query vector parameter matrix is ​​a trainable vector, and the video numerical vector V is the video numerical vector. v and video key vector K v For video features F v V is obtained by performing a linear transformation. v =F v W V ,K v =F v W K W V W is a trainable numerical vector parameter matrix. K is the trainable key vector parameter matrix.

[0047] A multimodal cross-fusion attention mechanism for stress-induced video is adopted, and the cross-attention fusion features are calculated using a multi-head attention mechanism. Where SoftMax is the normalized exponential function, d k For video key vector K v The dimension size is typically 2048. For video key vector K v The transpose of the vector has 4 attention heads. Attention fusion features F with residual information are obtained by fusing initial information using commonly used residual structures. ca The attention fusion feature F with residual information caThe position is fed into an element-wise feedforward network to obtain the attention fusion feature F, which contains residual information and local features. c ′ a The position-by-element feedforward network consists of two fully connected layers, a random deactivation layer, and a layer normalization structure. The input dimension is 2048, the hidden layer dimension is 2048, and the random deactivation probability is set to 0.1. Finally, it is compared with the original whole-body pressure feature F. p The final fused feature F is obtained after splicing. fusion .

[0048] F fusion feature F fusion The parameters are fed into the SMPL model parameter regression network, which consists of multiple fully connected layers, to obtain the final human pose and displacement parameters. These estimated human pose and displacement parameters can then drive the SMPL 3D parametric human model.

[0049] To demonstrate the advantages of this invention over other human pose estimation methods, pose estimation was performed and compared using common methods on the same batch of data. Method 1 uses a convolutional neural network and gated recurrent units for human pose estimation; Method 2 uses a convolutional neural network and a multilayer perceptron; Method 3 uses a visual converter to construct a general model for human pose estimation; Method 4 uses a convolutional neural network to estimate image optical flow information for human pose estimation; and Method 5 uses a convolutional neural network, a recurrent neural network, and a multilayer perceptron for human pose estimation. This invention measures the accuracy of human pose estimation for each model using average joint position error, per-model vertex position error, and acceleration error. The average joint position error and per-model vertex position error measures the root mean square error between the estimated human joints and skin vertices and the true values. Smaller errors in these two metrics indicate a smaller difference between the estimated and true values, meaning a more accurate human pose estimation. The acceleration error measures the temporal jitter of the joints; a smaller acceleration error indicates a smoother pose estimation. Table 1 compares this invention with other methods in terms of average joint position error, per-model vertex position error, and acceleration error.

[0050] Table 1 Comparison of human posture estimation indices between existing commonly used methods and the method of this invention

[0051]

[0052]

[0053] As shown in Table 1, compared with existing human posture estimation methods, the average position error per joint of this invention is reduced by 9.8-19.6 mm, the position error per model vertex is reduced by 10.0-24.3 mm, and the acceleration is reduced by 11.6-434.3 m / s².2 .

[0054] To demonstrate the ability of this invention to estimate human body displacement parameters relative to other methods, this embodiment compares the results using commonly used models on the same batch of data. This invention measures the human body displacement estimation performance of various models using global displacement error, global average joint position error, and global joint jitter error. Global displacement error refers to the error between the estimated global human body displacement and the true value; global average joint error refers to the root mean square error between the estimated and true values ​​of human body joints; both of these metrics measure the effectiveness of global human body displacement estimation, with smaller values ​​indicating more accurate estimation. Global joint jitter refers to the jitter error between adjacent frames of global human body joints; smaller values ​​indicate higher smoothness of human body displacement. Table 2 compares this invention with other methods in terms of global displacement error, global average joint position error, and global joint jitter error (Methods 1, 2, and 3 do not have global displacement estimation capabilities):

[0055] Table 2 Comparison of human displacement estimation indices between existing commonly used methods and the method of this invention

[0056] method Global displacement error Global average position error per joint Global joint jitter error Method 4 1193(1151.4↓) 141.2(80.4↓) 686(626↓) Method 5 1023(981.4↓) 75.6(14.8↓) 92(32↓) This invention 41.6 60.8 60

[0057] As shown in Table 2, compared with existing human body displacement estimation methods, the global displacement error of this invention is reduced by 981.4-1151.4 mm, the global average position error per joint is reduced by 14.8-80.4 mm, and the global joint jitter error is reduced by 32-626 m / s. 2 .

[0058] To demonstrate the effectiveness of this invention in human pose estimation, this embodiment performs a visual comparison with other methods on the same data, such as... Figure 3 As shown. In the plank exercise, this invention estimates that all four limbs are in contact with the ground, while Method 3 estimates that the feet are not in contact with the ground. Method 2 has an estimation error for the left foot, and Method 3 has some degree of estimation error for all four limbs. In the stepping exercise, this invention estimates that both feet are in contact with the ground, while Methods 1, 2, and 3 all have some degree of estimation error for foot contact, and Method 3 even has an estimation error of floating in the air.

[0059] To demonstrate the effectiveness of this invention in estimating human body displacement, this embodiment presents a visual comparison with other methods using the same data, such as... Figure 4 As shown. During the spinning motion, the human body displacement trajectory estimated by this invention has a very small error compared to the true value. Method 5 estimates the human body displacement with incorrect direction, while Method 4 fails to estimate the true human body displacement.

Claims

1. A human pose estimation device based on whole-body pressure and monocular video, characterized in that, The device includes: A video feature encoder is used to extract video features related to human actions from a monocular frontal view video frame sequence. A pressure feature encoder is used to extract whole-body pressure features related to human movement from a whole-body ground pressure frame sequence. The whole-body pressure features of human movement include the pressure exerted on the ground by the body parts in contact with the ground, as well as the relative positional relationship between the body parts in contact with the ground. The relative positional relationship between the body parts in contact with the ground refers to the distance or angular positional relationship between the body parts that exert pressure on the ground. A feature fusion unit is used to fuse the whole-body pressure features and the video features to obtain fused features; An action regressor is used to perform nonlinear regression calculations based on the fused features to obtain the posture and displacement parameters of the human parametric model.

2. The human pose estimation device based on whole-body pressure and monocular video according to claim 1, characterized in that, The whole-body pressure characteristics include the pressure exerted on the ground by the body parts in contact with the ground, as well as the relative positional relationship between the body parts in contact with the ground.

3. The human posture estimation device based on whole-body pressure and monocular video according to claim 1, characterized in that, The video feature encoder includes four convolutional layers. Each convolutional layer consists of a two-dimensional convolutional function, a batch normalization function, and a linear rectifier function. Each stage downsamples the feature map of the previous stage and fuses the downsampled feature map with the feature map of each previous stage to obtain the final video features fused from various scales.

4. The human pose estimation device based on whole-body pressure and monocular video according to claim 1, characterized in that, The pressure feature encoder includes an initial convolutional layer, a max pooling layer, a four-stage convolutional layer, an average pooling layer, and a fully connected layer. Each of the four-stage convolutional layers consists of two residual blocks, and each residual block consists of two convolutional layers. As the number of stages increases, the spatial resolution of the features decreases while the number of channels increases, thus achieving high-dimensional feature extraction.

5. The human pose estimation device based on whole-body pressure and monocular video according to claim 1, characterized in that, The feature fusion unit comprises a fully connected layer, a cross-attention mechanism, a normalized exponential function, and a position-by-element feedforward network. Specifically, the whole-body pressure feature is processed through a fully connected layer to obtain a pressure query vector, and the video features are processed through a fully connected layer to obtain a video key vector and a video numerical vector. A correlation matrix between the pressure query vector and the video key vector is calculated. This correlation matrix is ​​then processed through a normalized exponential function to obtain attention weights. These attention weights are then used to perform matrix multiplication on the video numerical vectors to obtain a cross-attention fusion feature. This cross-attention fusion feature is then residually connected to the whole-body pressure feature and the video feature to obtain an attention fusion feature with residual information. This attention fusion feature with residual information is input into a position-by-element feedforward network consisting of two fully connected layers, a random deactivation layer, and a layer normalization function to obtain an attention fusion feature with residual information and local features. Finally, the feature fusion unit concatenates the attention fusion feature with residual information and local features with the whole-body pressure feature to obtain the final fused feature.

6. The human pose estimation device based on whole-body pressure and monocular video according to claim 1, characterized in that, The motion regressor includes a fully connected layer and a randomly deactivated layer. The motion regressor performs nonlinear regression on the fused features to output human posture parameters and displacement parameters.

7. The method using the human pose estimation device based on whole-body pressure and monocular video as described in claim 1, characterized in that, The method includes the following steps: Step S1: Construct a network model for the human pose estimation device and train the network model; Step S2: Collect a monocular frontal view video frame sequence of human movements; input the video frame sequence into a video feature encoder to obtain monocular video features of human movements; Step S3: Collect the whole-body ground pressure frame sequence of human movement; input the whole-body ground pressure frame sequence into the pressure feature encoder to obtain the whole-body pressure features of human movement; The whole-body pressure characteristics of the human body movement include the pressure exerted on the ground by the body parts in contact with the ground, as well as the relative positional relationship between the body parts in contact with the ground; wherein the relative positional relationship between the body parts in contact with the ground refers to the distance or angular positional relationship between the body parts that exert pressure on the ground. Step S4: Input the whole-body pressure features of human movement and the monocular video features into the feature fusion unit to fuse the whole-body pressure features and video features, and output the fused features containing the whole-body pressure information of human movement and video information. Step S5: Input the fused features into the motion regressor to obtain the final human posture and displacement parameters.

8. The method according to claim 7, characterized in that, In step S1, the training loss function is: ,in, Let be the human pose loss function. The loss function for the three-dimensional joints of the human body. For the two-dimensional joint loss function of the human body, For human body displacement parameter loss function, This is the whole-body contact loss function.

9. The method according to claim 8, characterized in that, Human body two-dimensional joint loss function ,in, Human posture parameters and human body displacement parameters 3D joint coordinates of a controlled SMPL human 3D parametric model. True values ​​of human posture parameters and the true reference value of human body displacement parameters 3D joint coordinates of a controlled SMPL human 3D parametric model. This indicates orthographic projection.

10. The method according to claim 8, characterized in that, The whole-body contact loss function is , Human posture parameters and human body displacement parameters The coordinates of the three-dimensional joints of the SMPL human body three-dimensional parametric model in contact with the ground under control. True reference values ​​for human posture parameters and the true reference value of human body displacement parameters The coordinates of the three-dimensional joints of the SMPL human body three-dimensional parametric model in contact with the ground under control.

Citation Information

Patent Citations

  • Behavioral disorder detection method considering vision and plantar pressure multimode perception

    CN115601840A

  • Monocular video-based multi-stage human motion capture method and device, and medium

    CN116386141A