Calorific value map guidance-based time sequence smooth human motion capture method and system
By introducing heat value maps and double-branch timing attention blocks in human motion capture, combined with the graph attention mechanism, the problems of timing inconsistency and structural instability in the existing methods are solved, and three-dimensional human posture estimation with higher accuracy and timing smoothness are achieved.
Patent Information
- Application Number
- CN202510714371.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing video stream-based human motion capture methods have shortcomings in timing inconsistent, structural instability, and failure to fully utilize dynamic features and structural dependencies.
The timing smooth human motion capture method based on heat map guidance is adopted. By acquiring RGB video sequences and generating heat map sequences, the timing characteristics of the RGB video sequence and the heat map sequence are fused using a dual-branch timing attention block, and the structural dependence relationship between three-dimensional human joints is modeled through the intra-pose inference module using the graph attention mechanism.
It significantly improves the accuracy and timing smoothness of three-dimensional human posture estimation, can better cope with complex motion and perspective changes, and enhances structural consistency.
Smart Images

Figure CN120220254A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and human motion capture, and particularly relates to a method and system for temporal smoothing of human motion capture guided by a heat value map. Background Art
[0002] Three-dimensional human mesh recovery is one of the core technologies in the fields of computer vision, human-computer interaction, and intelligent monitoring. Accurate and stable three-dimensional human mesh recovery is of great significance for improving the naturalness of interactive systems, enhancing virtual reality experiences, and promoting medical rehabilitation and sports analysis.
[0003] With the development of deep learning, three-dimensional human mesh recovery methods based on single-frame images have achieved good performance. However, these methods mainly rely on single-frame features for prediction and fail to effectively utilize the temporal information of human motion, resulting in problems such as jitter and instability in the estimation results. In contrast, three-dimensional human mesh recovery methods based on video streams improve the stability of prediction to a certain extent by modeling temporal relationships. However, existing methods still have deficiencies in capturing long-term and short-term temporal dependencies and are less adaptable to complex motions and perspective changes. For example, some researchers use a pre-trained CNN network to extract static features for each frame and then train a temporal network to predict SMPL parameters. TCMR uses a Transformer to extract the temporal information of static features; MPS-Net proposes motion persistent attention to extract the temporal information of static features. However, these methods rely on static features extracted by a CNN and do not fully utilize the dynamic features of human motion, thus having certain limitations in capturing dynamic changes and details.
[0004] In addition, some researchers have made innovative attempts in SMPL parameter regression. For example, MAED pays attention to temporal information through a self-attention mechanism and proposes a KTD structure. This method regresses the pose information of child nodes from the parent node step by step based on the kinematic tree. However, this method does not fully consider the dynamic dependencies of joint points inside the kinematic structure, and all child nodes default to having the same weight coefficient as the parent node, which limits the model's ability to capture complex human postures.
[0005] In three-dimensional pose estimation based on video streams, significant achievements have been made in learning 2D pose information. However, how to use 2D pose information as a prior condition for the temporal three-dimensional human mesh recovery task and further improve the accuracy and temporal smoothness of pose estimation remains a challenge.
[0006] Therefore, how to improve the temporal smoothness and enhance the structural consistency while improving the accuracy of human pose estimation remains a key challenge in current research. Summary of the Invention
[0007] The present invention aims to solve the problems of inconsistent timing, unstable structure, and failure to fully utilize dynamic features and structural dependencies in existing human motion capture methods based on video streams, and provides a timing-smoothing human motion capture method and system guided by a heat map.
[0008] In a first aspect, the present invention provides a timing-smoothing human motion capture method guided by a heat map, including the following steps: S1: Obtain an RGB video sequence containing multiple frames of images; S2: Extract 2D human joint position information from each frame of the RGB video sequence, and generate a corresponding heat map sequence based on the 2D human joint position information; S3: Respectively perform feature extraction on the RGB video sequence and the heat map sequence to obtain the timing features of the RGB video sequence and the timing features of the heat map sequence; S4: Use a dual-branch timing attention block to fuse the timing features of the RGB video sequence and the timing features of the heat map sequence to obtain fused timing features, and the dual-branch timing attention block includes a multi-head self-attention mechanism and a multi-head cross-attention mechanism; S5: Regress SMPL model parameters according to the fused timing features, and the SMPL model parameters include pose parameters, body shape parameters, and camera parameters. Among them, the step of regressing the pose parameters includes using an intra-frame pose inference module to model the structural dependencies between the three-dimensional human joints corresponding to the pose parameters through a graph attention mechanism, and optimizing the pose parameters based on the structural dependencies; S6: Output three-dimensional human motion capture data according to the regressed SMPL model parameters.
[0009] As an optional implementation manner of the first aspect of the present application, the step S2 specifically includes: obtaining 2D human joint position information in the RGB video sequence through a 2D pose extractor ; based on the 2D human joint position information, use Gaussian distribution spots to generate a heat map sequence, and the value of the heat map sequence at a position is calculated according to the following formula: , where represents the value of the heat map sequence at the position , and represent the positions of each pixel on the grid, is the standard deviation controlling the diffusion range of the Gaussian distribution.
[0010] As an alternative implementation of the first aspect of the present application, in step S4, obtaining the fused temporal features using the dual-branch temporal attention block specifically includes: linearly projecting the temporal features of the RGB video sequence and the temporal features of the heat value map sequence into the same low-dimensional feature space respectively; adding learnable temporal position encodings to the temporal features of the mapped RGB video sequence and the temporal features of the heat value map sequence respectively; using the multi-head self-attention mechanism to calculate the temporal attention matrix of the temporal features of the mapped RGB video sequence and the temporal attention matrix of the temporal features of the mapped heat value map sequence to capture the temporal information within each modality; using the multi-head cross-attention mechanism to establish a potential association between the temporal features of the mapped RGB video sequence and the temporal features of the mapped heat value map sequence to enhance cross-modal information interaction; using a linear layer and an adaptive weight factor to perform weighted fusion on the temporal features of the fused RGB video sequence and the heat value map sequence to obtain the fused temporal features; wherein, the adaptive weight factor is calculated according to the following formula: , where the adaptive weight factor , represents the length of the time series, represents a learnable linear transformation matrix, represents the temporal features of the RGB video sequence and the temporal features of the heat value map sequence the concatenation operation between modalities; the fused temporal features are calculated according to the following formula: .
[0011] As an alternative implementation of the first aspect of the present application, in step S5, using the graph attention mechanism of the intra-frame pose inference module to model the structural dependencies between the three-dimensional human joints corresponding to the pose parameters specifically includes: regarding the three-dimensional human joints corresponding to the pose parameters as nodes in the graph; using the graph attention mechanism to learn a trainable weight matrix to transform the features of the nodes; using the graph attention mechanism to calculate the association strength between the joint nodes to obtain a graph attention matrix, and the association strength between node and node is calculated according to the following formula: , where, represents the transpose of the shared attention vector, represents the pose node i and k the concatenation operation; pruning the unconnected nodes according to the graph attention matrix.
[0012] As an alternative implementation of the first aspect of the present application, in step S5, when using the graph attention mechanism of the intra-frame pose inference module to model the structural dependency between the three-dimensional human joints corresponding to the pose parameters, it further includes: for each joint node , perform softmax function normalization on the set of corresponding parent nodes in the preset kinematic tree structure. The sum of the attention weights of the set of parent nodes is 1, and the normalized attention weights are calculated according to the following formula: ; Based on the topological relationship of the kinematic tree structure and the normalized attention weights, calculate the local pose parameters of the joint node . The local pose parameters are derived by calculating the weight distribution coefficient of the joint node relative to its set of parent nodes, and performing a linear combination of the global pose parameters of the parent nodes and the corresponding edge weights . The calculation method is as shown in the following formula: .
[0013] As an alternative implementation of the first aspect of the present application, in step S5, based on the structural dependency, optimizing the pose parameters specifically includes: using a multi-layer perceptron to combine the fused temporal features and the calculated local pose parameters to model the pose parameters of the final SMPL model.
[0014] As an alternative implementation of the first aspect of the present application, the loss function adopted during the training process of the method includes: a 2D joint position loss calculated based on the difference between the predicted 2D joint positions and the true 2D joint positions; a 3D joint position loss calculated based on the difference between the predicted 3D joint positions and the true 3D joint positions; an SMPL parameter loss calculated based on the difference between the predicted SMPL pose parameters and body shape parameters and the true SMPL parameters; a temporal smoothness loss imposed on the global fused features output by the dual-branch temporal attention block to penalize the drastic change of features between adjacent frames.
[0015] In a second aspect, an embodiment of the present application provides a temporal-smoothing human motion capture system guided by heat maps, including: An input module configured to obtain an RGB video sequence including multiple frames of images; A 2D pose extraction and heat map generation module configured to extract 2D human joint position information from each frame image in the RGB video sequence and generate a corresponding sequence of heat maps based on the 2D human joint position information; A feature extraction module, configured to perform feature extraction on the RGB video sequence and the heat value map sequence respectively, to obtain the temporal features of the RGB video sequence and the temporal features of the heat value map sequence; A dual-branch temporal attention module, configured to use dual-branch temporal attention blocks to fuse the temporal features of the RGB video sequence and the temporal features of the heat value map sequence, so as to obtain fused temporal features, and the dual-branch temporal attention blocks include a multi-head self-attention mechanism and a multi-head cross-attention mechanism; An SMPL regression module, configured to regress SMPL model parameters according to the fused temporal features, and the SMPL model parameters include pose parameters, body shape parameters, and camera parameters. Among them, the step of regressing the pose parameters includes using an intra-frame pose inference module to model the structural dependencies between the three-dimensional human joints corresponding to the pose parameters through a graph attention mechanism, and optimizing the pose parameters based on the structural dependencies; An output module, configured to output three-dimensional human motion capture data according to the regressed SMPL model parameters.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0017] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0018] Compared with the prior art, the beneficial effects of the present invention include but are not limited to: 1. By introducing 2D pose information and converting it into a heat value map as prior knowledge, the accuracy of joint position estimation is enhanced, providing a reliable basis for subsequent three-dimensional pose estimation.
[0019] 2. The proposed dual-branch temporal attention block can effectively fuse RGB features and heat value map features, capture the temporal information of both modalities at the same time, and improve the temporal consistency and feature expression ability through adaptive fusion.
[0020] 3. The proposed intra-frame pose inference module uses a graph attention mechanism to dynamically model the structural relationships of joint points within a single frame, overcoming the limitations of traditional methods that rely on a fixed kinematic tree structure, and enhancing the structural coherence and accuracy of pose prediction.
[0021] 4. Through the combination of the above technologies, the accuracy and temporal smoothness of three-dimensional human pose estimation are significantly improved, and it can better handle complex movements and perspective changes. Description of the Drawings
[0022] Figure 1 It is a flowchart of a proposed temporal smoothing human motion capture method guided by a heat value map; Figure 2 It is a schematic diagram of the overall technical framework of the proposed temporal smoothing human motion capture method guided by a heat value map; Figure 3 It is a structural display diagram of the proposed dual-branch temporal attention block (DBTA); Figure 4 It is a schematic diagram of iterative SMPL model regression based on a motion tree structure; Figure 5 It is a detailed display diagram of the proposed in-frame pose reasoning (IPR) module; Figure 6 It is a schematic diagram of the structure of a temporal smoothing human motion capture system guided by a heat value map provided by an embodiment of the present invention. Detailed Embodiments
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.
[0024] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.
[0025] Embodiment 1 Please refer to Figure 1 , which is an implementation flowchart of a temporal smoothing human motion capture method guided by a heat value map provided by an embodiment of the present invention. Figure 2 It is a schematic diagram of the overall technical framework of the proposed temporal smoothing human motion capture method guided by a heat value map. The method includes the following steps: S1: Obtain an RGB video sequence containing multiple frames of images.
[0026] S2: Extract the 2D human joint position information for each frame image in the RGB video sequence, and generate the corresponding heat map sequence based on the 2D human joint position information.
[0027] First, a RGB video containing multiple frames is given as input. For each frame in the video sequence, this method extracts its 2D human joint position information. Specifically, this method adopts an existing top-down method to obtain the 2D human joint position information, such as using a pre-trained 2D pose extractor to obtain the position information of human joints, and this information is represented as , where represents the number of human joints, depending on the selected 2D pose extractor.
[0028] Next, Gaussian distribution spots are constructed based on the extracted 2D human joint position information and propagated throughout the image through spatial diffusion to generate the corresponding heat map . For the given 2D pose joint position , the generation process of the heat map can be represented by formula (1): where represents the value of the heat map sequence at position , and represent the positions of each pixel on the grid, is the standard deviation that controls the diffusion range of the Gaussian distribution.
[0029] In the above way, this method converts the key feature points extracted from the 2D pose information into heat map image information, which is used as prior knowledge for subsequent modal feature fusion.
[0030] S3: Extract features from the RGB video sequence and the heat map sequence respectively to obtain the temporal features of the RGB video sequence and the temporal features of the heat map sequence.
[0031] In this embodiment, the temporal features of the RGB feature map and the heat map feature map are extracted through a convolutional neural network (CNN) and a pose network (Pose Net) respectively, and are represented as: , where , represents the number of channels of the feature map, represents the length of the time series.
[0032] Specifically, the generated heat map is then input into a pose network (Pose Net), which is a lightweight feature extraction network composed of convolutional layers and pooling layers, aiming to learn the key pose information in the heat map, so as to obtain the temporal features of the heat map sequence, represented as Meanwhile, the RGB video sequence extracts its temporal features through a Convolutional Neural Network (CNN), denoted as .
[0033] S4: Use a Dual-Branch Temporal Attention Block (DBTA) to fuse the temporal features of the RGB video sequence and the temporal features of the heat map sequence to obtain fused temporal features. The DBTA includes a Multi-Head Self-Attention (MHSA) mechanism and a Multi-Head Cross-Attention (MCA) mechanism.
[0034] In this embodiment, a Dual-Branch Temporal Attention Block (DBTA) is used to obtain fused temporal features, specifically including: linearly projecting the temporal features of the RGB video sequence and the temporal features of the heat map sequence into the same low-dimensional feature space respectively; adding learnable temporal position encodings to the temporal features of the projected RGB video sequence and the temporal features of the heat map sequence respectively; using the Multi-Head Self-Attention (MHSA) mechanism to calculate the temporal attention matrix of the temporal features of the projected RGB video sequence and the temporal attention matrix of the temporal features of the projected heat map sequence to capture the temporal information within each modality; using the Multi-Head Cross-Attention (MCA) mechanism to establish potential associations between the temporal features of the projected RGB video sequence and the temporal features of the projected heat map sequence to enhance cross-modal information interaction; using a linear layer and an adaptive weight factor to perform weighted fusion on the temporal features of the fused RGB video sequence and the temporal features of the heat map sequence to obtain fused temporal features.
[0035] Specifically,[[]] Figure 3 shows the model architecture of the proposed Dual-Branch Temporal Attention Block (DBTA), aiming to learn the temporal consistency between RGB images and heat maps. First, using the method of linear projection, the temporal features of the input RGB video sequence and the temporal features of the heat map sequence are mapped into the same low-dimensional feature space, aiming to ensure the dimensional alignment of the two features. Subsequently, the learnable temporal position encoding is added to and respectively, where represents the length of the time series, and is the dimension of the feature. The purpose of this step is to introduce temporal position information so that the model can capture the feature dynamics that change over time. The core of the Dual-Branch Temporal Attention Block is to establish an image relationship model across time frames. The DBTA adopts a Multi-Head Self-Attention (MHSA) mechanism to calculate the temporal attention matrix of to capture the temporal information between different time frames within the same batch.
[0036] Although and can capture certain temporal information within their respective feature spaces, they still lack the ability of direct cross-modal interaction and cannot fully integrate information from different modalities. Therefore, in order to further enhance the information interaction between different modal features, this method introduces a cross-attention mechanism (Multi-head Cross Attention, MCA) to establish and the potential correlation between the two modalities. Its mechanism is similar to MHSA. By using matrices to calculate the vectors of their respective modalities respectively, the correlation is established between different modal features, thus effectively enhancing the global feature representation ability. As shown in formula (2): where , L represents the depth of the DBTA block. Since and belong to different feature spaces respectively, direct fusion may lead to information loss. Therefore, this method uses a linear layer to fuse and through an adaptive weight factor to control the contribution degree of different features in the final representation. The calculation method of the weight is as shown in formula (3): where the adaptive weight factor , represents the length of the time series, represents the learnable linear transformation matrix, represents the temporal features of the RGB video sequence and the temporal features of the heat map sequence the concatenation operation between modalities. Finally, this method uses the calculated weights and to perform weighted summation on the two features to obtain a global feature representation with better temporal consistency and modal fusion ability , and its calculation method is as shown in formula (4): S5: Regress the SMPL model parameters based on the fused temporal features. The SMPL model parameters include pose parameters, shape parameters, and camera parameters. Among them, the steps of regressing the pose parameters include using the Intra-frame Pose Reasoning (IPR) module to model the structural dependencies between the three-dimensional human joints corresponding to the pose parameters through a graph attention mechanism, and optimizing the pose parameters based on the structural dependencies.
[0037] S6: Output the three-dimensional human motion capture data according to the regressed SMPL model parameters.
[0038] In this embodiment, using the Intra-frame Pose Reasoning (IPR) module graph attention mechanism to model the structural dependencies between the three-dimensional human joints corresponding to the pose parameters specifically includes: regarding the three-dimensional human joints corresponding to the pose parameters as nodes in the graph; using the graph attention mechanism to learn a trainable weight matrix to transform the features of the nodes; calculating the association strength between the joint nodes through the graph attention mechanism to obtain the graph attention matrix.
[0039] Further, using the Intra-frame Pose Reasoning (IPR) module graph attention mechanism to model the structural dependencies between the three-dimensional human joints corresponding to the pose parameters also includes: for each joint node , performing softmax function normalization on the corresponding set of parent nodes in the preset kinematic tree structure , and the sum of the attention weights of the set of parent nodes is 1; calculating the local pose parameters of the joint node based on the topological relationship of the kinematic tree structure and the normalized attention weights. The local pose parameters are derived by calculating the weight distribution coefficients of the joint node relative to its set of parent nodes, and performing a linear combination of the global pose parameters of the parent nodes and the corresponding edge weights .
[0040] In this embodiment, optimizing the pose parameters based on the structural dependencies specifically includes: using a multi-layer perceptron to model the final pose parameters of the SMPL model by combining the fused temporal features and the calculated local pose parameters.
[0041] In three-dimensional human motion capture, a regressor is used to obtain the SMPL parameters for recovering the pose and shape information of the human body from the image data. Here, represents the pose parameters of the SMPL joints, represents the shape parameters, represents the camera parameters for projecting 3D coordinates into the 2D space. Among them, the pose parameters It directly determines the rotation relationship of the skeletal joints. Therefore, its recovery accuracy greatly affects the overall accuracy of 3D human mesh recovery. Existing methods, such as HMR, usually adopt image features extracted by deep neural networks to directly regress the relative rotation parameters of joint points to predict the SMPL pose parameters of the human body, as shown in Equation (5): where represents the regression function, is represented by the 6D rotation of 24 joint points (including 23 joint points and a root node). However, although this end-to-end deep learning-based regression method can estimate the 3D information of the human body, it often ignores the mutual correlation between joint points when modeling the human body structure, especially the hierarchical dependence relationship of human motion. To address this issue, He et al. proposed a step-by-step regression method based on the kinematic tree. This method hierarchically models joint nodes according to the hierarchical structure of human motion to enhance the structural rationality of pose prediction. As Figure 4 shown, this method performs step-by-step regression with a specific parent node according to the human kinematic structure. For example, for the given joint point 7, regression is performed with joint 0 as its parent node, and the set of its parent nodes is defined as L = {0, 1, 4}. Therefore, for the regression of a current node, the specific regression method is as shown in Equation (6): where represents the regression function of the pose node, F represents the image feature, represents the k th pose parameter of the node.
[0042] Although the above step-by-step regression method based on the kinematic tree can utilize the hierarchical dependence of human motion to a certain extent, thereby enhancing the structural rationality of pose prediction, this method still has certain limitations. Specifically, this method only relies on the fixed parent-child node relationship for regression and fails to fully consider the dynamic correlation between different joint points within the kinematic structure. In addition, during the regression process, all child nodes are defaulted to have the same weight coefficient as the parent node, ignoring the importance differences of different joints in different motion states, which may limit the model's ability to capture complex human postures.
[0043] To overcome the above deficiencies and enhance the structural relationship between 24 joint points within a single frame, this method adopts a graph attention mechanism (GAT) to design weights for node connections in order to capture the dynamic dependency relationships between joint points. This method introduces the graph attention mechanism (GAT), allowing the model to dynamically adjust the association strength between joints, thereby enhancing the modeling ability for local pose relationships. Specifically, this paper proposes an in-frame pose reasoning (IPR) method that uses GAT to model the joint poses of each frame and dynamically learn the weight distribution between joint points to more accurately capture the human pose structure, such as Figure 5 As shown, this method treats each joint point as a node in the graph and transforms the features of each node by learning a trainable weight matrix to obtain its latest representation. On this basis, the graph attention matrix is responsible for measuring the association strength between joint points, and its calculation method is shown in formula (7): where represents the linear activation function, represents the shared attention vector of nodes in the graph, represents the pose node i and k concatenation operation. Through this calculation method, the model is not only restricted by the topological structure of the kinematic tree but can also flexibly capture the dynamic constraint relationships in different motion states. Further, this paper filters the pose nodes and prunes the unconnected nodes. Subsequently, for each node i perform softmax function normalization on its parent node set , and the sum of the attention weights of the parent node set is 1. The specific calculation method is shown in formula (8): Subsequently, based on the topological relationship of the edge connection weight matrix, this method calculates the weight distribution coefficient of the node relative to its parent node set, and linearly combines the global pose parameters of the parent node with the corresponding edge weight to derive the local pose parameter of the current node , realizing hierarchical human pose modeling. This process can be formally expressed as shown in formula (9): where represents the parent node set of node i , represents the 6D rotation parameter of the parent node.
[0044] Finally, the pose regression parameter Combining local pose parameters through a multi - layer perceptron (MLP) with the feature representation for modeling, and its formal expression is shown in Equation (10): In this embodiment, the method further includes a model training process. The loss functions used in the model training process include: a 2D joint position loss calculated based on the difference between the predicted 2D joint positions and the true 2D joint positions; a 3D joint position loss calculated based on the difference between the predicted 3D joint positions and the true 3D joint positions; an SMPL parameter loss calculated based on the difference between the predicted SMPL pose parameters and body shape parameters and the true SMPL parameters; a temporal smoothness loss imposed on the global fusion features output by the dual - branch temporal attention block (DBTA) to penalize the drastic changes in features between adjacent frames.
[0045] Specifically, the overall training process of the HG - HMR proposed in this paper is shown in Algorithm 1. In this method, following the method of previous studies, for the position losses of 2D and 3D joints, an L2 loss function is used for calculation. Specifically, the position loss of 2D joints and the position loss of 3D joints are shown in Equation (11): where and represent the predicted 2D and 3D joint positions respectively, represents the corresponding true joint position. For the SMPL regression parameters, this method uses the SMPL loss function , which comprehensively considers the differences in body shape and pose parameters, as shown in Equation (12): where represents the predicted joint parameters, represents the predicted body shape parameters, represents the true joint parameters, represents the true body shape parameters.
[0046] To further improve the stability of the model, prevent the jitter of the prediction results, and ensure that the generated motion sequences have good coherence, this method introduces a smoothness loss on the temporal feature map output by the model at the end. This loss function is shown in Equation (13): where represents thei Temporal feature map of the frame.
[0047] The training process of the method proposed in this study is as follows: Algorithm 1 Overall training process of the proposed method (HG-HMR) Input: Training set , 2D pose detector , CNN network , Pose-Net network , DBTA network block , SMPL regressor , GAT network .
[0048] 1. For each training sample (RGB video sequence) in the training dataset do; 2. For each frame in the RGB video sequence, extract its 2D joint points through formula (1) and generate the corresponding heat map sequence: ; 3. Perform feature encoding: ; 4. As shown in formulas (3) and (4), obtain the temporal global features (fused temporal features) of the image through DBTA: ; 5. Initially regress the global SMPL parameters (including pose parameters, body shape parameters, camera parameters) from the fused temporal features through formula (5): ; 6. As shown in formula (7), perform intra-frame pose reasoning (IPR) on the pose parameters to model the dependency relationship between joint points: ; 7. Calculate the normalized parent node attention weights based on the kinematic tree structure and the modeled dependency relationship, as shown in formula (8); 8. Parse the local pose parameters according to the pose parameters and attention weights of the parent node, as shown in formula (9): ; 9. Combine the fused temporal features and the local pose parameters, regress the final SMPL pose parameters through MLP, and calculate the total loss (including 2D and 3D joint point losses, SMPL parameter losses, smoothness losses) in combination with the body shape parameters and camera parameters, as shown in formulas (11), (12), and (13), and update the model parameters; 10. End.
[0049] Experimental results show that the proposed method has achieved excellent performance on multiple datasets, significantly improving the accuracy and temporal smoothness of 3D human pose estimation.
[0050] In summary, in this embodiment, the method and system for temporally smoothed human motion capture based on heat map guidance provided by the present invention effectively solve the deficiencies of existing methods by organically combining 2D pose priors, cross-modal temporal feature fusion, and in-frame pose reasoning based on graph attention, providing an effective solution for high-precision and temporally stable video stream human motion capture.
[0051] Embodiment 2 Please refer to Figure 6 , which shows a schematic structural diagram of a system for temporally smoothed human motion capture based on heat map guidance proposed in the second embodiment of the present application. The system includes the following key modules: Input module 100, configured to obtain an RGB video sequence containing multiple frames of images; 2D pose extraction and heat map generation module 200, configured to extract 2D human joint position information for each frame of image in the RGB video sequence and generate a corresponding heat map sequence based on the 2D human joint position information; Feature extraction module 300, configured to extract features from the RGB video sequence and the heat map sequence respectively to obtain the temporal features of the RGB video sequence and the temporal features of the heat map sequence; Dual-branch temporal attention module 400, configured to fuse the temporal features of the RGB video sequence and the temporal features of the heat map sequence using a dual-branch temporal attention block (DBTA) to obtain fused temporal features. The DBTA includes a multi-head self-attention (MHSA) mechanism and a multi-head cross-attention (MCA) mechanism; SMPL regression module 500, configured to regress SMPL model parameters according to the fused temporal features. The SMPL model parameters include pose parameters, body shape parameters, and camera parameters. Among them, the step of regressing the pose parameters includes using an in-frame pose reasoning (IPR) module to model the structural dependencies between the three-dimensional human joints corresponding to the pose parameters through a graph attention mechanism and optimizing the pose parameters based on the structural dependencies; Output module 600, configured to output three-dimensional human motion capture data according to the regressed SMPL model parameters.
[0052] A temporal smoothing human motion capture system based on calorific value map guidance in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.
[0053] A temporal smoothing human motion capture system based on calorific value map guidance in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0054] A temporal smoothing human motion capture system based on calorific value map guidance provided in an embodiment of the present application can implement Figure 1 each process implemented in a method embodiment of a temporal smoothing human motion capture method based on calorific value map guidance. To avoid repetition, it will not be elaborated here.
[0055] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of a temporal smoothing human motion capture method based on calorific value map guidance and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0056] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of a temporal smoothing human motion capture method based on calorific value map guidance and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0057] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.
[0058] It should be noted that in this text, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0059] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0060] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. A timing-smoothing human motion capture method guided by a calorific value map, characterized in that, Including the following steps: S1: Obtain an RGB video sequence containing multiple frames of images; S2: Extract 2D human joint point position information for each frame image in the RGB video sequence, and generate a corresponding heat map sequence based on the 2D human joint point position information; S3: Respectively perform feature extraction on the RGB video sequence and the heat map sequence to obtain the temporal features of the RGB video sequence and the temporal features of the heat map sequence; S4: Use a dual-branch temporal attention block to fuse the temporal features of the RGB video sequence and the temporal features of the heat map sequence to obtain fused temporal features. The dual-branch temporal attention block includes a multi-head self-attention mechanism and a multi-head cross-attention mechanism; S5: Regress SMPL model parameters according to the fused temporal features. The SMPL model parameters include pose parameters, body shape parameters, and camera parameters. Among them, the step of regressing the pose parameters includes using an intra-frame pose inference module to model the structural dependencies between the three-dimensional human joints corresponding to the pose parameters through a graph attention mechanism, and optimizing the pose parameters based on the structural dependencies; S6: Output three-dimensional human motion capture data according to the regressed SMPL model parameters.
2. The method according to claim 1, wherein The step S2 specifically includes: Obtain the 2D human joint position information in the RGB video sequence through a 2D pose extractor ; Based on the 2D human joint point position information, use Gaussian distribution spots to generate a heat map sequence. The value of the heat map sequence at a position is calculated according to the following formula: , wherein represents the value of the calorific value map sequence at the position , and represent the positions of each pixel on the grid is the standard deviation that controls the diffusion range of the Gaussian distribution 3. The method according to claim 1, wherein In the step S4, using the dual-branch temporal attention block to obtain the fused temporal features specifically includes: Linearly project the temporal features of the RGB video sequence and the temporal features of the heat map sequence into the same low-dimensional feature space respectively; Add learnable temporal position encodings to the temporal features of the mapped RGB video sequence and the temporal features of the heat map sequence respectively; Use the multi-head self-attention mechanism to calculate the temporal attention matrix of the temporal features of the mapped RGB video sequence and the temporal attention matrix of the temporal features of the mapped heat map sequence to capture the temporal information within each modality; Use the multi-head cross-attention mechanism to establish a potential association between the temporal features of the mapped RGB video sequence and the temporal features of the mapped heat map sequence to enhance cross-modal information interaction; Use a linear layer and an adaptive weight factor to perform weighted fusion on the temporal features of the fused RGB video sequence and the temporal features of the heat map sequence to obtain the fused temporal features; Among them, the adaptive weight factor is calculated according to the following formula: , where the adaptive weight factor , represents the length of the time series, represents a learnable linear transformation matrix, represents the temporal features of the RGB video sequence and the temporal features of the heat map sequence concatenation operation between modalities; The fused temporal features are calculated according to the following formula: 。 4. The method according to claim 1, wherein In the step S5, using the graph attention mechanism of the intra-frame pose inference module to model the structural dependencies between the three-dimensional human joints corresponding to the pose parameters specifically includes: Regard the three-dimensional human joints corresponding to the pose parameters as nodes in the graph; Use the graph attention mechanism to learn a trainable weight matrix to transform the features of the nodes; Calculate the association strength between joint nodes through the graph attention mechanism to obtain a graph attention matrix. The association strength between node and node is calculated according to the following formula: , Among them, represents the transpose of the shared attention vector, represents the pose node i and k is the concatenation operation; Prune the unconnected nodes according to the graph attention matrix.
5. The method according to claim 4, wherein In step S5, using the graph attention mechanism of the intra-frame pose inference module to model the structural dependency relationship between the three-dimensional human joints corresponding to the pose parameters further includes: For each joint node , perform softmax function normalization on the corresponding set of parent nodes in the preset kinematic tree structure, the sum of the attention weights of the set of parent nodes is 1, and the normalized attention weights are calculated according to the following formula: ; Calculate the local pose parameters of the joint node based on the topological relationship of the kinematic tree structure and the normalized attention weights where the local pose parameters are derived by calculating the weight distribution coefficients of the joint node relative to its set of parent nodes, and linearly combining the global pose parameters of the parent nodes with the corresponding edge weights . The calculation method is shown in the following formula: 。 6. The method according to claim 5, wherein In step S5, optimizing the pose parameters based on the structural dependency relationship specifically includes: Using a multi-layer perceptron to combine the fused temporal features and the calculated local pose parameters to model the pose parameters of the final SMPL model.
7. The method according to claim 1, wherein The loss function adopted during the training process of the method includes: A 2D joint position loss calculated based on the difference between the predicted 2D joint position and the true 2D joint position; A 3D joint position loss calculated based on the difference between the predicted 3D joint position and the true 3D joint position; An SMPL parameter loss calculated based on the difference between the predicted SMPL pose parameters and body shape parameters and the true SMPL parameters; A temporal smoothness loss applied to the globally fused features output by the dual-branch temporal attention block, which is used to penalize the drastic change of features between adjacent frames.
8. A time-series smoothing human motion capture system guided by a calorific value map, characterized in that, Including: An input module configured to obtain an RGB video sequence containing multiple frames of images; A 2D pose extraction and heatmap generation module configured to extract 2D human joint position information from each frame image in the RGB video sequence and generate a corresponding heatmap sequence based on the 2D human joint position information; A feature extraction module configured to perform feature extraction on the RGB video sequence and the heatmap sequence respectively to obtain the temporal features of the RGB video sequence and the temporal features of the heatmap sequence; A dual-branch temporal attention module configured to use a dual-branch temporal attention block to fuse the temporal features of the RGB video sequence and the temporal features of the heatmap sequence to obtain fused temporal features, and the dual-branch temporal attention block includes a multi-head self-attention mechanism and a multi-head cross-attention mechanism; An SMPL regression module configured to regress SMPL model parameters according to the fused temporal features, and the SMPL model parameters include pose parameters, body shape parameters, and camera parameters. Among them, the steps of regressing the pose parameters include using an intra-frame pose inference module to model the structural dependency relationship between the three-dimensional human joints corresponding to the pose parameters through a graph attention mechanism and optimizing the pose parameters based on the structural dependency relationship; An output module configured to output three-dimensional human motion capture data according to the regressed SMPL model parameters.
9. An electronic device, characterized in that, Including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for temporally smooth human action capture based on heatmap guidance as described in any one of claims 1-7 are implemented.
10. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of a method for temporally smooth human action capture based on heatmap guidance as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Teacher classroom video action recognition method based on space-time double-branch feature fusion
CN118247849A
Capture method based on video stream attitude simulation
CN118629085A
Three-dimensional human body posture estimation method and system based on alternating space-time encoder
CN118692137A
Video tag determination method, device, terminal, and storage medium
WO2021143624A1
Cited By
Human body posture recognition method and system based on double-attention structured position coding
CN121600558A