A Human Motion Capture Method Based on Mask-Aware Graph Convolution and Skeletal Priors
Through the combination of mask-perceptual graph convolution and bone prior, the problem of insufficient feature representation and lack of geometric consistency in monocular human motion capture technology is solved, and the robustness and accuracy in complex backgrounds are improved, and more accurate human motion capture is achieved.
Patent Information
- Application Number
- CN202510577887.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing monocular human motion capture technology has insufficient feature representation ability, lack of geometric consistency and poor robustness to interference, making it difficult to accurately capture three-dimensional human postures and shapes in complex backgrounds.
Using a method of combining mask-aware graph convolution network and skeleton prior, images and mask features are extracted through mask-aware graph convolution network, mask-aware adjacency matrix is constructed, and 3D joint position and body mesh are predicted by combining the skeleton and vertex information of the SMPL model.
It improves the accuracy and robustness of three-dimensional human motion capture, enhances the ability to express key features of the human body, ensures geometric consistency, and achieves more accurate human motion capture.
Smart Images

Figure CN120088331B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly relates to a human motion capture method based on mask-aware graph convolution and skeletal prior. Background Art
[0002] Human motion capture technology has broad application prospects in the fields of virtual reality (VR), augmented reality (AR), human-computer interaction, film and animation production, medical rehabilitation, etc. This technology aims to accurately acquire, record, and reconstruct the motion postures and shapes of the human body.
[0003] Traditional human motion capture methods usually rely on hardware devices such as optical marker points, inertial sensors, or multi-camera arrays. Although these methods can achieve high precision, they have the disadvantages of high cost, the need to wear special devices, strict environmental requirements, complex settings, etc., which limit their application in consumer-level scenarios and outdoor environments.
[0004] In recent years, the monocular vision human motion capture technology based on a single RGB image has attracted much attention due to its advantages of low cost, convenience, and the absence of special devices. Its goal is to directly estimate the three-dimensional joint positions and three-dimensional mesh models (such as the SMPL model) of the human body from a single image. However, monocular vision methods face inherent challenges: depth ambiguity (the same two-dimensional posture may correspond to multiple three-dimensional postures), self-occlusion and external occlusion, complex background interference, illumination changes, and the diversity of the human body itself.
[0005] Existing monocular methods can be roughly divided into two categories: optimization-based methods and regression-based methods. Optimization-based methods usually first detect two-dimensional joint points, and then use an optimization algorithm to fit the parameters of a statistical human model (such as SMPL) to match the two-dimensional observations. These methods are highly dependent on the accuracy of two-dimensional detection, and the optimization process may fall into a local optimum. Regression-based methods attempt to directly predict three-dimensional joint positions or SMPL model parameters from image features end-to-end. For example, methods such as HMR (HumanMesh Recovery) use an encoder to extract image features, and then use a regressor to predict SMPL parameters. Subsequent research has improved performance by introducing temporal information (for videos), improving the network structure (such as Transformer, GCN), or using stronger prior knowledge.
[0006] However, existing methods still have some limitations:
[0007] 1. Insufficient feature representation ability: Traditional CNN features may be difficult to effectively distinguish the foreground human body from the complex background, especially when there are interfering objects or the proportion of the human body area is small. Global features may ignore local details, resulting in inaccurate pose or shape estimation.
[0008] 2. Lack of geometric consistency: Directly regressing coordinates or parameters may result in generated poses that do not conform to human kinematic constraints, or problems such as local distortion and incoherence in the body mesh. How to effectively utilize the inherent structural priors of the human body remains a challenge.
[0009] 3. Poor robustness to interference: Factors such as background noise, occlusion, illumination changes, and scale changes of the human body in the image can all cause a significant decline in the performance of existing methods.
[0010] Some methods attempt to model the structural relationships between joint points using graph convolutional networks (GCNs), or enhance features using attention mechanisms, but often do not fully combine mask information to precisely focus on the human body region and suppress background interference. At the same time, there is still room for improvement in how to more effectively integrate skeletal prior information such as SMPL into the feature learning process to guide the recovery of geometric consistency.
[0011] Therefore, developing a new method that can effectively utilize image mask information to focus on the human body region, enhance feature representation ability, and combine skeletal structure priors to ensure geometric consistency, thereby improving the accuracy of monocular 3D human action capture, has important research significance and application value. Summary of the Invention
[0012] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a human action capture method and system based on mask-aware graph convolution and skeletal priors, aiming to improve the accuracy of recovering 3D human joint positions and body meshes from a single image.
[0013] To achieve the above purpose, the first aspect of the present invention provides a human action capture method, the method comprising:
[0014] a) Obtain a single RGB image and extract the human mask corresponding to the RGB image;
[0015] b) Through the feature encoding layer of the mask-aware graph convolution network, respectively extract features from the RGB image and the human mask to obtain image features and mask features, and respectively encode the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information;
[0016] c) Through the adjacency matrix construction layer of the mask-aware graph convolution network, construct a mask-aware adjacency matrix according to the image token sequence and the mask token sequence, and use the graph convolution network to update the image token sequence and the adjacency matrix to obtain enhanced image features;
[0017] d) Through the SMPL body fitting layer of the bone prior decoupling network, based on the skeleton and vertex information of the SMPL model as priors, combined with the enhanced image features, the three-dimensional joint position coordinates and body mesh vertex coordinates are predicted through the decoder structure.
[0018] As an alternative implementation manner of the first aspect of the present application, in step b), the steps of respectively encoding the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information include: using a convolution operation to reduce the channel dimension of the image features and the mask features to a preset dimension D; respectively reorganizing the downsampled image feature map and mask feature map into an image token sequence and a mask token sequence with a length of HW + 1, where represents the height of the image, represents the width of the image, and the token sequence contains a learned camera token; applying a linear projection operation to the image token sequence and the mask token sequence respectively; adding a position encoding P to the linearly projected image token sequence and mask token sequence to obtain an image token sequence and a mask token sequence containing spatial position information.
[0019] As an alternative implementation manner of the first aspect of the present application, after the steps of respectively encoding the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information in step b), the method further includes: parallelly inputting the image token sequence and the mask token sequence added with the position encoding P into a Transformer encoder; using the self-attention mechanism of the Transformer encoder to respectively model the spatial relationships inside the image token sequence and the mask token sequence, and outputting the encoded image token sequence and mask token sequence.
[0020] As an alternative implementation manner of the first aspect of the present application, in step c), the step of constructing a mask-aware adjacency matrix specifically includes: generating an initial mask-aware matrix by calculating the Hadamard product between the image token sequence and the mask token sequence; using a human parsing model to extract the original mask M as a screening condition to process the initial mask-aware matrix: filtering the connection relationships between tokens with a value of zero at the corresponding positions in the original mask M, and assigning weights controlled by a preset scaling factor β to the connection relationships in the background area to obtain a filtered matrix; performing weight normalization processing on the filtered matrix to generate a mask-aware adjacency matrix for the graph convolutional network.
[0021] As an alternative implementation of the first aspect of the present application, after step c), it further includes: calculating a mask constraint loss through the mask constraint loss layer of the mask-aware graph convolutional network; the calculation of the mask constraint loss includes: identifying token pairs with the same non-zero label value in the original mask M , and minimizing the Euclidean distance between the enhanced image features corresponding to the token pairs; the calculation formula of the mask constraint loss is: , where is the mask loss, represents the enhanced image feature of the and th tokens, represents the value of the original mask M at the position.
[0022] As an alternative implementation of the first aspect of the present application, step d) includes: initializing learnable joint features and vertex features, where the vertex features are obtained by downsampling and dimension elevation based on the vertices of the SMPL model; inputting the learnable joint features and vertex features into a decoder for feature modeling; the decoder includes a structure-aware multi-head self-attention layer and a cross-attention layer; the multi-head self-attention layer uses a predefined edge connection matrix based on the SMPL skeleton structure, and combines mask weighting to incorporate attention weights for calculation; the cross-attention layer uses the enhanced image features as key vectors and value vectors, and uses the joint features and vertex features as query vectors to perform cross-attention calculation.
[0023] As an alternative implementation of the first aspect of the present application, step d) further includes: regressing the joint features and vertex features output by the decoder to obtain initial three-dimensional joint coordinates and a rough vertex mesh; performing an upsampling operation on the rough vertex mesh to restore a fine SMPL parameterized body mesh.
[0024] As an alternative implementation of the first aspect of the present application, step d) further includes: calculating a total loss through the loss optimization layer of the bone prior decoupling network, where the total loss includes: a 2D joint projection loss for constraining the predicted two-dimensional projected joints and the true two-dimensional joints; a 3D geometric loss for constraining the geometric consistency between the predicted three-dimensional joint coordinates and the body mesh vertex coordinates and the true coordinates; and a mask constraint loss.
[0025] In a second aspect, an embodiment of the present application provides a human action capture system based on mask-aware graph convolution and bone prior, the system includes:
[0026] An image acquisition module, configured to acquire a single RGB image;
[0027] A mask extraction module for extracting a human mask corresponding to the RGB image;
[0028] The mask-aware graph convolution module includes:
[0029] A feature encoding unit for respectively extracting features from the RGB image and the human mask to obtain image features and mask features, and respectively encoding the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information; and
[0030] An adjacency matrix construction unit for constructing a mask-aware adjacency matrix according to the image token sequence and the mask token sequence, and using a graph convolution network to update the image token sequence and the adjacency matrix to obtain enhanced image features;
[0031] The bone prior decoupling module includes an SMPL body fitting unit for using the skeleton and vertex information of the SMPL model as a prior, combining the enhanced image features, and predicting the three-dimensional joint position coordinates and the body mesh vertex coordinates through a decoder structure.
[0032] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0033] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] (1) Improve accuracy and robustness: By introducing a mask-aware graph convolution network, using mask information to guide the graph convolution network to focus on the feature learning of the human body region, effectively suppressing the interference of complex backgrounds, and enhancing the expression ability of human key features. At the same time, mask-based pruning of the adjacency matrix reduces the impact of redundant information and improves the robustness of the model under interference such as occlusion and noise.
[0036] (2) Enhance geometric consistency: Combining the bone prior decoupling network (SMPL model), integrating the inherent skeleton structure and vertex topology information of the human body as prior conditions into the decoder. Through the structure-aware self-attention mechanism and the guidance of image features, the consistency between the predicted 3D joint points and vertex coordinates and the real human body geometry is optimized.
[0037] (3) More accurate motion capture: By comprehensively utilizing the visual features of images (enhanced by the mask-aware graph convolutional network) and the structural priors of the human body (integrated through the SMPL model), the model can not only estimate the accurate 3D joint positions but also recover the body mesh poses that better conform to the kinematic and anatomical constraints of the human body, thereby achieving more precise human motion capture. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 The flowchart of a human motion capture method based on mask-aware graph convolution and skeletal priors provided by the first embodiment of the present invention;
[0039] Figure 2 The overall model process diagram of a human motion capture method based on mask-aware graph convolution and skeletal priors provided by the first embodiment of the present invention;
[0040] Figure 3 The mask-aware graph convolution learning process diagram in the first embodiment of the present invention;
[0041] Figure 4 The three-dimensional pose and body recovery process diagram in the first embodiment of the present invention;
[0042] Figure 5 The schematic diagram of the human joint connection matrix of the SMPL model in the first embodiment of the present invention;
[0043] Figure 6 The structural schematic diagram of a human motion capture system based on mask-aware graph convolution and skeletal priors provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0045] Terms such as "first" and "second" in the description and claims of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the description and claims means at least one of the connected objects. The character " / " generally represents an "or" relationship between the associated objects before and after. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0046] Example 1
[0047] Please refer to Figure 1 , which is a flowchart of a human motion capture method based on mask-aware graph convolution and skeletal prior proposed in the first embodiment of the present invention. The method includes the following steps. Please refer to Figure 2 , which is an overall model process diagram of a human motion capture method based on mask-aware graph convolution and skeletal prior provided in the first embodiment of the present invention. The overall model architecture mainly includes a mask-aware graph convolution learning process (refer to steps S2, S3, and S4) and a skeletal prior decoupling process (refer to step S5).
[0048] S1: Obtain a single RGB image and extract the human mask corresponding to the RGB image.
[0049] It should be noted that, for the convenience of understanding the overall framework of this method, this method is described. Formally, for a given picture , its goal is to recover the three-dimensional joint position coordinates of this picture and the body mesh vertex coordinates , where K represents the number of three-dimensional human joint points, and V = 6890, representing the estimated human mesh vertices. In this framework, this method first uses an existing human parsing model to extract the mask information of the input single image and uses it as guiding information for subsequent feature modeling.
[0050] S2: Respectively perform feature extraction on the RGB image and the human mask to obtain image features and mask features;
[0051] S3: Respectively encode the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information;
[0052] S4: According to the image token sequence and the mask token sequence, construct a mask-aware adjacency matrix, and use a graph convolutional network to update the image token sequence and the adjacency matrix to obtain enhanced image features.
[0053] It should be noted that the input RGB image and the parsed mask information respectively extract their static features through a CNN-based backbone network. In addition, to maintain spatial consistency, this method adds position encoding to the features so that each Token has the ability to perceive position information. Subsequently, the image features and the mask features obtained through the encoder are jointly put into the mask-aware graph convolution learning network to establish the association information between token nodes.
[0054] In this embodiment, step S3 uses a convolution operation to reduce the channel dimension of the image features and the mask features to a preset dimension D; the downsampled image feature map and mask feature map are respectively reorganized into an image token sequence and a mask token sequence with a length of HW + 1, where represents the height of the image, represents the width of the image, and the token sequence contains a learned camera token label; linear projection operations are respectively applied to the image token sequence and the mask token sequence; position encoding P is added to the linearly projected image token sequence and mask token sequence to obtain an image token sequence and a mask token sequence containing spatial position information.
[0055] In this embodiment, after step S3, it further includes parallelly inputting the image token sequence and the mask token sequence added with position encoding P into a Transformer encoder; using the self-attention mechanism of the Transformer encoder to respectively perform spatial relationship modeling on the inside of the image token sequence and the mask token sequence, and outputting the encoded image token sequence and mask token sequence.
[0056] In this embodiment, the step of constructing the mask-aware adjacency matrix in step S4 specifically includes: generating an initial mask-aware matrix by calculating the Hadamard product between the encoded image feature sequence and the encoded mask feature sequence; using the original mask M extracted by the human parsing model as a filtering condition to process the initial mask-aware matrix: filtering the connection relationships between tokens with a value of zero at the corresponding positions in the original mask M, and assigning weights controlled by a preset scaling factor β to the connection relationships in the background area to obtain a filtered matrix; performing weight normalization processing on the filtered matrix to generate a mask-aware adjacency matrix for the graph convolutional network.
[0057] In this embodiment, after step S4, it further includes calculating a mask constraint loss; the calculation of the mask constraint loss includes: identifying token pairs with the same non-zero label value in the original mask M , and minimizing the Euclidean distance between the enhanced image features corresponding to the token pairs.
[0058] Specifically, the mask-aware graph convolutional learning network consists of the following three parts: a feature encoding layer, an adjacency matrix construction layer, and a mask constraint loss layer, and its purpose is to perform feature encoding learning of images, and the specific content is as follows:
[0059] 1. Feature encoding layer
[0060] From the given RGB image I, this method extracts information M related to human body parts by using an existing human parsing model. First, the input RGB image I and M are processed by a pre-trained convolutional neural network (CNN) to extract the static feature representation of the RGB image itself and the feature representation of its corresponding mask respectively: . Subsequently, this method uses convolution to reduce the dimension of the extracted image features in the channel dimension, reducing its feature dimension to D = 512. After completing the feature dimension reduction, these feature maps are reorganized into a series of token blocks of length HW + 1, which contains a learned camera token marker for regressing the representation form of the 3D pose. In the feature encoding stage, this method further introduces positional encoding P to retain the spatial position information between tokens, explicitly guiding the model to learn the spatial relationship between tokens. The feature encoding process is shown in Equation (1):
[0061] (1)
[0062] where represents the linear projection operation, and the image features , represent the positional encodings for RGB features and mask features respectively. After completing the positional encoding, the flattened token sequence is parallelly input into the Transformer encoder for further spatial modeling within the tokens.
[0063] 2. Adjacency Matrix Construction Layer
[0064] To effectively establish the complex relationship between RGB features and mask features, a common and intuitive method is based on feature fusion, that is, simply concatenating the two features and then restoring the original channel dimension through downsampling operations. However, this direct concatenation method may destroy the spatial structure of the original features, resulting in the model's difficulty in accurately capturing the dependencies between tokens in subsequent processing and affecting the ability to model human poses. Therefore, this method proposes a mask-aware graph convolutional learning network (MGCL), as Figure 3 shown. This method constructs a graph convolutional network (GCN) based on the image features to enhance the ability to model the relationship between each token. In this method, the mask feature The construction conditions for the adjacency matrix are as follows: each token is regarded as a node in the GCN, and the adjacency matrix is used to represent the association relationship between tokens. In this method, the mask perception matrix of each token in the RGB image and each token in the mask image is calculated through the Hadamard product (where ). Its calculation method is shown in formula (2):
[0065] (2)
[0066] where represents the Hadamard product operation, .
[0067] Since the matrix contains some tokens that are irrelevant to the human body main region, these irrelevant tokens may introduce redundant information, thus affecting the effectiveness of the model in the feature learning process. Therefore, this method performs a pruning process on the connection relationship of the matrix . Specifically, this method uses the initial mask as the screening condition and only retains the tokens in the regions where the mask value is not zero and the connection relationship between. For the associations between tokens with a mask value of zero, they will be filtered to reduce the interference of irrelevant backgrounds. At the same time, the matrix is used as the adjacency matrix of the GCN. This process can be shown by formula (3):
[0068] (3)
[0069] where is set to 0.1, which is used to control the ratio of the weight of the background region to the weight of the target region. represents the value of the mask M at the position. Subsequently, for the filtered matrix, this method performs weight normalization to ensure a more reasonable overall distribution of the weight matrix. This weight normalization process is shown in formula (4):
[0070] (4)
[0071] The main purpose of the above filtering method is to effectively reduce the interference of irrelevant background regions by focusing the calculation on the feature modeling of the target region. This strategy can ensure that the model focuses its attention on the key information related to the human body region rather than being affected by irrelevant background information.
[0072] 3. Mask Constraint Loss Layer
[0073] After obtaining the adjacency matrix the matrix will be used to represent the connection relationships between nodes in the GCN. In the GCN operation, a residual connection method is adopted to transmit features, thereby ensuring stable gradient propagation. Finally, after being processed by the GCN, the output enhanced image features will fuse information from the graph structure and the original input features, so as to effectively capture the dependencies and context information between tokens. This process is shown in Equation (5):
[0074] (5)
[0075] where represents the graph convolutional network, represents the learnable weight matrix of the -th layer in the graph convolutional network. To improve the local consistency of the model during the token feature learning process, this method imposes a masked region constraint on the output features, aiming to control the similarity within the feature region, thereby enhancing the model's learning ability for local region features. Specifically, this method uses the original mask as a constraint condition, and by minimizing the Euclidean distance between similar tokens, the model can better focus on and strengthen the features in the similar regions during the learning process. In this way, the model can maintain strong local consistency during the feature learning process, thereby more effectively capturing the region information related to the human body. The masked constraint loss is calculated as shown in Equation (6):
[0076] (6)
[0077] where represents the Euclidean distance between the -th and the -th tokens of
[0078] S5: Based on the skeleton and vertex information of the SMPL model as priors, combined with the enhanced image features, the three-dimensional joint position coordinates and body mesh vertex coordinates are predicted through the decoder structure.
[0079] In this embodiment, the learnable joint features and vertex features are initialized. The vertex features are obtained by downsampling and dimension elevation based on the vertices of the SMPL model. The learnable joint features and vertex features are input into the decoder for feature modeling. The decoder includes a structure-aware multi-head self-attention layer and a cross-attention layer. The multi-head self-attention layer uses a predefined edge connection matrix based on the SMPL skeleton structure and incorporates attention weight calculation by combining with masking weighting. The cross-attention layer uses the enhanced image features as key vectors and value vectors, and uses the joint features and vertex features as query vectors to perform cross-attention calculation.
[0080] Further, the joint features and vertex features output by the decoder are regressed to obtain the initial three-dimensional joint coordinates and the rough vertex mesh. An upsampling operation is performed on the rough vertex mesh to restore the fine SMPL parameterized body mesh.
[0081] In this embodiment, after step S5, calculating the total loss is further included. The total loss includes: a 2D joint projection loss for constraining the predicted two-dimensional projected joints and the ground-truth two-dimensional joints; a 3D geometric loss for constraining the geometric consistency between the predicted three-dimensional joint coordinates and the body mesh vertex coordinates and the ground-truth coordinates; and a masking constraint loss.
[0082] Specifically, the bone prior decoupling network consists of the following two parts: an SMPL body fitting layer and a loss optimization layer, and its purpose is to recover the pose information and body mesh coordinates of the three-dimensional human body from the image, and the content is as follows:
[0083] 1. SMPL body fitting layer
[0084] In the previous HMR method, the image features regress the camera parameters, pose, and body parameters of the human body through subsequent multiple linear layers, and then these parameters are applied to the SMPL model for human body recovery. This method effectively realizes the reconstruction of human geometry by extracting key information from the image and mapping it to the parameter space of the human body model. However, in recent years, some research works have embedded the parameterized features of SMPL into the neural network for end-to-end joint optimization. This method significantly improves the model's ability to model human geometry by directly optimizing the SMPL parameters in the network. Inspired by this, this study decides to use SMPL as the skeleton model to perform feature learning, and the parameter regression process is as Figure 4 shown. Specifically, this study uses the skeleton structure of SMPL as a prior condition for feature learning, and uses this as a guide to help the model better learn the geometric shape of the human body.
[0085] Set the learnable joint features and a set of learnable vertex features , where and respectively represent the number of joint points and vertices, denotes the feature dimension. These two features are jointly input into the decoder for modeling through a concatenation operation. To reduce the computational complexity and improve the adaptability of the model, first, this study downsamples the original number of vertices (6890 vertices) of the SMPL parametric human model to , and uses a linear layer to increase its feature dimension to 512. Then, a residual connection is designed as the bias guidance for vertex features to ensure that information can be more stable and effective during transmission.
[0086] In the design of the decoder, this study introduces a structure-aware multi-head self-attention mechanism (MHSA) and combines it with a predefined edge connection matrix for feature modeling, and its structure is as Figure 5 shown. To further improve the expression ability of the model, the edge connection matrix is incorporated into the calculation of attention weights through masked weighting. This process is shown in Equation (7):
[0087] (7)
[0088] where, d represents the channel dimension, represents the minimum threshold parameter between non-adjacent nodes. are respectively transformed by the corresponding projection matrices and represent the query vector, key vector, and value vector respectively. To further improve the feature fusion ability of the model, this study follows the common decoder architecture method and introduces a cross-attention layer on this basis. In this mechanism, the image features are used as the key vector K and value vector V, while the concatenated joint point features are used as the query vector Q. Through the cross-attention mechanism, the image features can effectively interact with the joint point features to guide and enhance the feature expression of human joint points.
[0089] 2. Loss Optimization Layer
[0090] Through the feature modeling of multiple decoder layers, the model regresses the output joint point features and vertex features, and then obtains the predicted 3D joint point coordinates . On this basis, an upsampling operation is gradually performed on the corresponding rough vertex mesh to restore a more refined SMPL parameter .
[0091] To align the input image with the 3D reconstructed body mesh, this method follows the existing method for 2D joint projection loss, and uses the camera parameters to project the estimated 3D joints to 2D to obtain joint points for supervision: , where , is a projection function. Geometric consistency calculation between the 3D coordinates for prediction and the ground-truth coordinates. It includes 3D joint points and 3D body mesh calculation: . The loss function is shown in Equation (8):
[0092] (8)
[0093] In summary, combined with the masked constraint loss , the total loss is shown in Equation (9):
[0094] (9)
[0095] where represent the weight values of 2D coordinates, 3D coordinates, and masked constraints in training, respectively.
[0096] The training process of the method proposed in this study is as follows:
[0097] Input: Training set , human parsing model , CNN network , graph convolutional network , encoder , MGCL network , SMPL parametric model , decoder .
[0098] 1. Use the human parsing model to extract the human part parsing graph of the training set: ;
[0099] 2. for in do;
[0100] 3. Respectively extract the features of image I and mask M through the CNN encoder: ;
[0101] 4. Perform feature encoding on through the Transformer encoder ;
[0102] 5. Construct the mask-aware matrix for through Equations (2, 3, 4);
[0103] 6. Use the graph convolutional network for feature update: ;
[0104] 7. Calculate the mask constraint loss through formula (6). : ;
[0105] 8. Downsample the SMPL parametric model and input it into the decoder for joint feature learning;
[0106] 9. The features of the image guide the learning of joint point features to regress the three-dimensional human pose information;
[0107] 10. Update the model parameters through the loss calculation of formula (9);
[0108] 11. End.
[0109] In summary, this embodiment proposes a human motion capture method based on mask-aware graph convolution and skeletal prior, which mainly includes two core processes: the mask-aware graph convolution learning (MGCL) process and the skeletal prior decoupling process. Specifically, MGCL extracts the static features of the RGB image and the mask, uses convolution to reduce the dimension to 512 dimensions and then flattens it into a token sequence, combines positional encoding to retain spatial information, and optimizes the feature representation through the Transformer encoder. Subsequently, a GCN adjacency matrix is constructed using the mask-aware matrix, and a mask constraint loss is imposed to enhance local consistency. In the skeletal prior decoupling process, the SMPL skeletal prior model is introduced, and its skeleton and vertex information are used as prior knowledge for feature learning. The decoder structure combines the multi-head self-attention and cross-attention mechanisms, uses a predefined adjacency matrix, and fuses the attention weights through mask weighting. Finally, under the guidance of the image features, the cross-attention mechanism is used to perform multi-modal data augmentation on the SMPL skeletal node features, thereby optimizing the geometric consistency of the three-dimensional coordinate prediction.
[0110] Embodiment 2
[0111] Please refer to Figure 6 , which shows the structural schematic diagram of a human motion capture system based on mask-aware graph convolution and skeletal prior proposed in the second embodiment of this application. The system includes the following key modules:
[0112] An image acquisition module 100, configured to acquire a single RGB image;
[0113] A mask extraction module 200, configured to extract the human mask corresponding to the RGB image;
[0114] A mask-aware graph convolution module 300, including:
[0115] A feature encoding unit 301 is configured to perform feature extraction on the RGB image and the human mask respectively to obtain an image feature and a mask feature, and encode the image feature and the mask feature respectively to obtain an image token sequence and a mask token sequence containing spatial position information; and
[0116] An adjacency matrix construction unit 302 is configured to construct a mask-aware adjacency matrix according to the image token sequence and the mask token sequence, and update the image token sequence by using a graph convolutional network to obtain an enhanced image feature;
[0117] A bone prior decoupling module 400 includes an SMPL body fitting unit 401, which is configured to use the skeleton and vertex information of the SMPL model as a prior, and combine the enhanced image feature to predict the three-dimensional joint position coordinates and body mesh vertex coordinates through a decoder structure.
[0118] A human action capture system based on mask-aware graph convolution and bone prior in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc., which are not specifically limited in the embodiments of the present application.
[0119] A human action capture system based on mask-aware graph convolution and bone prior in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0120] A human action capture system based on mask-aware graph convolution and bone prior provided in an embodiment of the present application can implement Figure 1 each process implemented by a human action capture method based on mask-aware graph convolution and bone prior in the method embodiment. To avoid repetition, it will not be elaborated here.
[0121] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above embodiment of the human action capture method based on mask-aware graph convolution and skeleton prior, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0122] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above embodiment of the human action capture method based on mask-aware graph convolution and skeleton prior, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0123] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0124] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be executed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0126] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A human motion capture method based on mask-aware graph convolution and skeletal prior, characterized in that The method includes: a) Obtain a single RGB image and extract the human mask corresponding to the RGB image; b) Through the feature encoding layer of the mask-aware graph convolutional network, extract features from the RGB image and the human mask respectively to obtain image features and mask features, and encode the image features and the mask features respectively to obtain an image token sequence and a mask token sequence containing spatial position information, specifically including: using a convolutional operation to reduce the channel dimension of the image features and the mask features to a preset dimension D; reorganizing the downsampled image feature map and mask feature map into an image token sequence and a mask token sequence with a length of HW + 1 respectively, where represents the height of the image, represents the width of the image, and the token sequence contains a learned camera token; applying a linear projection operation to the image token sequence and the mask token sequence respectively; adding a position encoding P to the linearly projected image token sequence and mask token sequence to obtain an image token sequence and a mask token sequence containing spatial position information; c) Through the adjacency matrix construction layer of the mask-aware graph convolutional network, construct a mask-aware adjacency matrix according to the image token sequence and the mask token sequence, and use the graph convolutional network to update the image token sequence and the adjacency matrix to obtain enhanced image features; d) Through the SMPL body fitting layer of the bone prior decoupling network, based on the skeleton and vertex information of the SMPL model as priors, combined with the enhanced image features, predict the three-dimensional joint position coordinates and body mesh vertex coordinates through the decoder structure, specifically including: initializing the learnable joint point features and vertex features, where the vertex features are obtained by downsampling and dimension elevation based on the vertices of the SMPL model; inputting the learnable joint point features and vertex features into the decoder for feature modeling; the decoder includes a structure-aware multi-head self-attention layer and a cross-attention layer; the multi-head self-attention layer uses a predefined edge connection matrix based on the SMPL skeleton structure and incorporates the attention weights by combining mask weighting; the cross-attention layer uses the enhanced image features as the key vector and value vector, and uses the joint point features and vertex features as the query vector to perform cross-attention calculation.
2. The method according to claim 1, characterized in that After the step of encoding the image features and the mask features respectively in step b) to obtain the image token sequence and the mask token sequence containing spatial position information, it further includes: Parallelly input the image token sequence and the mask token sequence after adding the position encoding P into the Transformer encoder; Use the self-attention mechanism of the Transformer encoder to respectively perform spatial relationship modeling inside the image token sequence and the mask token sequence, and output the encoded image token sequence and mask token sequence.
3. The method according to claim 1 or 2, characterized in that, In step c), the step of constructing the mask-aware adjacency matrix specifically includes: Generate an initial mask-aware matrix by calculating the Hadamard product between the image token sequence and the mask token sequence; Use the human parsing model to extract the original mask M as a screening condition to process the initial mask-aware matrix: filter the connection relationships between tokens with zero values at the corresponding positions in the original mask M, and assign weights controlled by a preset scaling factor β to the connection relationships in the background area to obtain a filtered matrix; Perform weight normalization processing on the filtered matrix to generate a mask-aware adjacency matrix for the graph convolutional network.
4. The method according to claim 3, wherein After step c), it further includes: Calculate the mask constraint loss through the mask constraint loss layer of the mask-aware graph convolutional network; The calculation of the masked constraint loss includes: identifying token pairs with the same non-zero label value in the original mask M , and minimizing the Euclidean distance between the enhanced image features corresponding to the token pairs; The calculation formula of the mask constraint loss is: Among them is the mask loss, represents the Euclidean distance between the th and th tokens of the enhanced image features, represents the value of the original mask M at the position.
5. The method according to claim 1, wherein Step d) further includes: Regress the joint point features and vertex features output by the decoder to obtain the initial three-dimensional joint point coordinates and the rough vertex grid; Perform an upsampling operation on the rough vertex grid to restore the fine SMPL parameterized body mesh.
6. The method according to claim 1 or 4 or 5, characterized in that, Step d) further includes: Calculating a total loss through a loss optimization layer of a bone prior decoupling network, where the total loss includes: A 2D joint projection loss for constraining the predicted two-dimensional projected joint points and the ground-truth two-dimensional joint points; A 3D geometric loss for constraining the geometric consistency between the predicted three-dimensional joint point coordinates and body mesh vertex coordinates and the ground-truth coordinates; And a mask constraint loss.
7. A human motion capture system based on mask-aware graph convolution and skeletal prior, characterized in that, The system includes: An image acquisition module for acquiring a single RGB image; A mask extraction module for extracting a human mask corresponding to the RGB image; A mask-aware graph convolution module, including: A feature encoding unit, configured to respectively extract features from the RGB image and the human mask to obtain image features and mask features, and respectively encode the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information, specifically including: using a convolution operation to reduce the channel dimension of the image features and the mask features to a preset dimension D; respectively reorganizing the downsampled image feature map and mask feature map into an image token sequence and a mask token sequence of length HW+1, where represents the height of the image, represents the width of the image, the token sequence contains a learned camera token; applying a linear projection operation to the image token sequence and the mask token sequence respectively; adding a position encoding P to the linearly projected image token sequence and mask token sequence to obtain an image token sequence and a mask token sequence containing spatial position information; and An adjacency matrix construction unit for constructing a mask-aware adjacency matrix according to the image token sequence and the mask token sequence, and updating the image token sequence and the adjacency matrix by using a graph convolution network to obtain enhanced image features; A bone prior decoupling module, including an SMPL body fitting unit for predicting three-dimensional joint position coordinates and body mesh vertex coordinates through a decoder structure based on the skeleton and vertex information of the SMPL model as a prior and in combination with the enhanced image features, specifically including: initializing learnable joint point features and vertex features, where the vertex features are obtained by downsampling and dimension elevation based on the vertices of the SMPL model; inputting the learnable joint point features and vertex features into the decoder for feature modeling; the decoder includes a structure-aware multi-head self-attention layer and a cross-attention layer; the multi-head self-attention layer uses a predefined edge connection matrix based on the SMPL skeleton structure and incorporates attention weights by mask weighting; the cross-attention layer uses the enhanced image features as key vectors and value vectors and uses the joint point features and vertex features as query vectors to perform cross-attention calculation.
8. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a human action capture method based on mask-aware graph convolution and bone prior as described in any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Three-dimensional human motion prediction method based on optimized graph convolutional network
CN116596959A
Underground dangerous behavior identification method based on multi-modal joint training network
CN118968614A