Human motion capture method based on mask perception graph convolution and skeleton prior

By introducing mask-perceptual graph convolution and bone prior information into monocular visual human motion capture technology, the problems of insufficient feature representation ability, lack of geometric consistency and poor robustness in the prior art are solved, and a three-dimensional human motion capture effect with higher accuracy and robustness are achieved.

CN120088331AActive Publication Date: 2025-06-03NANCHANG UNIV

Patent Information

Application Number
CN202510577887.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-03
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing monocular visual human body motion capture technology has problems such as insufficient feature representation ability, lack of geometric consistency and poor robustness to interference, making it difficult to achieve high-precision three-dimensional human body motion capture in complex environments.

Method used

The human motion capture method based on mask-aware graph convolution and bone prior is adopted. Image and mask features are extracted through mask-aware graph convolution network, combined with the skeletal prior information of the SMPL model, a mask-aware adjacency matrix is ​​constructed, and the three-dimensional joint position and body mesh are predicted through the graph convolution network and decoder structure.

Benefits of technology

It improves the accuracy and robustness of three-dimensional human motion capture, enhances the ability to express key features of the human body, and ensures geometric consistency, and is suitable for complex backgrounds and environments with a lot of interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088331A_ABST
    Figure CN120088331A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and discloses a human motion capture method based on mask perceptual graph convolution and skeleton prior, and the method comprises the steps: obtaining an input image and a corresponding human mask; extracting features of the image and the mask and coding; constructing a mask perception graph convolutional network, constructing an adjacent matrix of the graph convolutional network by using mask information, and applying mask constraint loss to enhance the expression ability of the human body region features of the image; a skeleton prior decoupling network is constructed, skeleton and vertex information of an SMPL model is used as prior, a cross attention mechanism is combined, and multi-modal data enhancement of SMPL skeleton node features is guided through image features; and finally outputting a three-dimensional joint position coordinate and a shape grid vertex coordinate. According to the method, the local feature consistency is enhanced through mask perception image convolution, the geometric consistency is improved in combination with skeleton prior, and the precision of three-dimensional human body motion capture under monocular vision is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly relates to a human action capture method based on masked perception graph convolution and skeletal prior knowledge. Background Art

[0002] Human action capture technology has broad application prospects in the fields of virtual reality (VR), augmented reality (AR), human-computer interaction, film animation production, medical rehabilitation, etc. This technology aims to accurately acquire, record, and reconstruct the motion postures and shapes of the human body.

[0003] Traditional human action capture methods usually rely on hardware devices such as optical marker points, inertial sensors, or multi-camera arrays. Although these methods can achieve high precision, they have disadvantages such as high cost, the need to wear special devices, strict environmental requirements, and complex settings, which limit their application in consumer-level scenarios and outdoor environments.

[0004] In recent years, monocular vision human action capture technology based on a single RGB image has attracted much attention due to its advantages of low cost, convenience, and the absence of special devices. Its goal is to directly estimate the three-dimensional joint positions and three-dimensional mesh models (such as the SMPL model) of the human body from a single image. However, monocular vision methods face inherent challenges: depth ambiguity (the same two-dimensional posture may correspond to multiple three-dimensional postures), self-occlusion and external occlusion, complex background interference, illumination changes, and the diversity of the human body itself.

[0005] Existing monocular methods can be roughly divided into two categories: optimization-based methods and regression-based methods. Optimization-based methods usually first detect two-dimensional joint points, and then use an optimization algorithm to fit the parameters of a statistical human model (such as SMPL) to match the two-dimensional observations. These methods are highly dependent on the accuracy of two-dimensional detection, and the optimization process may fall into a local optimum. Regression-based methods attempt to directly predict three-dimensional joint positions or SMPL model parameters from image features end-to-end. For example, methods such as HMR (HumanMesh Recovery) use an encoder to extract image features, and then use a regressor to predict SMPL parameters. Subsequent research has improved performance by introducing temporal information (for videos), improving network structures (such as Transformer, GCN), or using stronger prior knowledge.

[0006] However, existing methods still have some limitations: 1. Insufficient feature representation ability: Traditional CNN features may be difficult to effectively distinguish the foreground human body from the complex background, especially when there are interfering objects or the proportion of the human body area is small. Global features may ignore local details, resulting in inaccurate pose or shape estimation.

[0007] 2. Lack of geometric consistency: Directly regressing coordinates or parameters may result in generated poses that do not conform to human kinematic constraints, or local distortions and incoherences in the body mesh. How to effectively utilize the inherent structural priors of the human body remains a challenge.

[0008] 3. Poor robustness to interference: Factors such as background noise, occlusion, illumination changes, and scale variations of the human body in the image can all cause a significant decline in the performance of existing methods.

[0009] Some methods attempt to model the structural relationships between joint points using graph convolutional networks (GCNs), or enhance features using attention mechanisms, but often do not fully combine mask information to precisely focus on the human body region and suppress background interference. At the same time, there is still room for improvement in how to more effectively incorporate skeletal prior information such as SMPL into the feature learning process to guide the restoration of geometric consistency.

[0010] Therefore, developing a new method that can effectively utilize image mask information to focus on the human body region, enhance feature representation capabilities, and combine skeletal structure priors to ensure geometric consistency, thereby improving the accuracy of monocular 3D human motion capture, has important research significance and application value. Summary of the Invention

[0011] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a human motion capture method and system based on mask-aware graph convolution and skeletal priors, aiming to improve the accuracy of restoring 3D human joint positions and body meshes from a single image.

[0012] To achieve the above object, in the first aspect of the present invention, a human motion capture method is provided, the method comprising: a) Obtain a single RGB image and extract the human mask corresponding to the RGB image; b) Through the feature encoding layer of the mask-aware graph convolution network, respectively extract features from the RGB image and the human mask to obtain image features and mask features, and respectively encode the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information; c) Through the adjacency matrix construction layer of the mask-aware graph convolution network, construct a mask-aware adjacency matrix according to the image token sequence and the mask token sequence, and use the graph convolution network to update the image token sequence and the adjacency matrix to obtain enhanced image features; d) Through the SMPL body fitting layer of the skeletal prior decoupling network, based on the skeleton and vertex information of the SMPL model as priors, and in combination with the enhanced image features, predict the 3D joint position coordinates and body mesh vertex coordinates through a decoder structure.

[0013] As an alternative implementation of the first aspect of the present application, in step b), the steps of encoding the image feature and the mask feature respectively to obtain an image token sequence and a mask token sequence containing spatial position information include: using a convolution operation to reduce the channel dimension of the image feature and the mask feature to a preset dimension D; respectively reorganizing the downsampled image feature map and mask feature map into an image token sequence and a mask token sequence with a length of HW + 1, where represents the height of the image, represents the width of the image, and the token sequence contains a learned camera token; applying a linear projection operation to the image token sequence and the mask token sequence respectively; adding a position encoding P to the linearly projected image token sequence and mask token sequence to obtain an image token sequence and a mask token sequence containing spatial position information.

[0014] As an alternative implementation of the first aspect of the present application, after the steps of encoding the image feature and the mask feature respectively to obtain an image token sequence and a mask token sequence containing spatial position information in step b), the following steps are further included: parallelly inputting the image token sequence and the mask token sequence after adding the position encoding P into a Transformer encoder; using the self-attention mechanism of the Transformer encoder to respectively model the spatial relationships inside the image token sequence and the mask token sequence, and outputting the encoded image token sequence and mask token sequence.

[0015] As an alternative implementation of the first aspect of the present application, in step c), the step of constructing a mask-aware adjacency matrix specifically includes: generating an initial mask-aware matrix by calculating the Hadamard product between the image token sequence and the mask token sequence; using a human parsing model to extract the original mask M as a screening condition to process the initial mask-aware matrix: filtering the connection relationships between tokens with a value of zero at the corresponding positions in the original mask M, and assigning weights controlled by a preset scaling factor β to the connection relationships in the background area to obtain a filtered matrix; performing weight normalization processing on the filtered matrix to generate a mask-aware adjacency matrix for a graph convolutional network.

[0016] As an alternative implementation of the first aspect of the present application, after step c), the following steps are further included: calculating a mask constraint loss through a mask constraint loss layer of a mask-aware graph convolutional network; the calculation of the mask constraint loss includes: identifying token pairs with the same non-zero label value in the original mask M , and minimize the Euclidean distance between the enhanced image features corresponding to the said tokens; the calculation formula of the mask constraint loss is: , where is the mask loss, represents the enhanced image features of the and th tokens, represents the value of the original mask M at the position.

[0017] As an alternative implementation of the first aspect of the present application, step d) includes: initializing learnable joint features and vertex features, where the vertex features are obtained by downsampling and dimension elevation based on the vertices of the SMPL model; inputting the learnable joint features and vertex features into a decoder for feature modeling; the decoder includes a structure-aware multi-head self-attention layer and a cross-attention layer; the multi-head self-attention layer utilizes a predefined edge connection matrix based on the SMPL skeleton structure, and incorporates mask weighting into the calculation of attention weights; the cross-attention layer uses the enhanced image features as key vectors and value vectors, and uses the joint features and vertex features as query vectors to perform cross-attention calculation.

[0018] As an alternative implementation of the first aspect of the present application, step d) further includes: regressing the joint features and vertex features output by the decoder to obtain initial three-dimensional joint coordinates and a rough vertex mesh; performing an upsampling operation on the rough vertex mesh to restore a fine SMPL parameterized body mesh.

[0019] As an alternative implementation of the first aspect of the present application, step d) further includes: calculating the total loss through the loss optimization layer of the bone prior decoupling network, where the total loss includes: a 2D joint projection loss for constraining the predicted two-dimensional projected joints and the ground truth two-dimensional joints; a 3D geometric loss for constraining the geometric consistency between the predicted three-dimensional joint coordinates and the body mesh vertex coordinates and the ground truth coordinates; and a mask constraint loss.

[0020] In a second aspect, an embodiment of the present application provides a human motion capture system based on mask-aware graph convolution and bone prior, the system includes: An image acquisition module for acquiring a single RGB image; A mask extraction module for extracting the human mask corresponding to the RGB image; A mask-aware graph convolution module, including: A feature encoding unit, configured to respectively extract features from the RGB image and the human mask to obtain image features and mask features, and respectively encode the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information; and An adjacency matrix construction unit, configured to construct a mask-aware adjacency matrix according to the image token sequence and the mask token sequence, and update the image token sequence and the adjacency matrix by using a graph convolutional network to obtain enhanced image features; A skeletal prior decoupling module, including an SMPL body fitting unit, configured to use the skeleton and vertex information of the SMPL model as a prior, and combine the enhanced image features to predict three-dimensional joint position coordinates and body mesh vertex coordinates through a decoder structure.

[0021] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0022] In a fourth aspect, an embodiment of the present application provides a readable storage medium, where a program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0023] Compared with the prior art, the present invention has the following beneficial effects: (1) Improving accuracy and robustness: By introducing a mask-aware graph convolutional network, using mask information to guide the graph convolutional network to focus on feature learning in the human body region, effectively suppressing the interference of complex backgrounds, and enhancing the ability to express key human body features. At the same time, mask-based pruning of the adjacency matrix reduces the impact of redundant information and improves the robustness of the model under interferences such as occlusion and noise.

[0024] (2) Enhancing geometric consistency: Combining a skeletal prior decoupling network (SMPL model), integrating the inherent skeleton structure and vertex topology information of the human body as prior conditions into the decoder. Through a structure-aware self-attention mechanism and the guidance of image features, the consistency between the predicted 3D joint points and vertex coordinates and the true human body geometry is optimized.

[0025] (3) More accurate action capture: By comprehensively using the visual features of the image (enhanced by the mask-aware graph convolutional network) and the structural prior of the human body (integrated by the SMPL model), the model can not only estimate the accurate positions of 3D joint points, but also recover the body mesh poses that are more in line with human kinematic and anatomical constraints, thereby achieving more accurate human action capture. Description of the Drawings

[0026] Figure 1 Flow chart of a human action capture method based on mask-aware graph convolution and skeletal prior provided in the first embodiment of the present invention; Figure 2 Overall model process diagram of a human action capture method based on mask-aware graph convolution and skeletal prior provided in the first embodiment of the present invention; Figure 3 Mask-aware graph convolution learning process diagram in the first embodiment of the present invention; Figure 4 Three-dimensional pose and shape recovery process diagram in the first embodiment of the present invention; Figure 5 Schematic diagram of the human joint connection matrix of the SMPL model in the first embodiment of the present invention; Figure 6 Structural schematic diagram of a human action capture system based on mask-aware graph convolution and skeletal prior provided in the embodiments of the present invention. Specific implementation manners

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0028] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0029] Embodiment 1 Please refer to Figure 1 , which is a flow chart of a human action capture method based on mask-aware graph convolution and skeletal prior proposed in the first embodiment of the present invention. The method includes the following steps. Please refer to Figure 2, which is the overall model process diagram of a human motion capture method based on mask-aware graph convolution and skeletal prior provided by the first embodiment of the present invention. The overall model architecture mainly includes a mask-aware graph convolution learning process (refer to steps S2, S3, and S4) and a skeletal prior decoupling process (refer to step S5).

[0030] S1: Obtain a single RGB image and extract the human mask corresponding to the RGB image.

[0031] It should be noted that, for the convenience of understanding the overall framework of this method, this method is described. Formally, for a given picture , its aim is to recover the three-dimensional joint position coordinates of this picture and the body mesh vertex coordinates , where K represents the number of three-dimensional human joint points, and V = 6890, representing the estimated human mesh vertices. In this framework, this method first uses an existing human parsing model to extract the mask information of the input single image and uses it as guiding information for subsequent feature modeling.

[0032] S2: Respectively perform feature extraction on the RGB image and the human mask to obtain image features and mask features; S3: Respectively encode the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information; S4: According to the image token sequence and the mask token sequence, construct a mask-aware adjacency matrix, and use a graph convolutional network to update the image token sequence and the adjacency matrix to obtain enhanced image features.

[0033] It should be noted that the input RGB image and the parsed mask information respectively extract their static features through a backbone network based on CNN. In addition, in order to maintain spatial consistency, this method adds position encoding to the features so that each Token has the ability to perceive position information. Subsequently, the image features and the mask features are jointly put into the mask-aware graph convolution learning network to establish the association information between token nodes.

[0034] In this embodiment, step S3 uses a convolution operation to reduce the channel dimension of the image features and the mask features to a preset dimension D; the downsampled image feature map and mask feature map are respectively reorganized into an image token sequence and a mask token sequence with a length of HW + 1, where represents the height of the image, Denote the width of the image, and the token sequence contains a learned camera token; apply linear projection operations to the image token sequence and the mask token sequence respectively; add positional encoding P to the linearly projected image token sequence and mask token sequence to obtain an image token sequence and a mask token sequence containing spatial position information.

[0035] In this embodiment, after step S3, it further includes parallelly inputting the image token sequence and the mask token sequence after adding positional encoding P into the Transformer encoder; using the self-attention mechanism of the Transformer encoder to respectively model the spatial relationships inside the image token sequence and the mask token sequence, and outputting the encoded image token sequence and mask token sequence.

[0036] In this embodiment, the step of constructing the mask-aware adjacency matrix in step S4 specifically includes: generating an initial mask-aware matrix by calculating the Hadamard product between the encoded image feature sequence and the encoded mask feature sequence; using the original mask M extracted by the human parsing model as a screening condition to process the initial mask-aware matrix: filtering the connection relationships between tokens with zero corresponding position values in the original mask M, and assigning weights controlled by a preset scaling factor β to the connection relationships in the background area to obtain a filtered matrix; performing weight normalization processing on the filtered matrix to generate a mask-aware adjacency matrix for the graph convolutional network.

[0037] In this embodiment, after step S4, it further includes calculating the mask constraint loss; the calculation of the mask constraint loss includes: identifying token pairs with the same non-zero label value in the original mask M , and minimizing the Euclidean distance between the corresponding enhanced image features of the token pairs.

[0038] Specifically, the mask-aware graph convolutional learning network consists of the following three parts: a feature encoding layer, an adjacency matrix construction layer, and a mask constraint loss layer, and its purpose is to perform feature encoding learning of the image, and the specific content is as follows: 1. Feature encoding layer From the given RGB image I, this method extracts information M related to human body parts by using an existing human parsing model. First, the input RGB image I and M are processed by a pre-trained convolutional neural network (CNN) to respectively extract the static feature representation of the RGB image itself and the feature representation of its corresponding mask: . Subsequently, this method adopts Convolution reduces the dimensionality of the extracted image features in the channel dimension, reducing the feature dimension to D = 512. After completing the feature dimensionality reduction, these feature maps are reorganized into a series of token blocks of length HW + 1, which contain a learned camera token marker for representing the regression of three-dimensional poses. In the feature encoding stage, this method further introduces positional encoding P to preserve the spatial position information between tokens and explicitly guide the model to learn the spatial relationship between tokens. The feature encoding process is shown in Equation (1): (1) where represents the linear projection operation, and the image feature , represent the positional encodings for RGB features and mask features respectively. After completing the positional encoding, the flattened token sequence is input into the Transformer encoder in parallel for further spatial modeling within the tokens.

[0039] 2. Adjacency Matrix Construction Layer To effectively establish the complex relationship between RGB features and mask features, a common and intuitive method is based on feature fusion, that is, simply concatenating the two features and then restoring the original channel dimension through downsampling operations. However, this direct concatenation method may disrupt the spatial structure of the original features, making it difficult for the model to accurately capture the dependencies between tokens in subsequent processing and affecting the ability to model human poses. Therefore, this method proposes a mask-aware graph convolutional learning network (MGCL), as Figure 3 shown. This method constructs a graph convolutional network (GCN) based on the image feature to enhance the ability to model the relationship between each token. In this method, the mask feature is used as the condition for constructing the adjacency matrix, that is: each token is regarded as a node in the GCN, and the adjacency matrix is used to characterize the association relationship between tokens. This method calculates the mask-aware matrix of each token in the RGB image and each token in the mask image through the Hadamard product (where ). Its calculation method is shown in Equation (2): (2) where represents the Hadamard product operation, .

[0040] Since the matrix contains some tokens that are not related to the main body region of the human body. These irrelevant tokens may introduce redundant information, thus affecting the effectiveness of the model in the feature learning process. Therefore, this method performs pruning on the connection relationships of the matrix Specifically, this method uses the initial mask as a filtering condition and only retains the tokens in the regions where the mask values are not zero and between the connections. For the associations between tokens with zero mask values, they will be filtered to reduce the interference of irrelevant backgrounds. At the same time, the matrix is used as the adjacency matrix of the GCN. This process can be shown by formula (3): (3) where is set to 0.1, which is used to control the ratio of the weight of the background region to the weight of the target region. represents the value of the mask M at the position. Subsequently, for the filtered matrix, this method performs weight normalization to ensure a more reasonable overall distribution of the weight matrix. This weight normalization process is shown by formula (4): (4) The main purpose of the above filtering method is to effectively reduce the interference of irrelevant background regions by focusing the calculation on the feature modeling of the target region. This strategy can ensure that the model focuses on the key information related to the human body region rather than being affected by irrelevant background information.

[0041] 3. Mask Constraint Loss Layer After obtaining the adjacency matrix , this matrix will be used for the connection relationships between nodes in the GCN. In the GCN operation, the residual connection method is adopted to transfer features to ensure stable gradient propagation. Finally, after being processed by the GCN, the output enhanced image features will fuse the information from the graph structure and the original input features, thus being able to effectively capture the dependency relationships and context information between tokens. This process is shown by formula (5): (5) where represents the graph convolutional network, represents the The learnable weight matrix of the layer. To improve the local consistency of the model during the token feature learning process, this method imposes a masked region constraint on the output features, aiming to control the similarity within the feature region, thereby enhancing the model's ability to learn local region features. Specifically, this method uses the original mask as a constraint condition. By minimizing the Euclidean distance between similar tokens, the model can better focus on and strengthen the features of similar regions during the learning process. In this way, the model can maintain strong local consistency during the feature learning process, thus more effectively capturing the region information related to the human body. This masked constraint loss is calculated as shown in Equation (6): (6) where represents the and th tokens' Euclidean distance.

[0042] S5: Based on the skeleton and vertex information of the SMPL model as priors, combined with the enhanced image features, the three-dimensional joint position coordinates and body mesh vertex coordinates are predicted through the decoder structure.

[0043] In this embodiment, the learnable joint point features and vertex features are initialized. The vertex features are obtained by downsampling and dimension elevation based on the vertices of the SMPL model; the learnable joint point features and vertex features are input into the decoder for feature modeling; the decoder includes a structure-aware multi-head self-attention layer and a cross-attention layer; the multi-head self-attention layer uses a predefined edge connection matrix based on the SMPL skeleton structure and incorporates the attention weights calculation by combining with mask weighting; the cross-attention layer uses the enhanced image features as the key vectors and value vectors, and uses the joint point features and vertex features as the query vectors to perform cross-attention calculation.

[0044] Furthermore, the joint point features and vertex features output by the decoder are regressed to obtain the initial three-dimensional joint point coordinates and the rough vertex grid; an upsampling operation is performed on the rough vertex grid to restore the fine SMPL parameterized body mesh.

[0045] In this embodiment, after step S5, it further includes calculating the total loss, which includes: a 2D joint projection loss for constraining the predicted two-dimensional projected joint points and the ground truth two-dimensional joint points; a 3D geometric loss for constraining the geometric consistency between the predicted three-dimensional joint point coordinates and body mesh vertex coordinates and the ground truth coordinates; and a masked constraint loss.

[0046] Specifically, the bone prior decoupling network consists of the following two parts: the SMPL body fitting layer and the loss optimization layer, whose purpose is to recover the pose information and body mesh coordinates of a 3D human body from an image, as follows: 1. SMPL Body Fitting Layer In the previous HMR method, image features were used to regress the camera parameters, pose, and body parameters of a human body through subsequent multiple linear layers, and then these parameters were applied to the SMPL model for human body recovery. This method effectively achieved the reconstruction of human geometry by extracting key information from the image and mapping it to the parameter space of the human body model. However, in recent years, some research works have embedded the parametric features of SMPL into neural networks for end-to-end joint optimization. This method has significantly improved the model's ability to model human geometry by directly optimizing SMPL parameters in the network. Inspired by this, this study decided to use SMPL as the skeleton model to perform feature learning, and the parameter regression process is as Figure 4 shown. Specifically, this study used the skeleton structure of SMPL as a prior condition for feature learning to guide the model to better learn the geometric shape of the human body.

[0047] Set learnable joint point features and a set of learnable vertex features , where and represent the number of joint points and vertices respectively, and represents the feature dimension. These two types of features are input into the decoder for modeling through a concatenation operation. To reduce the computational complexity and improve the adaptability of the model, first, this study downsamples the original number of vertices (6890 vertices) of the SMPL parametric human body model to , and uses a linear layer to increase its feature dimension to 512. Then, a residual connection is designed as the bias guidance for vertex features to ensure that information can be more stable and effective during transmission.

[0048] In the design of the decoder, this study introduced a structure-aware multi-head self-attention mechanism (MHSA) and combined it with a predefined edge connection matrix for feature modeling, and its structure is as Figure 5 shown. To further improve the model's expressive ability, the edge connection matrix is incorporated into the calculation of attention weights through masked weighting. This process is shown in Equation (7): (7) where d represents the channel dimension, and represents the minimum threshold parameter between non-adjacent nodes. are respectively obtained by the corresponding projection matrices They are obtained through transformation and represent the query vector, key vector, and value vector respectively. To further improve the model's feature fusion ability, this study follows the common decoder architecture method and introduces a cross-attention layer on this basis. In this mechanism, the image features are used as the key vector K and value vector V, while the concatenated joint features are used as the query vector Q. Through the cross-attention mechanism, the image features can effectively interact with the joint features to guide and enhance the feature expression of human joints.

[0049] 2. Loss Optimization Layer Through the feature modeling of multiple decoder layers, the model regresses the output joint features and vertex features, and then obtains the predicted 3D joint coordinates. On this basis, an upsampling operation is gradually performed on the corresponding rough vertex mesh to recover more refined SMPL parameters. .

[0050] To align the input image with the 3D reconstructed body mesh, this method follows the existing method to perform 2D joint projection loss, and uses the camera parameters to project the estimated 3D joints to 2D to obtain the joint points for supervision: , where , is the projection function. It is used for calculating the geometric consistency between the predicted 3D coordinates and the ground truth coordinates. It includes the calculation of 3D joint points and 3D body mesh: . The loss function is shown in formula (8): (8) In summary, combined with the mask constraint loss , the total loss is shown in formula (9): (9) where represent the weight values of 2D coordinates, 3D coordinates, and mask constraints in training respectively.

[0051] The training process of the method proposed in this study is as follows: Input: Training set , human parsing model , CNN network , graph convolutional network , encoder , MGCL network , SMPL parameterization model , decoder .

[0052] 1. Use the human parsing model to extract the human part parsing diagram of the training set: ; 2. for in do; 3. Extract the features of image I and mask M respectively through the CNN encoder: ; 4. For After passing through the Transformer encoder Perform feature encoding; 5. Through formulas (2, 3, 4) for Construct a mask-aware matrix ; 6. Use the graph convolutional network for Feature update: ; 7. Calculate the mask constraint loss through formula (6) : ; 8. The SMPL parametric model performs downsampling and inputs it into the decoder for joint feature learning; 9. The features of the image Guide the learning of joint point features to regress the three-dimensional human pose information; 10. Update the model parameters through loss calculation in formula (9); 11. End.

[0053] In summary, this embodiment proposes a human action capture method based on mask-aware graph convolution and skeletal prior, which mainly includes two core processes: the mask-aware graph convolution learning (Mask-aware Graph ConvolutionLearning, MGCL) process and the skeletal prior decoupling process. Specifically, MGCL extracts the static features of RGB images and masks, flattens them into a token sequence after convolutional dimensionality reduction to 512 dimensions, combines positional encoding to retain spatial information, and optimizes the feature representation through the Transformer encoder. Subsequently, a mask-aware matrix is used to construct the GCN adjacency matrix, and a mask constraint loss is imposed to enhance local consistency. In the skeletal prior decoupling process, the SMPL skeletal prior model is introduced, and its skeleton and vertex information are used as prior knowledge for feature learning. The decoder structure combines multi-head self-attention and cross-attention mechanisms, uses a predefined adjacency matrix, and fuses the attention weights through mask weighting. Finally, under the guidance of image features, the cross-attention mechanism is used to perform multi-modal data augmentation on the SMPL skeletal node features, thereby optimizing the geometric consistency of three-dimensional coordinate prediction.

[0054] Embodiment 2 Please refer toFigure 6 As shown, it is a schematic structural diagram of a human motion capture system based on mask-aware graph convolution and skeletal prior proposed in the second embodiment of the present application. The system includes the following key modules: An image acquisition module 100, configured to acquire a single RGB image; A mask extraction module 200, configured to extract a human mask corresponding to the RGB image; A mask-aware graph convolution module 300, including: A feature encoding unit 301, configured to respectively perform feature extraction on the RGB image and the human mask to obtain image features and mask features, and respectively encode the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information; and An adjacency matrix construction unit 302, configured to construct a mask-aware adjacency matrix according to the image token sequence and the mask token sequence, and update the image token sequence by using a graph convolution network to obtain enhanced image features; A skeletal prior decoupling module 400, including an SMPL body fitting unit 401, configured to use the skeleton and vertex information of the SMPL model as a prior, and combine the enhanced image features to predict three-dimensional joint position coordinates and body mesh vertex coordinates through a decoder structure.

[0055] A human motion capture system based on mask-aware graph convolution and skeletal prior in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.

[0056] A human motion capture system based on mask-aware graph convolution and skeletal prior in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0057] The human motion capture system based on mask-aware graph convolution and skeletal prior provided by the embodiments of the present application can achieve Figure 1 each process implemented by a human motion capture method based on mask-aware graph convolution and skeletal prior in the method embodiments. To avoid repetition, it will not be elaborated here.

[0058] Optionally, the embodiments of the present application further provide an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned method embodiment of the human motion capture method based on mask-aware graph convolution and skeletal prior, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0059] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned method embodiment of the human motion capture method based on mask-aware graph convolution and skeletal prior, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0060] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0061] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0062] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0063] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A human motion capture method based on mask-aware graph convolution and skeleton prior, characterized in that: The method comprises: a) obtaining a single RGB image and extracting a human body mask corresponding to the RGB image; b) extracting features of the RGB image and the human body mask respectively through a feature encoding layer of a mask-aware graph convolutional network to obtain image features and mask features, and encoding the image features and the mask features respectively to obtain an image token sequence and a mask token sequence containing spatial position information; c) constructing a mask-aware adjacency matrix according to the image token sequence and the mask token sequence through an adjacency matrix construction layer of a mask-aware graph convolutional network, and updating the image token sequence and the adjacency matrix using a graph convolutional network to obtain enhanced image features; d) Through the SMPL body fitting layer of the skeleton prior decoupling network, based on the skeleton and vertex information of the SMPL model as prior, combined with the enhanced image features, the three-dimensional joint position coordinates and the body mesh vertex coordinates are predicted through the decoder structure.

2. The method according to claim 1, characterized in that In step b), the steps of respectively encoding the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information include: A convolution operation is used to reduce the channel dimension of the image features and mask features to a preset dimension D; The reduced image feature map and mask feature map are reorganized into image token sequences and mask token sequences of length HW+1, respectively. Indicates the height of the image. Indicates the width of the image. The token sequence contains a learned camera token tag. Apply linear projection operations to the image token sequence and mask token sequence respectively; Add position encoding P to the linearly projected image token sequence and mask token sequence to obtain image token sequence and mask token sequence containing spatial position information.

3. The method according to claim 2, characterized in that In step b), after respectively encoding the image features and the mask features to obtain an image token sequence and a mask token sequence containing spatial position information, the method further includes: The image token sequence and mask token sequence after adding the position code P are input into the Transformer encoder in parallel; The self-attention mechanism of the Transformer encoder is used to model the spatial relationship within the image token sequence and the mask token sequence, respectively, and the encoded image token sequence and mask token sequence are output.

4. The method according to claim 1 or 3, characterized in that: In step c), the step of constructing a mask-aware adjacency matrix specifically includes: Generate an initial mask perception matrix by calculating the Hadamard product between the image token sequence and the mask token sequence; The original mask M is extracted using the human body parsing model as a screening condition, and the initial mask perception matrix is ​​processed: the connection relationship between the tokens whose corresponding position values ​​in the original mask M are zero is filtered, and the connection relationship of the background area is given a weight controlled by a preset scale factor β to obtain a filtered matrix; The filtered matrix is ​​weight normalized to generate a mask-aware adjacency matrix for a graph convolutional network.

5. The method according to claim 4, characterized in that Step c) further comprises: The mask-constrained loss is calculated through the mask-constrained loss layer of the mask-aware graph convolutional network; The calculation of the mask constraint loss includes: identifying token pairs with the same non-zero label value in the original mask M , and minimize the Euclidean distance between the enhanced image features corresponding to the token pairs; The calculation formula of the mask constraint loss is: in is the mask loss, Represents enhanced image features No. and The Euclidean distance of tokens, Indicates that the original mask M is The value of the position.

6. The method according to claim 1, characterized in that Step d) comprises: Initialize learnable joint point features and vertex features, where the vertex features are obtained by downsampling and dimension enhancement based on the vertices of the SMPL model; Inputting the learnable joint point features and vertex features into a decoder for feature modeling; The decoder includes a structure-aware multi-head self-attention layer and a cross-attention layer; The multi-head self-attention layer uses a predefined edge connection matrix based on the SMPL skeleton structure and combines mask weighting into attention weight calculation; The cross-attention layer uses the enhanced image features as key vectors and value vectors, and uses the joint point features and vertex features as query vectors to perform cross-attention calculations.

7. The method according to claim 6, characterized in that Step d) further comprises: Regress the joint point features and vertex features output by the decoder to obtain the initial three-dimensional joint point coordinates and a rough vertex mesh; An upsampling operation is performed on the coarse vertex mesh to restore a fine SMPL parameterized body mesh.

8. The method according to claim 1, 5 or 7, characterized in that: Step d) further comprises: The total loss is calculated by the loss optimization layer of the skeleton prior decoupling network, and the total loss includes: 2D joint projection loss used to constrain the predicted 2D projected joint points and the true 2D joint points; 3D geometry loss for constraining the geometric consistency of predicted 3D joint point coordinates and body mesh vertex coordinates with the true value coordinates; and mask constrained loss.

9. A human motion capture system based on mask-aware graph convolution and skeleton prior, characterized in that: The system comprises: Image acquisition module, used to acquire a single RGB image; A mask extraction module, used to extract a human body mask corresponding to the RGB image; Mask-aware graph convolution module, including: A feature encoding unit, used to extract features from the RGB image and the human body mask respectively to obtain image features and mask features, and to encode the image features and the mask features respectively to obtain an image token sequence and a mask token sequence containing spatial position information; and An adjacency matrix construction unit, configured to construct a mask-aware adjacency matrix according to the image token sequence and the mask token sequence, and to update the image token sequence and the adjacency matrix using a graph convolutional network to obtain enhanced image features; The skeleton prior decoupling module includes an SMPL body fitting unit, which is used to obtain three-dimensional joint position coordinates and body mesh vertex coordinates through decoder structure prediction based on the skeleton and vertex information of the SMPL model as priors and the enhanced image features.

10. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a human motion capture method based on mask-aware graph convolution and skeleton prior as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Three-dimensional human motion prediction method based on optimized graph convolutional network

    CN116596959A

  • Underground dangerous behavior identification method based on multi-modal joint training network

    CN118968614A

  • Method for intervening and correcting posture of digital human generated by AIGC

    CN119091013A

  • Three-dimensional (3D) integrated teaching field system based on flipped platform and method for operating same

    US20240038086A1

Cited By

  • Motion capture method for unmarked dynamic shielding scene

    CN121861728A

  • An action capture method for a markerless dynamic occlusion scene

    CN121861728B

  • A method, system, device and storage medium for three-dimensional human pose estimation

    CN122574961A