Transform-based human body grid reconstruction method

Through the multi-view feature aggregation and posture optimization framework, the robustness problem of existing methods under complex posture and occlusion conditions is solved, efficient and accurate human body mesh reconstruction is achieved, and the reconstruction accuracy under extreme postures is improved.

CN120655858APending Publication Date: 2025-09-16HANGZHOU DIANZI UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510718433.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing human body mesh reconstruction methods are not robust enough under complex posture and occlusion conditions. Their reliance on single-view images results in low reconstruction accuracy, especially in extreme postures.

Method used

Through the multi-view feature aggregation and posture optimization framework, features from different viewpoints are used to compensate for occlusion information, combined with the posture optimization strategy, to decouple human body shape and posture information and generate high-precision meshes.

Benefits of technology

The mesh reconstruction quality in complex scenes and high-density occlusion is significantly improved, the reconstruction accuracy and robustness are improved, and the errors caused by occlusion and posture changes are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655858A_ABST
    Figure CN120655858A_ABST
Patent Text Reader

Abstract

The invention discloses a human body grid reconstruction method based on Transform, and the method comprises the steps: firstly obtaining a public human body data set, and carrying out the standardization of an image in the human body data set; secondly, images in the human body data set generate a plurality of visual angle features through two branches of a front view encoder and a visual angle conversion network respectively, optimization and enhancement are carried out respectively, the optimized and enhanced features are fused, and a multi-visual angle aggregation feature is generated; and then, in combination with the optimization of a Transform encoder, initial joint features are extracted from the images in the human body data set, and final posture features are generated. And finally, constructing a grid regression module, adding the multi-view aggregated features and the attitude features, and inputting the added features to the grid regression module to generate a final human body grid reconstruction result. According to the method, the feature information of different visual angles is fused, so that the precision and detail representation of human body grid reconstruction are effectively improved under the conditions of complex posture change and shielding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer graphics and deep learning, and specifically relates to a human body mesh reconstruction method based on Transformer. Background Art

[0002] In recent years, with the rapid development of computer vision and image processing technologies, human body mesh reconstruction has become a highly sought-after research area. Human body mesh reconstruction aims to recover the geometric structure of the human body from two-dimensional images or videos, thereby enabling accurate reconstruction of human form and posture analysis. This technology has important application prospects in fields such as virtual reality, augmented reality, and motion analysis. Due to the spatial differences between the source image and the target mesh, existing mesh reconstruction methods typically rely on paired data for training, using the source image and its corresponding target mesh as supervisory signals to guide the model's learning process.

[0003] Existing methods for human mesh reconstruction primarily include methods based on the SMPL (Skinned Multi-Person Linear model) human model and model-free methods. SMPL-based methods generate meshes by regressing human body parameters, while model-free methods directly regress mesh vertices from images without requiring a predefined model. SMPL-based methods map images to the model parameter space, while model-free methods rely on deep neural networks to directly recover the mesh structure.

[0004] When reconstructing a human body mesh, relying solely on joint information in an image to represent the pose is often insufficient to handle complex pose variations, especially when there are significant differences between the source image and the target mesh. To address this issue, some methods have introduced additional auxiliary information, such as deep features or full-body contour information, to provide more precise guidance and help recover a more accurate mesh structure.

[0005] Traditional methods for human mesh reconstruction mostly rely on paired source images and ground-truth meshes as supervisory signals for training. However, in recent years, some studies have proposed training methods that do not rely on explicit ground-truth meshes. By using techniques such as self-supervision or generative adversarial networks, mesh reconstruction can be performed without real annotated data. These methods can effectively train in the absence of annotated data.

[0006] Some methods use methods based on the SMPL human model and model-free methods. For example, Kolotouros et al. recover human shape and pose by regressing SMPL parameters. However, challenges remain when dealing with occlusion and complex movements, especially poor accuracy in extreme poses. To address this issue, Kim et al. proposed a point-guided feature sampling method that guides feature sampling based on the projection results of 3D mesh vertices. However, this method still suffers from insufficient local feature learning and cannot fully handle global dependencies, resulting in inconsistent details in some generated results in complex scenes. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this paper proposes a Transformer-based method for human mesh reconstruction. By aggregating features from different viewpoints and optimizing pose joints, this method decouples human shape and pose information, effectively improving reconstruction accuracy. This method addresses the robustness issues of existing methods in complex pose and occlusion conditions, eliminates reliance on single-view images, and achieves efficient and accurate human mesh reconstruction.

[0008] This paper uses a multi-view feature aggregation and pose optimization framework to address occlusion issues in human mesh reconstruction (HMR). This framework fuses multi-view information and uses features from different perspectives to compensate for information loss caused by occlusion, thereby enhancing the ability to recover from occluded areas. Simultaneously, combined with a pose optimization strategy, it can further improve reconstruction accuracy in scenes with large pose variations and reduce reconstruction errors caused by occlusion. In this way, the present invention significantly improves the quality of mesh reconstruction in complex scenes and under high-density occlusion.

[0009] The technical solution adopted by the present invention to achieve the above-mentioned purpose is a human body mesh reconstruction method based on Transformer, comprising the following steps:

[0010] Step 1: Use a publicly available human dataset containing full-body frontal images from various viewpoints. All images are annotated with ground truth (GT) annotations, including skeletal keypoints and joint information. These annotations are used for subsequent image processing and pose optimization. Next, all images undergo normalization, including resizing, background padding, and cropping, to ensure that each image has a consistent size format.

[0011] Step 2: View feature aggregation. For the images in the human body dataset (full-body frontal images), multiple view features are generated through the front view encoder and the view conversion network respectively. These features are optimized and enhanced respectively, and the optimized and enhanced features are fused to generate multi-view aggregated features.

[0012] The input full-body frontal image is first passed through the backbone network of the HRNet model to extract the initial feature representation. This initial feature is then fed into two processing branches in parallel: the front view modeling branch and the perspective conversion branch. In the front view modeling branch, the initial feature is input to the Transformer-based front view encoder (Front View Encoder). The encoder uses a multi-head self-attention mechanism to perform global modeling and structural perception on the front view features, generating a structured feature representation with contextual semantic information, denoted as X. front . This feature will be passed as guidance information to the front view guidance module (FOB) for further optimization. In the perspective conversion branch, the initial feature is input into the perspective conversion network (View Conversion Network). The network integrates a normal prediction module (Normal Prediction Network, NPN) to generate a unit normal map. Subsequently, the unit normal is dot-producted with the predefined perspective direction vector (corresponding to left view, right view, and rear view) pixel by pixel. The dot product result is used as a binary perspective mask matrix, and the binary perspective mask matrix is ​​dot-multiplied with the unit normal map to obtain the visible area feature. The visible area feature is convolutionally enhanced and up-sampled to finally obtain the feature map F under each perspective. i ∈

[0013] R H×W×C , where i includes three directions i∈{left,back,right}: left, right and back.

[0014] The predefined viewing direction vectors (corresponding to left view, right view, and rear view) are specifically: The predefined viewing direction vectors used in the present invention are intended to simulate the geometric perspective of analyzing the human body from different observation angles in three-dimensional space. Specifically, the left view, right view, and rear view directions are set to [-1, 0, 0],

[0015] The coordinates [1,0,0] and [0,0,-1] are defined based on a standard 3D Cartesian coordinate system, where the human body faces the positive z-axis by default, with left and right corresponding to the x-axis, and the top of the head facing the positive y-axis. This configuration aligns with the world coordinate system used by the SMPL model and conforms to the common perspective division conventions used in human modeling and rendering. By performing a pixel-by-pixel dot product of these direction vectors with the unit normal map, we can effectively determine whether the current pixel is facing the viewer at a specific viewpoint, which is then used for subsequent visibility determination and view mask construction.

[0016] The Normal Prediction Network (NPN) proposed in this paper uses initial features as input and adopts a shallow network structure consisting of four convolutional layers and one fully connected layer. The convolutional layers are used to gradually extract local spatial geometric information and are combined with the ReLU activation function to enhance nonlinear expression capabilities. Finally, the fully connected layer maps the features at each pixel position into a three-dimensional vector and normalizes it to output a unit normal map. This normal map maintains the same spatial resolution as the input image, with each pixel corresponding to a normalized three-dimensional normal direction. It is used to provide accurate local surface orientation information, giving the network the ability to perceive the target geometry.

[0017] Multiple directional features F from the view conversion network i The left view, right view, and rear view are fed into three separate side view decoders, each corresponding to a specific direction. This module uses a Transformer-based decoder structure to model features from different view angles. Each directional feature is used as input, combined with the guidance feature from the front view encoder as a query, and the cross-region contextual dependencies are captured through a cross-attention mechanism. Each decoder outputs an enhanced feature representation X for its corresponding direction. i and pass it as input to the multi-view optimization module (MIOB) for further feature fusion.

[0018] The front view guidance module (FOB) and the multi-view optimization module (MIOB) together constitute the key feature aggregation structure of this method. The FOB module first receives the front view feature X front The spatial dimensions are expanded by the Flatten operation, and layer normalization (LN) is applied to stabilize the feature distribution. Subsequently, the multi-head self-attention mechanism (MHA) is applied to model the global dependencies of the front view features, and the structure expression is enhanced by the multi-layer perceptron (MLP). The FOB output is then reshaped for subsequent use with the enhanced feature representation X. i Alignment and fusion in the spatial dimension to obtain the optimized front view feature X′ front The multi-view optimization module (MIOB) uses the FOB output as a guide (Query) to optimize each directional feature (left view, right view, and rear view). Each MIOB branch receives the corresponding directional feature X i As the key and value, Flatten, layer normalization, cross attention mechanism (MCA), re-normalization and MLP are performed in sequence, and then the channel dimension is adjusted and the feature is compressed through 1×1 convolution operation. The final output is the perspective perception feature representation after structure enhancement, which is recorded as

[0019] Finally, the optimized feature X′ output by FOB is front View-aware features of MIOB output By element-by-element summation, we can obtain the enhanced multi-view aggregation feature X fusion .

[0020] Step 3: Pose extraction and optimization (parallel auxiliary branch): extract initial joint features from the full-body frontal image and optimize its structural representation in combination with the Transformer encoder to enhance the modeling ability between key joints.

[0021] A parallel pose optimization path is introduced to enhance the structural modeling capabilities between key joints. Starting from the input image, this branch first obtains the initial joint feature representation through the joint extraction module. To improve the robustness of the model in occlusion scenarios, a random occlusion mask is constructed to simulate occlusion of the joint features. Specifically, the mask is fused with the initial joint feature representation through element-by-element multiplication to simulate the situation where some joints are invisible. Then, positional encoding is introduced to embed the position information of each joint in the spatial structure to assist the network in modeling the relative spatial relationship between key points. The joint features that fuse occlusion information and position encoding are then input into the joint optimization module to further model the structural dependencies between key points. The module as a whole consists of four parts: First, the joint optimization module flattens the joint features from the spatial structure to a sequence form through the Flatten operation, and then applies LayerNorm to stabilize the feature distribution. Next, the features are fed into the pose encoder, which consists of a multi-layer stacked Transformer encoder. It uses a multi-head self-attention mechanism to capture global context dependencies and strengthen the structural relationship between joints. The output of the Transformer encoder is fed into the MLP layer for feature conversion and mapping to generate the final pose feature X pose .

[0022] Step 4: Mesh regression, build a mesh regression module, add the multi-view aggregation features to the final posture features, and input them into the mesh regression module to generate the final mesh reconstruction result.

[0023] In the mesh regression stage, the aggregated multi-view features are first element-wise added to the optimized pose features to fuse the global information from different viewpoints with the detailed pose features. These fused features are then fed into the mesh regression module, whose primary task is to map these high-dimensional features into a 3D mesh space. This module typically employs fully connected layers or convolution-based structures to gradually map the fused features into vertex positions in a 3D coordinate system.

[0024] In order to optimize the accuracy of mesh reconstruction, four key loss terms are used: vertex loss 3D joint loss 2D joint loss and heatmap loss These loss terms each play a different role and together promote the accuracy improvement of the model in vertex positioning, joint estimation and key point detection. By measuring the Euclidean distance between the generated mesh vertices and the real vertices, the shape of the mesh is ensured to be consistent with the real data. By calculating the distance between the predicted 3D joints and the real joints, the spatial position of the joints is accurately optimized. By projecting the 3D joints onto the 2D plane, the positions of the 2D joints are optimized to enhance the robustness under multiple views. The pixel-level perspective support is used to improve the positioning accuracy of key points, especially in complex postures and occlusion situations, which can effectively improve the positioning accuracy of joints.

[0025] Beneficial effects of the present invention:

[0026] Existing methods for human body mesh reconstruction rely on single-view image data for mesh reconstruction, which often faces issues such as view dependency, occlusion, and pose variations. This paper proposes a multi-view aggregation compensation method that effectively compensates for occlusion and pose ambiguity by fusing feature information from different viewpoints. Unlike traditional methods, this method dynamically weights and fuses multi-view features, eliminating the reliance on a single viewpoint or complex viewpoint transformation processes, resulting in significant advantages in accuracy and robustness.

[0027] This paper proposes a pose optimization framework that uses deep learning to dynamically optimize joint positions and poses, eliminating mesh distortion and positioning errors caused by complex movements and rapid pose changes. This framework accurately optimizes poses when handling dynamic movements, improving the accuracy and detail of mesh reconstruction.

[0028] The feature fusion strategy in this invention specifically includes extracting features from different viewpoints, dynamically weighted fusion, and compensating for occlusion and pose ambiguity. This mechanism can significantly improve the performance of 3D mesh reconstruction in complex scenes, especially in complex backgrounds and challenging pose reconstruction tasks. This strategy is detailed in step 2 of the specific implementation method.

[0029] The core technologies of the pose optimization framework in this paper, including the multi-stage feedback mechanism during the pose optimization process, can effectively improve the accuracy and detail of human mesh reconstruction under complex pose changes and occlusions, and show significant advantages in handling extreme movements and dynamic changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic diagram of the process of the present invention;

[0031] Figure 2 Schematic diagram of the network structure of the present invention;

[0032] Figure 3 Network structure diagram of view feature aggregation;

[0033] Figure 4 A comparison chart of the present invention and other existing methods in the task of human body mesh reconstruction;

[0034] Figure 5 This is the result of testing the present invention on real-world human body images. DETAILED DESCRIPTION

[0035] To address the shortcomings of existing technologies, this paper proposes a human body mesh reconstruction method based on view feature aggregation and pose optimization. By aggregating features from different viewpoints and optimizing pose joints, this method decouples human body shape and pose information, effectively improving reconstruction accuracy. This method addresses the robustness issues of existing methods in complex pose and occlusion conditions, eliminates reliance on single-viewpoint images, and achieves efficient and accurate human body mesh reconstruction.

[0036] The technical solution adopted by the present invention to achieve the above-mentioned purpose is a human body mesh reconstruction method based on Transformer. The present invention is described in detail below with reference to the accompanying drawings and specific embodiments:

[0037] like Figure 1 As shown in Figure 2, the Transformer-based human body mesh reconstruction method includes the following steps:

[0038] Step 1: Use a public human dataset containing full-body frontal images to provide input for feature extraction and multi-view modeling.

[0039] To ensure accurate human mesh reconstruction, the present invention selected publicly available standard human image datasets, specifically Human3.6M and 3DPW. These datasets contain a large number of full-body images with annotated skeletal joint information. Using this data, the present invention can extract joint information and appearance features from single-view images, enabling accurate mesh reconstruction.

[0040] First, the present invention screened approximately 20,000 full-body images of 256×256 size from the two datasets. All images were annotated with GT annotations, including skeletal key points and joint information of the human body, which were used for subsequent image processing and pose optimization. Next, all images were normalized, including resizing, background filling, and cropping, to ensure that each image had a uniform size format (256×256). This step provided a standardized data foundation for subsequent feature extraction and mesh regression.

[0041] like Figure 2 Figure 2 shows the overall block diagram of the human body mesh reconstruction method based on multi-view feature aggregation and pose optimization in the present invention. The block diagram includes input: a single image; initial feature extraction (Backbone); view conversion (ViewConversion): view feature aggregation (View Feature Aggreation); pose extraction (Pose Estimation); pose optimization (Pose Optimization) and mesh regressor (Mesh Regression).

[0042] Steps 2 to 4 are the overall process of human body mesh reconstruction with multi-view feature aggregation and posture optimization.

[0043] Step 2: View feature aggregation: The initial features are input into the forward-looking encoder and the view conversion module respectively to generate multi-view features and perform fusion optimization.

[0044] like Figure 2 As shown in the figure, the input full-body frontal image is first extracted through the backbone network (Backbone) to extract the initial feature representation. The initial feature is then sent to two processing branches in parallel: the front view modeling branch and the perspective conversion branch. In the front view modeling branch, the initial feature is input to the Transformer-based front view encoder (Front View Encoder). The encoder uses a multi-head self-attention mechanism to perform global modeling and structural perception on the front view features, generating a structured feature representation with contextual semantic information, denoted as X. front. This feature will be passed as guidance information to the front view guidance module (FOB) for further optimization. During the perspective conversion process, a normal prediction module (Normal Prediction Network, NPN) is introduced to generate a pixel-level unit surface normal map. This module adopts a lightweight convolution structure, takes the initial feature map as input, and outputs the three-dimensional unit normal vector corresponding to each pixel. In order to achieve direction-aware perspective modeling, the geometric perception idea in SMPL-X is borrowed to define a fixed perspective direction vector for each target perspective (left view, right view, and rear view). Subsequently, for each pixel position, the dot product value between its unit normal and the target perspective direction is calculated, and a binary perspective mask is constructed based on this. The area where the dot product result is greater than zero is considered to be the visible area of ​​the current perspective, and the remaining areas are masked. The obtained perspective mask is used to filter the initial feature map, retaining only the salient areas under the current perspective. In order to further enhance the geometric structure features of these areas, a convolution layer is used on the masked feature map for local modeling and enhancement, and then an upsampling operation is performed to obtain a perspective feature map F aligned with the backbone feature. i ∈R H×W×C , where i∈

[0045] {left,back,right}.

[0046] Multiple directional features (including left view, right view, and rear view) from the view conversion network are fed into three independent side view decoders, each corresponding to a specific direction. This module uses a Transformer-based decoder structure to model features from different viewpoints. Each directional feature is used as input, combined with the guidance feature from the front view encoder as a query, and the cross-attention mechanism is used to capture cross-region contextual dependencies, enhancing the directional perception and structural expression capabilities of the viewpoint features. Each decoder outputs an enhanced feature representation X for its corresponding direction. i , where i∈{left,back,right}, and passes it as input to the multi-view optimization module (MIOB) for further feature fusion.

[0047] like Figure 3 As shown in the figure, the front view guidance module (FOB) and the multi-view optimization module (MIOB) together constitute the key feature aggregation structure of this method. The FOB module first receives the front view feature X frontThe spatial dimension is expanded by the Flatten operation, and layer normalization (LN) is applied to stabilize the feature distribution. Subsequently, the multi-head self-attention mechanism (MHA) is applied to model the global dependency of the front view features, and the structural expression is enhanced by the multi-layer perceptron (MLP). The FOB output is then reshaped to align and fuse with the subsequent view features in the spatial dimension to obtain the optimized front view feature X front The multi-view optimization module MIOB consists of three parallel branches, which process the left view, right view and back view directional features generated by the side view decoder respectively. Each MIOB branch receives its corresponding directional feature X i , where (i∈{left, back, right}), the following processing flow is performed: first, Flatten and position encoding are superimposed, followed by layer normalization, multi-head self-attention mechanism and cross attention mechanism (MCA) in sequence, where the query comes from the front view feature output by FOB, and the key and value come from the directional feature of the current view. This structure is used to capture the contextual interaction between the front view and the side view, and enhance the cross-view structure alignment capability. Finally, each branch outputs its structurally enhanced directional feature representation, denoted as The formula for multi-head attention is as follows:

[0048]

[0049] The formula for cross attention is as follows:

[0050] Q=X front W Q ,K=X i W K ,V=X i W V

[0051]

[0052] Where W Q 、W K and W V is the corresponding weight, and d is the dimension of the key vector.

[0053] Finally, the optimized feature X output by FOB is front View-aware features of MIOB output By summing up the elements one by one, we can get the enhanced aggregate feature X fusion .

[0054] Step 3: Pose extraction and optimization, extract the initial joint features from the full-body frontal image and optimize its structure in combination with the Transformer encoder.

[0055] To further improve the accuracy of human mesh reconstruction in complex scenes, independent pose extraction and optimization branches are designed, such as Figure 2 As shown in Figure 2, the overall network consists of a multi-view feature aggregation path and a pose optimization path, which respectively model the image's structural information and joint details. While multi-view feature aggregation can capture global spatial relationships and provide relatively complete information about the human body's structure, its ability to fine-grainedly model key local joints can be insufficient under conditions of drastic pose changes or partial occlusion, resulting in unstable or incomplete pose representations.

[0056] To this end, a parallel pose optimization path is introduced to enhance the structural modeling capabilities between key joints. Starting from the input image, this branch first obtains an initial joint feature representation through the joint extraction module. To improve the model's robustness in occlusion scenarios, a random occlusion mask is constructed to simulate occlusion of the joint features. Specifically, the mask is fused with the original joint features through an element-wise product, simulating the situation where some joints are invisible, thereby enabling the network to acquire occlusion awareness and structural completion capabilities. Furthermore, positional encoding is introduced to embed the spatial position information of each joint, assisting the network in modeling the relative spatial relationships between key points. The joint features, fused with occlusion information and position encoding, are then input to the joint refinement module to further model the structural dependencies between key points. This module consists of four parts: First, the joint features are flattened from a spatial structure into a sequence form through the Flatten operation, followed by layer normalization to stabilize the feature distribution. Next, the features are fed into the pose encoder, which consists of a multi-layer stacked Transformer encoder. It uses a multi-head self-attention mechanism to capture global context dependencies and strengthen the structural relationship between joints. The output of the Transformer encoder is fed into the MLP layer for feature conversion and mapping to generate the final pose feature X pose The attention formula is as follows:

[0057]

[0058] Step 4: Mesh regression, add the aggregated features to the final posture features and input them into the mesh regression module to generate the final mesh reconstruction result.

[0059] In the mesh regression stage, the aggregated multi-view features are first element-wise added to the optimized pose features to fuse global information from different views with detailed pose features. This fusion process ensures that the model fully utilizes multi-view information while preserving the accuracy of the optimized pose, resulting in a more accurate body mesh. This fusion overcomes the shortcomings of single-view information, especially in complex poses, occlusions, or incomplete viewports. The fused features are then input to the mesh regression module, whose primary task is to map these high-dimensional features into a 3D mesh space. The mesh regression module typically employs fully connected layers or convolution-based structures to gradually map the fused features into vertex positions in a 3D coordinate system. The resulting body mesh should be consistent with the shape and structure of the real-world data, ensuring that the relative positions of each keypoint and joint are accurately reconstructed.

[0060] In order to optimize the accuracy of mesh reconstruction, four key loss terms are used: vertex loss 3D joint loss 2D joint loss and heatmap loss These loss terms each play a different role and together promote the accuracy improvement of the model in vertex positioning, joint estimation and key point detection. By measuring the Euclidean distance between the generated mesh vertices and the real vertices, the shape of the mesh is ensured to be consistent with the real data. By calculating the distance between the predicted 3D joints and the real joints, the spatial position of the joints is accurately optimized. By projecting the 3D joints onto the 2D plane, the positions of the 2D joints are optimized to enhance the robustness under multiple views. By using pixel-level perspective support to improve key point positioning accuracy, especially in complex postures and occlusion situations, it can effectively improve joint positioning accuracy. The loss formula is as follows:

[0061]

[0062] Where V n represents the predicted vertex, Represents the true value of the vertex, and N is the total number of vertices:

[0063]

[0064] in represents the estimated joint position in 3D space, represents the actual joint position, represents the predicted joint position, and K represents the total number of joints.

[0065]

[0066] in Indicates that the estimated joint positions in three-dimensional space are mapped to the estimated joint positions in two-dimensional space, Represents the real 2D joint position, represents the predicted 2D space joint position, K represents the total number of joints;

[0067]

[0068] in Represents the true joint heat map, H n Represents the predicted joint heatmap.

[0069] Figure 4 The visualization comparison results of the method of the present invention and the existing methods PointHMR and POTTER in the human mesh reconstruction task are shown. Each row from left to right is the input image, PointHMR, POTTER and the reconstruction result of the present method. As can be seen from the figure, the present method shows stronger capabilities in posture restoration, structural coherence and detail recovery. In the bending action in the first row, the mesh generated by the present method is more natural, and the action restoration is more consistent with the real human body structure; in the stair-going scene in the third row, the present method can more accurately reconstruct the spatial relationship between the limbs and the torso, avoiding the deformation problem that occurs in other methods; and in the sitting scene in the last row, even in the presence of large occlusion and background interference, the present method still maintains a stable structural restoration effect. The overall results show that the method of the present invention exhibits superior mesh reconstruction capabilities in complex postures and occlusion environments.

[0070] also, Figure 5 The figure shows the results of testing the present invention on real-world human images. Given a set of human images collected online, the present invention generated corresponding reconstruction meshes. Analysis of the results shows that the present method can accurately reconstruct the shape and structure of the human body in different movements, demonstrating high accuracy and realism.

Claims

1. A human body mesh reconstruction method based on Transformer, characterized in that: The following steps are involved: Step 1: Obtain a public human dataset and standardize the images in the human dataset; Step 2: For the images in the human body dataset, multiple view features are generated through the front view encoder and the view conversion network respectively. These features are then optimized and enhanced respectively. The optimized and enhanced features are then fused to generate multi-view aggregate features. Step 3: Combined with Transformer encoder optimization, the initial joint features are extracted from the images in the human body dataset to generate the final posture features; Step 4: Construct a mesh regression module, add the multi-view aggregation features and posture features, and input them into the mesh regression module to generate the final human body mesh reconstruction result.

2. The human body mesh reconstruction method based on Transformer according to claim 1, characterized in that: In step 1, the human body dataset contains full-body frontal images from different perspectives. All images are annotated with true GT information, and the annotation information includes skeletal key points and joint information of the human body. The standardization processing includes size adjustment, background filling and cropping.

3. The human body mesh reconstruction method based on Transformer according to claim 2, characterized in that: The specific implementation process of step 2 is as follows: Step 2.1: The input full-body frontal image is first passed through the backbone network of the HRNet model to extract the initial feature representation; Step 2.2: The initial features are then fed into two processing branches in parallel: the front view modeling branch and the perspective conversion branch. In the front view modeling branch, the initial features are fed into the Transformer-based front view encoder to generate a structured feature representation with contextual semantic information. This feature representation is then passed as guidance information to the front view guidance module FOB for optimization. In the perspective conversion branch, the initial features are fed into the perspective conversion network to obtain the feature maps F at each perspective. i ,; Step 2.3: Feature maps F from multiple perspectives i , including left view, right view and rear view, will be fed into three independent side view decoders, each side view decoder corresponds to one direction; the three side view decoders use a Transformer-based decoder structure to model features under different perspectives; each direction feature is used as input, combined with the guidance feature from the front view encoder as the query, and the cross-region context dependency is captured through the cross-attention mechanism. Each decoder outputs the enhanced feature representation X of its corresponding direction. i and pass it as input to the multi-view optimization module MIOB for feature fusion; Step 2.4: Output the optimized feature X′ from FOB front View-aware features of MIOB output By element-by-element summation, we can obtain the enhanced multi-view aggregation feature X fusion .

4. The method for human body mesh reconstruction based on Transformer according to claim 3, characterized in that: In the front view modeling branch, the initial features are input into the Transformer-based front view encoder, which uses a multi-head self-attention mechanism to perform global modeling and structural perception on the front view features, generating a structured feature representation with contextual semantic information, denoted as X front ,This feature representation will be passed as guidance information to the front view guidance module FOB for optimization; In the perspective conversion branch, the initial features are input to the perspective conversion network. The perspective conversion network integrates a normal prediction module NPN to generate a unit normal map. The unit normal is dot-producted with the predefined perspective direction vector pixel by pixel. The dot product result is used as a binary perspective mask matrix. The binary perspective mask matrix is ​​dot-multiplied with the unit normal map to obtain the visible area feature. The visible area feature is then convolution-enhanced and up-sampled to obtain the feature map F at each perspective. i ∈R H×W×C , where i includes three directions: left view, right view and rear view.

5. The method for human body mesh reconstruction based on Transformer according to claim 4, characterized in that: The FOB is specifically implemented as follows: first receiving the front view feature X front The spatial dimension is expanded by Flatten operation, layer normalization is applied, and then the multi-head self-attention mechanism MHA is applied to model the global dependency of the front view features, and the structural expression is enhanced by the multi-layer perceptron MLP. The FOB output result is then deformed and reshaped for subsequent enhancement of feature representation X. i Alignment and fusion in the spatial dimension to obtain the optimized front view feature X′ front ; The MIOB is specifically implemented as follows: MIOB uses FOB output as a guide query to optimize each direction feature separately; each MIOB branch receives the corresponding X i As the key and value, Flatten, layer normalization, cross attention mechanism MCA, re-normalization and MLP are performed in sequence, and then the channel dimension adjustment and feature compression are performed through 1×1 convolution operation. The final output is the perspective perception feature representation after structure enhancement, which is recorded as 6. The method for human body mesh reconstruction based on Transformer according to claim 5, characterized in that: The normal prediction module NPN is specifically implemented as follows: a network structure consisting of four convolutional layers and one fully connected layer is adopted. The convolutional layer is used to gradually extract local spatial geometric information and is combined with the ReLU activation function to enhance the nonlinear expression ability; finally, the feature of each pixel position is mapped into a three-dimensional vector through the fully connected layer, and after normalization processing, a unit normal map is output. The unit normal map maintains the same spatial resolution as the input map. Each pixel corresponds to a normalized three-dimensional normal direction, providing local surface orientation information and giving the network the ability to perceive the target geometric structure.

7. The method for human body mesh reconstruction based on Transformer according to claim 6, characterized in that: The specific implementation process of step 3 is as follows: A parallel posture optimization path is introduced to enhance the structural modeling capability between key joints. The posture optimization path starts with the input image. First, the initial joint feature representation is obtained through the joint extraction module, and a random occlusion mask is constructed. The mask and the initial joint feature representation are fused through element-by-element multiplication to simulate the situation where some joints are invisible and perform occlusion simulation on the joint features. Then, position encoding is introduced to embed the position information of each joint in the spatial structure to assist the network in modeling the relative spatial relationship between key points. The joint features that have been fused with occlusion information and position encoding are then input to the joint optimization module. The joint optimization module first flattens the joint features from the spatial structure to a sequence form through the Flatten operation, and then applies the regularization LayerNorm to stabilize the feature distribution. Next, the features are fed into the posture encoder, which consists of a multi-layer stacked Transformer encoder. The multi-head self-attention mechanism is used to capture global context dependencies. The output of the Transformer encoder is passed to the MLP layer for feature conversion and mapping to generate the final posture feature X. pose .

8. The method for human body mesh reconstruction based on Transformer according to claim 7, characterized in that: The specific implementation process of step 4 is as follows: First, the aggregated multi-view features and the optimized posture features are added element by element to fuse the global information from different viewpoints and the detailed features of the posture. The fused features are input into the grid regression module, which uses a fully connected layer or a convolution-based structure to gradually map the fused features into vertex positions in a three-dimensional coordinate system; and constructs a loss function to optimize accuracy.

9. The method for human body mesh reconstruction based on Transformer according to claim 8, characterized in that: The specific implementation process of constructing the loss function is as follows: Four key loss terms are used: vertex loss 3D joint loss 2D joint loss and heatmap loss Vertex Loss By measuring the Euclidean distance between the generated mesh vertices and the real vertices, the shape of the mesh is ensured to be consistent with the real data; 3D joint loss By calculating the distance between the predicted three-dimensional joint and the real joint, the spatial position of the joint is accurately optimized; the two-dimensional joint loss By projecting the 3D joints onto the 2D plane, the position of the 2D joints is optimized to enhance the robustness under multiple views; heat map loss The key point positioning accuracy is improved through pixel-level perspective support; the loss formula is as follows: Where V n represents the predicted vertex, represents the true value of the vertex, and N is the total number of vertices; in represents the estimated joint position in 3D space, represents the actual joint position, represents the predicted joint position, K represents the total number of joints; in Indicates that the estimated joint positions in three-dimensional space are mapped to the estimated joint positions in two-dimensional space, Represents the real 2D joint position, represents the predicted 2D space joint position, K represents the total number of joints; in Represents the true joint heat map, H n Represents the predicted joint heatmap.

Citation Information

Cited By

  • Human body grid construction method, device and equipment for shielded large scene and medium

    CN120877334A

  • Human mesh construction method and device for large scene with occlusion, equipment and medium

    CN120877334B

  • 3D human body posture and grid reconstruction method and system

    CN121616789A