Monocular dressing human body reconstruction method for perspective distortion image
Through virtual camera perspective transformation and multi-scale attention enhancement network, combined with pseudo-multi-view feature fusion, the problem of human body reconstruction in perspective distorted images is solved, and high-precision clothing body reconstruction is achieved.
Patent Information
- Application Number
- CN202510682870.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-08-19
AI Technical Summary
The existing methods have problems of nonlinear geometric distortion, depth information loss and occlusion ambiguity in the reconstruction of the monocular clothing body of the perspective image, making it difficult to achieve high-precision reconstruction of the clothing body.
Using fusion distortion correction, three-dimensional geometric representation and pseudo-multi-view constraints, a multi-scale attention enhancement network is designed, and a high-precision human body reconstruction model is generated by combining Fourier transform and pseudo-multi-view feature fusion.
It effectively alleviates the impact of perspective distortion, improves the geometric accuracy and detail fidelity of three-dimensional human body reconstruction, and improves the reconstruction quality under perspective distortion images.
Smart Images

Figure CN120510299A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of three-dimensional vision technology, and in particular relates to a monocular clothed human body reconstruction method for perspective-distorted images. Background Art
[0002] With the rapid development of 3D modeling and virtual reality technologies, demand for their applications in medical visualization, digital twins, and other fields is growing. 3D human reconstruction based on monocular perspective images has become a cutting-edge topic in computer vision and graphics due to its significant advantages in cost and convenience. In this context, recovering a geometrically sound and detailed human model from a single perspective image has become a key challenge in achieving high-fidelity digital human construction and large-scale application. However, 3D human reconstruction from monocular perspective images faces significant technical bottlenecks. First, the perspective projection mechanism introduces nonlinear geometric distortion. The core characteristic of perspective projection is that objects appear larger when closer to the camera than when projected onto the image plane. This means that objects projected onto the image plane larger when closer to the camera (such as with a mobile phone selfie) appear larger when the hands are closer to the camera than the torso, despite the fact that the hand-to-toss ratio in real 3D space is normal. This nonlinear distortion can lead to distorted body proportions when inferring 3D models from 2D images. Second, the lack of depth information and occlusion ambiguity exacerbate the reconstruction challenge. Common orthogonal projections or weak perspective projections are usually scaled to average depth information, which is effective when shooting at a distance, but cannot handle perspective distortion when shooting at close range. Although parametric models and deep learning techniques have improved reconstruction efficiency, mainstream methods still have two major limitations: First, most studies on human body reconstruction from perspective images focus on reconstruction of naked bodies, and lack the ability to restore details such as wrinkles and textures in clothed states; second, the end-to-end framework has difficulty in explicitly modeling the coupling relationship between perspective distortion and human body geometry, resulting in limited reconstruction quality of the model in perspective scenes. To address the problems of limb proportion distortion, missing details, and occlusion ambiguity caused by nonlinear projection distortion in monocular perspective images, the present invention proposes a clothed human body reconstruction method that integrates distortion correction, three-dimensional geometric representation, and pseudo-multi-view constraints to achieve high-precision clothed human body reconstruction in perspective-distorted images.
[0003] In recent years, research on reconstructing a 3D human body from a single image has gradually become a hot topic. Based on the different reconstruction methods, existing research can be mainly divided into two categories: explicit reconstruction and implicit reconstruction. Explicit surface-based 3D reconstruction directly defines the topological structure and spatial coordinates of the surface geometry through a parameterized human body model, decoupling the human body shape and posture into low-dimensional interpretable parameters. The SMPL model (Skinned Multi-Person Linear Model) is a typical example. It decomposes the human body geometry into posture parameters (joint rotations) and shape parameters (body proportions) through linear skinning animation, and superimposes vertex offsets to simulate clothing deformation. For example, Huang et al. (Huang Z, Xu Y, Lassner C, et al. ARCH: Animatable Reconstruction of Clothed Humans [C]. In Proc. IEEE / CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2020.) introduced a hierarchical network structure to jointly optimize human posture parameters and clothing deformation parameters. Through a hierarchical modeling strategy, they decoupled human posture, shape, and clothing deformation into multiple subtasks, optimized them separately, and then globally integrated them. This design not only captures the local impact of human joint motion on clothing deformation, but also maintains the overall smoothness and continuity of the clothing through a global optimization mechanism. Saito et al. (Saito S, Yang J, MaQ, et al. SCANimate: Weakly Supervised Learning of Skinned Clothed AvatarNetworks[C]. In Proc. IEEE / CVF Conf. Computer Vision and Pattern Recognition(CVPR), June 2021.) combined motion capture data with physical simulation technology to incorporate cloth dynamics into the prediction of vertex offsets, significantly improving the naturalness of clothing. This dynamic modeling capability enables SCANimate to generate more realistic clothing deformation effects, such as the fluttering trajectory of a long skirt during a rotation or the swaying effect of a loose coat while running. At the same time, a weakly supervised learning framework was introduced to guide the network to learn complex cloth deformation patterns through a small amount of labeled data, thereby reducing data dependence while maintaining high-precision reconstruction effects. Point cloud or voxel-based methods directly predict discrete three-dimensional points or voxel grids from images, partially alleviating the limitations of fixed models through topology-independent representations. Unlike parametric models, these methods do not rely on predefined human body templates, but directly represent 3D geometry through point clouds or voxels, which can handle more complex shape variations.For example, Charles R. Qi et al. (Qi CR, Yi L, Su H, et al. PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space[J]. Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2017.) employed a hierarchical point cloud feature learning framework to capture both local geometric details and global structural information, demonstrating strong robustness when dealing with occluded regions. However, such methods typically require significant computational resources and face challenges in generating smooth surfaces. The advantage of explicit methods lies in the strong constraints of anatomical priors: the standard topology provided by SMPL avoids reconstruction distortions. However, their limitations are also significant. First, fixed topology makes it difficult to model the complex deformations of loose clothing. For example, the swaying of a long skirt requires non-rigid deformation modeling, but explicit models, with their fixed number of vertices and connectivity, cannot flexibly capture dynamic wrinkles. Second, capturing high-frequency details relies on precise prediction of vertex offsets, but limited by the input image resolution, local details are easily smoothed or lost. Implicit surface-based 3D reconstruction techniques implicitly define the human body surface geometry through continuous functions, breaking through the topological limitations of traditional explicit modeling and enabling direct representation of clothed human forms of arbitrary complexity. These methods learn surface distribution characteristics in 3D space through implicit geometric fields. Their core advantage lies in modeling high-frequency geometric details and nonlinear deformations in the form of continuous functions, while avoiding the drawbacks of fixed mesh resolution in explicit parameterized models. Pixel-aligned implicit function methods, such as Saito et al. (Saito S, Simon T, Saragih J, et al. PIFuHD: Multi-level pixel-aligned implicit function for high-resolution 3D human digitization [C]. In Proc. IEEE / CVFConf. Computer Vision and Pattern Recognition (CVPR), 2020.), map high-resolution images to 3D space through a multi-scale feature fusion mechanism, achieving high-precision human reconstruction. The key lies in designing a pixel-level aligned implicit field encoder that associates each 3D point with the semantic and geometric features of a local region in the image.However, purely implicit methods are prone to limb topology errors (such as finger adhesion or foot penetration) in the absence of explicit human priors. To this end, Xiu et al. (Xiu Y, Yang J, Tzionas ea, Dimitrios. ICON: Implicit Clothed Humans Obtained from Normals[C]. In Proc. IEEE / CVFConf. Computer Vision and Pattern Recognition (CVPR), 2022.) proposed to jointly optimize the anatomical constraints of the parameterized human model with the implicit field, and drive the implicit surface deformation through the explicit template. This not only retains the flexibility of implicit modeling, but also avoids non-physically reasonable phenomena such as limb breakage. In particular, it can effectively maintain the kinematic consistency of human joints when dealing with loose clothing. Xiu et al. (Xiu Y, Yang J, Cao X, et al. ECON: Explicit Clothed Humans Optimized via Normal Integration [C]. In Proc. IEEE / CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2023: 2-6.) integrated an explicit parameterized model with an implicit detail completion network to restore local deformation of clothing while maintaining the rationality of the main structure. The advantages of implicit methods lie in their topology independence and detail fidelity, but they face three major challenges: First, inference requires dense sampling of spatial points, which is computationally expensive. For example, a single reconstruction requires evaluating the occupancy probability of millions of spatial points, making real-time applications difficult. Second, they rely heavily on large-scale 3D scan data for training, while real human scan data is scarce and the annotation cost is high. Existing methods mostly rely on synthetic data, resulting in insufficient generalization to real-world occlusions or low-resolution inputs. Third, the lack of anatomical prior constraints can easily lead to the generation of non-physically plausible geometries (such as limb fractures or joint dislocations) due to data bias. This is especially true when the input image has perspective distortion or occlusions; the implicit field may incorrectly infer the surface morphology of the occluded area. Although explicit and implicit methods have their own unique geometric representations, they both suffer from the following core blind spots:
[0004] (1) Simplified projection model assumptions: Existing methods generally assume that the input image uses orthogonal or weak perspective projection, ignore the real camera parameters, and simplify the depth scale to a single scaling factor, resulting in the inability to model perspective distortion (such as the near-large-far-small effect). When human body parts are distributed along the camera optical axis (such as arms extended forward), projection errors will cause limb proportion disproportion, such as abnormal enlargement of the palm size or distortion of the leg length, seriously affecting geometric rationality. (2) Dataset bias: Mainstream datasets are mostly based on synthetic data or rectified real images, using orthogonal or weak perspective projection, and lack perspective distortion samples. For example, the THuman2.0 dataset collects data through a multi-view camera array, but the default projection parameters assume long-distance shooting, which cannot cover the geometric distortion characteristics of close-range scenes, resulting in insufficient generalization of the model in real scenes. (3) Distortion amplification effect: When the input image has perspective distortion, explicit methods may amplify the local vertex offset due to projection errors, resulting in limb proportion distortion; implicit methods may generate surface convexity or topological errors due to abnormal implicit field density in the distorted area. Therefore, to overcome this bottleneck, it is necessary to explicitly integrate real camera parameters into the geometric representation, build a training framework that adapts to perspective scenes, and overcome the inherent problems of explicit and implicit representations through hybrid representation (explicit-implicit joint modeling). Perspective distortion is a geometric deformation phenomenon caused by the perspective projection characteristics of the camera imaging model.
[0005] In human mesh reconstruction, perspective distortion can cause the following problems:
[0006] 1) Proportional distortion: Errors in estimating the length of body parts 2) Joint position deviation: Due to projection ambiguity, the same 2D key point may correspond to multiple 3D postures. The current technical routes for solving perspective distortion can be divided into two categories: explicit correction methods based on geometric calibration and adaptive correction methods based on deep learning. The geometric calibration method explicitly corrects perspective distortion by accurately estimating camera parameters (such as focal length, principal point, and distortion coefficient). Its research mainly focuses on camera calibration algorithms and projection model optimization. Zhang Zhengyou's calibration method (Zhang Z.A Flexible New Technique for Camera Calibration[J].IEEE Transactions on Pattern Analysis and Machine Intelligence, 2000, 22(11):1330-1334.) is a milestone work in the field of camera calibration. It calculates the camera intrinsic parameter matrix and distortion parameters by using the pixel coordinates and physical coordinates of the checkerboard corner points, laying a theoretical foundation for subsequent research. On this basis, ( J, Silvén O. A Four-step Camera Calibration Procedure with Implicit Image Correction [J]. Proc. IEEE / CVF Conf. Computer Vision and Pattern Recognition (CVPR), 1997: 1106-1112.) et al. proposed an improved calibration method that improves calibration accuracy, particularly robustness in large distortion scenarios, by introducing nonlinear optimization techniques. Although geometric calibration methods offer high accuracy in camera parameter estimation, their reliance on calibration plates and scene assumptions limits their application in monocular image reconstruction. Furthermore, deep learning technology offers a new approach to perspective distortion correction, implicitly learning the mapping between camera parameters and human body geometry through a data-driven approach. Wang et al. (WANG W, GE Y, MEI H, et al. Zolly: Zoom focal length correctly for perspective-distorted human mesh reconstruction supplementary material[C] / / Proceedings of the IEEE International Conference on Computer Vision(ICCV),2023) proposed a dynamic focal length optimization module that infers focal length by analyzing human anatomical proportions, significantly reducing reconstruction scale error. By introducing a human scale prior, this method can adaptively adjust focal length parameters without relying on a calibration plate, making it particularly suitable for dynamic scene reconstruction using a monocular camera. Kocabas et al. (Kocabas M, Huang C-H, Hilliges O, et al. PARE: Part Attention Regressor for 3D Human Body Estimation[C]. In IEEE / CVF International Conference on Computer Vision(ICCV),2021.) designed an independent camera parameter prediction branch to generate physically plausible projection parameters through adversarial learning.Li et al. (LiM, Chen Z, Liu Y, et al. CLIFF: Carrying Location Information in Full Frames into Human Pose and Shape Estimation[C]. In European Conference on Computer Vision (ECCV), 2022.) proposed a focal length conditional feature fusion module, which dynamically adjusts the frequency domain response characteristics of the convolution kernel by encoding camera focal length information, enabling the network to perceive the perspective scaling effect at different shooting distances. Wei et al. (W. Jiang, K. M. Yi, G. Samei, O. Tuzel, and A. Ranjan, “Neuman: Neural human radiance field from a single video,” in Computer Vision-ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XXXII. Springer, 2022, pp. 402-418.) utilize neural radiance fields to jointly optimize the in-camera trajectory of participating human motion. Differentiable volume rendering is used to compensate for lens distortion and the nonlinear effects of perspective projection, achieving real-time distortion correction in mobile video. The coupled effects of perspective distortion and clothing deformation on the reconstruction of 3D clothed humans and their decoupling methods have become a research challenge in computer vision. Current neural parameterized models based on monocular images typically assume that projection distortion (caused by camera pose) and clothing deformation (caused by human motion or fabric properties) follow a linear superposition relationship in 2D image space. However, in real-world imaging, these two types of deformations couple through the nonlinear geometric transformation of perspective projection, leading to ill-posed optimization problems when estimating reconstruction model parameters. For example, localized fabric wrinkles caused by elbow flexion can produce similar 2D pixel displacement patterns under perspective projection as foreshortening of the limb caused by the camera's pitch angle. Because monocular methods lack multi-viewpoint geometric constraints and depth priors, it is difficult for the model to separate the physical generation mechanisms of the two deformations from a single observation, leading to ambiguity in pose estimation and clothing reconstruction.
[0007] In the current work of clothed human body reconstruction from perspective-distorted images, the following problems still need to be solved:
[0008] 1) The perspective projection mechanism leads to nonlinear geometric distortion. 2) Depth information is missing and occlusion ambiguity exists. 3) The conflict between perspective distortion processing and the details of clothing. Existing human reconstruction work from perspective images focuses on recovering the human pose and basic shape, while ignoring the complex geometry and appearance changes introduced by clothing. These three challenges reveal the bottlenecks of current methods from the perspective of the underlying assumptions of geometric modeling, the inherent limitations of monocular data, and the trade-off between detail preservation and global correction.
[0009] Therefore, to address the above problems, the present invention proposes a monocular clothed human body reconstruction method that integrates distortion correction, three-dimensional geometric representation and pseudo multi-view constraints. Summary of the Invention
[0010] (1) Technical problems to be solved by the present invention:
[0011] The purpose of this invention is to address the challenges faced by existing methods in the task of monocularly reconstructing a clothed human body from perspective images, such as the inherent characteristics of perspective cameras where objects appear larger near and smaller far away, as well as depth ambiguity, which lead to unsatisfactory reconstruction results. This paper proposes a 3D human body reconstruction framework that integrates distortion correction, 3D geometric representation, and pseudo multi-view constraints.
[0012] (2) In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:
[0013] A monocular clothed human body reconstruction method for perspective-distorted images includes the following steps:
[0014] S1, performing average block division and perspective distortion correction processing on the input image to generate a processed image block set;
[0015] S2. Design a multi-scale attention enhancement network to process the image blocks generated in S1 and complete the feature extraction of the image blocks;
[0016] S3, linearly map the vertex coordinates of the SMPL human body model and combine it with Fourier transform analysis to generate a fusion feature representation containing visual appearance and 3D geometric information;
[0017] S4, based on the virtual camera optical center position and differentiable sampling, pseudo multi-view features are generated, and multi-view feature fusion is achieved by combining distance weighted processing;
[0018] S5. Generate spatial sampling points through a hybrid sampling strategy, use implicit functions to predict the signed distance values of spatial points, and combine with optimizer training to finally generate a high-precision human body reconstruction model.
[0019] Preferably, the S1 specifically includes the following contents:
[0020] S1.1. Divide the input image into p×p non-overlapping sub-blocks. Calculate the coordinate position of the center point of each sub-block in the original image, which is recorded as the optical center position. The specific calculation formula is as follows:
[0021]
[0022] Among them, x ij and y ij Respectively represent the horizontal and vertical coordinates of the optical center position; W represents the width of the image; H represents the height of the image; i is the row index, j is the column index, which are 0, 1, 2, ..., p / 2 respectively;
[0023] S1.2. Define a virtual camera. Generate the virtual camera internal parameters corresponding to the real camera by constructing an affine transformation matrix. Use the virtual camera internal parameters to reproject the image block. Calculate the camera rotation matrix based on the offset of the virtual camera center relative to the original image center. The specific calculation formula is as follows:
[0024]
[0025] Among them, K represents the intrinsic parameter matrix of the real camera; K virt Represents the intrinsic parameter matrix of the virtual camera; Represents the rotation matrix from the center of the real camera to a certain viewing angle of the virtual camera; u is the horizontal coordinate of the pixel; v is the vertical coordinate of the pixel, and s is the scaling factor, which is used to control the scaling degree of the virtual camera to the target area.
[0026] S1.3. After the calculation is completed, a set of image blocks after distortion correction is output.
[0027] Preferably, the S2 specifically includes the following contents:
[0028] S2.1. Use 3×3 convolution and batch normalization to extract initial features from the input image and generate a basic feature map.
[0029] S2.2, expand the receptive field by using 7×7 convolution with a stride of 2, and construct a multi-scale hierarchical feature pyramid by combining the improved convolutional block structure;
[0030] S2.3 stacks four cascaded hourglass processing units, each of which uses a symmetrical codec structure. In the encoding stage, a downsampling module with a 7×7 convolution kernel is used to extract deep semantic features, and in the decoding stage, a bilinear interpolation upsampling module is used to restore spatial details.
[0031] S2.4. Residual connections are introduced in the encoding and decoding process to alleviate gradient disappearance, and a channel attention module is embedded to adaptively calibrate feature channel weights and output features capable of capturing multi-scale information.
[0032] Preferably, the S3 specifically includes the following contents:
[0033] S3.1. Linearly map the mesh vertex coordinates of the SMPL human body model from the original model space to the normalized cube space to eliminate scale differences.
[0034] S3.2. Project the three-dimensional surface onto a two-dimensional grid plane, and simultaneously record the depth distribution information within the projection area;
[0035] S3.3. Perform Fourier series integration on the occupancy function of each pixel along the depth direction and calculate the first 32 order coefficients to construct frequency domain features.
[0036] S3.4. Integrate spatial-frequency domain information and output a multi-channel image with a resolution of 512×512 and a channel dimension of 32 to achieve lightweight representation of three-dimensional human body geometry.
[0037] Preferably, the S4 specifically includes the following contents:
[0038] S4.1. Assume that the camera intrinsic parameter matrix corresponding to the original input image is K0, divide the image plane evenly into M×M sub-regions, and use the center point of each sub-region as the virtual principal point coordinate. Recorded as the virtual camera optical center position, generating N = M 2 The corresponding virtual camera intrinsic parameter matrix {K i}, to simulate different viewing angles;
[0039] S4.2. Applying the virtual camera intrinsic parameters to the projected sampling points, where significant pixel displacement occurs in the near-field region due to the principal point offset.
[0040] S4.3. Calculate the initial fusion weight based on the distance between the projected sampling point and the image center. Perform weight normalization by analyzing the geometric distribution characteristics of the 3D projection points on the multi-view image plane, and dynamically calculate the contribution weight of each view feature.
[0041] S4.4. Realize multi-view feature fusion through dynamic weighted averaging.
[0042] Preferably, the S5 specifically includes the following contents:
[0043] S5.1. A hybrid sampling strategy is used to generate a large number of spatial sampling points in a standard cubic space. Based on the initial coarse reconstruction model and prior knowledge of the human body, a dynamic narrowband region is set outside the surface, and high-density sampling is performed within the dynamic narrowband region.
[0044] S5.2. Identify detail areas and increase the density of sampling points within a certain space around key points to enhance sampling.
[0045] S5.3. Predict the signed distance values of the sampling points through a neural network, and use the optimizer to train the model, ultimately outputting a three-dimensional human body model with reasonable human topology and clothing details.
[0046] (3) The beneficial effects of the present invention include the following:
[0047] (1) Aiming at the inherent distortion effect and depth ambiguity of perspective projection, the present invention proposes a three-dimensional reconstruction method for clothed human bodies that integrates perspective distortion correction. Most existing methods are based on orthogonal projection or weak perspective projection assumptions, and it is difficult to achieve accurate geometric reconstruction in scenes with significant perspective distortion. The present invention alleviates image distortion by using virtual camera perspective transformation. On this basis, the present invention designs a multi-scale attention enhancement network based on the hourglass structure, and uses the hourglass model and channel attention mechanism to realize multi-scale feature extraction, so that the model can adapt to perspective distorted images.
[0048] (2) Existing technologies are limited by the two-dimensional projection plane and fail to explicitly model three-dimensional geometric continuity. This paper proposes a three-dimensional human body reconstruction method that integrates distortion correction, three-dimensional geometric representation, and pseudo-multi-view constraints. First, a three-dimensional geometric representation is introduced to align the two-dimensional perspective correction features with the three-dimensional geometric representation, and efficient three-dimensional geometric modeling is achieved through frequency domain representation. Second, to address the information loss problem of single-view input, a pseudo-multi-view module is proposed. Through the optical center movement strategy and pseudo-multi-view allocation mechanism, the network's robustness to perspective distortion is enhanced.
[0049] (3) To address the problem of adhesion of human body parts in complex postures, a hybrid sampling strategy that integrates surface sampling, spatial sampling, and local enhancement is proposed. Based on the skeletal semantic information of the SMPL model, the vertices of parts such as the arms are guided to offset along the normal direction. By calculating the distance from the external spatial point to the human body surface as a supervision signal, it is ensured that the human body model generated by the implicit field conforms to the anatomical structure.
[0050] (4) The monocular clothed human reconstruction method for perspective-distorted images proposed in this paper has achieved optimal results on the public datasets THuman2.0 and CustomHumans. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flowchart of a monocular clothed human body reconstruction method for perspective-distorted images proposed by the present invention;
[0052] Figure 2 This is the network diagram of the multi-scale attention enhancement network based on the hourglass structure proposed by the present invention;
[0053] Figure 3This is a schematic diagram of the human body reconstruction results of the monocular clothed human body reconstruction method for perspective-distorted images proposed in the present invention on the THuman2.0 and CustomHumans test sets;
[0054] Figure 4 This is a comparison chart of the results of single-person reconstruction on perspective-distorted images using the method proposed in the present invention and the prior art. DETAILED DESCRIPTION
[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0056] This paper proposes a monocular clothed human reconstruction method for perspective-distorted images. This method decouples the correlation between image position and perspective distortion through virtual viewpoint transformation, providing accurate image features for subsequent reconstruction. Furthermore, to more accurately model the continuous structure of 3D shapes, the method introduces the Skinned Multi-Person Linear Model_Fourier Occupancy Field (SMPL_FOF) representation of the SMPL human model to characterize 3D geometry. Furthermore, to address the issue of missing input information for single images, a pseudo-multi-view module is proposed. By simulating multi-view data augmentation, the network learns a feature representation that is robust to perspective distortion. To address the problem of body part adhesion in complex poses, a hybrid sampling strategy is proposed that combines surface sampling, spatial sampling, and local enhancement. Based on the skeletal semantics of the SMPL model, vertices of parts such as the arms are offset along the normal direction. The distance from external spatial points to the human surface is calculated as a supervisory signal to ensure that the human model generated by the implicit field conforms to the anatomical structure. To validate the effectiveness of the designed module, extensive comparative and ablation experiments were conducted on the THuman2.0 and CustomHuman datasets. Experimental results demonstrate that the proposed algorithm exhibits significant advantages in both reconstruction accuracy and robustness. The following describes the proposed monocular clothed human reconstruction method for perspective-distorted images, using accompanying figures and specific examples. The details are as follows.
[0057] Example 1:
[0058] The present invention proposes a monocular clothed human body reconstruction method for perspective-distorted images, comprising the following steps:
[0059] S1. Perform perspective distortion correction on the input 512×512 resolution RGBA format image, evenly divide the image into multiple square sub-blocks of equal size, and calculate the coordinates of the center point of each sub-block as the optical center position. Generate the corresponding virtual camera internal parameters by constructing an affine transformation matrix containing image scaling, position translation, and tilt transformation parameters. Use the virtual camera parameters to reproject the image block, where the camera rotation matrix is calculated based on the offset of the virtual optical center relative to the center of the original image. Finally, output a set of image blocks after distortion correction, which specifically includes the following contents:
[0060] S101. Evenly divide the image block into p×p non-overlapping sub-blocks. Determine the coordinate position of the center point of each sub-block in the original image, that is, the optical center position is:
[0061]
[0062] Among them, x ij and y ij Respectively represent the horizontal and vertical coordinates of the optical center position; W represents the width of the image; H represents the height of the image; i is the row index, j is the column index, which are 0, 1, 2, ..., p / 2 respectively;
[0063] S102. A virtual camera is defined, and the optical center of the virtual camera corresponds to the region of interest. The mapping from the original image to the cropped region is implemented using the following formula:
[0064]
[0065] Among them, K represents the intrinsic parameter matrix of the real camera; K virt Represents the intrinsic parameter matrix of the virtual camera; Represents the rotation matrix from the center of the real camera to a certain viewing angle of the virtual camera; u is the horizontal coordinate of the pixel; v is the vertical coordinate of the pixel, and s is the scaling factor, which is used to control the scaling degree of the virtual camera to the target area.
[0066] S2. During the feature extraction phase, a multi-scale attention-enhanced network is used. This network employs a four-stage hourglass structure, with each stage consisting of a 7×7 convolutional kernel downsampling module and a bilinear interpolation upsampling module. A channel-wise attention mechanism is incorporated into the network decoding process to dynamically adjust the weight coefficients of each feature channel. By integrating the global average features and attention weights through a cross-block interaction module, the final output features are capable of capturing multi-scale information, specifically including the following:
[0067] S201, use 3×3 convolution and batch normalization to perform initial feature extraction on the input image and generate a 64-channel basic feature map;
[0068] S202, expand the receptive field through 7×7 convolution with a stride of 2, and build a multi-scale hierarchical feature pyramid with an improved convolutional block structure;
[0069] S203, stacking four cascaded hourglass processing units, each unit adopts a symmetrical encoding and decoding structure, extracting deep semantic features through downsampling in the encoding stage, and restoring spatial details through upsampling in the decoding stage;
[0070] S204. Introduce residual connections in the encoding and decoding process to alleviate gradient disappearance, and embed a channel attention module to adaptively calibrate feature channel weights to enhance key feature expression.
[0071] S3. During the construction of the 3D geometric representation, the vertex coordinates of the SMPL human body model are first linearly mapped into a standardized cubic space. For each pixel unit, the 3D occupancy function is Fourier transformed along the depth direction, retaining the real part of the first 32 low-frequency coefficients. The generated frequency domain feature map is concatenated with the 2D image in the channel dimension to form a fused feature representation that combines both visual appearance and 3D geometric information. Specifically, it includes the following:
[0072] S301. Linearly map the SMPL mesh vertices from the original model space to the standardized cube space to eliminate scale differences.
[0073] S302, mapping the three-dimensional surface to a two-dimensional grid plane by projection, and simultaneously recording depth distribution information within the projection area;
[0074] S303 , performing Fourier series integration on the occupancy function of each pixel along the depth direction, calculating the first 32 order coefficients to construct frequency domain features.
[0075] S304 , integrating spatial-frequency domain information, outputting a multi-channel image with a resolution of 512×512 and a channel dimension of 32, and realizing a lightweight two-dimensional representation of three-dimensional human body geometry.
[0076] S4. In the pseudo multi-view feature fusion stage, a 4×4 uniformly distributed virtual camera optical center position is set on the imaging plane to generate the corresponding virtual camera intrinsic parameter matrix. The virtual camera intrinsic parameter matrix is applied to the projected two-dimensional sampling points. The initial fusion weight is calculated based on the distance between the sampling point and the image center. The weight is normalized based on the projection point visibility and surface normal consistency conditions. Finally, the effective fusion of multi-view features is achieved through weighted averaging. The specific contents include the following:
[0077] S401, assuming that the camera intrinsic parameter matrix corresponding to the original input image is K0, the image plane is evenly divided into M×M sub-regions, and the center point of each sub-region is used as the virtual principal point coordinate. Generate N=M 2The virtual camera intrinsic parameter matrix {K i}, and applies it to the sample points after projection. The near view area has significant pixel displacement due to the principal point offset, while the distant view area has less deformation.
[0078] S402: By analyzing the geometric distribution characteristics of the 3D projection points on the multi-view image plane, the contribution weight of each view feature is dynamically calculated to achieve robust feature fusion.
[0079] S5. During the 3D reconstruction optimization phase, a hybrid sampling strategy is used within the standard cubic space to generate a large number of spatial sampling points. 30% of these sampling points are concentrated near the human body surface, and triple-density sampling is used for detailed areas such as the hands. The signed distance values of the spatial points are predicted using an implicit function, and the L1 norm is used as the loss function for optimization. Using an optimizer with an initial learning rate of 0.001, after 15 training cycles, the final output is a 3D human body model with a reasonable human topology and clothing details, including the following:
[0080] S501. Based on the initial coarse reconstruction model and prior knowledge of the human body, a dynamic narrowband area is set outside the surface, and high-density sampling is performed in the area.
[0081] S502: Identify the key point area of the hand and increase the sampling density in the spherical space around the key point to three times that of other areas.
[0082] S503: predict the signed distance values of the sampling points through a neural network, and finally output a three-dimensional human body model with a reasonable human body topology and clothing details.
[0083] Example 2:
[0084] Based on Example 1, but with the following differences: Figure 1-4 The specific implementation process of the monocular clothed human body reconstruction method for perspective-distorted images proposed in this invention is as follows:
[0085] (1) Data processing:
[0086] This paper uses the THuman2.0 and CustomHumans datasets as benchmarks. The training set contains 381 THuman2.0 models, and the test set consists of 145 THuman2.0 models and all CustomHumans dataset models. Each model renders 256 images (evenly sampled along the y-axis). All images are rendered using a perspective camera and a precomputed radiosity renderer at a resolution of 512×512.
[0087] (2) Perspective distortion correction processing:
[0088] During the training process, the input image is first transformed using a virtual perspective to decouple the correlation between image position and perspective distortion, thereby alleviating image distortion. On this basis, a multi-scale attention enhancement network is designed, which uses the hourglass model and channel attention mechanism to extract block features for reconstructing a full-body human mesh with complete geometric information.
[0089] (3) Multi-scale Attention Enhancement Network Based on Hourglass Structure
[0090] The multi-scale attention enhancement network model architecture consists of three main components: image patch preprocessing, a cross-patch interaction module, and an hourglass module. The image patch preprocessing first uses a 3×3 convolution kernel with batch normalization to perform an initial feature transformation on the input image, generating a basic feature representation with 64 channels. This is followed by spatial downsampling through a 7×7 convolutional layer with a large receptive field. Combined with an improved convolutional block structure, a hierarchical feature pyramid with multi-scale representation capabilities is gradually constructed. To further enhance the network's ability to model complex features, the system stacks four cascaded hourglass processing units. Each hourglass unit adopts a symmetrical encoder-decoder structure. During the encoding phase, deep semantic features are extracted through continuous downsampling operations, while during the decoding phase, spatial details are restored through precise upsampling. A residual connection mechanism is introduced to alleviate the vanishing gradient problem, and a channel attention module is embedded to adaptively recalibrate the weight distribution of feature channels, thereby enhancing the representation of key features and suppressing redundant features. The result is a deep network architecture with powerful feature extraction and reconstruction capabilities.
[0091] (3) Pseudo-multi-view strategy:
[0092] The present invention proposes a pseudo-multi-view data generation method based on virtual camera principal point offset. By setting four symmetrically distributed virtual camera principal point coordinates (128,128), (128,384), (384,384) and (384,128) on the imaging plane, the imaging geometry changes under different perspectives are effectively simulated. This design causes the optical center of each virtual camera to produce different degrees of offset relative to the original image center (256,256), and naturally introduces different degrees of image distortion effects through perspective projection transformation. This method can not only efficiently simulate multi-view and multi-distortion data from a single input image, but also maintain geometric consistency between images, providing richer and more accurate geometric constraint information for subsequent feature extraction. Through this controllable virtual perspective generation mechanism, the present invention can enhance the model's adaptability to different perspectives and distortion conditions at the data level, thereby improving the robustness of feature representation under real complex imaging conditions.
[0093] like Figure 1Figure 2 shows the monocular clothed human reconstruction method for perspective-distorted images proposed in this paper. Block processing and virtual camera perspective transformation are used to eliminate uneven distortion and positional dependence between image blocks. The SMPL_FOF three-dimensional geometric representation is introduced to decompose the three-dimensional occupancy field into a two-dimensional Fourier series along the line of sight, and the image is aligned with the Fourier coefficients of each pixel in the SMPL_FOF. Pseudo-multi-view images with different perspective distortions are generated by virtually moving the camera's optical center, simulating the geometric constraints of real multi-view imaging and enhancing the network's robustness to perspective distortion.
[0094] like Figure 2 As shown in the figure, the specific implementation details of the multi-scale attention enhancement network based on the hourglass structure proposed in this invention are demonstrated. It restores spatial details through precise upsampling, introduces a residual connection mechanism to alleviate the gradient vanishing problem, and embeds a channel attention module to adaptively calibrate the weight distribution of feature channels, thereby achieving enhanced expression of key features and suppression of redundant features, and finally forming a deep network architecture with powerful feature extraction and reconstruction capabilities.
[0095] like Figure 3 As shown in the figure, the human body reconstruction results of the present invention on the THuman2.0 and CustomHumans test sets are demonstrated. The results fully demonstrate that the three-dimensional clothed human body reconstruction method under perspective images proposed by the present invention takes into account the distortion effect of the perspective image itself, ensuring the consistency of the distorted image features in the subsequent reconstruction network, thereby reconstructing a more accurate three-dimensional human body model.
[0096] like Figure 4 As shown, the qualitative results of the present invention are compared with the current mainstream single-person reconstruction method. It can be seen that the present invention is relatively robust to various postures, achieves fine reconstruction results in the visible area, and also reconstructs relatively reasonable results in the invisible area.
[0097] Table 1
[0098]
[0099] Table 1 lists the comparison of the quantitative results of the present invention and the current mainstream single-person reconstruction methods on the THuman2.0 dataset; ECON was proposed by Xiu et al. (Xiu Y, Yang J, Cao X, et al. ECON: Explicit Clothed humans Optimized via Normal integration [C]. In Proc. IEEE / CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2023: 2-6.) in 2023, SITH was proposed by Ho HI et al. (Ho HI, Song J, Hilliges O. SITH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion [C]. In Proc. IEEE / CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2024.) in 2024, and FOFX was proposed by Feng et al. (Feng Q, Liu Y, Lai YK, et al. FOF-X: Towards Real-time Detailed Human Reconstruction from a Single Image[C]. In ArXiv:2412.05961, 2024.) was proposed in 2024. The quantitative evaluation metrics used are chamfer distance, point-to-surface distance, and normal consistency, and are used to assess the accuracy of single-image human reconstruction. The dataset is the THuman2.0 dataset. As can be seen, the reconstruction results of our method achieve the best results compared to the current state-of-the-art methods.
[0100] Table 2
[0101]
[0102] Table 2 lists the comparison of the quantitative results of the present invention and the current mainstream multi-person reconstruction methods on the CustomHumans dataset; ECON was proposed by Xiu et al. (Xiu Y, Yang J, Cao X, et al. ECON: Explicit Clothed humans Optimized via Normal integration [C]. In Proc. IEEE / CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2023: 2-6.) in 2023, SITH was proposed by Ho HI et al. (Ho HI, Song J, Hilliges O. SiTH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion [C]. In Proc. IEEE / CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2024.) in 2024, and FOFX was proposed by Feng et al. (Feng Q, Liu Y, Lai YK, et al. FOF-X: Towards Real-time Detailed Human Reconstruction from a Single Image[C]. In ArXiv:2412.05961, 2024.) was proposed in 2024. The quantitative evaluation metrics used are chamfer distance, point-to-surface distance, and normal consistency, and are used to assess the accuracy of single-image human reconstruction. The dataset is the CustomHumans dataset. As can be seen, the reconstruction results of our method achieve the best results compared to the current state-of-the-art methods.
[0103] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes based on the technical solutions and improved concepts of the present invention within the technical scope disclosed by the present invention, and these changes should be covered by the scope of protection of the present invention.
Claims
1. A monocular clothed human reconstruction method for perspective-distorted images, characterized by: The steps include: S1. Perform block division and perspective distortion correction on the input image to generate a set of processed image blocks; S2. Design a multi-scale attention enhancement network to process the image blocks generated in S1 and complete the feature extraction of the image blocks; S3, linearly map the vertex coordinates of the SMPL model and combine it with Fourier transform analysis to generate a fusion feature representation containing visual appearance and 3D geometric information; S4, based on the virtual camera optical center position and differentiable sampling, pseudo multi-view features are generated, and multi-view feature fusion is achieved by combining distance weighted processing; S5. Generate spatial sampling points through a hybrid sampling strategy, use implicit functions to predict the signed distance values of spatial points, and combine with optimizer training to finally generate a high-precision human body reconstruction model.
2. The monocular clothed human body reconstruction method for perspective-distorted images according to claim 1, characterized in that: The S1 specifically includes the following contents: S1.
1. Divide the input image into p×p non-overlapping sub-blocks. Calculate the coordinate position of the center point of each sub-block in the original image, which is recorded as the optical center position. The specific calculation formula is as follows: Among them, x ij and y ij Respectively represent the horizontal and vertical coordinates of the optical center position; W represents the width of the image; H represents the height of the image; i is the row index, j is the column index, which are 0, 1, 2, ..., p / 2 respectively; S1.
2. Define a virtual camera. By constructing an affine transformation matrix, generate the virtual camera internal parameters corresponding to the real camera. Use the virtual camera internal parameters to reproject the image block. Calculate the camera rotation matrix based on the offset of the virtual camera center relative to the original image center. The specific calculation formula is as follows: Among them, K represents the intrinsic parameter matrix of the real camera; K virt Represents the intrinsic parameter matrix of the virtual camera; Represents the rotation matrix from the center of the real camera to a certain viewing angle of the virtual camera; u is the horizontal coordinate of the pixel; v is the vertical coordinate of the pixel, and s is the scaling factor, which is used to control the scaling degree of the virtual camera to the target area; S1.
3. After the calculation is completed, a set of image blocks after distortion correction is output.
3. The monocular clothed human body reconstruction method for perspective-distorted images according to claim 2, characterized in that: The S2 specifically includes the following contents: S2.
1. Use 3×3 convolution and batch normalization to extract initial features from the input image and generate a basic feature map. S2.2, expand the receptive field by using 7×7 convolution with a stride of 2, and construct a multi-scale hierarchical feature pyramid by combining the improved convolutional block structure; S2.3 stacks four cascaded hourglass processing units, each of which uses a symmetrical codec structure. In the encoding stage, a downsampling module with a 7×7 convolution kernel is used to extract deep semantic features, and in the decoding stage, a bilinear interpolation upsampling module is used to restore spatial details. S2.
4. Residual connections are introduced in the encoding and decoding process to alleviate gradient disappearance, and a channel attention module is embedded to adaptively calibrate feature channel weights and output features capable of capturing multi-scale information.
4. The monocular clothed human body reconstruction method for perspective-distorted images according to claim 3, characterized in that: The S3 specifically includes the following contents: S3.
1. Linearly map the mesh vertex coordinates of the SMPL human body model from the original model space to the normalized cube space to eliminate scale differences. S3.
2. Project the three-dimensional surface onto a two-dimensional grid plane, and simultaneously record the depth distribution information within the projection area; S3.
3. Perform Fourier series integration on the occupancy function of each pixel along the depth direction and calculate the first 32 order coefficients to construct frequency domain features. S3.
4. Integrate spatial-frequency domain information and output a multi-channel image with a resolution of 512×512 and a channel dimension of 32 to achieve lightweight representation of three-dimensional human body geometry.
5. The monocular clothed human body reconstruction method for perspective-distorted images according to claim 4, characterized in that: The S4 specifically includes the following contents: S4.
1. Assume that the camera intrinsic parameter matrix corresponding to the original input image is K0, divide the image plane evenly into M×M sub-regions, and use the center point of each sub-region as the virtual principal point coordinate. Recorded as the virtual camera optical center position, generating N = M 2 The corresponding virtual camera intrinsic parameter matrix {K i }, to simulate different viewing angles; S4.
2. Applying the virtual camera intrinsic parameters to the projected sampling points, where significant pixel displacement occurs in the near-field region due to the principal point offset. S4.
3. Calculate the initial fusion weight based on the distance between the projected sampling point and the image center. Perform weight normalization by analyzing the geometric distribution characteristics of the 3D projection points on the multi-view image plane, and dynamically calculate the contribution weight of each view feature. S4.
4. Realize multi-view feature fusion through dynamic weighted averaging.
6. The monocular clothed human body reconstruction method for perspective-distorted images according to claim 5, characterized in that: The S5 specifically includes the following contents: S5.
1. A hybrid sampling strategy is used to generate a large number of spatial sampling points in a standard cubic space. Based on the initial coarse reconstruction model and prior knowledge of the human body, a dynamic narrowband region is set outside the surface, and high-density sampling is performed within the dynamic narrowband region. S5.
2. Identify detail areas and increase the density of sampling points within a certain space around key points to enhance sampling. S5.
3. Predict the signed distance values of the sampling points through a neural network, and use the optimizer to train the model, ultimately outputting a three-dimensional human body model with reasonable human topology and clothing details.
Citation Information
Cited By
Method and system for reconstructing an animatable three-dimensional human body model based on a single occluded image
CN122492946A