A garment three-dimensional reconstruction method based on geometric prior and generated image assistance

By combining real and synthetic image datasets, extracting multimodal prior information and optimizing model parameters, the dependence of existing 3D clothing reconstruction methods on high-cost 3D scanning data is resolved, achieving higher accuracy and better generalization in 3D clothing reconstruction.

CN122391482APending Publication Date: 2026-07-14CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV OF POSTS & TELECOMM
Filing Date
2026-04-14
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing 3D clothing reconstruction methods rely on high-cost 3D scanning data, which limits the model's generalization ability and robustness. Furthermore, the training and inference costs are high, making it difficult to effectively handle complex clothing topological changes and large pose variations.

Method used

A method based on geometric priors and generated images is adopted. By constructing a mixed dataset of real and synthetic images, multimodal prior information is extracted. Combined with a clothing generation network and a geometric enhancement module, the model parameters are optimized using a loss function to achieve 3D clothing reconstruction.

Benefits of technology

Without relying on additional 3D scanning data, the model's ability to represent complex clothing shapes and its reconstruction accuracy have been improved, thus enhancing its generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391482A_ABST
    Figure CN122391482A_ABST
Patent Text Reader

Abstract

The present application belongs to the field of augmented reality, computer vision, computer graphics, and relates to a kind of garment three-dimensional reconstruction method based on geometric prior and generated image auxiliary, comprising: obtaining image, inputting image into trained garment three-dimensional reconstruction model, and obtaining detailed three-dimensional garment grid;Garment three-dimensional reconstruction model includes: multi-modal prior information extraction module, garment generation network and geometry enhancement module;The present application extracts prior information that can reflect geometric structure from synthetic image, i.e.semantic segmentation map and surface normal map, which is jointly trained with real data set, and combined with real data set to construct garment statistical model, and according to garment statistical model, principal component coefficient is restored to garment grid, to make up for the problem of insufficient three-dimensional supervision signal, without relying on additional three-dimensional scanning grid data, improve the representation ability of model to complex garment shape, thereby improve the precision and generalization performance of three-dimensional garment reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of augmented reality, computer vision, and computer graphics, and relates to a method for 3D reconstruction of clothing based on geometric priors and generated image assistance. Background Technology

[0002] 3D clothing reconstruction based on a single image aims to recover a 3D clothing mesh model that matches the appearance of the input image from a single 2D image. This task has significant application value in fields such as virtual try-on, digital human modeling, and computer vision. Existing technologies can be mainly divided into methods based on implicit representation and methods based on explicit representation, depending on the different 3D representation methods.

[0003] Implicit representation-based methods typically utilize continuous functions such as occupancy fields or signed distance fields (SDFs) to represent 3D geometry. In 3D clothing reconstruction tasks, these methods usually rely on relatively accurate prior information about human geometry and require large-scale, high-quality 3D scan mesh data as supervision. However, the high cost and limited scale of 3D scan data acquisition restrict the generalization ability and robustness of the models in practical applications, while also incurring significant training and inference costs. Nevertheless, implicit representation methods have certain advantages in depicting geometric details and can handle complex clothing structures well.

[0004] Explicit representation-based methods directly construct 3D models using discrete geometric representations (such as mesh templates or 3D Gaussian representations). These methods typically use a predefined template as a basis, predicting the deformation parameters of each vertex or representation unit to recover the garment shape consistent with the input image. Compared to implicit methods, explicit representation methods are more structurally intuitive and generally exhibit better stability and robustness due to the introduction of geometric topological constraints. However, limited by the expressive power of templates, these methods have certain limitations when handling complex garment topological changes or significant pose variations.

[0005] In addition, both of the above methods usually rely on 3D scanned mesh data for training. Due to the high cost of acquiring this type of data and the complexity of the production process, its data size is much smaller than that of 2D image data, which to some extent restricts the further improvement of model performance and its application in real-world scenarios. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention employs a clothing 3D reconstruction method based on geometric priors and generated images, comprising: acquiring an image; inputting the image into a trained clothing 3D reconstruction model to obtain a detailed 3D clothing mesh; the clothing 3D reconstruction model includes: a multimodal prior information extraction module, a clothing generation network, and a geometric enhancement module; the training process of the clothing 3D reconstruction model includes:

[0007] S1. Obtain the real dataset; the real dataset includes 3D scanned clothing meshes of various clothing categories and their rendered real images; construct statistical models of various clothing categories based on the 3D scanned clothing meshes of various clothing categories;

[0008] S2. Construct a synthetic image dataset based on a text-driven image generation model; combine the real dataset and the synthetic image dataset to obtain a hybrid image dataset;

[0009] S3. Input the images from the mixed image dataset into the multimodal prior information extraction module to obtain the multimodal prior information of the images; the multimodal prior information includes: SMPL parameters, camera parameters, normal map and semantic segmentation map;

[0010] S4. Input the image and its SMPL parameters into the clothing generation network to obtain the principal component coefficients of the image and the predicted clothing category;

[0011] S5. Select the clothing statistical model corresponding to the clothing category predicted from the image. Based on the principal component coefficients of the image and the clothing statistical model Construct the initial 3D clothing mesh of the image;

[0012] S6. Combine the image and its principal component coefficients with the clothing statistical model. The initial 3D clothing mesh, SMPL parameters, camera parameters, and semantic segmentation map are input into the geometry enhancement module to obtain a detailed 3D clothing mesh.

[0013] S7. Calculate the loss function value based on the principal component coefficients of the image, the predicted clothing category, the initial 3D clothing mesh, the detailed 3D clothing mesh, the normal map, and the semantic segmentation map. Update the parameters of the clothing 3D reconstruction model based on the loss function value. When the loss function value is minimized, the trained clothing 3D reconstruction model is obtained.

[0014] Beneficial effects:

[0015] This invention proposes a 3D clothing reconstruction method using synthetic images without real 3D annotations for training. This method extracts prior information reflecting the geometric structure (i.e., semantic segmentation map and surface normal map) from the synthetic image, trains it jointly with a real dataset, and constructs a clothing statistical model based on the real dataset. The principal component coefficients are then restored to clothing mesh based on the clothing statistical model to compensate for the lack of 3D supervision signals. Without relying on additional 3D scan mesh data, this method improves the model's ability to represent complex clothing shapes, thereby enhancing the accuracy and generalization performance of 3D clothing reconstruction. Attached Figure Description

[0016] Figure 1 A flowchart of a method for 3D reconstruction of clothing based on geometric priors and generated image assistance provided in an embodiment of the present invention;

[0017] Figure 2 A flowchart illustrating the generation process of a synthetic dataset provided in this embodiment of the invention;

[0018] Figure 3 This is a diagram of the network structure for clothing generation provided in an embodiment of the present invention.

[0019] Figure 4 This is a comparative schematic diagram of the present invention and other methods provided for embodiments of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] like Figure 1 As shown, this embodiment of the invention employs a clothing 3D reconstruction method based on geometric priors and generated image assistance, including: acquiring an image, inputting the image into a trained clothing 3D reconstruction model to obtain a detailed 3D clothing mesh; the clothing 3D reconstruction model includes: a multimodal prior information extraction module, a clothing generation network, and a geometric enhancement module; the training process of the clothing 3D reconstruction model includes:

[0022] S1. Obtain the real dataset; the real dataset includes 3D scanned clothing meshes of various clothing categories and their rendered real images; construct statistical models of various clothing categories based on the 3D scanned clothing meshes of various clothing categories;

[0023] The real datasets are CTD, SIZER, and Deep3DGarment datasets, which include real 3D scanned clothing meshes and their rendered real images.

[0024] Statistical models for various clothing categories include:

[0025] S11. Obtain clothing templates for various clothing categories. Using the registration method, align the 3D scanned clothing meshes of various clothing categories in the real dataset to the clothing templates of the corresponding clothing categories to obtain 3D registered clothing meshes for various clothing categories.

[0026] The clothing templates are D3GNet clothing templates, including T-shirts, shirts, pants, and shorts, with 1954, 2468, 1180, and 678 vertices respectively. The 3D scanned clothing meshes of each category in the real image datasets (CTD, SIZER, and Deep3DGarment datasets) are aligned to the corresponding templates in the D3GNet clothing templates through non-rigid deformation using the D3GNet registration method, resulting in 3D registered clothing meshes for each clothing category.

[0027] S12. Principal component analysis is performed on the 3D registered clothing meshes for various clothing categories to obtain statistical models for each category of clothing.

[0028] Principal component analysis (PCA) of the 3D registered clothing mesh includes: expanding the vertices of the 3D registered clothing mesh into a one-dimensional vector in index order, i.e., "x0,y0,z0,x1,y1,z1,x2,y2,z2.....", and performing PCA on the one-dimensional vector to obtain the clothing statistical model (i.e., 64-dimensional PCA parameters).

[0029] S2. Construct a synthetic image dataset based on a text-driven image generation model; combine the real dataset and the synthetic image dataset to obtain a hybrid image dataset;

[0030] like Figure 2 As shown, the synthetic image dataset for 3D clothing reconstruction, built using a text-driven image generation model, includes:

[0031] The system features a pre-defined multi-dimensional attribute space, including character attributes, clothing attributes, and environmental attributes. Character attributes cover different age groups (children, youth, middle-aged, and elderly), genders, and various body types (such as thin, medium, strong, and obese), and include multiple perspectives such as front, back, and side views. Clothing attributes are defined in the form of "top and bottom combinations," with a total of 38 subcategories of clothing combinations. Tops include shirts, T-shirts, and other types, while bottoms include trousers, shorts, and other types. Environmental attributes include various indoor and outdoor scenes, totaling more than 90 background types.

[0032] By combining the attributes of the aforementioned multidimensional attribute space to construct multiple text descriptions, a text-driven image generation model is obtained. Each text description is then input into the text-driven image generation model to generate a 1024×1024 resolution image of a dressed human body, thus obtaining a synthetic image dataset. This ensures the diversity of the data in terms of appearance, posture, and scene.

[0033] The text-driven image generation model is Stable Diffusion v3.

[0034] S3. Input the images from the mixed image dataset into the multimodal prior information extraction module to obtain the multimodal prior information of the images; the multimodal prior information includes: SMPL parameters and camera parameters. Normal graph and semantic segmentation graph; SMPL parameters include: morphological parameters and attitude parameters ;

[0035] The multimodal prior information extraction module includes: a normal inference module, a mask inference module, and a SMPL inference module. Image processing within the multimodal prior information extraction module includes: inputting the image into the normal inference module to obtain a normal map; inputting the image into the mask inference module to obtain a semantic segmentation map; and inputting the image into the SMPL inference module to obtain SMPL parameters and camera parameters. .

[0036] The normal inference module and the mask inference module are based on the Sapiens model, which is Meta's 2D human vision base model and outputs semantic segmentation map and normal map. The SMPL inference module is based on the PyMAF model, which is an SMPL parameter regressor that takes an image as input and outputs SPML parameters and camera parameters.

[0037] This invention employs a hybrid training strategy combining real and synthetic data to enhance the model's generalization ability. To unify the spatial representation of data from different sources, the pre-trained PyMAF model acquires SMPL parameters and camera parameters, thereby mapping different data to a consistent projection space.

[0038] S4. Input the image and its SMPL parameters into the clothing generation network to obtain the principal component coefficients z of the clothing in the image and the predicted clothing category;

[0039] like Figure 3 As shown, the clothing generation network includes an encoder, a decoder, and a classifier; the clothing generation network processes the image and its SMPL parameters as follows:

[0040] S41. Input the image into the encoder to obtain the high-level semantic features g of the image;

[0041] S42. Input the high-level semantic features g of the image into the classifier to obtain the predicted clothing category of the image; wherein, the classifier contains two branches, top and bottom, for predicting the clothing category;

[0042] S43. Integrate high-level semantic features and morphological parameters of the image After being stitched together, the images are input into a decoder to obtain the principal component coefficients z of the clothing in the image, which are 64-dimensional PCA parameters.

[0043] S5. Select the clothing statistical model corresponding to the clothing category predicted by the image from the clothing statistical models for each clothing category. Based on the principal component coefficients z of the image and the clothing statistical model Construct the initial 3D clothing mesh of the image;

[0044] Initial 3D clothing mesh for generating images include:

[0045] in, The principal component coefficient z of the image is the th The principal component vector corresponding to dimension , For clothing statistical models The Middle The principal component basis vectors corresponding to each dimension.

[0046] S6. Combine the image and its principal component coefficients with the clothing statistical model. The initial 3D clothing mesh, SMPL parameters, camera parameters, and semantic segmentation map are input into the geometry enhancement module to obtain a detailed 3D clothing mesh.

[0047] The processing steps of the geometry enhancement module include:

[0048] S61. Use SMPL parameters to perform pose transformation on the initial 3D clothing mesh of the image to obtain the 3D clothing mesh after pose transformation.

[0049] S62. Calculate the keypoint loss (Landmark Loss) based on the image, its camera parameters, and the 3D clothing mesh after pose changes. Calculate the boundary loss (Boundary Loss) based on the semantic segmentation map of the image and the 3D clothing mesh after pose changes.

[0050] Calculating keypoint loss includes:

[0051] Step 1: Select representative regions on the 3D clothing mesh after the pose change, and combine the vertices of the representative regions to obtain a set of representative vertices;

[0052] For tops, shirts should be chosen around the apex of the neck, shoulders, and elbows, while T-shirts should exclude the elbow area, retaining only the apex of the neck and shoulders. For bottoms, trousers should be chosen around the apex of the hips and knees, while shorts should only focus on the hip area.

[0053] Since the order of the vertices is the same, the indices of these vertices can be determined in advance on the clothing template.

[0054] Step 2: Calculate the mean of the spatial coordinates of all vertices in the representative vertex set to obtain the spatial coordinates of the 3D virtual joints;

[0055] Step 3: Using camera parameters The spatial coordinates of 3D virtual joints are projected onto a 2D image plane to obtain the projected coordinates. ;

[0056] Step 4: Obtain the joint inference model. Input the image into the joint inference model to obtain the ground truth joint coordinates. Calculate the projected coordinates Image ground truth key coordinates Between Distance, obtaining key point loss :

[0057]

[0058] Preferably, the key-point reasoning model is the OpenPose model.

[0059] Calculating the boundary loss includes:

[0060] Step 1, through Neighborhood detection algorithms extract boundary points from the semantic segmentation map of an image to obtain a set of ground truth boundaries. ;in, This represents the number of boundary points extracted from the semantic segmentation map.

[0061] For each pixel in the semantic segmentation map If there are pixels in its neighborhood that do not belong to the mask, then the pixel is determined to be a pixel. These are boundary points.

[0062] Step 2: Extract the boundary vertices of the 3D clothing mesh after pose changes, and project the extracted boundary vertices onto the 2D image plane to obtain the predicted boundary set. ; The number of boundary vertices extracted;

[0063] Step 3: For the truth boundary set For each point i in the prediction boundary set Find the N points i that are closest in distance (e.g., Manhattan distance) to obtain the set of points i. Preferably, N is 10;

[0064] Step 4: Calculate the truth boundary set Each point i in the set and its neighborhood The boundary loss is obtained by calculating the distance to each point in the boundary. :

[0065]

[0066] S63. Update the principal component coefficients z of the image based on the keypoint loss and boundary loss, and then update the principal component coefficients... and clothing statistical models Construct the aligned 3D clothing mesh;

[0067] Aligned 3D clothing mesh include:

[0068] in, The updated principal component coefficients The Middle The principal component vector corresponding to dimension.

[0069] The backpropagation update of the initial 3D clothing mesh based on keypoint loss and boundary loss includes: calculating the total loss L based on the keypoint loss and boundary loss, and obtaining the gradient of the principal component coefficients z using backpropagation based on the total loss L. Based on gradient The principal component coefficients z are updated using the Adam optimizer.

[0070] By applying keypoint loss and boundary loss simultaneously, the model is forced to maintain contour alignment while achieving precise coupling with the human skeletal structure, thereby correcting internal distortions caused by pose estimation bias and providing a more accurate starting shape for subsequent detail sculpting.

[0071] S64. Perform detail processing on the aligned 3D clothing mesh to restore the surface details of the clothing mesh and obtain a detailed 3D clothing mesh.

[0072] Detail processing of the aligned 3D clothing mesh includes:

[0073] S641. Subdivide the aligned 3D clothing mesh to obtain a subdivided 3D clothing mesh.

[0074] The subdivision process increases the number of facets through an interpolation algorithm, thereby expanding the deformation degrees of freedom of the mesh.

[0075] S642 represents each vertex in the subdivided 3D clothing mesh. Configure a local affine transformation matrix ;

[0076] S643, By optimizing the affine transformation matrix Drive vertex Displacement in three-dimensional space enables the reconstruction of geometric surface details, resulting in a detailed three-dimensional clothing mesh.

[0077] The above-mentioned detail processing procedure is existing and adopts the detail processing procedure in the literature "ICON: Implicit Clothed humans Obtained from Normals".

[0078] S7. Calculate the loss function value based on the principal component coefficients of the image, the predicted clothing category, the initial 3D clothing mesh, the detailed 3D clothing mesh, the normal map, and the semantic segmentation map. Update the parameters of the clothing 3D reconstruction model based on the loss function value. When the loss function value is minimized, the trained clothing 3D reconstruction model is obtained.

[0079] During the joint training process, this invention designs multiple loss functions to constrain the model.

[0080] Cross-entropy loss is introduced into the classifier branch to supervise the superpack and underpack categories:

[0081]

[0082] in, and These represent the ground truth label and predicted probability of the image's upper category, respectively. and These represent the ground truth label and predicted probability of the lower category of the image, respectively. For the number of categories of top garments, Let be the number of categories downloaded, and c be the index of each category. This loss ensures that the encoder can extract semantic information with category discrimination, preventing the introduction of noise due to category misclassification.

[0083] When training on generated images, due to the lack of direct guidance from 3D ground truth, this invention designs a boundary loss to drive the alignment of geometry on a 2D projection plane. The implementation of this loss includes:

[0084] pass Neighborhood detection algorithms extract boundaries from the semantic segmentation map of the input image, that is, for each pixel within the semantic segmentation map... If there are pixels in its neighborhood that do not belong to the semantic segmentation map, then the pixel is determined to be a non-segmented pixel. Using these as boundary points, we ultimately obtain the set of truth boundaries. .

[0085] Projecting the boundary vertices of the initial 3D clothing mesh onto the image plane yields the predicted boundary pixel set. ;

[0086] For the set of truth boundaries For each point i in the predicted boundary pixel set Find the nearest Manhattan Given points, obtain the set. ;

[0087] Boundary loss The calculation is as follows:

[0088]

[0089] Since boundary loss only constrains geometric alignment on the two-dimensional projection plane, simply optimizing the edges may lead to irregular geometric distortions within the mesh. To maintain the physical plausibility of the geometric surface, this invention also introduces Laplacian Smoothing Loss, which prevents severe unevenness or folding of the surface by constraining the consistency between mesh vertices and their neighborhood centers. Its formula is as follows:

[0090]

[0091] in, This represents the coordinates of vertex i in the initial 3D clothing mesh. Let i represent the set of adjacent vertices of vertex i in the initial 3D clothing mesh.

[0092] To further ensure that the generated garment surface has continuous and reasonable curvature, this framework also employs normal consistency loss. This loss function, by constraining the angle between the normal vectors of adjacent triangular facets, prevents unnatural distortions in the mesh during alignment. The loss function is expressed as:

[0093]

[0094] in, and For the initial 3D clothing mesh Adjacent faces that share an edge , They are dough pieces , The corresponding unit normal vector.

[0095] Given that the boundary loss is mainly corrected and Axial deviation, for depth ( To ensure accuracy in the (axis) direction, this invention corrects for errors using a strong supervisory signal provided by real scan data. For real image data, this invention applies a force to the 64-dimensional PCA parameters z obtained from regression. loss Directly align predicted coefficients with true coefficients This ensures the model's prediction accuracy in both depth and global dimensions.

[0096]

[0097] True coefficients The calculation process includes: aligning the 3D scanned clothing mesh corresponding to the real image to the corresponding template in the D3GNet clothing template using the D3GNet registration method through non-rigid deformation, thus obtaining a 3D registered clothing mesh; performing principal component analysis on the 3D registered clothing mesh to compress the high-dimensional vertex representation into 64-dimensional parameters, thus obtaining the clothing statistical model, i.e., the true coefficients. .

[0098] Finally, to prevent the model from producing extreme shapes that defy physical common sense when fitting complex edges, this invention applies a regularization loss to the 64-dimensional PCA parameter z of the regression. (Regularization Loss). This is achieved by performing a regularization loss on the coefficient vector. Regularization forces the model to make predictions within a reasonable distribution of the statistical model, avoiding topological distortions.

[0099]

[0100] To incorporate high-frequency details in the image, this invention utilizes each vertex of a detailed 3D clothing mesh. unit normal vector Calculate its color mapping in RGB space :

[0101]

[0102] Through this mapping, the range of values ​​for the normal vector is changed from... Linear transformation to image pixels space.

[0103] Subsequently, a differentiable renderer is used to map the colors of all vertices i. Generate prediction normal map And calculate the pixel-level relationship between it and the truth normal map. loss:

[0104]

[0105] In this process, to ensure the physical realism of the mesh surface is maintained while generating details and to prevent unnatural tearing or self-intersection, Laplacian smoothing loss and normal consistency loss are introduced as regularization terms in this stage. After multiple rounds of iterative optimization, the transformation matrix of the vertices tends to stabilize, and finally a 3D clothing model with surface details is obtained.

[0106] Laplacian smoothing loss for detailed 3D clothing mesh:

[0107]

[0108] in, This represents the coordinates of vertex i in the detailed 3D clothing mesh. This represents the set of adjacent vertices of vertex i in a detailed 3D clothing mesh.

[0109] Loss of normal consistency in detailed 3D clothing mesh:

[0110]

[0111] in, For detailed 3D clothing mesh Adjacent faces that share an edge They are dough pieces The corresponding unit normal vector.

[0112] In summary, this invention optimizes the entire network through end-to-end training by constructing a joint loss function. Specifically, the loss function for synthesized images... :

[0113]

[0114] Loss function for real images :

[0115]

[0116]

[0117] in, The classification loss is based on the predicted clothing category. The boundary loss is based on the initial 3D clothing mesh and semantic segmentation map. For the Laplacian smoothing loss based on the initial 3D clothing mesh, For the normal consistency loss based on the initial 3D clothing mesh and normal map, The L1 loss is based on the depth direction of the initial 3D clothing mesh. The regularization loss is based on the principal component coefficients z. For high-frequency detail loss based on detailed 3D clothing mesh, For Laplacian smoothing loss based on detailed 3D clothing mesh, For the normal consistency loss based on detailed 3D clothing mesh and normal map, , , , , , , These are the weighting coefficients for each type of loss.

[0118] Figure 4 The results are compared with those of ICON, ECON, SIFU, PIFUHD, etc., with GHNet representing the results of this method. The method of this invention provides more accurate details and forms compared to these methods, and the layering of upper and lower garments is more distinct.

[0119] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for 3D reconstruction of clothing based on geometric priors and generative image assistance, characterized in that, include: Acquire images and input them into a trained 3D clothing reconstruction model to obtain a detailed 3D clothing mesh. The 3D clothing reconstruction model includes: a multimodal prior information extraction module, a clothing generation network, and a geometry enhancement module; the training process of the 3D clothing reconstruction model includes: S1. Obtain the real dataset; the real dataset includes 3D scanned clothing meshes of various clothing categories and their rendered real images; construct statistical models of various clothing categories based on the 3D scanned clothing meshes of various clothing categories; S2. Construct a synthetic image dataset based on a text-driven image generation model; combine the real dataset and the synthetic image dataset to obtain a hybrid image dataset; S3. Input the images from the mixed image dataset into the multimodal prior information extraction module to obtain the multimodal prior information of the images; the multimodal prior information includes: SMPL parameters, camera parameters, normal map and semantic segmentation map; S4. Input the image and its SMPL parameters into the clothing generation network to obtain the principal component coefficients of the image and the predicted clothing category; S5. Select the clothing statistical model corresponding to the clothing category predicted from the image. Based on the principal component coefficients of the image and the clothing statistical model Construct the initial 3D clothing mesh of the image; S6. Combine the image and its principal component coefficients with the clothing statistical model. The initial 3D clothing mesh, SMPL parameters, camera parameters, and semantic segmentation map are input into the geometry enhancement module to obtain a detailed 3D clothing mesh. S7. Calculate the loss function value based on the principal component coefficients of the image, the predicted clothing category, the initial 3D clothing mesh, the detailed 3D clothing mesh, the normal map, and the semantic segmentation map. Update the parameters of the clothing 3D reconstruction model based on the loss function value. When the loss function value is minimized, the trained clothing 3D reconstruction model is obtained.

2. The method for 3D reconstruction of clothing based on geometric priors and generated image assistance according to claim 1, characterized in that, Statistical models for various clothing categories include: Obtain clothing templates for various clothing categories. Using a registration method, 3D scanned clothing meshes for each clothing category in the real dataset are aligned to the corresponding clothing templates to obtain 3D registered clothing meshes for each clothing category. Principal component analysis is then performed on the 3D registered clothing meshes for each clothing category to obtain statistical models for each clothing category.

3. The method for 3D reconstruction of clothing based on geometric priors and generated image assistance according to claim 1, characterized in that, The construction of a synthetic image dataset based on a text-driven image generation model includes: a pre-defined multi-dimensional attribute space, which includes: character attributes, clothing attributes, and environmental attributes; constructing multiple text descriptions based on the multi-dimensional attribute space; obtaining a text-driven image generation model; and inputting each text description into the text-driven image generation model to obtain the synthetic image dataset.

4. The method for 3D reconstruction of clothing based on geometric priors and generated image assistance according to claim 1, characterized in that, SMPL parameters include: morphological parameters and pose parameters; the clothing generation network includes: encoder, decoder, and classifier; the clothing generation network processes the image and its SMPL parameters as follows: S41. Input the image into the encoder to obtain the high-level semantic features of the image; S42. Input the high-level semantic features of the image into the classifier to obtain the predicted clothing category of the image; S43. After concatenating the high-level semantic features and morphological parameters of the image, input the result into the decoder to obtain the principal component coefficients of the image.

5. The method for 3D reconstruction of clothing based on geometric priors and generated image assistance according to claim 1, characterized in that, Initial 3D clothing model for generating images include: in, The principal component coefficient z of the image is the th The principal component vector corresponding to dimension , For clothing statistical models The Middle The principal component basis vectors corresponding to each dimension.

6. The method for 3D reconstruction of clothing based on geometric priors and generated image assistance according to claim 1, characterized in that, The processing steps of the geometry enhancement module include: S61. Use SMPL parameters to perform pose transformation on the initial 3D clothing mesh of the image to obtain the 3D clothing mesh after pose transformation. S62. Calculate the key point loss based on the image, its camera parameters, and the 3D clothing mesh after pose changes; calculate the boundary loss based on the semantic segmentation map of the image and the 3D clothing mesh after pose changes. S63. Update the principal component coefficients z of the image based on keypoint loss and boundary loss, and then use the updated principal component coefficients and the clothing statistical model. Construct the aligned 3D clothing mesh; S64. Perform detail processing on the aligned 3D clothing mesh to obtain a detailed 3D clothing mesh.

7. The method for 3D reconstruction of clothing based on geometric priors and generated image assistance according to claim 6, characterized in that, Calculating keypoint loss includes: Select representative regions on the 3D clothing mesh after the pose change, and combine the vertices of the representative regions to obtain a set of representative vertices; Calculate the mean of the spatial coordinates of all vertices in the representative vertex set to obtain the spatial coordinates of the 3D virtual joints; Using camera parameters The spatial coordinates of the 3D virtual joints are projected onto the two-dimensional image plane to obtain the projected coordinates; Obtain the key point inference model, input the image into the key point inference model, and obtain the image ground truth key point coordinates; The distance between the projected coordinates and the ground truth key point coordinates in the image is calculated to obtain the key point loss.

8. The method for 3D reconstruction of clothing based on geometric priors and generated image assistance according to claim 6, characterized in that, Calculating the boundary loss includes: Boundary points are extracted from the semantic segmentation map of the image to obtain the set of truth boundaries; Extract the boundary vertices of the 3D clothing mesh after the pose change, and project the extracted boundary vertices onto the 2D image plane to obtain the predicted boundary set; For each point i in the truth boundary set, find the k nearest points in the prediction boundary set to construct a neighborhood set for point i. ; Calculate the set of truth boundaries for each point i and its neighborhood set. The boundary loss is calculated by finding the distance to each point in the boundary.

9. The method for 3D reconstruction of clothing based on geometric priors and generated image assistance according to claim 1, characterized in that, Loss function of synthesized image : Loss function for real images : Boundary loss based on the initial 3D clothing mesh and semantic segmentation map, For the Laplacian smoothing loss based on the initial 3D clothing mesh, For the normal consistency loss based on the initial 3D clothing mesh and normal map, The L1 loss is based on the depth direction of the initial 3D clothing mesh. For regularization loss based on principal component coefficients, For high-frequency detail loss based on detailed 3D clothing mesh, For Laplacian smoothing loss based on detailed 3D clothing mesh, For the normal consistency loss based on detailed 3D clothing mesh and normal map, , , , , , , These are the weighting coefficients for each type of loss.

10. The method for 3D reconstruction of clothing based on geometric priors and generated image assistance according to claim 9, characterized in that, The calculation process for high-frequency detail loss in detailed 3D clothing meshes includes: Calculate the vertices of the detailed 3D clothing mesh unit normal vector Calculate each vertex unit normal vector Color mapping in RGB space ; Using a differentiable renderer based on all vertices Color mapping Generate prediction normal map ; Calculate the predicted normal map The pixel-level L2 loss between the image and the normal map yields the high-frequency detail loss of the detailed 3D clothing mesh.