A text-driven based method for generating a three-dimensional clothed human

CN122473365BActive Publication Date: 2026-09-22WUHAN TEXTILE UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610926258.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-22
Estimated Expiration
2046-06-25

AI Technical Summary

Technical Problem

该方法在单视图输入条件下具备一定的重建能力,但其输入形式限定为单张图像,缺乏对多视角图像与法向图协同生成的支持,在多视角几何一致性保持与服装细粒度纹理还原方面存在一定局限

Benefits of technology

(1)本发明将用户输入的纹理信息与服装类型信息统一转化为固定格式的结构化文本,涵盖面料质感、图案样式、色彩分布、服装品类、版型设计与穿着场景等属性字段。这种标准化编码方式有效减少了自由文本描述中因表达多样性引入的语义歧义,为后续属性解耦与条件生成提供了规范化、稳定的输入基础,有助于提升文本描述与生成结果之间的对应一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473365B_ABST
    Figure CN122473365B_ABST
Patent Text Reader

Abstract

The present application provides a kind of three-dimensional dressed human generation method based on text driving, it is related to three-dimensional generation technical field, including: obtaining the texture information and clothing type information input by user, conversion is standardized text;Standardized text is input text coding module, adopts attribute decoupling strategy to carry out semantic analysis, obtains uniform condition point;With uniform condition point driving condition generation model, obtains two-dimensional dressed human image;Two-dimensional dressed human image and view point parameter are input double-domain heterogeneous denoising model, after reverse denoising sampling, obtain multi-view RGB image and corresponding normal map;Multi-view RGB image, corresponding normal map and camera parameter are obtained point-level fusion feature using multi-view feature aggregation strategy;Point-level fusion feature is input neural implicit field module and carries out three-dimensional implicit reconstruction, obtains three-dimensional dressed human.The present application can realize attribute controllable, cross-angle consistent high-quality three-dimensional dressed human generation under text condition input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D generation technology, and in particular to a text-driven method for generating 3D dressed human figures. Background Technology

[0002] 3D dressed human figures are a crucial foundational data format for scenarios such as virtual try-on, digital human e-commerce, game and film asset production, and online clothing design and pattern simulation. Compared to static 2D images, 3D dressed human figures need to simultaneously depict the human form, clothing pattern structure, and fabric folds and texture details, while maintaining geometric and appearance consistency across multiple viewpoints. Existing industrial processes typically rely on manual modeling, scanning reconstruction, or physically based clothing simulation. While these methods offer high accuracy, they are highly dependent on equipment conditions, calibration processes, human experience, and production cycles, making it difficult to meet the practical needs of rapid iteration of massive SKUs and personalized content generation. In recent years, deep learning-based 3D generation technologies have made continuous progress in text-to-image conversion, multi-view consistency modeling, and implicit field reconstruction, driving the automated production of 3D assets. However, there are still obvious technical bottlenecks in the composite object of the dressed human body: the high degree of coupling between clothing material, structure and texture attributes makes it difficult to stably parse the text semantic conditions and accurately map them to the generated results; at the same time, in the processing link from 2D generation to 3D reconstruction, problems such as inconsistent appearance from multiple perspectives, geometric drift, unstable normal estimation and loss of fine-grained texture information are prominent, which in turn restricts the usability and editability of the generated results.

[0003] Chinese patent CN120047617A discloses a single-view method based on pixel and voxel feature fusion. Figure 3 This paper proposes a method for reconstructing a 3D dressed human body. This method constructs a dual-branch feature fusion network combining pixel and voxel branches to extract spatial features of the human body and clothing from a single-view input, thereby reconstructing the 3D dressed human body. While this method possesses certain reconstruction capabilities under single-view input conditions, its input format is limited to a single image, lacking support for the collaborative generation of multi-view images and normal maps. This results in limitations in maintaining multi-view geometric consistency and restoring fine-grained textures of clothing. Summary of the Invention

[0004] In view of this, the present invention provides a text-driven method for generating 3D dressed human bodies. By standardizing and encoding the clothing attribute information input by the user, decoupling the attributes of the text semantics and unifying the condition representation, and using a dual-domain heterogeneous denoising model to collaboratively generate multi-view RGB images and normal maps, and completing 3D reconstruction through a neural topological structure field, a high-quality 3D dressed human body generation with controllable attributes and consistent cross-viewpoints can be achieved under text condition input.

[0005] The technical solution of this invention is implemented as follows: This invention provides a text-driven method for generating 3D dressed human figures, comprising: S1. Obtain the texture information and clothing type information input by the user, and convert the texture information and clothing type information into standardized text using a fixed format template; S2. Input the standardized text into the text encoding module, and use the attribute decoupling strategy to perform semantic parsing on the material, structure and texture attributes to obtain unified condition points; use the unified condition points to drive the condition generation model, and obtain a two-dimensional dressed human body image through joint loss constraint training and guided sampling. S3. Input the two-dimensional dressed human body image and viewpoint parameters into the dual-domain heterogeneous denoising model, extract the shared conditional features through conditional encoding, and then send them to the RGB domain modulation branch and the normal domain modulation branch for gated fusion. After reverse denoising sampling, obtain the multi-view RGB image and the corresponding normal map. S4. The multi-view RGB image, corresponding normal map and camera parameters are used to obtain point-level fusion features by multi-view feature aggregation strategy; the point-level fusion features are input into the neural implicit field module for three-dimensional implicit reconstruction, and after differentiable rendering consistency optimization and mesh extraction, a three-dimensional dressed human body is obtained.

[0006] Preferably, in step S1, the texture information and clothing type information are converted into standardized text in the format A [T_style][T_pattern], where A represents the unit of measurement, [T_style] represents the clothing style attribute field, and [T_pattern] represents the clothing texture pattern attribute field; the texture information includes fabric texture, pattern style, and color distribution, and the clothing type information includes clothing category, pattern design, and wearing scenario.

[0007] Preferably, step S2 specifically includes: S21. Input the standardized text into the text encoding module, perform vMF hybrid decomposition and hierarchical weighting on the material, structure and texture attributes, and obtain the unified condition point c. S22, Derivation of structural conditions from the unified condition point c Appearance conditions The conditional denoising network is injected and trained with a joint loss that includes a score matching term, a Sinkhorn semantic alignment term, and an HSIC decoupling term. S23. After training, conditional sampling is performed using inverse stochastic differential equations to obtain preliminary generated results. After refinement by probability flow ordinary differential equations, the energy function guides the correction of conditional score estimation to obtain a two-dimensional clothing human image.

[0008] Preferably, step S21 specifically includes: The standardized text T is input into the text encoding module, and after L2 normalization and direction mapping, the direction representation data of the word sequence is obtained. ,in Let L represent the directional representation of the i-th lexical unit, and L represent the total number of lexical units. Regarding the By attribute category Perform vMF hybrid decomposition, where m represents material, s represents structure, and p represents texture, and calculate the responsibility degree. : ; Where j represents the prototype index under attribute k. This indicates the number of prototypes for attribute k. Indicates mixed weights, This represents the concentration parameter of the vMF distribution. Represents the prototype direction vector. This indicates transpose, and j' indicates the summation index. Represents an exponential function; Based on the above Perform direction-intensity weighted convergence to obtain attribute decoupling semantic vectors. And calculate the attribute confidence weights. ; The anchor point of manifold with property k As hierarchical weights, the unified condition points are solved in hyperbolic space: ; Where argmin represents the independent variable that minimizes the objective function, and x represents a candidate manifold point. Represents attributes manifold anchor points, This represents the geodesic distance in hyperbolic space.

[0009] Preferably, step S22 specifically includes: Clean two-dimensional images of dressed human bodies Construct a continuous forward perturbation process to generate noisy samples. and target score data ,in This represents the conditional probability density of the forward perturbation process; Project the unified condition point c onto the structural semantic direction center vector. With the semantic direction center vector of appearance The structural conditions are derived from this. Appearance conditions ,Will and As an independent conditional injection conditional denoising network, conditional score estimates are obtained. ; The conditional denoising network is trained using the following joint loss: ; Where L represents the total training loss, Let t represent expectation, and t represent time. Represents the time weighting function. Represents the noise scaling function. Represents the alignment of image embedding distribution With conditional embedding distribution The Sinkhorn divergence term, The HSIC decoupling term represents the statistical independence of structural and aesthetic constraints. This represents the loss weighting coefficient. This represents the square of the L2 norm.

[0010] Preferably, step S23 specifically includes: Preliminary generation results were obtained by conditional sampling using inverse stochastic differential equations: ; Where x represents the state in the reverse generation process. Represents the diffusion coefficient function. This represents the conditional score estimation function. Represents a reverse-time Wiener process; The trajectory is refined using probability flow ordinary differential equations: ; in, This represents the derivative term of the probability flow ordinary differential equation; With energy function Guided correction of conditional scores: ; in This indicates the score after bootstrapping. This represents the guiding strength coefficient. This represents the gradient operator with respect to the variable x. It consists of structural consistency constraints and appearance consistency constraints, used to penalize and or The inconsistent generation state yields a two-dimensional image of a dressed human body.

[0011] Preferably, step S3 specifically includes: The two-dimensional dressed human image and viewpoint parameter set The conditional coding layer of the input dual-domain heterogeneous denoising model is used to fuse and encode image features and viewpoint conditional vectors to obtain dual-domain shared conditional features. Where M represents the number of viewpoints, This represents the m-th viewpoint; The The data are fed into the RGB domain modulation branch and the normal domain modulation branch respectively to obtain the intermediate features in the RGB domain. Intermediate features of the normal domain ; Regarding the Introducing a direction-selective convolution operator to enhance the modeling of the orientation field of clothing folds, resulting in enhanced orientation normal features. Then, a gating fusion structure is used to... and Cross-domain fusion is performed to obtain dual-domain fusion features. ; The They are fed into the RGB domain score prediction head respectively. With normal domain score prediction head The combined features are the dual-domain denoising driving features. ; The After dual-domain inverse denoising sampling, a multi-view RGB image set is obtained. and the corresponding set of normal diagrams ,in, Indicate viewpoint Generate RGB images, Indicate viewpoint The generated normal graph.

[0012] Preferably, the direction-selective convolution operator pairs The enhancement calculation is as follows: ; in Learnable directional convolution kernels, Indicates along direction Anisotropic convolution operations, These are the local principal direction parameters predicted from clothing category features; The gated fusion structure uses RGB domain features as its backbone and introduces normal domain geometric information through gating to achieve asymmetric modulation and fusion of geometry and appearance. ; ; in This represents the Sigmoid activation function. This represents a learnable linear mapping matrix. Represents a cross-domain gated matrix. This represents element-wise multiplication. This indicates feature splicing.

[0013] Preferably, step S4 specifically includes: The multi-view RGB image set Corresponding normal graph set and camera parameter set The organization is a multi-view observation matrix, in which Representing viewpoints The camera's intrinsic parameter matrix, extrinsic parameter rotation matrix, and extrinsic parameter translation vector; A multi-view projection sampling and weighted aggregation strategy is employed to evaluate the bilinear sampling point features of each 3D sampling point X from various viewpoints using a scoring function. Calculate viewpoint fusion weights Weighted aggregation yields point-level fusion features. ,in This indicates that the 3D sampling point X is at the viewpoint. Projection sampling features; The point-level fusion features Input neural implicit field module, output implicit geometric field With topological modulation field And by the implicit geometric field Conductor density With differentiable normal : ; ; in Represents the scaling factor. This represents the gradient operator with respect to the three-dimensional coordinate X. Represents the L2 norm, Represents a smoothed positive function; Based on the above , , and Differentiable volume rendering and multi-view consistency constraint optimization are performed, and the three-dimensional dressed human body is obtained by isosurface extraction and mesh refinement.

[0014] Preferably, the optimization of differentiable volume rendering and multi-view consistency constraints specifically includes: Based on volume density Transmittance along rays and volume rendering weight Where s represents the ray sampling depth and p represents the pixel coordinates at viewpoint v. This represents the three-dimensional sampling point of the ray at depth s for pixel p under viewpoint v; With respect to color radiation function With differentiable normal Integrating along the ray yields the predicted RGB color of pixel p at viewpoint v. With prediction normal ; The multi-view joint optimization loss is defined as follows: ; ; ; in, For RGB consistency loss, For loss of normal consistency, This represents the actual observed RGB color of pixel p at viewpoint v. Indicates the corresponding actual observed normal. Represents the normal consistency weight, Indicates the weight of the regularization term. This represents the smoothing regularization term of the field function; By minimizing The network parameters of the neural implicit field module are updated, and the three-dimensional dressed human body is obtained by isosurface extraction and mesh refinement.

[0015] The present invention has the following advantages over the prior art: (1) This invention transforms the texture information and clothing type information input by the user into structured text with a fixed format, covering attribute fields such as fabric texture, pattern style, color distribution, clothing category, pattern design and wearing scene. This standardized encoding method effectively reduces the semantic ambiguity introduced by the diversity of expression in free text description, and provides a standardized and stable input basis for subsequent attribute decoupling and condition generation, which helps to improve the correspondence between text description and generation results.

[0016] (2) This invention introduces a hierarchical semantic parsing strategy in the text encoding stage, decoupling and representing the three types of attributes—clothing material, structure, and texture—and fusing them into unified condition points in the hyperbolic manifold space. These condition points then constrain different generation dimensions of the denoising network using structural and appearance conditions. Compared to directly using global text embedding, this strategy can more precisely capture the fine-grained semantic differences in clothing attributes, making the generated two-dimensional dressed human body image more accurate in responding to the input text in terms of clothing structure and appearance texture.

[0017] (3) This invention designs a dual-branch denoising structure that coordinates the processing of the RGB domain and the normal domain. The normal domain branch models the directional geometric features of clothing folds and achieves asymmetric interaction between the two features through a cross-domain gating fusion mechanism. This design enables the multi-view RGB images and their corresponding normal maps to constrain each other during the generation process, which helps to suppress appearance drift and instability in normal estimation between different viewpoints, and provides a relatively consistent geometric and appearance input for subsequent 3D reconstruction.

[0018] (4) The 3D reconstruction module of this invention uses multi-view RGB images and normal maps as joint inputs. It obtains implicit geometric fields and topological modulation fields through multi-view feature aggregation and neural implicit field inference, and performs multi-view joint optimization in conjunction with differentiable rendering constraints. Compared with reconstruction methods that rely solely on a single view or utilize only RGB information, introducing normal map supervision signals helps to constrain the accuracy of local geometry of curved surfaces, and has a certain improvement effect on the reconstruction of clothing fold details and thin-layer structures. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of the method of the present invention; Figure 2 A flowchart for generating unified condition points in this invention; Figure 3 This is a flowchart of the dual-domain heterogeneous noise reduction process of the present invention; Figure 4 This is a schematic diagram of the neural topology field reconstruction model of the present invention; Figure 5 This is a diagram illustrating the effects of the present invention. Detailed Implementation

[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0022] like Figure 1 As shown, this invention provides a text-driven method for generating a 3D dressed human body, comprising the following steps: S1. Obtain the texture information and clothing type information input by the user, and convert the texture information and clothing type information into standardized text using a fixed format template; S2. Input the standardized text into the text encoding module, and use the attribute decoupling strategy to perform semantic parsing on the material, structure and texture attributes to obtain unified condition points; use the unified condition points to drive the condition generation model, and obtain a two-dimensional dressed human body image through joint loss constraint training and guided sampling. S3. Input the two-dimensional dressed human body image and viewpoint parameters into the dual-domain heterogeneous denoising model, extract the shared conditional features through conditional encoding, and then send them to the RGB domain modulation branch and the normal domain modulation branch for gated fusion. After reverse denoising sampling, obtain the multi-view RGB image and the corresponding normal map. S4. The multi-view RGB image, corresponding normal map and camera parameters are used to obtain point-level fusion features by multi-view feature aggregation strategy; the point-level fusion features are input into the neural implicit field module for three-dimensional implicit reconstruction, and after differentiable rendering consistency optimization and mesh extraction, a three-dimensional dressed human body is obtained.

[0023] In one embodiment of the present invention, step S1 includes: The system acquires user-inputted texture and clothing type information. Texture information includes fabric texture (such as the feel and luster of cotton, linen, silk, and knitwear), pattern style (such as checks, stripes, prints, and solid colors), and color distribution (such as main color tone and color area division). Clothing type information includes clothing category (such as tops, skirts, and coats), silhouette design (such as fitted, loose, and waist-cinching silhouettes), and wearing occasion (such as casual, formal business, and sports / outdoor). This information is formatted and concatenated according to the template "A [T_style][T_pattern]" to generate standardized text. The "A" field contains the corresponding unit of measurement, the "[T_style]" field carries the clothing style attribute, and the "[T_pattern]" field carries the clothing texture pattern attribute. This standardized text format unifies multi-granularity free input into a fixed-structure string, providing a standardized input for subsequent text encoding modules.

[0024] Specifically, this invention utilizes templates to transform user-inputted texture and clothing type information into standardized text input in a fixed format. By structurally encoding and parsing fields such as fabric texture, pattern style, color distribution, clothing category, pattern design, and wearing scenario in the standardized text, computable conditional information about the appearance and structural attributes of the clothing can be obtained. This conditional information varies between different clothing styles and texture patterns because changes in material texture, pattern, and pattern structure directly alter the visual appearance and geometric structural features of the clothing, thus providing stable and controllable semantic constraints for subsequent text-driven two-dimensional clothing human body generation.

[0025] like Figure 2 As shown, in one embodiment of the present invention, step S2 includes: S21. Perform hierarchical semantic parsing and attribute decoupling on standardized text to finally generate unified condition points c for subsequent denoising networks; S22. Calculate the total training loss L; S23. Generate a two-dimensional image of a dressed human body using reverse denoising sampling.

[0026] Specifically, by designing personalized fine-tuning strategies and semantic condition injection mechanisms, this invention can extract the most critical clothing material attributes, structural attributes, and fine-grained texture details from standardized text, and enhance and constrain these features in the denoising generation network. This suppresses text-irrelevant noise and redundant visual patterns, reduces crosstalk between material and structure attributes, and highlights core representations such as fabric texture, pattern, and silhouette. This results in generating two-dimensional clothing images that are more consistent with the text description, have clearer details, and are structurally more stable, providing more reliable input for subsequent multi-view RGB image and corresponding normal map generation.

[0027] In one embodiment of the present invention, step S21 includes: S211. Input the standardized text into the text encoding and manifold representation module to obtain the directional representation data of the word sequence. .

[0028] in Represents a set; Representing the The directional information of each lexical unit in the attribute space; L represents the total number of lexical units; the text encoding and manifold representation module adopts a hybrid architecture of pre-trained language model fine-tuning + directional mapping enhancement. The basic encoding layer uses the CLIP text encoder as the basic backbone network. Its cross-modal semantic representation ability learned in the pre-training stage can effectively capture the fine-grained semantic differences of clothing attributes (material, structure, texture); for the clothing domain characteristics of this invention, fine-tuning is performed on the clothing attribute annotation dataset (containing 10,000 clothing description text-attribute label pairs), and the parameters of the first 12 layers of the encoder are frozen. Only the last 6 layers and the output projection layer are trained to adapt to the semantic distribution of the clothing domain. The process of generating the directional representation data of the lexical sequence is to perform directional mapping on the lexical embeddings (768-dimensional) output by the fine-tuned CLIP encoder: First, the original embedding of each lexical is L2 normalized (mapped to a unit sphere) to obtain a normalized vector; then, a lightweight fully connected layer (1024 hidden layer dimensions, GELU activation function) is used to map the normalized vector to the attribute semantic space, keeping the output dimension at 768, finally obtaining the directional representation of the lexical. .

[0029] S212. The lexical orientation representation data are assigned prototypes according to three attributes: material (m), structure (s), and texture (p). The responsibility weight of each lexical on each attribute prototype is calculated using the von Mises-Fisher (vMF) hybrid model. The calculation formula is as follows: ; Where j represents the prototype index under attribute k. This indicates the number of prototypes for attribute k. Indicates mixed weights, This represents the concentration parameter of the vMF distribution. Represents the prototype direction vector. This indicates transpose, and j' indicates the summation index. This represents an exponential function.

[0030] Based on the above Perform direction-intensity weighted convergence to obtain attribute decoupling semantic vectors. And calculate the attribute confidence weights. The calculation formula is as follows: ; in, This represents the confidence weight of attribute k; Indicates the index of the attribute summation; Represents attributes The number of prototypes; L represents the lexical length; j represents the prototype index; This represents the responsibility weight assigned to the j-th prototype by the i-th term under attribute category k.

[0031] S213, the above anchor point of manifold with property k As a hierarchical weight, the weighted mean of the unified condition point c is calculated using the following formula: ; Where argmin represents the independent variable that minimizes the objective function, and x represents a candidate manifold point. Representing attributes manifold anchor points, This represents the geodesic distance in hyperbolic space.

[0032] Specifically, the process begins by inputting standardized text into the text encoding and manifold representation module to obtain directional representation data of the lexical sequence. Next, this data undergoes a hybrid VMF decomposition based on three attributes: material, structure, and texture, calculating the responsibility level data. Weighted convergence is then used to obtain attribute-decoupled semantic vectors, along with the calculation of attribute confidence weights. Finally, using the attribute confidence weights and manifold anchors as hierarchical weights, a unified condition point is obtained by minimizing the objective function in hyperbolic space. This entire process, through a series of mathematical operations, completes the attribute decoupling and quantification representation of text semantics, providing accurate conditional input for subsequent generation of two-dimensional dressed human images.

[0033] In one embodiment of the present invention, step S22 includes: S221. Clean two-dimensional image of a dressed human body Construct a forward perturbation process, input the continuous forward perturbation module, and generate noisy sample data. The target score data is obtained from the logarithmic gradient of the forward perturbation kernel. ;in Represents the variable Operators for finding the gradient; This represents the conditional probability density of the forward perturbation process.

[0034] S222, Derive structural condition data from the unified condition point c. With appearance condition data The calculation formula is as follows: ; ; in, The vector transpose of the point c under the unified condition; Represents the structural semantic direction center vector; This represents the center vector of the semantic direction of appearance.

[0035] The and Injected as an independent conditional branch into the denoising network, the conditional score estimation data is obtained. .

[0036] S223, The conditional score estimation data The image-conditional semantic alignment is fused, and a structure-appearance decoupling constraint is added. The total training loss is defined and calculated as follows: ; in, t represents expectation; t represents time. Indicates a clean sample; This represents noisy sample data at time t; This represents the time weighting function; Represents the square of the L2 norm; Indicates structural conditions Appearance conditions Conditional score estimation; Represents the noise scaling function; Indicates the loss weighting coefficient; Indicates the distribution used to align image embeddings. With conditional embedding distribution The Sinkhorn divergence term; This represents the HSIC decoupling term used to constrain the statistical independence of structural and appearance conditions.

[0037] Specifically, firstly, a forward perturbation process is constructed on a clean 2D dressed human image to generate noisy sample data and obtain target score data. Next, structural condition data and appearance condition data are derived from unified condition points. The noisy sample data and the two types of condition data are then input into a conditional injection denoising network to obtain conditional score estimation data. Finally, the conditional score estimation data is fused with image-conditional semantic alignment, adding structure-appearance decoupling constraints, and defining the total training loss. The entire process, through data perturbation, conditional injection, and loss constraints, provides a training optimization basis for the accurate generation of 2D dressed human images.

[0038] In one embodiment of the present invention, step S23 includes: Conditional sampling was performed using inverse stochastic differential equations to obtain preliminary generation results: ; Where x represents the state in the reverse generation process; Represents the diffusion coefficient function. This represents the conditional score estimation function. This represents a reverse-time Wiener process.

[0039] The trajectory is refined using probability flow ordinary differential equations: ; A conditional guidance mechanism is introduced to adjust the direction of the generated trajectory: ; in, This indicates the score after booting; This represents a conditional score estimate; Indicates the guiding strength coefficient; This represents the gradient operator with respect to the variable x; This represents the energy function, which is composed of structural consistency constraints and appearance consistency constraints, and is used to penalize and constrain structural conditions. or appearance conditions Inconsistent generation state.

[0040] Specifically, the process begins with conditional sampling using inverse stochastic differential equations. This is combined with an inverse time Wiener process, diffusion coefficient function, conditional fraction estimation function, and structural and appearance conditions to obtain preliminary generated results. Next, probabilistic flow ordinary differential equations are used to refine the preliminary results, and derivative calculations further optimize the generated trajectory. Finally, the conditional fractions of the generation process are guided and corrected using a guiding intensity coefficient, variable gradient operator, and an energy function containing structural and appearance consistency constraints to obtain the final two-dimensional dressed human image. This progressive process of sampling, refinement, and guided correction ensures that the generated two-dimensional dressed human image conforms to structural and appearance constraints, improving image quality and text matching accuracy.

[0041] like Figure 3 As shown, in one embodiment of the present invention, step S3 includes: The two-dimensional dressed human image and viewpoint parameter set The conditional coding layer of the input dual-domain heterogeneous denoising model is used to fuse and encode image features and viewpoint conditional vectors to obtain dual-domain shared conditional features. Where M represents the number of viewpoints, This represents the m-th viewpoint.

[0042] The conditional coding layer employs a Transformer-based image encoder (such as the conditional injection structure of Zero123++) to... Feature extraction is performed, and the viewpoint parameters (azimuth, elevation, and distance) are encoded into continuous vectors and then fused with image features through cross-attention in the channel dimension, so that the conditional features carry both appearance semantics and viewpoint geometric information.

[0043] Dual-domain sharing conditional features The calculation formula is expressed as follows: ; ; ; in, Represents a two-dimensional image of a dressed human body; Represents the image feature coding layer; This represents the feature vector or feature map of the input image; Represents a set of viewpoints; Indicates the number of viewpoints; Indicates the first One perspective; Indicates the viewpoint coding layer; Represents the viewpoint condition vector; Indicates splicing; This indicates a conditional fusion coding layer; This represents the dual-domain shared conditional output matrix.

[0044] The The data are fed into the RGB domain modulation branch and the normal domain modulation branch respectively to obtain the intermediate features in the RGB domain. Intermediate features of the normal domain The calculation expression is as follows: ; ; in, Indicates the RGB domain modulation branch; This represents the normal domain modulation branch.

[0045] Regarding the Introducing a direction-selective convolution operator to enhance the modeling of the orientation field of clothing folds, resulting in enhanced orientation normal features. Its calculation form is: ; in Learnable directional convolution kernels, Indicates along direction Anisotropic convolution operations, These are the local principal direction parameters predicted from clothing category features; Then adopt a gating fusion structure to and Cross-domain integration and heterogeneous interaction fusion layer A geometric modulation-gated fusion structure is adopted, with RGB domain features as the main representation. Normal domain geometric information is introduced through gating to achieve asymmetric modulation fusion of geometry and appearance, resulting in dual-domain fused features. The calculation formula is as follows: ; ; in This represents the Sigmoid activation function. This represents a learnable linear mapping matrix. Represents a cross-domain gated matrix. This represents element-wise multiplication. This indicates feature splicing.

[0046] The They are fed into the RGB domain score prediction head respectively. With normal domain score prediction head The dual-domain denoising driving output matrix is ​​obtained. The calculation formula is as follows: ; ; ; in, This indicates the RGB field score prediction header; Represents the normal domain score prediction header; This represents the RGB domain conditional score output matrix; The output matrix represents the normal domain conditional score. This represents the dual-domain denoising driver output matrix; Indicates splicing.

[0047] The matrix The images are fed into a dual-domain inverse denoising sampling layer to obtain a multi-view RGB image set. and the corresponding set of normal diagrams The calculation formula is as follows: ; ; in, This represents the RGB domain inverse denoising sampling operator; The normal domain inverse denoising sampling operator is represented; m represents the viewpoint; This represents the m-th viewpoint; Indicate viewpoint Generate RGB images; Indicate viewpoint The generated normal graph.

[0048] Specifically, by designing a dual-domain heterogeneous denoising model for clothing, this invention can use a two-dimensional clothing human image generated step-by-step as prior input, establish denoising generation branches in the RGB appearance domain and the normal geometry domain respectively, and achieve collaborative generation through shared conditional features and cross-domain consistency constraints. This allows for the simultaneous output of normal maps strictly corresponding to each viewpoint while generating multi-view RGB images, reducing texture drift and geometric inconsistencies between multiple views, improving appearance continuity and surface normal stability under viewpoint switching, and thus providing higher quality and more consistent multi-view observation data for subsequent 3D reconstruction based on structural fields.

[0049] like Figure 4 As shown, in one embodiment of the present invention, step S4 includes: S41, assemble the multi-view RGB image set Corresponding normal graph set and camera parameter set Organized as input matrix The calculation formula is as follows: ; in, Represents the multi-view observation input matrix; This represents the organization stacking operator; m represents the viewpoint; Indicate viewpoint Generate RGB images; Indicate viewpoint The generated normal graph; Represents the camera intrinsic parameter matrix; This represents the camera extrinsic rotation matrix; M represents the camera's extrinsic translation vector; M represents the number of viewpoints. This represents the m-th viewpoint.

[0050] The input matrix The data is fed into the ray parameterization, multi-view projection sampling, and viewpoint fusion modules to obtain the point-level fusion feature matrix. The calculation formula is as follows: ; ; ; ; Where X represents a three-dimensional sampling point; This indicates that the 3D point X is at the viewpoint. The projected pixel coordinates; Represents the bilinear sampling operator; Indicate viewpoint Point features; Represents the scoring function; Indicates the viewpoint fusion weights; Represents an exponential function; This represents point-level fusion features; Indicates the organization stacking operator; This represents the point-level fusion feature matrix.

[0051] The point-level fusion feature matrix The structure field encoding input matrix is ​​obtained by grouping and packaging the corresponding ray sampling point sequences. The calculation formula is as follows: ; in, This indicates the organization stacking operator; p represents the pixel coordinate; Indicates the sampling depth of the k-th ray; This represents the ray sampling point of the pixel at viewpoint p; Indicates point-level fusion features; subscript These represent the viewpoint index, pixel index, and depth index, respectively. This represents the structure field encoding input matrix.

[0052] Specifically, the process begins with the first step of indexing and aligning the multi-view RGB image set, corresponding normal map set, and camera parameter set (including intrinsic parameter matrix, extrinsic parameter rotation matrix, and extrinsic parameter translation vector), integrating them into a multi-view observation input matrix using an organization stacking operator. The second step feeds this input matrix into the ray parameterization, multi-view projection sampling, and viewpoint fusion module. First, a bilinear sampling operator extracts the point features of the 3D sampling points at each viewpoint. Then, a scoring function and an exponential function are used to calculate the viewpoint fusion weights. Finally, a weighted sum is obtained to acquire the point-level fusion features, which are then processed by an organization stacking operator to obtain the point-level fusion feature matrix. The third step combines the point-level fusion feature matrix with the corresponding ray sampling point sequence using a grouping and packaging operator to obtain the structure field encoding input matrix. This entire process, through data integration, multi-view feature fusion, and grouping and packaging, provides standardized, high-quality input data for subsequent neural topology structure field encoding, laying the foundation for 3D reconstruction.

[0053] S42. Input the structure field encoding into the matrix. as input matrix The data is fed into the neural topological structure field encoding module (using a Transformer-based structure field precoding network) to obtain the intermediate field representation matrix. The above The implicit geometric field is obtained by feeding it into the neural topology field backbone module. With topological modulation field The calculation formula is as follows: ; ; in, This represents the multi-view feature sampling and fusion function; X represents a 3D point; This represents the projected coordinates of a 3D point X from various viewpoints. This represents the core module of the neural topology field; This represents network parameters.

[0054] The implicit geometric field Topological modulation field and the volume density derived from implicit geometric fields With differentiable normal Organize the data to obtain the field output matrix. The calculation formula is as follows: ; ; ; in, Indicates volume density; Represents a smoothed positive function; Indicates the scaling factor; Indicates the unit normal direction; This represents the gradient operator with respect to the three-dimensional coordinate X; Represents the L2 norm; Represents implicit geometric field data; Represents topological modulation field data; Represents volume density field data; Represents differentiable normal field data; Indicates the organization stacking operator; This represents the field output matrix.

[0055] Specifically, the first step involves feeding the structure field encoding input matrix into the structure field precoding module to obtain the field embedding matrix, which is generated by the structure field precoding module after processing the structure field encoding input matrix. The second step involves feeding the field embedding matrix into the neural topology structure field backbone module, combining the 3D points, field embedding vectors, and network parameters to obtain the implicit geometric field and the topological modulation field, which are obtained by processing the 3D points and field embedding vectors through the neural topology structure field backbone module. The third step involves deriving the volume density based on the implicit geometric field using a smoothing positive function and a scaling coefficient, deriving the unit normal using a gradient operator and a L2 norm, and then integrating the implicit geometric field, topological modulation field, volume density field, and differentiable normal field data through an organization stacking operator to obtain the field output matrix. Specifically, the volume density is obtained by processing the negative product of the scale coefficient and the implicit geometric field value using a smoothing positive function; the unit normal is obtained by dividing the implicit geometric field value processed by the 3D coordinate gradient operator by the L2 norm of that processing result; and the field output matrix is ​​generated by integrating the implicit geometric field data, topological modulation field data, volume density field data, and differentiable normal field data using an organization stacking operator. The entire process, through encoding, field modeling, and data organization, constructs a complete field representation containing geometric, topological, density, and normal information, providing core data support for subsequent rendering and optimization.

[0056] S43, output the field matrix as input matrix Rendering weights of ray-based structures based on volume density Transmittance is obtained by negativeing ​​the integral of the volume density over the ray from zero to the current depth and then applying it using an exponential function. The volume rendering weight is calculated by multiplying the transmittance by the volume density of the current ray point, as shown in the following formula: ; ; Where s represents the ray sampling depth, and p represents the pixel coordinates at viewpoint v. This represents the 3D sampling point at depth s of the ray from pixel p at viewpoint v.

[0057] Utilizing volume rendering weights Integrating the color and normal vectors separately yields the pixel prediction RGB image for each viewpoint. With pixel prediction normal map The calculation formula is as follows: ; ; ; in, Represents the pixel prediction RGB image of viewpoint v; Represents the pixel prediction normal map of viewpoint v; This represents the output of the color radiation function; Indicates a differentiable normal; Indicates volume rendering weight; Indicates the organization stacking operator; Represents the integral operator; Represents an integral infinitesimal element; This represents the set of predicted RGB images from viewpoint v=1 to v=M; This represents the set of predicted normal maps from viewpoint v=1 to v=M; This represents the predicted observation matrix.

[0058] The prediction observation matrix For multi-view observations, consistency constraints are applied, and a multi-view joint optimization loss function is defined, the calculation formula of which is as follows: ; ; ; in, For RGB consistency loss, For loss of normal consistency, This represents the actual observed RGB color of pixel p at viewpoint v. Indicates the corresponding actual observed normal. Represents the normal consistency weight, Indicates the weight of the regularization term. This represents the smoothing regularization term of the field function.

[0059] Network parameters are minimized by the Update the network parameters to obtain optimized network parameters. This leads to the optimized parameter output matrix. The calculation formula is as follows: ; in, Indicates network parameters; Represents implicit geometric field values; Indicates the topological modulation field value; Indicates the organization stacking operator; This represents the optimized parameter output matrix.

[0060] Specifically, the first step uses the field output matrix as input, constructs volume rendering weights along the ray based on volume density, and obtains the transmittance by processing the negative value of the integral result of the volume density along the ray from zero to the current depth using an exponential function. The volume rendering weight is obtained by multiplying the transmittance by the volume density of the current ray point. The second step uses the volume rendering weight to integrate the color and normal from zero to infinity respectively, obtaining the pixel predicted RGB image and pixel predicted normal image for each viewpoint. Then, the set of predicted RGB images and the set of predicted normal images for all viewpoints are integrated by the organization stacking operator to form the prediction observation matrix. The third step applies consistency constraints to the prediction observation matrix and multi-view observations, updates the network parameters according to the constraint results, and then integrates the updated network parameters, implicit geometric field data, and topological modulation field data by the organization stacking operator to obtain the optimized parameter output matrix. The entire process, through rendering to generate prediction data and consistency constraints to optimize parameters, ensures the accuracy and reliability of 3D reconstruction, providing a guarantee for the final output of a high-quality 3D dressed human body.

[0061] S44. Output the optimized parameter matrix. The data is fed into the isosurface extraction and meshing module, where the Marching Cubes algorithm is used to process the implicit geometry. Isosurface extraction is performed at the zero isosurface to generate a triangular mesh, and the topological modulation field is then applied. The encoded clothing topology modulation information is projected onto the mesh vertices to obtain the final high-quality 3D dressed human body mesh output.

[0062] The neural topological structure field reconstruction model jointly models the implicit geometric field and the topological modulation field, and simultaneously optimizes the consistency of RGB color and normal in a volume rendering manner. This allows the reconstructed 3D human body to be effectively constrained from multi-view and multi-modal observation in terms of both clothing geometric details (such as folds and seams) and surface texture, thus improving the problem of rough geometric surfaces when relying solely on RGB supervision.

[0063] This invention can use multi-view RGB images and their corresponding normal maps as joint observation inputs, simultaneously modeling surface shape and topological modulation information in an implicit geometric field, and optimizing field parameters through differentiable rendering consistency constraints. This fully utilizes the appearance texture cues provided by RGB and the local geometric constraints provided by the normal map, suppressing geometric drift, surface fragmentation, and topological artifacts during reconstruction, improving the geometric fidelity and normal continuity of detailed areas (such as folds, collars, and cuffs), and ultimately outputting a high-quality 3D dressed human model with complete structure and clear details.

[0064] like Figure 5The diagram illustrates the processing steps of two typical implementation examples of the method of this invention. The input for the left example is the texture and style information of a green velvet wrap dress, while the input for the right example is the texture and style information of a rose gold sequined fitted long dress. Both examples undergo S1 standardized text conversion, S2 two-dimensional dressed human image generation, S3 multi-view RGB image and normal map generation, and finally S4 reconstruction to obtain a complete three-dimensional dressed human model. The output results are shown below. Figure 5 As shown at the bottom. Two sets of examples demonstrate that the method of the present invention can adapt to clothing inputs of different styles, fabrics, and designs, generating three-dimensional dressed human figures with clear geometric structures and accurate reproduction of texture details.

[0065] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A text-driven method for generating 3D dressed human figures, characterized in that, include: S1. Obtain the texture information and clothing type information input by the user, and convert the texture information and clothing type information into standardized text using a fixed format template; S2. Input the standardized text into the text encoding module, and use the attribute decoupling strategy to perform semantic parsing on the material, structure and texture attributes to obtain unified condition points; use the unified condition points to drive the condition generation model, and obtain a two-dimensional dressed human body image through joint loss constraint training and guided sampling. S3. Input the two-dimensional dressed human body image and viewpoint parameters into the dual-domain heterogeneous denoising model, extract the shared conditional features through conditional encoding, and then send them to the RGB domain modulation branch and the normal domain modulation branch for gated fusion. After reverse denoising sampling, obtain the multi-view RGB image and the corresponding normal map. Step S3 specifically includes: The two-dimensional dressed human image and viewpoint parameter set The conditional coding layer of the input dual-domain heterogeneous denoising model is used to fuse and encode image features and viewpoint conditional vectors to obtain dual-domain shared conditional features. Where M represents the number of viewpoints, This represents the m-th viewpoint; The The data are fed into the RGB domain modulation branch and the normal domain modulation branch respectively to obtain the intermediate features in the RGB domain. Intermediate features of the normal domain ; Regarding the Introducing a direction-selective convolution operator to enhance the modeling of the orientation field of clothing folds, resulting in enhanced orientation normal features. Then, a gating fusion structure is used to... and Cross-domain fusion is performed to obtain dual-domain fusion features. ; The They are fed into the RGB domain score prediction head respectively. With normal domain score prediction head The combined features are the dual-domain denoising driving features. ; The After dual-domain inverse denoising sampling, a multi-view RGB image set is obtained. and the corresponding set of normal diagrams ,in, Indicate viewpoint Generate RGB images, Indicate viewpoint The generated normal graph; S4. The multi-view RGB image, corresponding normal map and camera parameters are used to obtain point-level fusion features by multi-view feature aggregation strategy; the point-level fusion features are input into the neural implicit field module for three-dimensional implicit reconstruction, and after differentiable rendering consistency optimization and mesh extraction, a three-dimensional dressed human body is obtained.

2. The text-driven three-dimensional dressed human body generation method according to claim 1, characterized in that, In step S1, the texture information and clothing type information are converted into standardized text in the format A [T_style][T_pattern], where A represents the unit of measurement, [T_style] represents the clothing style attribute field, and [T_pattern] represents the clothing texture pattern attribute field. The texture information includes fabric texture, pattern style, and color distribution, and the clothing type information includes clothing category, pattern design, and wearing scenario.

3. The text-driven three-dimensional dressed human body generation method according to claim 1, characterized in that, Step S2 specifically includes: S21. Input the standardized text into the text encoding module, perform vMF hybrid decomposition and hierarchical weighting on the material, structure and texture attributes, and obtain the unified condition point c. S22, Derivation of structural conditions from the unified condition point c Appearance conditions The conditional denoising network is injected and trained with a joint loss that includes a score matching term, a Sinkhorn semantic alignment term, and an HSIC decoupling term. S23. After training, conditional sampling is performed using inverse stochastic differential equations to obtain preliminary generated results. After refinement by probability flow ordinary differential equations, the energy function guides the correction of conditional score estimation to obtain a two-dimensional clothing human image.

4. The text-driven three-dimensional dressed human body generation method according to claim 3, characterized in that, Step S22 specifically includes: Clean two-dimensional images of dressed human bodies Construct a continuous forward perturbation process to generate noisy samples. and target score data ,in This represents the conditional probability density of the forward perturbation process; Project the unified condition point c onto the structural semantic direction center vector. With the semantic direction center vector of appearance The structural conditions are derived from this. Appearance conditions ,Will and As an independent conditional injection conditional denoising network, conditional score estimates are obtained. ; The conditional denoising network is trained using the following joint loss: ; Where L represents the total training loss, Let t represent expectation, and t represent time. Represents the time weighting function. Represents the noise scaling function. Represents the alignment of image embedding distribution With conditional embedding distribution The Sinkhorn divergence term, The HSIC decoupling term represents the statistical independence of structural and aesthetic constraints. This represents the loss weighting coefficient. This represents the square of the L2 norm.

5. The text-driven three-dimensional dressed human body generation method according to claim 3, characterized in that, Step S23 specifically includes: Preliminary generation results were obtained by conditional sampling using inverse stochastic differential equations: ; Where x represents the state in the reverse generation process. Represents the diffusion coefficient function. This represents the conditional score estimation function. Represents a reverse-time Wiener process; The trajectory is refined using probability flow ordinary differential equations: ; in, This represents the derivative term of the probability flow ordinary differential equation; With energy function Guided correction of conditional scores: ; in This indicates the score after bootstrapping. This represents the guiding strength coefficient. This represents the gradient operator with respect to the variable x. It consists of structural consistency constraints and appearance consistency constraints, used to penalize and or The inconsistent generation state yields a two-dimensional image of a dressed human body.

6. The text-driven three-dimensional dressed human body generation method according to claim 1, characterized in that, Step S4 specifically includes: The multi-view RGB image set Corresponding normal graph set and camera parameter set The organization is a multi-view observation matrix, in which Representing viewpoints The camera's intrinsic parameter matrix, extrinsic parameter rotation matrix, and extrinsic parameter translation vector; A multi-view projection sampling and weighted aggregation strategy is employed to evaluate the bilinear sampling point features of each 3D sampling point X from various viewpoints using a scoring function. Calculate viewpoint fusion weights Weighted aggregation yields point-level fusion features. ,in This indicates that the 3D sampling point X is at the viewpoint. Projection sampling features; The point-level fusion features Input neural implicit field module, output implicit geometric field With topological modulation field And by the implicit geometric field Conductor density With differentiable normal : ; ; in Represents the scaling factor. This represents the gradient operator with respect to the three-dimensional coordinate X. Represents the L2 norm, Represents a smoothed positive function; Based on the above , , and Differentiable volume rendering and multi-view consistency constraint optimization are performed, and the three-dimensional dressed human body is obtained by isosurface extraction and mesh refinement.

7. The text-driven three-dimensional dressed human body generation method according to claim 6, characterized in that, The optimization of differentiable volume rendering and multi-view consistency constraints specifically includes: Based on volume density Transmittance along rays and volume rendering weight Where s represents the ray sampling depth and p represents the pixel coordinates at viewpoint v. This represents the three-dimensional sampling point of the ray at depth s for pixel p under viewpoint v; With respect to color radiation function With differentiable normal Integrating along the ray yields the predicted RGB color of pixel p at viewpoint v. With prediction normal ; The multi-view joint optimization loss is defined as follows: ; ; ; in, For RGB consistency loss, For loss of normal consistency, This represents the actual observed RGB color of pixel p at viewpoint v. Indicates the corresponding actual observed normal. Represents the normal consistency weight, Indicates the weight of the regularization term. This represents the smoothing regularization term of the field function; By minimizing The network parameters of the neural implicit field module are updated, and the three-dimensional dressed human body is obtained by isosurface extraction and mesh refinement.

Citation Information

Patent Citations

  • Single-view three-dimensional dressed human body reconstruction method based on pixel and voxel feature fusion

    CN120047617A

  • Garment generation model training method, digital human garment generation method and hardware

    CN120912737A

  • Automatic clothing modeling method based on two-dimensional clothing image and three-dimensional human body model

    CN121883728A