Virtual garment consistent texture asset generation method and system based on garment region semantics and visibility reprojection constraints
Patent Information
- Application Number
- CN202610788794.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-09-04
AI Technical Summary
(1)已有服装几何基础上的一致纹理生成能力不足:现有文本驱动服装生成方法通常从文本直接生成完整人体与服装资产,重点在于整体生成流程;而在实际服装资产制作过程中,服装几何往往已经由建模、扫描、仿真或几何细化方法获得
本发明在已有服装网格基础上,从多个视角生成纹理,并保证不同视角下同一服装表面区域的颜色、图案、边界和局部细节保持一致,实现基于已有服装几何生成跨视角一致纹理;将服装前片、后片、袖子、领口、袖口、裙摆等区域标签引入纹理生成过程,使不同区域的纹理生成更加符合服装结构,并避免纹理跨区混乱的问题,实现利用服装区域语义约束纹理生成;利用深度图和相机内外参数,将一个视角下的像素反投影到三维服装表面,再投影到另一个视角,并根据深度一致性判断该点是否可见,从而只对有效可见区域计算纹理一致性约束,实现可靠的可见性重投影约束;在袖口、领口、前后片交界、裙摆边缘等区域边界处保持纹理连续,减少颜色突变、图案错位和接缝明显,在服装区域边界处实现连续融合;最终将服装网格、材质贴图、区域标签和贴图映射关系统一导出,形成面向 CG 渲染和后续仿真流程的虚拟服装一致纹理资产,实现可复用。
Smart Images

Figure CN122695151A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital clothing design technology, and more specifically, to a method and system for generating consistent texture assets for virtual clothing based on clothing region semantics and visibility reprojection constraints. Background Technology
[0002] 3D virtual clothing assets are widely used in scenarios such as digital humans, virtual try-on, clothing design, game animation, film and television production, and fabric physics simulation. Unlike ordinary 3D objects, virtual clothing assets not only need to have accurate clothing geometry, but also stable surface textures, continuous material maps, and asset structures that can be used in subsequent rendering or simulation processes. Currently, related technologies mainly include the following categories.
[0003] (1) Text-driven 3D clothing asset generation technology This type of technology typically includes a text parsing module, a human body or clothing template selection module, a clothing geometry generation module, a texture generation module, and an asset output module.
[0004] Its basic working process is as follows: First, semantic parsing is performed on the text description input by the user to obtain control information such as clothing category, human body shape, clothing style, color or material; then, a human body model and clothing template are selected according to the parsing results; next, the clothing template is geometrically deformed to make the clothing fit the human body or target shape; finally, the appearance of the clothing is generated through the texture generation module, and the human body asset with clothing is output.
[0005] These methods can generate relatively complete virtual clothing or human body assets with clothing, but they usually focus more on the overall generation chain from text to complete assets, and lack specific design for consistent texture generation based on existing clothing geometry, clothing region boundary processing and texture asset export.
[0006] (2) Single-view or multi-view three-dimensional clothing reconstruction technology This type of technology typically includes an image input module, a viewpoint generation module, a depth or geometry prediction module, a texture restoration module, and a 3D reconstruction module.
[0007] The basic working process is as follows: input one or more clothing images, fill in the invisible areas through depth prediction, viewpoint synthesis, image transformation, or multi-view generation models, and then restore the three-dimensional geometry and surface appearance of the clothing. Some methods also jointly predict RGB images and depth information to enhance multi-view consistency.
[0008] The focus of this type of technology is to recover clothing geometry from images or improve the consistency of appearance from new perspectives. However, its output is usually still mainly reconstructed meshes, images from several perspectives, or visual results, lacking a complete process of further mapping, fusing, and organizing multi-view texture results into material map assets that can be directly called.
[0009] (3) General 3D texture generation technology This type of technology typically includes a 3D model input module, a conditional image generation module, a texture generation module, a projection blending module, and a texture output module.
[0010] Its basic working process is as follows: given a 3D model, render depth maps, normal maps or other geometric condition images from multiple perspectives, then use the generative model to generate texture images from different perspectives, and finally project these texture results back onto the surface of the 3D model or UV space to form texture maps.
[0011] While this type of technology can generate textures for general 3D objects, it still has shortcomings in clothing scenarios. Clothing has a distinct regional structure, such as the front piece, back piece, sleeves, neckline, cuffs, and hem. Different regions need to maintain local semantic differences while also ensuring texture continuity at boundaries. Ordinary 3D texture generation methods typically treat clothing as a general 3D object, failing to fully utilize the semantics of clothing regions and lacking a continuous blending mechanism for the boundaries of regions such as the neckline, cuffs, hem, and the junction of the front and back pieces.
[0012] (4) Material mapping construction and virtual asset export technology This type of technology typically includes UV mapping modules, texture blending modules, material texture generation modules, and asset file export modules.
[0013] Its basic working process is as follows: mapping the two-dimensional texture image onto the surface of the three-dimensional clothing mesh, then constructing material information such as basic color map, normal map, roughness map, transparency map or displacement map according to rendering requirements, and finally exporting it as a virtual clothing asset that can be read by the rendering engine or asset management system.
[0014] This type of technology can improve the rendering usability of clothing models, but it usually relies on manual texture editing or seam repair, and lacks an integrated solution that is closely integrated with multi-view texture generation, clothing area labeling, and visibility reprojection processes.
[0015] Despite the progress made by the aforementioned technologies in 3D clothing generation, clothing reconstruction, texture generation, and asset export, the following shortcomings still exist: (1) Insufficient ability to generate consistent textures based on existing clothing geometry: Existing text-driven clothing generation methods typically generate complete human and clothing assets directly from text, focusing on the overall generation process; however, in actual clothing asset production, clothing geometry is often already obtained through modeling, scanning, simulation, or geometric refinement methods. What is needed at this time is to generate stable, continuous, and reusable texture assets based on existing clothing meshes, rather than regenerating the entire set of clothing geometry. Therefore, existing technologies lack a method specifically for generating consistent texture assets for virtual clothing based on existing clothing geometry.
[0016] (2) Per-view texture generation is prone to inconsistencies: In existing texture generation methods, texture results from different viewpoints are often independent or only weakly constrained. When there are folds, cuffs, collars, skirt edges, or occluded areas on the garment surface, problems such as inconsistent color styles, misaligned patterns, obvious boundary seams, and local texture breaks can easily occur between different viewpoints. Therefore, existing technologies cannot guarantee the continuity and consistency of texture on the same garment surface under multiple viewpoints.
[0017] (3) Lack of explicit utilization of garment area semantics: Different areas of garments have clear structures and semantic functions. For example, the front and back pieces are usually the main display areas, the cuffs and necklines have obvious boundaries, and the skirt area is prone to folds and swaying structures. Existing general 3D texture generation methods usually only use depth maps, normal maps, or text descriptions as conditions, and rarely use garment area tags to control the texture generation of different areas. Therefore, the generated textures are prone to regional semantic confusion, such as textures crossing the cuff boundaries, inconsistent texture styles between the front and back pieces, and pattern breaks at the boundaries of the skirt and body.
[0018] (4) Insufficient cross-view correspondence and occlusion handling: In multi-view texture generation, to constrain texture consistency under different views, it is necessary to know which pixel in one view corresponds to which pixel in another view. Existing methods sometimes only use approximate alignment or image-level consistency constraints, without explicitly using depth maps, camera parameters, and occlusion judgment to establish reliable cross-view correspondence. When clothing has self-occlusion, sleeves occluding the body of the garment, skirt hem occluding the back piece, or deep folds, if visibility judgment is not performed, incorrect cross-view consistency constraints will instead cause texture misalignment and boundary contamination.
[0019] (5) Insufficient continuity of texture at the boundaries of the regions: The most problematic areas for clothing texture are usually at the boundaries of the regions, such as cuffs, necklines, side seams, hem edges, the junction of the front and back pieces, and areas covered by folds. Existing texture blending methods mostly use ordinary weighted average or simple projection blending, which lacks continuity constraints for the boundaries of clothing regions, thus easily causing color abrupt changes, pattern breaks, and obvious seams.
[0020] (6) Texture generation results are difficult to directly convert into usable texture assets: Although many methods can generate images of clothing appearance, these results usually remain at the level of two-dimensional view images, without further using visibility judgment, surface projection and texture fusion to uniformly map multi-view textures onto the clothing mesh surface or UV space. Therefore, existing technologies often cannot directly generate material map sets such as basic color maps, normal maps, and roughness maps that can be called by the rendering system.
[0021] How to generate cross-view consistent textures based on existing clothing geometry, utilize semantic constraints of clothing regions to generate textures, establish reliable visibility reprojection constraints, achieve continuous fusion at the boundaries of clothing regions, and form reusable virtual clothing consistent texture assets are urgent problems to be solved. Summary of the Invention
[0022] The purpose of this invention is to provide a method and system for generating consistent virtual clothing texture assets based on clothing region semantics and visibility reprojection constraints, which can improve the cross-view consistency of clothing textures, the fit between textures and clothing region structures, the continuity of region boundaries, and the usability of clothing assets in the rendering process.
[0023] This invention provides a virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints, including a clothing geometry and region label input module, a multi-view region conditional rendering module, a region semantic-guided consistent texture generation module, a visibility reprojection and region boundary fusion module, and a material mapping and texture asset export module. The clothing geometry and region label input module is used to obtain the clothing geometry representation based on the clothing mesh, UV coordinates and clothing region labels; The multi-view region conditional rendering module is used to perform conditional rendering from multiple viewing angles based on the clothing geometric representation and the set of camera viewpoints, and generate a multi-view conditional image. The region semantically guided consistent texture generation module is used to generate clothing texture images from multiple perspectives based on multi-view conditional images, clothing region labels, and overall material or style text. The visibility reprojection and region boundary fusion module is used to establish cross-view correspondence and cross-view consistency loss based on depth map and camera parameters. Based on the cross-view consistency loss and region boundary continuity loss, the clothing texture images from multiple viewpoints are projected and fused onto the three-dimensional clothing mesh surface to obtain the fused three-dimensional surface texture. The material mapping and texture asset export module is used to construct a consistent texture asset for virtual clothing based on the fused 3D surface texture.
[0024] Furthermore, the expression for the geometric representation of the clothing is:
[0025]
[0026]
[0027]
[0028]
[0029] in, This represents the input clothing geometry mesh; Represents the set of vertices of the clothing mesh. Let be the i-th grid vertex, and N be the number of grid vertices; Represents a set of triangular facets. Let K be the j-th triangular facet, and K be the number of triangular facets. Represents the set of UV coordinates corresponding to a vertex. This represents the UV coordinates corresponding to the i-th vertex; This represents a set of clothing area labels. This is the label for the j-th clothing area.
[0030] Furthermore, the process of generating clothing texture images from multiple perspectives includes: Generate region-level prompts based on clothing area labels and overall material or style text; Using multi-view conditional images, region-level cue words, and overall material or style text as conditions, a conditional diffusion generation model is used to generate clothing texture images from multiple perspectives.
[0031] Furthermore, the optimization objective of the conditional diffusion generation model is:
[0032] in, The optimization objective for the conditional diffusion generation model; Let represent the mathematical expectation calculated after joint sampling of the diffusion time step t, the added noise ε, and the viewpoint number b; t represents the diffusion time step. This represents noise added to the texture image; This represents a denoising network; Represents the noisy texture image at the t-th diffusion time step; This represents the multi-view conditional image from the b-th perspective, where p represents the overall material or style text. Indicates a region-level prompt.
[0033] Furthermore, the formula for calculating the three-dimensional surface texture is as follows:
[0034]
[0035]
[0036] in, B represents the blended texture color of surface point s; B represents the number of viewpoints. Let be the visibility function, used to indicate whether a surface point s is visible from the b-th viewpoint. Otherwise ; This represents the fusion weight of the b-th viewpoint for surface point s; Represents the texture color or feature of a pixel in viewpoint b; This represents the projection function corresponding to the b-th viewpoint; To prevent extremely small constants with a denominator of zero; Let be the projected position of point s from the b-th viewpoint; This indicates taking the maximum value; The normal vector of surface point s; This indicates the direction of observation from the b-th perspective; This indicates the confidence or sharpness score of the texture generation for the corresponding pixel from that viewpoint. Indicates the boundary security score; These are the weighting coefficients.
[0037] Furthermore, the expression for the cross-perspective correspondence is:
[0038]
[0039]
[0040] in, This represents the homogeneous pixel coordinates of a point on the surface of a 3D garment at viewpoint q. Let q represent the camera intrinsic parameters, rotation matrix, and translation vector, respectively; s represents a point on the 3D clothing surface. It is a rotation matrix; This represents a 3D point in the camera coordinate system corresponding to a pixel. It is a translation vector; This is the depth value; The pixel position at viewpoint b; This is the camera intrinsic parameter matrix; These are homogeneous coordinates.
[0041] Furthermore, the formula for calculating the cross-perspective consistency loss is as follows:
[0042]
[0043] in, For cross-view consistency loss; B represents the number of views; The pixel position at viewpoint b; This represents the candidate common visible region for viewpoints b and q. For visibility; Indicates consistency weight; Represents pixels in viewpoint b Texture color or features at the location; This represents the texture color or feature at the corresponding pixel after reprojection onto the viewpoint q; This represents the first norm, used to sum the absolute values of the components of the two texture color or texture feature difference vectors within the brackets; The predicted depth of a pixel after it has been projected onto the viewpoint q; This represents the depth value at this location in the target's perspective depth map. This represents the depth tolerance threshold.
[0044] Furthermore, the formula for calculating the continuity loss at the region boundary is as follows:
[0045] in, For continuous loss at the regional boundary; For any pair of adjacent regions; This is a group of adjacent clothing areas; and These represent texture sampling points at the boundary of two adjacent regions; Indicates that point s is in the region The blended texture colors at the boundary This indicates that point s is in the adjacent region. The blended texture color at the boundary.
[0046] Furthermore, the process of constructing a consistent texture asset for virtual clothing includes: The blended 3D surface texture is mapped onto the UV plane to obtain the base color map, as shown in the formula:
[0047]
[0048] in, Represents UV coordinates The base color texture value at that location; Represents the UV plane coordinates; This represents the blended texture color of surface point s; This represents the UV coordinates corresponding to surface point s; Based on the aforementioned base color map, construct a set of material maps, as shown in the formula:
[0049] in, A collection of material textures; This represents the base color map value. This represents the normal map. Represents roughness map, Indicates transparency map, This indicates a height map or a displacement map; Based on the aforementioned material map set, a consistent texture asset for the virtual clothing is obtained, as shown in the formula:
[0050] in, This represents the final exported texture asset; Represents the geometric mesh of the garment; This represents the UV coordinates corresponding to the clothing mesh; Represents a collection of material textures; This refers to the material description file; This indicates region labels, texture index relationships, or subsequent simulation interface information.
[0051] The present invention also provides a method for generating virtual clothing consistent texture assets based on clothing region semantics and visibility reprojection constraints, which generates virtual clothing consistent texture assets using the above-mentioned virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints.
[0052] The method and system for generating consistent texture assets of virtual clothing based on clothing region semantics and visibility reprojection constraints provided by this invention have the following beneficial effects: This invention generates textures from multiple perspectives based on existing clothing meshes, ensuring consistency in color, pattern, boundary, and local details across the same clothing surface area from different perspectives. This achieves cross-view consistent texture generation based on existing clothing geometry. It incorporates region labels such as the front and back pieces, sleeves, neckline, cuffs, and hem into the texture generation process, making texture generation in different regions more consistent with the clothing structure and avoiding cross-regional texture confusion. This achieves texture generation using semantic constraints of clothing regions. Using depth maps and camera intrinsic and extrinsic parameters, pixels from one perspective are back-projected onto the 3D clothing surface and then onto another perspective. Visibility is determined based on depth consistency, allowing texture consistency constraints to be calculated only for validly visible areas, achieving reliable visibility reprojection constraints. Texture continuity is maintained at region boundaries such as cuffs, necklines, front and back piece junctions, and hem edges, reducing color abrupt changes, pattern misalignment, and obvious seams, achieving continuous blending at clothing region boundaries. Finally, the clothing mesh, material maps, region labels, and texture mapping relationships are uniformly exported to form a virtual clothing consistent texture asset for CG rendering and subsequent simulation processes, enabling reusability.
[0053] Specifically, this invention establishes an effective correspondence between different viewpoints by using visibility reprojection constraints, depth maps, and camera parameters, and applies consistency constraints only to common visible areas, thereby reducing color drift, pattern misalignment, and local texture breakage caused by independent generation per viewpoint, and improving the cross-view consistency of virtual clothing textures.
[0054] This invention utilizes clothing area tags to generate area tag maps and area-level prompts, enabling the texture generation process to distinguish different clothing areas such as the front piece, back piece, sleeves, neckline, cuffs, and hem, thereby reducing texture cross-area confusion and inconsistency in regional semantics, and improving the matching degree between texture and clothing area structure.
[0055] This invention introduces boundary continuity constraints for areas such as cuffs, necklines, hems, and the junction of front and back pieces, making the color and texture transitions between adjacent areas more natural, reducing problems such as color abrupt changes, pattern breaks, and obvious seams at the boundaries, and reducing seams and breaks at the boundaries of clothing areas.
[0056] This invention utilizes conditions such as normal maps, depth maps, light and dark gray model maps, contour maps, and region label maps to make the generated texture more consistent with clothing folds, boundaries, surface normals, and region structures, thereby improving the texture's fit on the three-dimensional clothing surface and enhancing the fit between the texture and the clothing geometry.
[0057] This invention uses visibility projection and texture fusion to uniformly map texture results from multiple perspectives onto the surface of a clothing mesh or the UV space, and further constructs a set of material maps such as basic color, normal, and roughness, thereby upgrading two-dimensional texture results to asset-level texture generation.
[0058] The output of this invention not only includes clothing geometry, but also material textures, region labels and mapping relationships that can be called by the rendering system. It can be directly entered into the rendering process of digital humans, virtual try-on, game animation or clothing display, thus improving the rendering usability of virtual clothing assets.
[0059] This invention integrates multi-view texture generation, visibility reprojection, region boundary blending, material map construction, and asset export into a unified process, reducing the workload of manually reorganizing textures, repairing seams, and maintaining mapping relationships across different software, and lowering subsequent manual processing costs.
[0060] In its extended implementation, this invention can output area labels, boundary constraints, or simulation parameter interfaces, providing an input basis for subsequent fabric physics simulation processes and retaining the interface basis for extending to simulable clothing assets.
[0061] In summary, this invention does not generate clothing geometry from scratch, nor does it directly generate complete clothing geometry from text, nor does it perform clothing physics simulation independently. Instead, it addresses existing clothing geometry, including existing clothing meshes, fixed-topology clothing templates, reconstructed clothing meshes, or high-fidelity refined clothing meshes. Through clothing region conditional rendering, region semantic-guided texture generation, visibility reprojection constraints, region boundary continuity blending, and material map export, it generates virtual clothing consistent texture assets that can be used for computer graphics rendering and subsequent simulation processes. This solves problems such as inconsistent clothing textures across multiple viewpoints, semantic confusion in clothing region textures, discontinuous textures at region boundaries, and difficulty in directly reusing texture assets. It can improve the cross-viewpoint consistency of clothing textures, the fit between textures and clothing region structures, the continuity of region boundaries, and the usability of clothing assets in the rendering process. Attached Figure Description
[0062] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a block diagram of the virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints provided by the present invention. Figure 2 This is a flowchart of the method for generating consistent texture assets of virtual clothing based on clothing region semantics and visibility reprojection constraints provided by the present invention. Detailed Implementation
[0063] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0064] Figure 1 This diagram illustrates a virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints, as shown in this embodiment. In this embodiment, the virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints includes: a clothing geometry and region label input module, a multi-view region conditional rendering module, a region semantic-guided consistent texture generation module, a visibility reprojection and region boundary fusion module, and a material mapping and texture asset export module. The clothing geometry and region label input module is used to obtain the clothing geometry representation based on the clothing mesh, UV coordinates and clothing region labels; The multi-view region conditional rendering module is used to perform conditional rendering from multiple viewing angles based on the clothing geometric representation and the set of camera viewpoints, and generate a multi-view conditional image. The region semantically guided consistent texture generation module is used to generate clothing texture images from multiple perspectives based on multi-view conditional images, clothing region labels, and overall material or style text. The visibility reprojection and region boundary fusion module is used to establish cross-view correspondence and cross-view consistency loss based on depth map and camera parameters. Based on the cross-view consistency loss and region boundary continuity loss, the clothing texture images from multiple viewpoints are projected and fused onto the three-dimensional clothing mesh surface to obtain the fused three-dimensional surface texture. The material mapping and texture asset export module is used to construct a consistent texture asset for virtual clothing based on the fused 3D surface texture.
[0065] In one exemplary embodiment, the expression for the clothing geometry is:
[0066]
[0067]
[0068]
[0069]
[0070] in, This represents the input clothing geometry mesh; Represents the set of vertices of the clothing mesh. Let be the i-th grid vertex, and N be the number of grid vertices; Represents a set of triangular facets. Let K be the j-th triangular facet, and K be the number of triangular facets. Represents the set of UV coordinates corresponding to a vertex. This represents the UV coordinates corresponding to the i-th vertex; This represents a set of clothing area labels. This is the label for the j-th clothing area.
[0071] In one exemplary embodiment, the multi-view conditional image includes a normal map, a depth map, a grayscale map, a contour map, and a region label map; In one exemplary embodiment, the region label map is rendered from clothing region labels, specifically including: at any viewpoint in the camera viewpoint set, projecting the clothing region label corresponding to each triangular facet in the clothing geometry representation onto a two-dimensional image plane to obtain a pixel-level region label map, which is used to indicate the clothing region to which each pixel belongs; As an exemplary embodiment, the garment area includes at least a front piece, a back piece, sleeves, a neckline, cuffs, and a skirt hem; In one exemplary embodiment, the process of generating clothing texture images from multiple viewpoints includes: Generate region-level prompts based on clothing area labels and overall material or style text; Using multi-view conditional images, region-level cue words, and overall material or style text as conditions, a conditional diffusion generation model is used to generate clothing texture images from multiple perspectives.
[0072] In one exemplary embodiment, the optimization objective of the conditional diffusion generation model is:
[0073] in, The optimization objective for the conditional diffusion generation model; This represents the noise added at diffusion time step t. And the mathematical expectation calculated after joint sampling of viewpoint number b; This indicates the added noise; This represents a denoising network; Represents the noisy texture image at the t-th diffusion time step; This represents the multi-view conditional image from the b-th perspective, where p represents the overall material or style text. Indicates a region-level prompt.
[0074] As an exemplary embodiment, the method for generating region-level prompt words utilizes a large language model; In one exemplary embodiment, the formula for calculating the three-dimensional surface texture is:
[0075]
[0076]
[0077] in, B represents the blended texture color of surface point s; B represents the number of viewpoints. Let be the visibility function, used to indicate whether a surface point s is visible from the b-th viewpoint. Otherwise ; This represents the fusion weight of the b-th viewpoint for surface point s; Represents the texture color or feature of a pixel in viewpoint b; This represents the projection function corresponding to the b-th viewpoint; To prevent extremely small constants with a denominator of zero; Let be the projected position of point s from the b-th viewpoint; This indicates taking the maximum value; The normal vector of surface point s; This indicates the direction of observation from the b-th perspective; This indicates the confidence or sharpness score of the texture generation for the corresponding pixel from that viewpoint. Indicates the boundary security score; These are the weighting coefficients.
[0078] In one exemplary embodiment, the expression for the cross-perspective correspondence is:
[0079]
[0080]
[0081] in, This represents the homogeneous pixel coordinates of a point on the surface of a 3D garment at viewpoint q. Let q represent the camera intrinsic parameters, rotation matrix, and translation vector, respectively; s represents a point on the 3D clothing surface. It is a rotation matrix; This represents a 3D point in the camera coordinate system corresponding to a pixel. It is a translation vector; This is the depth value; The pixel position at viewpoint b; This is the camera intrinsic parameter matrix; These are homogeneous coordinates.
[0082] In one exemplary embodiment, the formula for calculating the cross-view consistency loss is:
[0083]
[0084] in, For cross-view consistency loss; B represents the number of views; The pixel position at viewpoint b; This represents the candidate common visible region for viewpoints b and q. This is a visibility flag; used to indicate the pixel at viewpoint b. Whether the corresponding 3D surface points are visible at viewpoint q; Indicates consistency weight; Represents pixels in viewpoint b Texture color or features at the location; Represents pixels The corresponding 3D surface point is reprojected to the pixel position after the viewpoint q; This represents the texture color or feature at the corresponding pixel after reprojection onto the viewpoint q; This represents the first norm, used to sum the absolute values of the components of the two texture color or texture feature difference vectors within the parentheses; The predicted depth of a pixel after it has been projected onto the viewpoint q; This represents the depth value at this location in the target's perspective depth map. This represents the depth tolerance threshold.
[0085] In one exemplary embodiment, the formula for calculating the continuity loss at the region boundary is:
[0086] in, For continuous loss at the regional boundary; For any pair of adjacent regions; This is a group of adjacent clothing areas; and These represent texture sampling points at the boundary of two adjacent regions; Indicates that point s is in the region The blended texture colors at the boundary This indicates that point s is in the adjacent region. The blended texture color at the boundary.
[0087] In one exemplary embodiment, the process of constructing a consistent texture asset for virtual clothing includes: The blended 3D surface texture is mapped onto the UV plane to obtain the base color map, as shown in the formula:
[0088]
[0089] in, Represents UV coordinates The base color texture value at that location; Represents the UV plane coordinates; This represents the blended texture color of surface point s; This represents the UV coordinates corresponding to surface point s; Based on the aforementioned base color map, construct a set of material maps, as shown in the formula:
[0090] in, A collection of material textures; This represents the base color map value. This represents the normal map. Represents roughness map, Indicates transparency map, This indicates a height map or a displacement map; Based on the aforementioned material map set, a consistent texture asset for the virtual clothing is obtained, as shown in the formula:
[0091] in, This represents the final exported texture asset; Represents the geometric mesh of the garment; This represents the UV coordinates corresponding to the clothing mesh; Represents a collection of material textures; This refers to the material description file; This indicates region labels, texture index relationships, or subsequent simulation interface information.
[0092] In one exemplary embodiment, during the training or optimization phase, the overall objective function is:
[0093] in, This represents the diffusion denoising loss. This represents the cross-view consistency loss based on visibility reprojection. Indicates continuous loss at the regional boundary. This indicates a texture smoothing or regularization term. These represent the weighting coefficients of the corresponding loss terms.
[0094] Through the above objective function, the present invention can simultaneously constrain texture generation quality, cross-view consistency, region boundary continuity, and overall texture smoothness.
[0095] like Figure 2As shown, this embodiment provides a method for generating consistent texture assets of virtual clothing based on clothing region semantics and visibility reprojection constraints. The method generates consistent texture assets of virtual clothing using the aforementioned consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints.
[0096] In some embodiments, the above-described virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints can also be implemented in the following ways.
[0097] In this embodiment, the virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints mainly includes: clothing geometry and region label input module, multi-view region conditional rendering module, region semantic-guided consistent texture generation module, visibility reprojection and region boundary fusion module, and material mapping and texture asset export module.
[0098] The system comprises several modules: a clothing geometry and region label input module to receive clothing meshes, UV coordinates, and region labels for texture generation; a multi-view region conditional rendering module to generate conditional images such as normal maps, depth maps, grayscale maps, contour maps, and region label maps from multiple viewing perspectives; a region semantic-guided consistent texture generation module to generate clothing textures with consistent style and reasonable regional structure from different perspectives based on overall text, region labels, and multi-view geometric conditions; a visibility reprojection and region boundary fusion module to establish cross-view correspondence based on depth maps and camera parameters, and to perform consistency constraints and fusion at visible regions and region boundaries; and a material mapping and texture asset export module to convert the fusion results into material maps such as basic color maps, normal maps, and roughness maps, and to export them together with the clothing mesh as a consistent texture asset for virtual clothing.
[0099] 1. Module input-output relationship The input-output relationships of each module in this invention are as follows.
[0100] (1) Clothing geometry and region label input module: Input clothing mesh, UV coordinates, clothing region labels and optional material text descriptions; output standardized clothing geometry representation.
[0101] (2) Multi-view area condition rendering module: Input clothing geometry, area labels and camera view set; Output multi-view normal map, depth map, light and dark gray model map, outline map and area label map.
[0102] (3) Region semantic guidance consistent texture generation module: input multi-view conditional image, overall material text and region prompt words; output clothing texture images from multiple perspectives.
[0103] (4) Visibility reprojection and region boundary fusion module: Input multi-view texture image, depth map, camera parameters, region label map and visibility relationship; output fused 3D surface texture or UV texture.
[0104] (5) Material Mapping and Texture Asset Export Module: Input blended texture, clothing mesh, UV coordinates and region labels; output base color map, optional material map and virtual clothing consistent texture asset package.
[0105] 2. Clothing geometry and area label input The input to this invention is existing garment geometry data. This garment geometry can originate from 3D modeling, garment scanning, garment reconstruction, garment simulation, or the result of fixed topology geometry refinement. In one embodiment, the garment mesh is represented as: (1) in, This represents the input clothing geometry mesh; Represents the set of vertices of the clothing mesh. ; Represents a set of triangular facets; Represents the set of UV coordinates corresponding to a vertex; The set of garment area labels represents a collection of garment area labels that can be predefined on the triangular faces or vertices of the garment grid to identify different structural areas of the garment. For example, the triangular faces can be divided into front piece, back piece, left sleeve, right sleeve, neckline, cuff, hem, waist, edge area, or decorative area, etc.
[0106] Unlike conventional 3D texture generation methods, this invention not only uses clothing geometric meshes as input, but also explicitly uses clothing region tags. This region label is used to subsequently generate region label maps, region-level prompts, and region boundary continuity constraints, ensuring that the generated texture conforms to the structural semantics of the clothing itself.
[0107] 3. Conditional rendering of multi-view areas To ensure the generated texture matches the clothing geometry, this invention first performs conditional rendering of the clothing mesh from multiple perspectives. Let the set of multi-view cameras be: (2) Where B represents the number of viewpoints. This represents the camera parameters for the b-th viewpoint. Camera parameters include the camera intrinsic matrix, rotation matrix, translation vector, field of view, and image resolution.
[0108] For each perspective For clothing mesh Render the image to obtain the corresponding geometric condition image: (3) in, This represents a conditional rendering function. Let represent the set of conditional images from the b-th viewpoint.
[0109] In one implementation, Including normal map, depth map, light and dark grayscale map, contour map, and region label map, it can be represented as: (4) in, A normal diagram is used to describe the orientation of a surface. A depth map is used to describe spatial distances; This represents a grayscale image, used to preserve the light and shadow structures formed by wrinkles and surface undulations; This represents a silhouette diagram used to constrain cuffs, collars, and boundary areas. This represents a region label image, used to distinguish different parts of the clothing.
[0110] Area label map By the area labels on the clothing mesh Rendered. Specifically, from the perspective... Next, the region label corresponding to each triangular facet is projected onto the two-dimensional image plane to obtain a pixel-level region label map. This region label map is used to indicate whether each pixel belongs to the front piece, back piece, sleeve, neckline, cuff, skirt, or other garment area.
[0111] With the aforementioned multi-view geometric conditions, the subsequent texture generation process not only relies on text or color references, but is also constrained by clothing geometric boundaries, surface normals, fold undulations, and region structures, thereby reducing the problem of texture mismatch with geometry.
[0112] 4. Region-based semantic-guided consistent texture generation After obtaining the multi-view geometric condition image, this invention inputs it into the multi-view... Figure 1 The texture generation module generates clothing texture images from multiple viewpoints. Let the texture image generated from the b-th viewpoint be... The multi-view texture result is then represented as: (5) in, This represents the set of textures generated from all viewpoints. In one implementation, the system also receives an overall material text description p, such as "gray wool coat," "light-colored cotton dress," or "blue denim jacket." To make the texture generation for different clothing areas more consistent with the area structure, this invention generates area-level prompts based on the overall material text and area labels: (6) in, This indicates the region-level prompt word corresponding to region r. This represents the function for generating area prompts, where p represents the overall material or style text, and r represents the clothing area label.
[0113] For example, when the overall text is "gray wool coat", the following area prompts can be generated: Front panel area: Gray wool front panel with continuous texture; Back section: Gray wool back panel, style consistent with the front panel; Cuff area: Gray wool cuffs with clear edges; Neckline area: Gray wool collar with a clear outline and texture consistent with the body of the garment.
[0114] By using region-level cue words, the texture generation model can generate more reasonable textures for different clothing areas while maintaining overall style consistency.
[0115] In one implementation, the region semantically guided consistent texture generation module employs a conditional diffusion generation model. The diffusion model generates textures based on multi-view geometric conditional images. Overall text p and region-level prompts As a condition, the generated result must simultaneously satisfy the requirements of geometric structure, region semantics, and material style. Its training or optimization objective can be expressed as: (7) in, This represents the noisy texture image at the t-th diffusion time step. This indicates the added noise. This represents a denoising network. Let p represent the geometric condition image at the b-th viewpoint, and p represent the overall text texture. Indicates a region-level prompt.
[0116] The function of this module is to generate clothing textures based on multi-view geometry, so that the clothing textures under different viewpoints are consistent in color, pattern, style and boundary areas, rather than being generated independently for each viewpoint.
[0117] 5. Visibility reprojection constraint To establish a reliable pixel correspondence between different viewpoints, this invention introduces a visibility reprojection constraint. This constraint uses a depth map and camera parameters to backproject a pixel from one viewpoint onto the surface of a 3D garment, then onto another viewpoint, and determines whether the point is visible based on depth consistency.
[0118] Let the pixel position under viewpoint b be Its homogeneous coordinates are Depth value The camera intrinsic parameter matrix is: The rotation matrix is The translation vector is First, the pixels Back projection onto the camera coordinate system of viewpoint b: (8) in, This represents the 3D point in the camera coordinate system corresponding to that pixel.
[0119] Then, transform the point to the world coordinate system: (9) Here, s represents a point on the three-dimensional clothing surface. Next, this three-dimensional point is projected onto another viewpoint q: (10) in, Let represent the camera intrinsic parameters, rotation matrix, and translation vector of the viewpoint q, respectively. This represents the homogeneous pixel coordinates of the 3D point at viewpoint q. To determine whether the point is visible in the target viewpoint q, this invention introduces a depth consistency judgment. Let the predicted depth of the point projected onto viewpoint q be... The depth value of the target's perspective depth map at this location is The visibility flag is then defined as: (11) in, This represents the depth tolerance threshold. When... When, it means that the 3D surface point corresponding to the pixel is visible in the viewpoint q; when When this condition is met, it indicates that the point is occluded or its projection is invalid, and it does not participate in the cross-view consistency constraint. Based on the above visibility reprojection relationship, the cross-view consistency loss of this invention can be expressed as: (12) in, This represents the candidate commonly visible region for viewpoints b and q. Represents the consistency weight. Represents the pixels in viewpoint b Texture color or features at the location, This represents the texture color or feature at the corresponding pixel after reprojection to viewpoint q. Equations 8 to 12 illustrate how this invention establishes a reliable cross-view texture correspondence through depth backprojection, 3D coordinate transformation, target viewpoint projection, and occlusion judgment. Compared to methods that only perform image-level consistency constraints, this invention can avoid incorrectly including occluded areas in consistency constraints, thereby improving the stability of multi-view texture generation.
[0120] 6. Continuous integration of regional boundaries Multiview Figure 1 The texture generation module obtains two-dimensional texture images from different perspectives. In order to create a usable virtual clothing asset, these texture results need to be projected and fused onto the surface of the three-dimensional clothing mesh, while ensuring texture continuity at the boundaries of the clothing area.
[0121] Let s be a point on the surface of the garment. The projected position of this point from the b-th viewpoint is: (13) in, Let represent the projection function corresponding to the b-th viewpoint. If surface point s is visible from the b-th viewpoint, then the visibility function is denoted as . Otherwise This invention calculates the final texture color of point s on the clothing surface using a weighted fusion method: (14) in, This represents the blended texture color of surface point s. This represents the fusion weight of the b-th viewpoint for that surface point. To prevent extremely small constants with a denominator of zero, and to avoid overly ambiguous fusion weights, this invention provides a specific weight calculation method: (15) in, This represents the normal vector of surface point s. This indicates the direction of observation from the b-th perspective. This represents the confidence or sharpness score of the texture generation for the corresponding pixel at that viewpoint. Represents the boundary security weights. These are the weighting coefficients.
[0122] In this weighting system, the closer the surface normal is to the current viewing angle, the higher the weight; the clearer the generated texture, the higher the weight; and the farther away from the occlusion boundary, image boundary, or region boundary, the higher the boundary safety weight. This method prioritizes viewing angles with better perspectives, less occlusion, and higher generation quality, thereby reducing texture seams and local misalignments.
[0123] Furthermore, this invention introduces boundary continuity constraints for the boundaries of clothing areas. Let the set of adjacent clothing areas be... Where any pair of adjacent regions is represented as .set up and Let represent the texture sampling points at the boundary of two adjacent regions, respectively. Then, the region boundary continuity loss can be expressed as: (16) in, Indicates the region The blended texture colors at the boundary Indicates adjacent regions The blended texture color at the boundary.
[0124] This constraint is used to reduce color abruptness and texture breaks at the boundaries of garment areas. For example, at the junction of sleeves and body, the side seams of the front and back pieces, the neckline, and the hem, boundary continuity constraints can make texture transitions more natural.
[0125] 7. Material Mapping and Asset Export After completing the surface texture blending, this invention maps the blended surface texture onto the UV plane based on the UV coordinates of the clothing mesh to obtain a base color map: (17) in, Represents the UV plane coordinates; Represents UV coordinates The base color texture value at that location; This represents the UV coordinates corresponding to surface point s.
[0126] Based on this, the system can further construct a set of material textures: (18) in, This represents the base color texture. This represents the normal map. Represents roughness map, Indicates transparency map, This indicates a height map or a displacement map.
[0127] In one implementation, the base color map is obtained by directly mapping the blended surface color onto the UV plane; the normal map can be obtained from the surface normal of the clothing mesh, the result of local texture detail enhancement, or the prediction of the generation model; the roughness map can be obtained by mapping based on the material text description, region label, or preset material table; the transparency map can be generated based on the clothing cutout, tulle, or semi-transparent region label; the height map or displacement map can be obtained based on texture details, gray model light and dark structure, or local undulation estimation.
[0128] The final exported consistent texture asset for virtual clothing can be represented as: (19) in, This represents the final exported virtual clothing consistent texture asset package; Represents the geometric mesh of the garment; This represents the UV coordinates corresponding to the clothing mesh; Represents a collection of material textures; This refers to the material description file; This indicates region labels, texture index relationships, or subsequent simulation interface information.
[0129] In one implementation, the asset export format can be OBJ, FBX, glTF, USD, or other 3D asset formats; texture maps can be exported as image formats such as PNG, JPG, TIFF, and EXR. The exported assets can be used for digital human rendering, virtual try-on, clothing display, game animation, and subsequent simulation processes.
[0130] 8. Overall Optimization Goals During the training or optimization phase, the overall objective function of this invention can be expressed as:
[0131] in, This represents the diffusion denoising loss. This represents the cross-view consistency loss based on visibility reprojection. Indicates continuous loss at the regional boundary. This indicates a texture smoothing or regularization term. These represent the weighting coefficients of the corresponding loss terms.
[0132] Through the above objective function, the present invention can simultaneously constrain texture generation quality, cross-view consistency, region boundary continuity, and overall texture smoothness.
[0133] In some embodiments, the above-described method for generating consistent texture assets of virtual clothing based on clothing region semantics and visibility reprojection constraints can also be implemented in the following ways.
[0134] In this embodiment, the method for generating consistent texture assets for virtual clothing based on clothing region semantics and visibility reprojection constraints includes the following steps: Step S1: Input the clothing geometry to be generated for the texture. The clothing geometry includes clothing mesh vertices, faces, UV coordinates, and clothing region labels.
[0135] Step S2: Perform multi-view area condition rendering on the clothing mesh based on multiple preset perspectives to generate one or more of the following: normal map, depth map, light and dark gray model map, outline map, and area label map.
[0136] Step S3: Generate region-level prompts based on the overall material text and clothing region labels, so that different clothing regions have corresponding semantic control conditions.
[0137] Step S4: Input the multi-view conditional image, overall text, and region-level prompts into the region semantic-guided consistent texture generation module to generate clothing texture images from multiple perspectives.
[0138] Step S5: Based on the depth map and camera parameters, backproject the pixels from one viewpoint onto the 3D clothing surface, then project them onto another viewpoint, and determine whether the point is visible by using depth consistency.
[0139] Step S6: Calculate cross-view texture consistency constraints only for clothing surface areas that are commonly visible from different viewpoints to reduce color drift and pattern misalignment.
[0140] Step S7: Based on the visibility relationship of the clothing surface, surface normal, generation confidence and boundary safety weight, perform weighted fusion of textures generated from different viewpoints.
[0141] Step S8: Apply continuity constraints to the boundaries of adjacent areas of the garment to reduce texture breaks and color abrupt changes at the junctions of cuffs, necklines, skirts, and front and back pieces.
[0142] Step S9: Map the blended surface texture onto the UV plane to build a base color map, and further generate normal maps, roughness maps, transparency maps and height maps.
[0143] Step S10: Package and export the clothing mesh, UV coordinates, material maps, material description files, and region tags to form a consistent texture asset for the virtual clothing.
[0144] In some embodiments, the above-described method for generating consistent texture assets of virtual clothing based on clothing region semantics and visibility reprojection constraints can also be implemented in the following ways.
[0145] In one specific implementation, a geometrically reconstructed or refined dress mesh is input, comprising vertices, triangles, UV coordinates, and region labels. Region labels include front, back, neckline, cuffs, hem, and edge regions.
[0146] The system renders the clothing mesh from six directions: front view, back view, left view, right view, oblique front view, and oblique back view, obtaining a normal map, depth map, grayscale map, outline map, and region label map for each viewpoint. The region label map is rendered from the predefined region labels of each triangular facet in the clothing mesh.
[0147] Subsequently, the system receives a text description of the material, such as "light-colored cotton dress." Based on this text and region labels, the system generates region-level cue words, such as "light-colored cotton front panel, continuous texture," "light-colored cotton skirt, natural boundaries," and "light-colored cotton neckline, clear outline." These region-level cue words, along with the multi-view conditional image, are input into the texture generation model to generate clothing texture images from six perspectives.
[0148] During the cross-view consistency constraint phase, the system uses depth maps and camera parameters from each viewpoint to backproject pixels from one viewpoint onto points on the 3D clothing surface, and then projects them onto other viewpoints. It then determines whether the surface point is visible in the target viewpoint based on depth consistency. Only when the surface point is visible in both viewpoints does the system apply consistency constraints to its texture color or texture features.
[0149] During the texture fusion stage, the system calculates the fusion weight for each viewpoint based on the angle between the surface normal and the viewpoint, the generation confidence, and the boundary safety weight, and performs weighted fusion on the texture results under the visible viewpoint. At the same time, for the neckline, cuffs, hem, and the junction of the front and back pieces, the system uses the continuous loss constraint of the region boundary to constrain the texture differences on both sides of the boundary, thereby reducing color abrupt changes and seams.
[0150] Finally, based on the UV coordinates of the clothing mesh, the system unfolds the blended surface texture into a base color map, and can generate a normal map based on the surface normals and texture details, and a roughness map based on the material text and region labels. The asset export module packages and exports the clothing mesh, UV coordinates, base color map, normal map, roughness map, material description file, and region labels to obtain a consistent texture asset for virtual clothing that can be used for CG rendering.
[0151] It should be noted that the number of multiple views in this invention is not limited to six; it can also be four, eight, twelve, or the number of views can be automatically determined according to the clothing structure.
[0152] It should be noted that the clothing area label of the present invention is not limited to the front piece, back piece, sleeve, neckline, cuff and hem, but can also be extended to the waist, hat, pocket, placket, decorative piece, pleated area or other areas depending on the type of clothing.
[0153] It should be noted that the geometric condition image of the present invention is not limited to normal map, depth map and grayscale map, but may also include curvature map, boundary map, occlusion map, region label map or UV visibility map.
[0154] It should be noted that the texture generation model of the present invention is preferably a conditional diffusion model, but other image generation models with conditional generation capabilities may also be used.
[0155] It should be noted that the texture fusion method of the present invention is not limited to weighted average, and may also employ Poisson fusion, graph cut optimization, neural texture field fusion, or local selection based on confidence.
[0156] It should be noted that the material texture of the present invention is not limited to the basic color texture, and can also generate normal maps, roughness maps, transparency maps, metallic maps, height maps or displacement maps according to the requirements of the target platform.
[0157] It should be noted that the export format of this invention is not limited to OBJ, FBX, glTF or USD, and can also be exported in the corresponding format according to the requirements of different rendering engines, virtual fitting systems or digital human platforms.
[0158] It should be noted that, based on existing clothing geometry, this invention uses clothing region labels and region-level cue words to guide multi-view texture generation, and establishes a reliable texture correspondence between different viewpoints through visibility reprojection using depth maps and camera parameters. At the same time, it continuously blends the boundaries of clothing regions, ultimately generating a consistent virtual clothing texture asset that can be directly used in the CG rendering process.
[0159] It should be noted that this invention does not focus on generating clothing geometry from scratch, nor does it focus on predicting fabric physical parameters independently. Simulation-related content mainly exists as subsequent extension interfaces, such as region labels, boundary constraints, or simulation parameter interfaces.
[0160] It should be noted that, compared to general 3D texture generation methods, this invention does not target arbitrary 3D objects, but rather virtual clothing meshes with well-defined regional structures. This invention not only utilizes multi-view geometry for texture generation, but also introduces clothing region labels, region-level cue words, occlusion visibility judgment, and region boundary continuity constraints, enabling the generated texture to maintain more stable structural consistency across clothing regions such as cuffs, necklines, hems, front panels, and back panels. Compared to methods that only output multi-view images, this invention further generates directly exportable clothing texture assets through visibility reprojection and UV surface fusion.
[0161] It should be noted that this invention uses existing clothing geometry as input, rather than generating clothing geometry from scratch, and focuses on solving the problem of texture consistency completion of existing clothing assets; it uses clothing region tags to generate region label maps and region-level prompts, giving different clothing regions clear semantic constraints during texture generation; it establishes visibility reprojection relationships through depth maps, camera intrinsic and extrinsic parameters, and applies cross-view texture consistency constraints only to commonly visible areas; it reduces color abruptness and texture breaks at cuffs, collars, hems, and the junction of front and back pieces through region boundary continuity constraints; and it transforms multi-view texture results into reusable material map assets through surface projection, visibility fusion, and UV mapping.
[0162] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints, characterized in that, It includes modules for clothing geometry and region label input, multi-view region conditional rendering, region semantic-guided consistent texture generation, visibility reprojection and region boundary fusion, and material mapping and texture asset export. The clothing geometry and region label input module is used to obtain the clothing geometry representation based on the clothing mesh, UV coordinates and clothing region labels; The multi-view region conditional rendering module is used to perform conditional rendering from multiple viewing angles based on the clothing geometric representation and the set of camera viewpoints, and generate a multi-view conditional image. The region semantically guided consistent texture generation module is used to generate clothing texture images from multiple perspectives based on multi-view conditional images, clothing region labels, and overall material or style text. The visibility reprojection and region boundary fusion module is used to establish cross-view correspondence and cross-view consistency loss based on depth map and camera parameters. Based on the cross-view consistency loss and region boundary continuity loss, the clothing texture images from multiple viewpoints are projected and fused onto the three-dimensional clothing mesh surface to obtain the fused three-dimensional surface texture. The material mapping and texture asset export module is used to construct a consistent texture asset for virtual clothing based on the fused 3D surface texture.
2. The virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints according to claim 1, characterized in that, The expression for the geometric representation of the clothing is: in, This represents the input clothing geometry mesh; Represents the set of vertices of the clothing mesh. Let be the i-th grid vertex, and N be the number of grid vertices; Represents a set of triangular facets. Let K be the j-th triangular facet, and K be the number of triangular facets. Represents the set of UV coordinates corresponding to a vertex. This represents the UV coordinates corresponding to the i-th vertex; This represents a set of clothing area labels. This is the label for the j-th clothing area.
3. The virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints according to claim 1, characterized in that, The process of generating clothing texture images from multiple perspectives includes: Generate region-level prompts based on clothing area labels and overall material or style text; Using multi-view conditional images, region-level cue words, and overall material or style text as conditions, a conditional diffusion generation model is used to generate clothing texture images from multiple perspectives.
4. The virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints according to claim 3, characterized in that, The optimization objective of the conditional diffusion generation model is: in, The optimization objective for the conditional diffusion generation model; Let represent the mathematical expectation calculated after joint sampling of diffusion time step t, noise ε, and viewpoint number b; t represents the diffusion time step. This represents noise added to the texture image; This represents the noise predicted by the denoising network. Represents the noisy texture image at the t-th diffusion time step; This represents the multi-view conditional image from the b-th perspective, where p represents the overall material or style text. Indicates a region-level prompt.
5. The virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints according to claim 1, characterized in that, The formula for calculating the three-dimensional surface texture is: in, B represents the blended texture color of surface point s; B represents the number of viewpoints. Let be the visibility function, used to indicate whether a surface point s is visible from the b-th viewpoint. Otherwise ; This represents the fusion weight of the b-th viewpoint for surface point s; Represents the texture color or feature of a pixel in viewpoint b; This represents the projection function corresponding to the b-th viewpoint; To prevent extremely small constants with a denominator of zero; Let be the projected position of point s from the b-th viewpoint; This indicates taking the maximum value; The normal vector of surface point s; This indicates the direction of observation from the b-th perspective; This indicates the confidence or sharpness score of the texture generation for the corresponding pixel from that viewpoint. Indicates the boundary security score; These are the weighting coefficients.
6. The virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints according to claim 1, characterized in that, The expression for the cross-perspective correspondence is: in, This represents the homogeneous pixel coordinates of a point on the surface of a 3D garment at viewpoint q. Let q represent the camera intrinsic parameters, rotation matrix, and translation vector, respectively; s represents a point on the 3D clothing surface. It is a rotation matrix; This represents a 3D point in the camera coordinate system corresponding to a pixel. It is a translation vector; This is the depth value; The pixel position at viewpoint b; This is the camera intrinsic parameter matrix; These are homogeneous coordinates.
7. The virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints according to claim 1, characterized in that, The formula for calculating the cross-perspective consistency loss is as follows: in, For cross-view consistency loss; B represents the number of views; b and q represent different view numbers, and q is not equal to b; The pixel position at viewpoint b; This represents the candidate common visible region for viewpoints b and q. This is a visibility flag used to indicate the pixel position at viewpoint b. Whether the corresponding 3D surface points are visible at viewpoint q; Indicates consistency weight; Represents pixels in viewpoint b Texture color or features at the location; Represents pixels The corresponding 3D surface point is reprojected to the pixel position after the viewpoint q; This represents the texture color or feature at the corresponding pixel after reprojection onto the viewpoint q; The L1 norm is used to measure the absolute difference between two texture colors or texture features. The predicted depth of a pixel after it has been projected onto the viewpoint q; This represents the depth value at this location in the target's perspective depth map. This represents the depth tolerance threshold.
8. The virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints according to claim 1, characterized in that, The formula for calculating the continuous loss at the boundary of the region is as follows: in, For continuous loss at the regional boundary; For any pair of adjacent regions; This is a group of adjacent clothing areas; and These represent texture sampling points at the boundary of two adjacent regions; Indicates that point s is in the region The blended texture colors at the boundary This indicates that point s is in the adjacent region. The blended texture color at the boundary.
9. The virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints according to claim 1, characterized in that, The process of constructing consistent texture assets for virtual clothing includes: The blended 3D surface texture is mapped onto the UV plane to obtain the base color map, as shown in the formula: in, Represents UV coordinates The base color texture value at that location; Represents the UV plane coordinates; This represents the blended texture color of surface point s; This represents the UV coordinates corresponding to surface point s; Based on the aforementioned base color map, construct a set of material maps, as shown in the formula: in, A collection of material textures; This represents the base color map value. This represents the normal map. Represents roughness map, Indicates transparency map, This indicates a height map or a displacement map; Based on the aforementioned material map set, a consistent texture asset for the virtual clothing is obtained, as shown in the formula: in, This represents the final exported texture asset; Represents the geometric mesh of the garment; This represents the UV coordinates corresponding to the clothing mesh; Represents a collection of material textures; This refers to the material description file; This indicates region labels, texture index relationships, or subsequent simulation interface information.
10. A method for generating consistent texture assets for virtual clothing based on clothing region semantics and visibility reprojection constraints, characterized in that, Virtual clothing consistent texture assets are generated using the virtual clothing consistent texture asset generation system based on clothing region semantics and visibility reprojection constraints as described in any one of claims 1-9.