Method for optimizing texture quality of generative 3D model based on texture expansion graph
By constructing a multi-view generation model based on texture expansion maps and UV space image restoration technology, the generative 3D model texture is optimized, the problem of inconsistency between texture and reference image is solved, the generation of high-quality multi-view texture is achieved, and the detail and integrity of the texture are improved.
Patent Information
- Application Number
- CN202510829923.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
Smart Images

Figure CN120707611A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a method for optimizing the texture quality of a generated 3D model based on a texture expansion graph. Background Art
[0002] With the rapid development of computer vision and 3D modeling technologies, generating 3D models with realistic textures has become a core requirement for the digital transformation of many industries. Realistic 3D reconstruction requires not only geometric accuracy but also high-quality texture mapping and multi-view consistency. Enhancing the texture quality of 3D models is not only a technical upgrade but also provides underlying support for data analysis, automated decision-making, and user experience optimization. Its potential application value is increasingly prominent in areas such as the Metaverse and Industry 4.0. However, existing technologies for 3D model texture generation and optimization still face many challenges.
[0003] Existing generation technologies are already able to generate textures that are semantically close to the target content based on text descriptions. However, when a reference image is input as a constraint, existing texture generation methods often find it difficult to fully preserve the style features and local details of the reference image, resulting in significant differences in color, texture distribution, or local structure between the generated texture and the reference image, making it difficult to meet application requirements for high fidelity or consistency. This is especially true in scenarios where the generated texture needs to be applied to three-dimensional models or multi-view rendering. In addition, existing diffusion models that generate multiple views of objects based on reference images still face challenges in maintaining geometric consistency with the reference image. Due to the lack of clear three-dimensional structural constraints, such methods are prone to inconsistencies between perspectives during the generation process, resulting in the generated image not matching the reference image in terms of geometry. When the generated texture is pasted back onto the three-dimensional model, problems such as obvious artifacts often occur. The texture quality of the 3D model obtained in the generative 3D reconstruction task is low. When the texture generation model performs texture generation tasks for the 3D model based on the reference image, the texture is often inconsistent with the reference image. When the multi-view generation model is used to generate multiple views based on the reference image, the generated multiple views are inconsistent with the geometry of the 3D model. When performing texture mapping, the texture of the 3D model will have artifacts, obvious texture fusion seams, and multiple views. Figure 1 Insufficient consistency. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes a method for optimizing the texture quality of a generative 3D model based on a texture expansion map, which relates to the field of computer vision. Rendering a 3D model obtains a multimodal image set under multiple perspectives; obtaining a training data set and training a multi-view generation model; using the multi-view generation model to generate an RGB image set; back-projecting high-quality RGB images under six perspectives onto the surface of the 3D model, and then mapping the color of the 3D point to the UV space through UV mapping to obtain a multi-perspective texture map and a mask map of the UV space under six perspectives; fusing the multi-perspective texture maps to obtain a rough texture map; optimizing the rough texture map through an image restoration model to obtain a high-quality complete texture expansion map. The present invention optimizes the texture quality of a generative 3D model through a multi-view generation model and image restoration technology in UV space, and effectively solves the problem of artifacts and distortion caused by the inconsistency between the image content and the geometry of the three-dimensional model when generating high-quality multi-views based on images.
[0005] A method for optimizing the texture quality of a generative 3D model based on a texture unfolding graph, the specific technical solution includes the following steps:
[0006] Step 1: Render the 3D model to obtain a multimodal image set from multiple perspectives; the multimodal image set includes six images, each image including a textured RGB image, a depth map, a normal map, and a position map; the six images are acquired from six perspectives, with only one photo captured from each perspective;
[0007] The 3D model is placed at an upright viewing angle, the pitch angle is set to 0°, and the azimuth angles are respectively 0°, 45°, 90°, 180°, 270°, and 315°, to obtain 6 images under the six viewing angles;
[0008] Step 2: Obtain a training dataset and train a geometrically controllable multi-view generation model based on the diffusion model;
[0009] Step 3: Generate an RGB image set using the trained multi-view generative model; the RGB image set includes RGB images with high resolution and high texture clarity under six viewing angles; the RGB images with high resolution and high texture clarity under the six viewing angles are high-quality RGB images; the high-quality RGB images are geometrically consistent with the 3D model and have the same fidelity as the reference image; the reference image is a random RGB image; the RGB image rendered under random camera poses is a random RGB image;
[0010] Step 4: The high-quality RGB image is back-projected and mapped to obtain UV space information.
[0011] Step 5: Fuse the multi-view texture maps to obtain a rough texture map;
[0012] Step 6: Optimize the rough texture map through the image restoration model to obtain a high-quality complete texture unfolded map.
[0013] Furthermore, in step 2, the multi-view generation model uses the Stable Diffusion model as the basic model; the Stable Diffusion model includes a conditional encoder and a parallel attention mechanism module; the conditional encoder and the parallel attention mechanism module work together to ensure the multi-view Figure 1 The conditional encoder encodes the depth map and position map into multi-scale spatial features, and injects the multi-scale spatial features into different layers of the Stable Diffusion model as geometric control information to guide the generation of multi-views. The original Stable Diffusion model organizes the spatial self-attention layer and the text cross-attention layer in a serial manner. The parallel attention mechanism module converts the serial architecture into a parallel architecture by retaining the pre-trained weights of the original two-layer architecture of the Stable Diffusion model, constructs the multi-view attention layer and the image cross-attention layer according to the network structure of the spatial self-attention layer, and initializes the weights of the multi-view attention layer and the image cross-attention layer to 0. In order to utilize the prior information of the spatial attention layer, the parallel attention mechanism module reorganizes the order of different layers and organizes the spatial self-attention layer, the text cross-attention layer and the image cross-attention layer in parallel. The outputs of the spatial self-attention layer, the text cross-attention layer and the image cross-attention layer are the input of the text cross-attention layer, thereby ensuring that the new attention layer fully inherits the prior knowledge of the pre-trained self-attention layer and realizes efficient learning with geometric knowledge.
[0014] Furthermore, in step 2, the process of obtaining the training data set is as follows:
[0015] First, high-quality 3D models containing single objects are screened from public 3D datasets;
[0016] Then, data with occlusion or complex background is eliminated;
[0017] Next, for each object, we render multiple RGB images, depth images, position images, and normal images under fixed camera poses, and also render random RGB images under random camera poses.
[0018] Finally, the random RGB image is used as the reference image, the image obtained by splicing the depth map and position map under fixed camera pose is used as the control information, and the 6 RGB images under fixed camera pose are used as supervision images in the training process.
[0019] Furthermore, in step 3, the RGB image set generation process is as follows:
[0020] First, the depth maps and position maps of the six images are spliced in the channel dimension to obtain a multi-channel control map;
[0021] Then, the reference image and the multi-channel control image are input into the multi-view generation model;
[0022] Finally, the multi-view generation model is based on the prior knowledge of the pre-trained diffusion model, using a conditional encoder and a parallel attention mechanism. Under the geometric information constraints of the depth map and position map, it outputs six RGB images of the 3D model with a pitch angle of 0° and azimuth angles of 0°, 45°, 90°, 180°, 270°, and 315° respectively. The six RGB images are a set of RGB images, and the six RGB images are geometrically consistent with the 3D model, thus avoiding artifacts caused by directly back-projecting multiple views back to the 3D model when the geometry is inconsistent.
[0023] Furthermore, in step 4, the UV space information includes multi-view texture maps, mask maps and weight information of the UV space under six viewing angles;
[0024] The UV space information acquisition method is as follows:
[0025] First, based on the depth information in the depth map, the RGB images under six perspectives are back-projected onto the 3D model surface, and the color information of the 3D points is mapped to the UV space through UV mapping to generate the corresponding texture map;
[0026] Secondly, the weight information is generated by the angle deviation between the normal map and the camera viewing direction at each viewing angle, and the weight information in the 2D space is converted to the UV space through UV mapping;
[0027] Finally, the visibility information of each pixel in the texture map at the current viewing angle is recorded. The visibility information is the mask map.
[0028] Furthermore, in step 5, the fusion of multi-view texture maps adopts an iterative weighted averaging method;
[0029] The fusion process of the iterative weighted average method for multiple perspective texture maps is as follows:
[0030] Based on the weight information of the UV space under six viewing angles, the texture map under each viewing angle, the UV space weight information and the visibility information of each pixel in the UV space are iteratively accumulated to finally generate a fused rough texture map.
[0031] Furthermore, in step 6, the steps of optimizing the rough texture map by the image restoration model are as follows:
[0032] Combining visibility information from multiple perspectives and using image restoration technology, the missing parts in the rough texture map are filled, seams are eliminated, and discontinuities are eliminated based on the geometric and color consistency constraints between multi-view textures to achieve consistency with the surrounding texture.
[0033] The technical effects of the present invention are as follows:
[0034] By constructing a geometrically controllable multi-view generation model, the present invention overcomes the problem of artifacts and distortion caused by geometric inconsistencies between images and 3D models when generating high-quality multi-views based on images. To improve the problem of obvious texture seams in multi-views, the present invention introduces an iterative texture image restoration optimization. First, a rough texture map is initially constructed using the multi-views generated by the multi-view generation model. Then, image restoration technology is applied in the UV space to finally generate a smooth and complete texture map.
[0035] The present invention optimizes the texture quality of generative 3D models by combining a geometrically controllable multi-view generation model with an image restoration technology based on UV space. The geometrically controllable multi-view generation model effectively solves the artifacts and distortion problems caused by the geometric inconsistency between the image content and the three-dimensional model when generating high-quality multi-views based on images. The weighted average strategy is used to achieve the fusion of multiple views in UV space. A rough texture is quickly constructed in the initial stage, and then the rough texture map is optimized and filled through image restoration technology, ensuring that the final texture map is improved in detail richness and completeness. In 3D generation tasks, a large reconstructed model can generate its corresponding 3D object through a single image. The texture optimization method proposed in the present invention optimizes the texture of 3D objects and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Flowchart of the overall texture optimization of the present invention. DETAILED DESCRIPTION
[0037] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0038] A method for optimizing the texture quality of a generative 3D model based on texture expansion graph, the specific technical solution is as follows Figure 1 As shown, the following steps are included:
[0039] Step 1: Use a public rendering algorithm (Blender) to render images of the input 3D object at a pitch angle of 0° and azimuth angles of 0°, 45°, 90°, 180°, 270°, and 315°. The images include RGB images, depth images, position images, and normal images; render 6 RGB images with random camera poses as reference images.
[0040] Input a 3D model and render it using a public rendering tool (Blender) RGB image below i , depth map d i and location map p i , where c i Represents the parameters of the i-th camera perspective, i = (1, 2, 3, ..., 6), 6 represents the parameters of the orthographic perspective camera with a pitch angle of 0° and azimuth angles of 0°, 45°, 90°, 180°, 270° and 315° respectively, in addition to the rendered RGB image r i , depth map d i and location map p i Images under three types of modalities, and then rendering 6 RGB images of the 3D model with random camera perspectives Among them I ref_i Represents the i-th RGB reference image; the rendering process of the 3D model in the four modal images of RGB image, depth map, position map and normal map under the camera parameter c1 is expressed as R:(M,c1)→r1,d1,p1, where M represents the 3D model, c1 is the parameter of the first camera perspective, the parameters include camera intrinsic parameters and camera extrinsic parameters, and R is a rendering process, which means projecting the 3D model onto four different modal 2D images of RGB image r1, depth map d1, position map p1 and normal map n1. For the subsequent camera perspective c i =(i=1,2,3,...,6), repeat the rendering process R.
[0041] For each viewpoint, the following data is generated:
[0042] Textured RGB image: The resolution is set to 512x512, recording the color information of the object surface;
[0043] Depth map: records the distance from each pixel to the camera center, normalized to the range [0,1];
[0044] Position map: records the coordinates (x, y, z) of each pixel in three-dimensional space and stores them in three-channel form;
[0045] Normal map: records the normal vector (n x ,n y ,n z ), normalized to the range [0,1];
[0046] The generated RGB image, depth map, position map, and normal map are saved as input to the subsequent multi-view generation model.
[0047] Step 2: Obtain a training dataset and train a geometrically controllable multi-view generation model:
[0048] Filter high-quality data from public 3D datasets (ShapeNet or Objaverse) and retain 3D models that contain only single objects without occlusion or complex backgrounds.
[0049] For each 3D model in the screened 3D dataset, select the 6 RGB images, depth images, and position images of the specified viewing angles obtained by rendering in step 1, and use the 6 RGB images of random viewing angles as the reference image set;
[0050] The multi-view generation model includes a conditional encoder that encodes geometric information and a parallel attention mechanism added to the Stable Diffusion base model;
[0051] The conditional encoder consists of a convolutional network containing feature extraction blocks and downsampling layers;
[0052] Use the Stable Diffusion model as the basic image-to-text diffusion model;
[0053] Build a network architecture based on the T2I (Text-to-Image) pre-trained model:
[0054] The Stable Diffusion basic model includes a spatial self-attention layer and a text cross-attention layer. The network architecture of the spatial self-attention layer is copied to create a multi-view attention layer (Multi-ViewAttention) and an image cross-attention layer (Image Cross-Attention). The weight information of the original spatial self-attention layer of the Stable Diffusion model is retained, and the spatial self-attention layer, multi-view attention layer, and image cross-attention layer are organized in parallel. The above three layers are then arranged in series with the original text cross-attention layer of the Stable Diffusion model to ensure that the newly added attention layer inherits the prior knowledge of the pre-trained model and improve the efficiency of geometric consistency learning.
[0055] During training, we train the parameters of the conditional encoder, multi-view attention layer, and image cross attention.
[0056] Using the mean squared error loss L, the pre-trained autoencoder ε(·) is used to perform the diffusion process in the latent space by Randomly select a reference image I ref , encode it into the latent space to get the encoded image z0, add noise to the encoded image z0, and get the encoded image z at time step t t , denoising network∈ θ Reverse denoising is achieved by predicting the added noise;
[0057]
[0058] Where ∈~N(0,1) indicates that the ∈ noise distribution obeys the normal distribution of N(0,1), condition indicates the control information (geometric information of the splicing of RGB image, depth map and position map), t indicates the diffusion time step, ∈ indicates the actual noise distribution when the diffusion time step is t, ∈ θ (·) indicates that at the diffusion time step t, the denoising network is based on the control information condition and the encoded image z t Prediction of noise; E represents the expectation function.
[0059] Step 3: The single RGB image rendered in step 1 is input into the trained multi-view generation model as a reference image. The depth map and position map are used as control information. The multi-view generation model outputs multi-view images that are consistent with the geometric information of the 3D object and have the same texture and RGB.
[0060] Input the RGB image, depth map, and position map rendered in step 1 to the multi-view generation model. The depth map and position map are spliced into 6-channel data according to the channel dimension and passed through the conditional encoder as control information; so that the model generates a multi-view RGB image consistent with the geometric information of the depth map and position map; sample the texture image I1 in the case of the depth map d1 and the position map p1, I1 = S(I ref_i ,d1,p1), where S is a geometrically controllable multi-view generation model. According to the depth map d1 and the position map p1, the position map p1 provides geometric constraints and generates the corresponding texture image I1. In order to ensure the multi-view Figure 1 Consistent, multi-view generation model generates 6 camera perspectives simultaneously Texture image I i .
[0061] Step 4: The high-quality RGB image is back-projected and mapped to obtain UV space information.
[0062] The UV space information includes multi-view texture maps, mask maps and weight information of the UV space under six viewing angles;
[0063] The UV space information acquisition method is as follows:
[0064] First, based on the depth information in the depth map, the RGB images under six perspectives are back-projected onto the 3D model surface, and the color information of the 3D points is mapped to the UV space through UV mapping to generate the corresponding texture map;
[0065] Secondly, the weight information is generated by the angle deviation between the normal map and the camera viewing direction at each viewing angle, and the weight information in the 2D space is converted to the UV space through UV mapping;
[0066] Finally, the visibility information of each pixel in the texture map at the current viewing angle is recorded. The visibility information is the mask map.
[0067] Step 5: Fuse the multi-view texture maps to get a rough texture map:
[0068] The fusion method of the multiple perspective texture maps is an iterative weighted average method;
[0069] The fusion process of the iterative weighted average method for multiple perspective texture maps is as follows:
[0070] Based on the weight information of the UV space under six viewing angles, the texture map under each viewing angle, the UV space weight information and the visibility information of each pixel in the UV space are iteratively accumulated to finally generate a fused rough texture map;
[0071] The process of optimizing textures in UV space consists of two stages:
[0072] Through preliminary texture mapping, the multi-view generation model generates 6 RGB images for the 3D model, which are back-projected onto the 3D model surface to obtain texture maps in UV space under 6 viewing angles. The weight information and visibility information of each pixel are then fused to generate a preliminary texture map.
[0073] In the first stage, according to each view i RGB image I i , normal map n i , the normal map n i The angle information with the camera view is used as the weight information W i ; Use 2D space and UV space mapping to get the current view i Texture map in UV space i , visible area mask i , weight information
[0074]
[0075] Where mapping(·) represents the process of mapping 2D space to UV space, which is to use the relationship between UV space coordinates and 2D space coordinates to realize RGB image I i Convert to texture map i , weight information W i Weight information converted to UV space And the visible area mask in UV space i ;
[0076] In the second stage, combine the texture map texture of the previous stepi , weight information Visible area mask i , by accumulating the texture and weight information of the visible area at each perspective to fuse the texture:
[0077]
[0078] Through iterative accumulation, a rough texture map is finally obtained. coarse .
[0079] Step 6: Optimize the rough texture map in UV space through the image restoration model (paint3d) to obtain a high-quality complete texture expansion map;
[0080] Combining visibility information from multiple perspectives and utilizing image restoration technology, based on the geometric and color consistency constraints between multi-view textures, the missing parts in the rough texture map are filled, seams are eliminated, and discontinuities are eliminated to achieve consistency with the surrounding textures; thereby optimizing the rough texture map in UV space and achieving coverage.
[0081] Combined with UV space position map texture pos , location map texture pos Represents the 2D representation of 3D object coordinate information, using image restoration technology (Paint3D) to restore rough textures coarse Repair, expand and improve the boundary areas of the UV map, repair lighting artifacts, incomplete areas and missing high-frequency details in the rough texture, and obtain the final complete and smooth texture map texture fine :
[0082] texture fine =S UV (texture coarse ,texture pos )
[0083] Among them, S UV Representing a multi-view generative model in UV space; using texture maps fine As 3D model texture information.
[0084] The present invention proposes an innovative texture optimization method. By constructing a geometrically controllable multi-view generation model, it effectively solves the artifact and distortion problems caused by the geometric inconsistency between the image content and the three-dimensional model when generating high-quality multi-views based on images. Combined with texture fusion and repair strategies, the texture map is gradually improved from coarse to fine. In the initial stage, the basic texture is quickly constructed, and then the details are supplemented through iterative optimization. Finally, the quality of the texture map is improved by image repair technology, ensuring that the final texture map is improved in detail richness and completeness. In generative three-dimensional tasks, a large reconstruction model can generate a corresponding 3D object based on a single image. The texture optimization method proposed in the present invention can optimize the texture quality of generative 3D objects, making them have high fidelity and has broad application prospects.
Claims
1. A method for optimizing the texture quality of a generative 3D model based on a texture unfolding graph, characterized in that: The optimization method comprises the following steps: Step 1: Render the 3D model to obtain a multimodal image set from multiple perspectives; the multimodal image set includes six images, each image including a textured RGB image, a depth map, a normal map, and a position map; the six images are acquired from six perspectives, with only one photo captured from each perspective; The 3D model is placed at an upright viewing angle, the pitch angle is set to 0°, and the azimuth angles are respectively 0°, 45°, 90°, 180°, 270°, and 315°, to obtain 6 images under the six viewing angles; Step 2: Obtain a training dataset and train a geometrically controllable multi-view generation model based on the diffusion model; Step 3: Generate an RGB image set using the trained multi-view generative model; the RGB image set includes RGB images with high resolution and high texture clarity under six viewing angles; the RGB images with high resolution and high texture clarity under the six viewing angles are high-quality RGB images; the high-quality RGB images are geometrically consistent with the 3D model and have the same fidelity as the reference image; the reference image is a random RGB image; the RGB image rendered under random camera poses is a random RGB image; Step 4: The high-quality RGB image is back-projected and mapped to obtain UV space information. Step 5: Fuse the multi-view texture maps to obtain a rough texture map; Step 6: Optimize the rough texture map through the image restoration model to obtain a high-quality complete texture unfolded map.
2. The method for optimizing the texture quality of a generative 3D model based on a texture expansion graph according to claim 1, characterized in that: In step 2, the multi-view generation model uses Stable Diffusion as the basic model; the Stable Diffusion model includes a conditional encoder and a parallel attention mechanism module; the conditional encoder and the parallel attention mechanism module work together to ensure the multi-view consistency and geometric consistency between multi-view images, thereby generating high-quality multi-view images; the conditional encoder encodes the depth map and the position map into multi-scale spatial features, and injects the multi-scale spatial features as geometric control information into different layers of the Stable Diffusion model to guide the generation of multi-views; the original Stable Diffusion model organizes the spatial self-attention layer and the text cross-attention layer in a serial manner, and the parallel attention mechanism module retains Stable Diffusion by The pre-trained weights of the original two-layer architecture of the Diffusion model transform the serial architecture into a parallel architecture, construct the multi-view attention layer and the image cross-attention layer according to the network structure of the spatial self-attention layer, and initialize the weights of the multi-view attention layer and the image cross-attention layer to 0; in order to utilize the prior information of the spatial attention layer, the parallel attention mechanism module reorganizes the order of different layers and organizes the spatial self-attention layer, text cross-attention layer and image cross-attention layer in parallel; the outputs of the spatial self-attention layer, text cross-attention layer and image cross-attention layer are the input of the text cross-attention layer, thereby ensuring that the new attention layer fully inherits the prior knowledge of the pre-trained self-attention layer and realizes efficient learning with geometric knowledge.
3. The method for optimizing the texture quality of a generative 3D model based on a texture expansion graph according to claim 1, characterized in that: In step 2, the process of obtaining the training data set is as follows: First, high-quality 3D models containing single objects are screened from public 3D datasets; Then, data with occlusion or complex background is eliminated; Next, for each object, we render multiple RGB images, depth images, position images, and normal images under fixed camera poses, and also render random RGB images under random camera poses. Finally, the random RGB image is used as the reference image, the image obtained by splicing the depth map and position map under fixed camera pose is used as the control information, and the 6 RGB images under fixed camera pose are used as supervision images in the training process.
4. The method for optimizing the texture quality of a generative 3D model based on a texture expansion graph according to claim 1, wherein: In step 3, the RGB image set generation process is as follows: First, the depth maps and position maps of the six images are spliced in the channel dimension to obtain a multi-channel control map; Then, the reference image and the multi-channel control image are input into the multi-view generation model; Finally, the multi-view generation model is based on the prior knowledge of the pre-trained diffusion model, using a conditional encoder and a parallel attention mechanism. Under the geometric information constraints of the depth map and position map, it outputs six RGB images of the 3D model with a pitch angle of 0° and azimuth angles of 0°, 45°, 90°, 180°, 270°, and 315° respectively. The six RGB images are a set of RGB images, and the six RGB images are geometrically consistent with the 3D model, thus avoiding artifacts caused by directly back-projecting multiple views back to the 3D model when the geometry is inconsistent.
5. The method for optimizing the texture quality of a generative 3D model based on a texture expansion graph according to claim 1, characterized in that: In step 4, the UV space information includes multi-view texture maps, mask maps and weight information of the UV space under six viewing angles; The UV space information acquisition method is as follows: First, based on the depth information in the depth map, the RGB images under six perspectives are back-projected onto the 3D model surface, and the color information of the 3D points is mapped to the UV space through UV mapping to generate the corresponding texture map; Secondly, the weight information is generated by the angle deviation between the normal map and the camera viewing direction at each viewing angle, and the weight information in the 2D space is converted to the UV space through UV mapping; Finally, the visibility information of each pixel in the texture map at the current viewing angle is recorded. The visibility information is the mask map.
6. The method for optimizing the texture quality of a generative 3D model based on a texture expansion graph according to claim 1, characterized in that: In step 5, the fused multi-view texture map is obtained by using an iterative weighted average method; The fusion process of the iterative weighted average method for multiple perspective texture maps is as follows: Based on the weight information of the UV space under six viewing angles, the texture map under each viewing angle, the UV space weight information and the visibility information of each pixel in the UV space are iteratively accumulated to finally generate a fused rough texture map.
7. The method for optimizing the texture quality of a generative 3D model based on a texture expansion graph according to claim 1, characterized in that: In step 6, the steps of optimizing the rough texture map by the image restoration model are as follows: Combining visibility information from multiple perspectives and using image restoration technology, the missing parts in the rough texture map are filled, seams are eliminated, and discontinuities are eliminated based on the geometric and color consistency constraints between multi-view textures to achieve consistency with the surrounding texture.
8. An electronic device, characterized in that: include: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to perform the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program code, and the program code can be called by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
3D training data processing system
CN121837561A
Three-dimensional texture super-resolution method and system based on multi-view fusion
CN122048672A