Three-dimensional model generation method based on new visual angle texture correction

By using cross-domain diffusion model and three-dimensional Gaussian splattering representation in the three-dimensional model generation, combined with new perspective texture correction and adaptive density control strategies, the problem of insufficient texture details of the three-dimensional model in the existing technology is solved, and high-quality three-dimensional model generation is achieved.

CN120182533APending Publication Date: 2025-06-20TONGJI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510142459.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing three-dimensional model generation method has shortcomings in the details module of new perspective textures, which makes it difficult for the generated texture to present rich and detailed features of the reference image in new perspectives outside the reference perspective.

Method used

The basic three-dimensional patch model is generated through the cross-domain diffusion model, and the three-dimensional Gaussian splattering representation is initialized on it. The three-dimensional Gaussian splattering characterization is sampled from different new perspectives, the rendered image and correction mask are obtained, and the diffusion model is integrated into the reference image information for correction. Combining the loss function of different perspectives and adaptive density control strategies, the three-dimensional Gaussian splattering characterization is optimized to generate a high-quality three-dimensional model.

Benefits of technology

It significantly improves the texture details quality of the three-dimensional model from a new perspective, realizes high-quality three-dimensional model generation, and provides better solutions for practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182533A_ABST
    Figure CN120182533A_ABST
Patent Text Reader

Abstract

The invention relates to a three-dimensional model generation method based on new visual angle texture correction, and the method comprises the steps: generating a basic three-dimensional patch model through a cross-domain diffusion model, and initializing three-dimensional Gaussian splash representation on the three-dimensional patch model; sampling the three-dimensional Gaussian splash representation from different new view angles, and obtaining a rendered image and a correction mask of the three-dimensional Gaussian splash representation under the new view angles; the rendering image of the three-dimensional Gaussian splash representation under the new visual angle is fused into reference image information by using a diffusion model for correction, and a corrected image is obtained; projecting the reference image to the new view angle by using the three-dimensional patch model to obtain a reference image projection; and jointly optimizing three-dimensional Gaussian splash representation by utilizing the corrected image and the reference image projection and combining loss functions of different visual angles and a self-adaptive density control strategy, and obtaining a three-dimensional model. Compared with the prior art, the three-dimensional model with rich texture details generated by the scheme of the invention has excellent visual effect and keeps geometric rationality of the model at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three - dimensional model generation, and more particularly to a method for generating a three - dimensional model based on new - view texture correction. Background Art

[0002] In recent years, three - dimensional models have been widely used in people's lives. From education, architecture to film and television, three - dimensional models play an irreplaceable important role in various fields. However, for a long time, the production of three - dimensional models has been extremely complex. Even for professionals with rich experience, it is not an easy task to create a fine and realistic three - dimensional model. Therefore, people have been exploring methods to automatically generate three - dimensional models using artificial intelligence means. Until recent years, with the rapid development of diffusion models and differentiable representations, by extracting prior knowledge from diffusion models with a carefully designed loss function and then continuously optimizing the differentiable three - dimensional representation method, people have finally achieved remarkable breakthroughs in the field of three - dimensional model generation.

[0003] However, the current mainstream methods still have problems in generating new - view textures of three - dimensional models. Most of the generated textures are blurred in details and have a lower resolution compared to the reference image, resulting in difficulty in presenting the rich and detailed textures of the reference image in new views other than the reference view. This indicates that when the current generation method controls the model generation through images, it does not fully explore the detailed description information of objects in the reference image, and there is still a great deal of room for improvement in the generation quality. Therefore, how to use the rich information in the reference image to improve the visual effect of the generation model has become a very important issue. If the texture of the new view can be corrected specifically and the information of the reference image can be fully used to guide the generation process of the three - dimensional model, the texture detail quality of the model can be significantly improved, and high - quality three - dimensional model generation can be achieved.

[0004] In summary, the current three - dimensional model generation methods have deficiencies in texture details, which limit their applications in the real industry. Summary of the Invention

[0005] The purpose of the present invention is to overcome the defect that the existing three - dimensional model generation technology does not fully explore the detailed description information of objects in the reference image, and to provide a method for generating a three - dimensional model based on new - view texture correction.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] A method for generating a three - dimensional model based on new - view texture correction, the steps include:

[0008] Generate a basic three-dimensional patch model through a cross-domain diffusion model, and initialize a three-dimensional Gaussian splash representation on the three-dimensional patch model. The three-dimensional Gaussian splash representation contains the spatial position information of the three-dimensional patch model;

[0009] Sample the three-dimensional Gaussian splash representation from different new perspectives to obtain the rendered image and correction mask of the three-dimensional Gaussian splash representation from the new perspectives;

[0010] Use the diffusion model to incorporate the reference image information into the rendered image of the three-dimensional Gaussian splash representation obtained by sampling to obtain a corrected image;

[0011] Project the reference image onto the new perspective using the three-dimensional patch model to obtain a reference image projection;

[0012] Use the corrected image and the reference image projection, combined with the loss functions of different perspectives and the adaptive density control strategy, to jointly optimize the three-dimensional Gaussian splash representation and obtain a three-dimensional model.

[0013] A method for generating a three-dimensional model based on new perspective texture correction, characterized in that the steps include:

[0014] Generate a basic three-dimensional patch model through a cross-domain diffusion model, and initialize a three-dimensional Gaussian splash representation on the three-dimensional patch model. The three-dimensional Gaussian splash representation contains the spatial position information of the three-dimensional patch model;

[0015] Sample the three-dimensional Gaussian splash representation from different new perspectives to obtain the rendered image and correction mask of the three-dimensional Gaussian splash representation from the new perspectives;

[0016] Use the diffusion model to incorporate the reference image information into the rendered image of the three-dimensional Gaussian splash representation obtained by sampling to obtain a corrected image;

[0017] Project the reference image onto the new perspective using the three-dimensional patch model to obtain a reference image projection;

[0018] Use the corrected image and the reference image projection, combined with the loss functions of different perspectives, to jointly optimize the three-dimensional Gaussian splash representation to obtain a three-dimensional model.

[0019] As a preferred technical solution, the process of generating a basic three-dimensional patch model through the cross-domain diffusion model is as follows:

[0020] Input a given single-view image into the cross-domain diffusion model;

[0021] The cross-domain diffusion model marks different domains through a domain switcher and uses it as an additional input to the diffusion model to control the category of the generated image;

[0022] The diffusion model adopts a multi-view diffusion scheme to generate multi-view images. It propagates information between the normal domain and the color image domain through multi-view attention and cross-domain attention mechanisms to obtain multi-view normal maps and color images.

[0023] After obtaining the multi-view images, by sampling pixels and the corresponding world space rays, the implicit neural signed distance field is optimized to extract the 3D geometry, and then a three-dimensional patch model is obtained.

[0024] As a preferred technical solution, the three-dimensional Gaussian splash representation consists of multiple three-dimensional Gaussian distributions. The attributes of each three-dimensional Gaussian distribution include: center, opacity, scaling matrix, rotation matrix, and color. The specific process of initializing the three-dimensional Gaussian splash representation on the three-dimensional patch model is as follows:

[0025] Define a local coordinate system for each triangle in the three-dimensional patch model. Initialize the three-dimensional Gaussian distribution in the local coordinate system and bind it to the corresponding triangle in the three-dimensional patch model.

[0026] During initialization, the center position of the Gaussian point corresponds to the origin of the local coordinate system, and the opacity is initialized based on an activation function. The scaling matrix and rotation matrix are initialized in the form of a unit Gaussian distribution, and the exponential activation function is used to ensure the smoothness of the gradient. The color is based on the mean of the vertex colors of the patch and is represented using spherical harmonics.

[0027] As a preferred technical solution, the process of obtaining the rendered image and calibration mask of the three-dimensional Gaussian splash representation from a new perspective is as follows:

[0028] By arranging a series of cameras uniformly around the three-dimensional Gaussian splash representation, render the image from a new perspective at the corresponding angles.

[0029] Rotate the camera so that the camera is perpendicular to the difference region between every pair of adjacent camera views, and obtain the view of the difference region from the facing angle to get the calibration mask.

[0030] As a preferred technical solution, the process of obtaining the calibration image is as follows:

[0031] For the rendered image of the three-dimensional Gaussian splash representation from a new perspective, use the DDIM inversion technique for denoising processing to generate a deterministic intermediate noise image containing the original image information.

[0032] Use a denoising diffusion model to denoise the intermediate noise image of the rendered image from a new perspective. The denoising diffusion model respectively uses ControlNet to introduce additional depth information, uses IP-Adapter to introduce reference image information, and uses a mutual self-attention mechanism to perform feature layer replacement based on the feature information during the denoising process of the reference image.

[0033] Mix the denoising result with the intermediate noisy image according to the correction mask, and guide the denoising process of the correction area through the information of the non-correction area;

[0034] Repeat the above denoising process until the noisy image is converted into a corrected image.

[0035] As a preferred technical solution, the ControlNet reuses the encoding layer of the pre-trained denoising diffusion model as the backbone to learn a set of different conditional controls; add trainable copies of the encoding block and the intermediate block to the denoising diffusion model, and connect the output of the trainable copy to the corresponding decoding block and intermediate block of the original denoising diffusion model through zero convolution, and inject additional depth information parameters into the decoding block and intermediate block of the original denoising diffusion model.

[0036] As a preferred technical solution, the IP-Adapter sets a cross-attention layer for processing image features in each cross-attention layer of the original denoising diffusion model, separates the cross-attention layers of text features and image features through a decoupled cross-attention mechanism, and inputs the text features and image features into their respective cross-attention layers respectively; uses a pre-trained CLIP image encoder model to extract image features from the image prompt and convert them into image embeddings; uses a pre-trained projection network to map the image embeddings to a feature space with the same dimension as the text features of the pre-trained model; merges the output of the image cross-attention into the output of the text cross-attention to form the final cross-attention result.

[0037] As a preferred technical solution, the denoising diffusion model adopts a mutual self-attention mechanism to utilize the feature information in the reference graph denoising process: first, add noise to the reference graph by DDIM inversion and then denoise it, and save the feature information of all self-attention layers in the reference graph denoising process; when denoising the rendered image of the three-dimensional Gaussian splash representation from a new perspective, use the feature information of the self-attention layer saved in the reference graph denoising process to replace the feature information of the corresponding self-attention layer in the new perspective image, and implement the replacement operation of the feature layer.

[0038] As a preferred technical solution, the acquisition process of the reference graph projection is as follows:

[0039] Use a three-dimensional patch model to project the reference graph to a new perspective, identify the visible patches, and project the vertices of the visible patches onto the reference graph;

[0040] Use the reference graph as the texture of the visible part model, re-render the visible patches from different new perspectives, and remove the occluded areas in the new perspective to obtain an accurate reference graph projection as the visual information reference in the new perspective;

[0041] Process the edges of the calibrated image and the reference map projection, blur the edges of both, and achieve a smooth transition between the two.

[0042] As a preferred technical solution, the specific process of optimizing the three-dimensional Gaussian splash representation is as follows:

[0043] Use the obtained calibrated image and reference map projection to train and optimize the three-dimensional Gaussian splash representation;

[0044] Adopt an adaptive density control strategy, supplement the under-reconstructed area by cloning Gaussian points and placing the cloned Gaussian points in the direction of the view space position gradient, and split the over-reconstructed area, and randomly initialize the positions of the newly added Gaussian points according to the Gaussian distribution;

[0045] Introduce a binding inheritance strategy to make the newly added Gaussian points inherit the original binding relationship and bind to the same patch as the old Gaussian points;

[0046] Adopt multiple loss functions to train and optimize the three-dimensional Gaussian splash representation from different aspects, so that the three-dimensional Gaussian model can effectively integrate the information from the calibrated image and the reference map, expressed as:

[0047] L = L rgb + L mask + L alpha + L position + L scaling

[0048] L rgb = MSE(I rep · M rep , I gs · M rep ) + MSE(I pro · M pro , I gs · M pro )

[0049] L mask = MSE(M gs , M mesh )

[0050] L alpha = ||ReLU(∈ alpha - α)||2

[0051] L position = ||ReLU(∈ position - μ)||2

[0052] L scaling = ||ReLU(∈ scaling - s)||2

[0053] Among them, Lrgb , L mask , L alpha , L position , L scaling respectively represent color loss, mask loss, opacity loss, position loss, and size loss; I rep , M rep represent the corrected image and the corrected mask; I pro , M pro represent the reference image projection and its mask; I gs , M gs represent the image of 3D Gaussian rendering and the occluded area; M mesh represents the occluded area of the 3D patch model; ∈ alpha , ∈ position , ∈ scaling respectively represent the trigger thresholds for opacity, position, and size loss; α, μ, s respectively represent the opacity, central position, and scaling matrix attributes of the Gaussian points;

[0054] Through the comprehensive above loss functions, the training optimization of the 3D Gaussian splash representation is completed to obtain a 3D model.

[0055] Compared with the prior art, the present invention has the following beneficial effects:

[0056] 1) The present invention first generates a basic 3D patch model with reasonable geometric structure, initializes a differentiable representation on it, and endows the model with stronger visual performance capabilities. Then, samples the differentiable representation from different new perspectives, and corrects the sampled views using a diffusion model to improve the texture details therein. And projects the reference image onto the new perspective using the patch model. Finally, uses these images to optimize the differentiable representation, thereby obtaining a 3D model with excellent visual effects. The present invention further corrects the texture of the new perspective of the 3D model using the information of the reference image, so that the generated model can still maintain high-quality texture details from the new perspective, which will provide a new opportunity for the wide application of 3D model generation methods in practical applications.

[0057] 2) The present invention improves the rendered image using a new perspective texture correction method. The sampled views are corrected using a diffusion model to improve the texture details therein. Combining ControlNet and IP-Adapter technologies, additional depth information and reference image information are introduced, and a mutual self-attention mechanism is used to utilize the feature information during the denoising process of the reference image. ControlNet endows the model with in-depth control capabilities without destroying the model's capabilities; IP-Adapter realizes the image prompting ability of the text-to-image diffusion model; the mutual self-attention mechanism better restores the detailed texture in the reference image in the corrected image.

[0058] 3) The present invention utilizes the acquired calibration image and reference image projection data to train and optimize the three-dimensional Gaussian representation. During the optimization process, an adaptive density control strategy and a binding inheritance strategy are introduced, and multiple loss functions are adopted to train and optimize the three-dimensional Gaussian representation from different aspects, prompting the three-dimensional Gaussian model to effectively integrate the information from the calibration image and the reference image, while obtaining excellent visual effects and maintaining the geometric rationality of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flowchart of a three-dimensional model generation method based on new perspective texture calibration according to the present invention;

[0060] Figure 2 is an overall flowchart of the optimization of the three-dimensional Gaussian splash representation based on new perspective texture calibration according to the present invention;

[0061] Figure 3 is a flowchart of the new perspective texture calibration algorithm in the optimization of the three-dimensional Gaussian splash representation based on new perspective texture calibration in the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and provides a detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0063] Embodiment 1

[0064] The object of the present invention is to improve the texture detail quality of the generation model, and a three-dimensional model generation method based on new perspective texture calibration is proposed. First, a basic three-dimensional patch model with reasonable geometric structure is generated, and a three-dimensional Gaussian splash representation is initialized on it, endowing the model with stronger visual performance. Then, the three-dimensional Gaussian splash representation is sampled from different new perspectives, and the sampled views are corrected using a diffusion model to improve the texture details therein. And the reference image is projected to the new perspective using the three-dimensional patch model. Finally, these images are used to optimize the three-dimensional Gaussian splash representation, so as to obtain a three-dimensional model with excellent visual effects.

[0065] As Figure 1 shown, the steps of the three-dimensional model generation method based on new perspective texture calibration according to the present invention include:

[0066] S1. Generation of a basic model based on a cross-domain diffusion model, generating a basic three-dimensional patch model as a reference for the basic geometric structure in the subsequent generation process.

[0067] S2. Initialization of 3D Gaussian splash representation based on the patch model: Initialize the 3D Gaussian splash representation based on the 3D patch model, and incorporate the spatial position information of the 3D patch model into the 3D Gaussian splash representation to constrain the position and shape of the 3D Gaussian points during subsequent optimization.

[0068] S3. Optimization of 3D Gaussian splash representation based on novel view texture correction: Sample from the 3D Gaussian splash representation, use the method of novel view texture correction to improve its texture details, and combine with reference image projection to optimize the 3D Gaussian splash representation.

[0069] The specific implementation of each step is as follows:

[0070] S1. Generation of the basic model based on the cross-domain diffusion model using the Wonder3D method to achieve the automatic conversion from a single-view image to a high-fidelity 3D patch model. Cross-domain diffusion is achieved by introducing a domain switcher. The domain switcher is responsible for marking different domains and is used as an additional input when inputting into the diffusion model to control the category of the generated image. A multi-view diffusion scheme is adopted to generate multi-view images, and information is propagated between the normal domain and the color image domain with the help of multi-view attention and cross-domain attention mechanisms to ensure the consistency between different images. After obtaining the multi-view images, a high-quality 3D patch model is extracted by sampling pixels and their corresponding world space rays to optimize an implicit neural signed distance field.

[0071] Specifically, given an image, input it into the diffusion model with the introduced domain switcher, and add multi-view attention and cross-domain attention mechanisms to propagate information between the normal domain and the color image domain, thereby obtaining highly consistent multi-view normal maps and color images. Then, the 3D geometric shape is accurately extracted by optimizing the implicit neural signed distance field. The optimization functions used include normal loss, color loss, mask loss, eikonal regularization term, sparsity regularization term, smoothness regularization term, etc.

[0072] S2. Initialization of 3D Gaussian splash representation based on the patch model: Initialize the 3D Gaussian splash representation on the 3D patch model through the conversion of basic components. The 3D Gaussian splash representation consists of a large number of 3D Gaussian distributions. Each 3D Gaussian distribution, also known as a 3D Gaussian point, is defined by five key attributes: center, opacity, scaling matrix, rotation matrix, and color. When initializing the 3D Gaussian splash representation, each triangle in the 3D patch model defines a local coordinate system, and then 3D Gaussian points are initialized in the local coordinate system and the generated 3D Gaussian points are bound to the triangle.

[0073] Specifically, each triangle in the 3D patch model defines a local coordinate system, in which 3D Gaussian points are initialized and bound to the corresponding triangles. The establishment of the local coordinate system involves the calculation of the translation matrix, rotation matrix, and scaling factor. The translation matrix is determined by calculating the centroid position of the triangle. The calculation of the rotation matrix is based on the edge vectors and normal vectors of the triangle. The scaling factor is determined by calculating the mean of the triangle side length and its corresponding height. The key attributes of the Gaussian points include the center, opacity, scaling matrix, rotation matrix, and color, which jointly determine the distribution and appearance of the Gaussian points in space. During initialization, the center position of the Gaussian point corresponds to the origin of the local coordinate system, and the opacity is initialized to 0.5 based on the Sigmoid activation function. The scaling matrix and rotation matrix are initialized in the form of a unit Gaussian distribution and ensure the smoothness of the gradient through the exponential activation function. The color is based on the mean of the patch vertex colors and is represented using spherical harmonics.

[0074] S3. The overall process of optimizing the 3D Gaussian splash representation based on new view texture correction is as follows Figure 2 shown, which includes four main steps: obtaining the rendered image and the corrected mask, correcting the new view texture, obtaining the reference map projection, and optimizing the 3D Gaussian splash representation. The new view texture correction strategy corrects the sampled new view texture through a diffusion model, which can incorporate the reference map information into the generation process, adjust the texture specifically, and locally enhance the details of the texture, thereby improving the fineness of the model.

[0075] S3.1. Obtain the rendered image and the corrected mask. By evenly arranging cameras around the 3D Gaussian splash representation, the image in the new view is rendered at the corresponding angles. For the acquisition of the corrected mask, through the idea of camera pairing, the newly revealed parts of the model are focused on. For the difference region between each pair of adjacent camera views, the camera is rotated to face the difference region, and the view of the difference region part is obtained from the facing angle to get the initial mask. To enhance the supervision information, the initial mask is dilated and the range of the dilated mask is restricted to obtain the corrected mask, providing richer training information for the subsequent steps.

[0076] S3.2. Use the new view texture correction method to improve the visual effect of the sampled and rendered new view image. The specific process is as follows Figure 3 shown. This process starts from the reference image and the rendered image of the 3D Gaussian splash representation in the new view. First, the two images are denoised through the DDIM inversion method to generate a deterministic intermediate noise image containing the original image information, that is, the intermediate noise image of the reference image and the intermediate noise image x of the rendered image of the 3D Gaussian splash representation in the new view t, some information from the original rendered image is retained in this noise. Then, denoising is performed on the intermediate noise image x of the three-dimensional Gaussian splash representation for the rendered image from a new perspective. During this process, ControlNet and IP-Adapter are used to introduce additional depth information and reference image information, and the mutual self-attention mechanism is combined to finely control the denoising process. These control conditions not only help us retain the important texture features of the reference image but also enhance the detail expressiveness of the corrected area. t Denoising is carried out, and during this period, ControlNet and IP-Adapter are used to introduce additional depth information and reference image information, and the mutual self-attention mechanism is combined to finely control the denoising process. These control conditions not only help us retain the important texture features of the reference image but also enhance the detail expressiveness of the corrected area.

[0077] Subsequently, according to the correction mask, the denoising result is mixed with the intermediate noise image, and the denoising process of the corrected area is guided by the information in the non-corrected area. Mix with the denoising result according to the correction mask to obtain x t-1 , and the formula is as follows:

[0078]

[0079] where x t-1 represents the intermediate noise image at time step t - 1; represents the intermediate noise image at time step t; represents the result of denoising the intermediate noise image at time step t, and the equivalent time step is t - 1; M represents the correction mask.

[0080] Repeat the above denoising process to gradually convert the noise image into a corrected image.

[0081] The implementation of each part in texture correction is as follows:

[0082] ControlNet realizes conditional control by injecting additional parameters into the pre-trained diffusion model. These parameters are from the encoding layer and intermediate layer of the pre-trained diffusion model and are connected to the backbone network of the model through zero convolution, so as to endow the model with new control capabilities without destroying the capabilities of the pre-trained model. Specifically, by reusing the robust and powerful encoding layer of the pre-trained Stable Diffusion model as the backbone to learn a set of different conditional controls. Add trainable copies of the encoding block and intermediate block to the diffusion model, and connect their outputs to the corresponding decoding block and intermediate block of the original diffusion model through zero convolution, injecting additional parameters into it, so as to control the generation process using additional conditional information without affecting the generation ability of the pre-trained model.

[0083] For the IP-Adapter, a cross-attention layer specifically for image features is added. The text features and image features are respectively input into their own cross-attention layers, and the output of the image cross-attention is merged into the output of the text cross-attention to form the final cross-attention result, realizing the image prompting ability of the pre-trained text-to-image diffusion model. Specifically, through the decoupled cross-attention mechanism, the cross-attention layers for separating text features and image features are separated, realizing the image prompting ability of the pre-trained text-to-image diffusion model. First, use the pre-trained CLIP image encoder model to extract image features from the image prompt and convert them into image embeddings. Subsequently, use a small pre-trained projection network to map these image embeddings into the feature space with the same dimension as the text features of the pre-trained model. Next, in each cross-attention layer of the original diffusion model, a new cross-attention layer specifically responsible for processing image features is added. The text features and image features are processed separately, and the outputs of the two are fused to obtain the final attention fusion result.

[0084] The mutual self-attention mechanism reuses the feature information in the reference image denoising process. When denoising the rendered view of the three-dimensional Gaussian at a new perspective, the feature information in the reference image denoising process is fused. In the specific operation, first perform DDIM inversion on the reference image to add noise and then denoise it, and save the feature information of all self-attention layers during the denoising process; then when denoising the three-dimensional Gaussian image, use the feature information of the reference image to replace the feature information of the corresponding layer in the new perspective image to achieve the replacement operation of the feature layer, and better restore the detailed texture in the reference image in the corrected image.

[0085] S3.3. Project the reference image to a new perspective using a three-dimensional patch model. By identifying the visible patches and projecting their vertices onto the reference image, the reference image can be used as the texture for this part of the model. Render them from different new perspectives as the visual information reference in the new perspective. Finally, process the edges between the corrected image and the reference image projection, blur the edges of both, and achieve a smooth transition between the two.

[0086] Specifically, to project the reference image to different perspectives, extract the visible patches and map their vertices onto the reference image, and use the reference image as the texture for the visible patch part of the model; render these visible patches from different new perspectives and remove the occluded areas in the new perspective to obtain an accurate reference image projection. At the same time, blur the edges between the corrected image and the projection.

[0087] S3.4. After obtaining the calibration images and reference map projections required for training, various measures are used to optimize the three-dimensional Gaussian splash representation. An adaptive density control strategy and a binding inheritance strategy are introduced, and multiple loss functions are employed to train and optimize the three-dimensional Gaussian splash representation from different aspects, enabling the three-dimensional Gaussian model to effectively integrate information from the calibration images and the reference map, while achieving excellent visual effects and maintaining the geometric rationality of the model.

[0088] Specifically, first, the adaptive density control strategy is adopted. The under-reconstructed regions are supplemented by cloning Gaussian points and placing them in the direction of the view space position gradient, and the over-reconstructed regions are split, and the positions of the newly added Gaussian points are randomly initialized according to the Gaussian distribution. At the same time, the binding inheritance strategy is introduced to make the newly added Gaussian points inherit the original binding relationship and be bound to the same patch as the old Gaussian points. In addition, multiple loss functions are used, which can be generally expressed as:

[0089] L = L rgb + L mask + L alpha + L position + L scaling

[0090] L rgb = MSE(I rep · M rep , I gs · M rep ) + MSE(I pro · M pro , I gs · M pro )

[0091] L mask = MSE(M gs , M mesh )

[0092] L alpha = ||ReLU(∈ alpha - α)||2

[0093] L position = ||ReLU(∈ position - μ)||2

[0094] L scaling = ||ReLU(∈ scaling - s)||2

[0095] Among them, L rgb , L mask , L alpha , L position , L scalingrepresent color loss, mask loss, opacity loss, position loss, and size loss respectively; I rep , M rep represent the corrected image and the corrected mask; I pro , M pro represent the reference image projection and its mask; I gs , M gs represent the image and the occlusion area of the 3D Gaussian rendering; M mesh represents the occlusion area of the 3D patch model; ∈ alpha , ∈ position , ∈ scaling represent the trigger thresholds of opacity, position, and size loss respectively; α, μ, s represent the opacity, center position, and scaling matrix attributes of the Gaussian points. Finally, by integrating the above loss functions, the training optimization of the 3D Gaussian splash representation is completed, and a 3D model with excellent visual effects is obtained.

[0096] The present invention proposes a method for generating a 3D model based on novel view texture correction. This method uses the novel view texture correction method and differentiable representation to optimize the texture of the 3D model, and introduces various measures to control the correction process using additional image information, realizing a 3D model with rich texture details and providing a high-quality model source for related industries such as education and film.

[0097] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A method for generating a three-dimensional model based on new perspective texture correction, characterized in that the steps include: Generate a basic three-dimensional patch model through a cross-domain diffusion model, initialize a three-dimensional Gaussian splash representation on the three-dimensional patch model, and the three-dimensional Gaussian splash representation contains spatial position information of the three-dimensional patch model; Sampling the three-dimensional Gaussian splash representation from different new perspectives to obtain a rendered image and a corrected mask of the three-dimensional Gaussian splash representation at the new perspective; The rendered image of the sampled three-dimensional Gaussian splash representation at a new perspective is corrected by incorporating the reference image information using a diffusion model to obtain a corrected image. Projecting the reference image to a new viewing angle using a three-dimensional patch model to obtain a projection of the reference image; The three-dimensional Gaussian splash representation is optimized by using the projection of the corrected image and the reference image combined with the loss functions of different perspectives to obtain a three-dimensional model.

2. The method for generating a three-dimensional model based on new perspective texture correction according to claim 1, characterized in that: The basic three-dimensional patch model is generated by the cross-domain diffusion model, and the specific process is as follows: Input the given single view image into the cross-domain diffusion model; The cross-domain diffusion model marks different domains through a domain switcher and serves as an additional input of the diffusion model to control the category of the generated image; The diffusion model adopts a multi-view diffusion scheme to generate multi-view images, propagates information between the normal domain and the color image domain through multi-view attention and cross-domain attention mechanisms, and obtains multi-view normal maps and color images; After obtaining the multi-view images, the implicit neural signed distance field is optimized by sampling pixels and the world space rays corresponding to the pixels, and the 3D geometric shapes are extracted to obtain the 3D patch model.

3. The method for generating a three-dimensional model based on new perspective texture correction according to claim 1, characterized in that: The three-dimensional Gaussian splash representation is composed of multiple three-dimensional Gaussian distributions, and the attributes of each three-dimensional Gaussian distribution include: center, opacity, scaling matrix, rotation matrix and color; the specific process of initializing the three-dimensional Gaussian splash representation on the three-dimensional patch model is as follows: A local coordinate system is defined for each triangle in the three-dimensional patch model, and a three-dimensional Gaussian distribution is initialized in the local coordinate system and bound to the triangle in the corresponding three-dimensional patch model; During initialization, the center position of the Gaussian point corresponds to the origin of the local coordinate system, and the opacity is initialized based on the activation function; the scaling matrix and the rotation matrix are initialized in the form of a unit Gaussian distribution, and the smoothness of the gradient is ensured by the exponential activation function; the color is based on the mean of the vertex colors of the patch and is represented by spherical harmonics.

4. The method for generating a three-dimensional model based on new perspective texture correction according to claim 1, characterized in that: The process of obtaining the rendered image and the correction mask of the three-dimensional Gaussian splash representation at a new perspective is as follows: By evenly arranging a series of cameras around the three-dimensional Gaussian splash representation, an image under a new perspective is rendered at a corresponding angle; The camera is rotated so that it is perpendicular to the difference area of ​​each pair of adjacent camera perspectives, and the view of the difference area is obtained from the opposite angle to obtain the correction mask.

5. The method for generating a three-dimensional model based on new perspective texture correction according to claim 1, characterized in that: The process of obtaining the corrected image is as follows: For the rendered image of the three-dimensional Gaussian splash representation under the new perspective, the DDIM inversion technology is used to perform noise processing to generate a deterministic intermediate noise image containing the original image information; A denoising diffusion model is used to denoise the intermediate noise image of the rendered image under the new perspective. The denoising diffusion model uses ControlNet to introduce additional depth information, uses IP-Adapter to introduce reference image information, and uses a mutual self-attention mechanism to replace the feature layer based on the feature information in the reference image denoising process; The denoising result is mixed with the intermediate noise image according to the correction mask, and the denoising process of the correction area is guided by the information of the non-correction area; Repeat the above denoising process until the noisy image is converted into a corrected image.

6. The method for generating a three-dimensional model based on new perspective texture correction according to claim 5, characterized in that: The ControlNet reuses the encoding layer of the pre-trained denoising diffusion model as the backbone to learn a set of different conditional controls; adds trainable copies of the encoding block and the intermediate block in the denoising diffusion model, and connects the output of the trainable copy to the corresponding decoding block and the intermediate block of the original denoising diffusion model through zero convolution, injecting additional depth information parameters into the decoding block and the intermediate block of the original denoising diffusion model.

7. The method for generating a three-dimensional model based on new perspective texture correction according to claim 5, characterized in that: The IP-Adapter sets a cross-attention layer for processing image features in each cross-attention layer in the original denoising diffusion model, separates the cross-attention layers of text features and image features through a decoupled cross-attention mechanism, and inputs the text features and image features into their respective cross-attention layers; uses a pre-trained CLIP image encoder model to extract image features from image prompts and converts them into image embeddings; uses a pre-trained projection network to map the image embeddings to a feature space with the same dimension as the text features of the pre-trained model; and merges the output of the image cross-attention into the output of the text cross-attention to form a final cross-attention result.

8. The method for generating a three-dimensional model based on new perspective texture correction according to claim 5, characterized in that: The denoising diffusion model adopts a mutual self-attention mechanism to utilize the feature information in the reference image denoising process: first, the reference image is denoised after DDIM inversion and noise is added, and the feature information of all self-attention layers in the reference image denoising process is saved; when denoising the rendered image of the three-dimensional Gaussian splash representation at a new perspective, the feature information of the self-attention layer saved in the reference image denoising process is used to replace the feature information of the corresponding self-attention layer in the new perspective image, thereby realizing the feature layer replacement operation.

9. The method for generating a three-dimensional model based on new perspective texture correction according to claim 1, characterized in that: The process of obtaining the reference image projection is as follows: The reference image is projected to a new perspective using a three-dimensional face model, the visible face is identified, and the vertices of the visible face are projected onto the reference image; Use the reference image as the texture of the visible part of the model, re-render the visible face from different new perspectives, and remove the occluded areas in the new perspective to obtain an accurate reference image projection as a reference for visual information under the new perspective; The edges of the correction image and the reference image projection are processed to blur the edges of the two and achieve a smooth transition between the two.

10. The method for generating a three-dimensional model based on new perspective texture correction according to claim 1, characterized in that: The specific process of optimizing the three-dimensional Gaussian splash characterization is as follows: Using the acquired rectified image and reference image projection to train and optimize the three-dimensional Gaussian splatter representation; Adopting an adaptive density control strategy, the under-reconstructed area is supplemented by cloning Gaussian points and placing them in the direction of the view space position gradient. The over-reconstructed area is split and the positions of the newly added Gaussian points are randomly initialized according to the Gaussian distribution. Introduce a binding inheritance strategy so that newly added Gaussian points inherit the original binding relationship and are bound to the same face as the old Gaussian points; A variety of loss functions are used to train and optimize the three-dimensional Gaussian splash representation from different aspects, so that the three-dimensional Gaussian model can effectively integrate the information from the rectified image and the reference image, which can be expressed as: L=L rgb +L mask +L alpha +L position +L scaling L rgb =MSE(I rep ·M rep ,I gs ·M rep )+MSE(I pro ·M pro ,I gs ·M pro ) L mask =MSE(M gs ,M mesh ) L alpha =||ReLU(∈ alpha -α)||2 L position =||ReLU(∈ position -μ)||2 L scaling =||ReLU(∈ scaling -s)||2 Among them, L rgb , L mask , L alpha , L position , L scaling I represents color loss, mask loss, opacity loss, position loss, and size loss respectively; rep 、M rep I represents the correction image and correction mask; pro 、M pro Represents the reference image projection and its mask; I gs 、M gs Represents the image and occluded area of ​​3D Gaussian rendering; M mesh It represents the occluded area of ​​the three-dimensional patch model; ∈ alpha ,∈ position ,∈ scaling They represent the triggering thresholds of opacity, position and size loss respectively; α, μ and S represent the opacity, center position and scaling matrix properties of the Gaussian point respectively; By combining the above loss functions, the training optimization of the three-dimensional Gaussian splash representation is completed to obtain a three-dimensional model.

Citation Information

Cited By

  • Non-rigid three-dimensional editing method and system based on cross-modal attention guidance

    CN120655802A

  • A non-rigid 3D editing method and system based on cross-modal attention guidance

    CN120655802B