Universal 3D intelligent modeling method for multi-view coupling constraint

By combining pre-trained diffusion models with fine-tuned multi-view generation diffusion models, the three-dimensional representation is optimized, and the geometry and texture of the three-dimensional grid is fine-tuned using multi-view coupling constraints, the problems of poor generalization capabilities and inconsistent geometry of multi-view geometry are solved, and efficient and accurate three-dimensional modeling generation is achieved.

CN120088427APending Publication Date: 2025-06-03HARBIN INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510097700.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing three-dimensional modeling generation methods have poor generalization capabilities, easy to overfit the generation results, and easy to cause problems of geometric inconsistency in multiple perspectives.

Method used

The diffusion model of the image generated by pretrained text and fine-tuned multi-view generation diffusion model optimizes the learnable three-dimensional representation, combined with multi-view coupling constraints, respectively fine-tune the geometry and surface texture of the triangle mesh.

Benefits of technology

It significantly enhances the versatility and applicability of three-dimensional modeling, realizes multi-view geometric consistency, improves modeling efficiency and quality, and gets rid of the dependence on large-scale text-3D datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088427A_ABST
    Figure CN120088427A_ABST
Patent Text Reader

Abstract

The invention provides a universal 3D intelligent modeling method for multi-view coupling constraint, belongs to the technical field of three-dimensional simulation modeling, and solves the problems that an existing intelligent three-dimensional modeling generation method is poor in generalization ability, a generation result is easy to over-fit, and multi-view geometric inconsistency is easy to occur. Comprising the following steps: optimizing learnable three-dimensional representation through a pre-trained text generation image diffusion model and a fine-adjusted multi-view generation diffusion model, and generating initial three-dimensional representation which is completely consistent with text semantics; according to the transparency prediction information of each point in the preliminary three-dimensional representation and the multi-view rendering graph of the preliminary three-dimensional representation, extracting and converting into a rough three-dimensional grid representation containing geometric structure and texture information; a geometry and texture decoupling optimization method is adopted, multi-view coupling constraints are combined, the geometrical shape and the surface texture of the triangular mesh are finely adjusted and optimized respectively, and a three-dimensional triangular mesh model with high-quality geometrical morphology and realistic texture is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a general 3D intelligent modeling method with multi-perspective coupling constraints, belonging to the technical field of three-dimensional simulation modeling. Background Art

[0002] With the rapid development of digital technology and the increasing demands of architectural design and industrial manufacturing, three-dimensional modeling technology is gradually becoming an important tool in fields such as film production, game development, virtual reality (VR), architectural design, and industrial design. In these application scenarios, the quality, complexity, and generation efficiency of models directly affect the effectiveness and cost of the design and production processes. However, existing three-dimensional modeling methods still face many challenges in practical applications.

[0003] Traditional three-dimensional modeling mainly relies on professional design software, such as Maya, 3ds Max, Rhino, and Revit, etc. These tools support the full process of design from geometric shape construction to detailed texture addition by providing designers with rich modeling functions. However, these tools usually require designers to spend a lot of time manually adjusting model details, which not only has low work efficiency but also highly depends on the experience and skill level of designers, making it difficult to achieve large-scale and rapid generation of diverse models. In addition, frequent adjustments and trial-and-error in the design process significantly extend the design cycle, and the innovation and consistency of design results are often difficult to guarantee. How to get rid of the dependence on manual experience, generate high-quality three-dimensional models with both rationality and innovation, and achieve efficient and accurate modeling optimization in multiple fields and scenarios has become the core issue in promoting the intelligent and automated development of three-dimensional modeling.

[0004] In recent years, with the rapid development of machine learning and deep learning technologies, three-dimensional modeling generation methods based on artificial intelligence have gradually emerged. The application of generative models (such as generative adversarial network GAN, variational autoencoder VAE, and diffusion model) makes it possible to learn the latent sample distribution from massive data. These methods achieve efficient and automated generation of three-dimensional models by analyzing the geometric features and distribution laws of three-dimensional models. Currently, three-dimensional modeling methods based on artificial intelligence mainly include the following three approaches:

[0005] (1) Direct modeling method based on three-dimensional generative model: Use a large-scale text-three-dimensional model dataset to train a generative model, and directly generate a three-dimensional model according to the text input.

[0006] (2) Three-dimensional modeling method based on single-perspective optimization: Randomly perform single-perspective optimization on a learnable three-dimensional representation through a pre-trained two-dimensional diffusion model, and gradually optimize the three-dimensional representation to achieve generation from text to three-dimensional model.

[0007] (3) Fine-tuned multi-view joint optimization 3D modeling method: The multi-view diffusion model is obtained by fine-tuning the pre-trained 2D diffusion model, and its multi-view view output is used to perform multi-view joint optimization on the 3D representation, thereby achieving geometrically consistent 3D model generation.

[0008] Although these methods have significantly improved the efficiency of 3D modeling, there are still some problems that need to be solved. For example, the multi-view joint optimization method of directly training 3D generative models and fine-tuning is constrained by the lack of high-quality, large-scale text-3D datasets, resulting in poor generalization ability and easy overfitting of the generated results; while the method based on single-view optimization is prone to multi-view geometric inconsistency. Therefore, developing a general, efficient 3D modeling generation method that can ensure multi-view geometric consistency has become an important research direction in this field. Summary of the invention

[0009] In order to solve the problems of poor generalization ability, easy overfitting of generated results and easy multi-view geometric inconsistency in existing 3D modeling generation methods, the present invention proposes a universal 3D intelligent modeling method with multi-view coupling constraints.

[0010] The technical solution adopted by the present invention to solve the above-mentioned problem is: the present invention comprises the following steps:

[0011] Step 1: Optimize the learnable 3D representation through the pre-trained text-to-image diffusion model and the fine-tuned multi-view generative diffusion model to generate a preliminary 3D representation that is fully consistent with the text semantics;

[0012] Step 2: Based on the generated preliminary 3D representation, extract and convert the rough 3D mesh representation containing geometric structure and texture information according to the transparency prediction information of each point in the preliminary 3D representation and the multi-view rendering of the preliminary 3D representation;

[0013] Step 3: Use the geometry and texture decoupling optimization method, combined with multi-view coupling constraints, to fine-tune and optimize the geometry and surface texture of the triangular mesh respectively, and generate a three-dimensional triangular mesh model with high-quality geometry and realistic texture.

[0014] Preferably, step 1 specifically includes:

[0015] Step 1.1: Randomly initialize the parameters of the learnable 3D representation model, including the parameters that control the geometric structure and texture properties of the model;

[0016] Step 1.2: Randomly select four mutually orthogonal horizontal viewpoints in the learnable 3D representation with the same elevation angle and camera parameters, render each viewpoint by rasterization rendering method, and generate four corresponding rendering views;

[0017] Step 1.3: Perform diffusion denoising on the four rendered views, combine the text prompt and input it into the pre-trained diffusion model, output the multi-view target rendered images that meet the text conditions, and perform supervised learning on the learnable 3D representation;

[0018] Step 1.4: Repeat Steps 1.2 and 1.3 until convergence or reaching the preset number of iterations, complete the optimization of the learnable 3D representation, and generate a preliminary 3D representation that is exactly the same as the text semantics;

[0019] The rendering expression for each pixel is:

[0020]

[0021]

[0022] In Formulas (1) and (2), is the color rendered for each pixel along the ray r in the rendered image, and T i represents the cumulative transmittance of the learnable 3D representation at the query point, and σ i represents the opacity of the learnable 3D representation at the query point, and c i represents the color information of the learnable 3D representation at the query point;

[0023] The calculation formula for the rendered view is:

[0024]

[0025] In Formula (3), θ is the parameter of the learnable 3D representation; v i is the information of the rendering view; g is the differentiable rendering function, is the 3D representation at the view v i and the rendered image obtained under this view.

[0026] Preferably, Step 1.3 specifically includes:

[0027] Step 1.3.1: Perform step-by-step diffusion denoising on the generated rendered views. Input the randomly single-view rendered image with noise added at the t-th step and the text prompt into the pre-trained text-to-image diffusion model to predict the noise view at the (t - 1)-th step of the single view, and input the multi-view rendered image with noise added at the t-th step and the text prompt into the fine-tuned multi-view generation diffusion model to predict the noise view at the (t - 1)-th step of the multi-view;

[0028] Step 1.3.2: Use the noise view at the (t-1)-th step of the single view and the noise view at the (t-1)-th step of the multi-view as the target views at the (t-1)-th step of the four rendered views. Calculate the loss function of the learnable three-dimensional representation model parameter θ and perform supervised learning based on the target views at the (t-1)-th step of the four rendered views and the pre-trained text generation image diffusion model with the LoRA module added, so as to update the three-dimensional representation parameter θ;

[0029] The expression for adding noise to each rendered image is:

[0030]

[0031] In formula (4), is the image generated after adding noise according to the diffusion time step t, α t and σ t are hyperparameters, and ∈ is the noise randomly sampled from the multi-dimensional normal distribution;

[0032] The calculation formula for the gradient of the loss function of the learnable three-dimensional representation model parameter θ is:

[0033]

[0034]

[0035] In formulas (5) and (6), are four orthogonal views, is the noise-rendered image of the corresponding view, y is the input text condition, is the diffusion time step of the diffusion model, ∈ is the random noise added during the diffusion process, ∈ φ is the pre-trained text generation image diffusion model, ∈ φ parameters are fixed during the optimization process, s is the classification guidance parameter, λ is the regularization parameter, ∈ M represents the fine-tuned multi-view generation diffusion model, ∈ M parameters are fixed during the optimization process, is the pre-trained text generation image diffusion model with the LoRA module added, and the LoRA module parameters are alternately updated with the learnable three-dimensional representation model parameter θ during the optimization process;

[0036] The expression for optimizing the LoRA module parameters is:

[0037]

[0038] Preferably, step 2 specifically includes:

[0039] Step 2.1: Map the optimized learnable three-dimensional representation to (-1,1)3 In the cube space, the cube space is divided into 16 3 sub-blocks;

[0040] Step 2.2: Each sub-block is further divided into 8 3 grid cells, and the transparency information at the center position of each grid cell is calculated, so as to convert the cube space of (-1, 1) 3 into 256 3 grid cells, and the transparency information of each grid cell is obtained;

[0041] Step 2.3: Set a threshold β to screen the transparency information, convert the grid cell values into binary form to obtain a rough voxel grid, calculate the signed distance function value of each cell in the voxel grid through the distance_transform_edt function in the SciPy library in Python, use the initialized deformable tetrahedral grid, and combine the MarchingTetrahedra algorithm to divide the cube space of (-1, 1) 3 into tetrahedral elements, assign SDF values to the vertices of the tetrahedral elements, and obtain the distance from the corresponding vertices to the surface of the modeled three-dimensional object through the SDF values;

[0042] Step 2.4: According to the multi-view strategy, combine the optimized preliminary three-dimensional representation to initialize the mesh texture of the triangular mesh to generate a rough three-dimensional mesh representation containing geometric structure and texture information;

[0043] The calculation formula of the transparency information is:

[0044] d(x) = f θ (x) (8);

[0045] In formula (8), d(x) is the transparency information of the grid cell with x as the central coordinate, and f θ is the function for calculating the transparency information at the coordinate point x of the three-dimensional representation.

[0046] Preferably, in step 2.4, initializing the mesh texture of the triangular mesh in combination with the optimized preliminary three-dimensional representation includes:

[0047] Step 2.4.1: For the surface texture of the triangular mesh, the spatial position is encoded using a hash grid and simulated in combination with a single-layer MLP network. Here, the MLP network is a multi-layer perceptron network. The input of the MLP network is the representation of the three-dimensional coordinates of any point in space after position encoding by the hash network, and the output is the color of the three-dimensional modeling object at the corresponding position in space. Based on the generated triangular mesh and the color of the surface sampling points output by the MLP, under the condition of a given rendering perspective, the initial rendering of the triangular mesh is generated using the nvdiffrast library of Python;

[0048] Step 2.4.2: According to the multi-view strategy, different camera views are selected to render the preliminary three-dimensional representation and the triangular mesh after grid texture initialization respectively. The rendering of the preliminary three-dimensional representation is used as the target label and the loss is calculated with the grid rendering;

[0049] Step 2.4.3: Through gradient backpropagation, the parameters of the MLP network and the hash grid encoding are optimized to realize the grid texture initialization of the preliminary three-dimensional representation rendering. Here, during the grid texture initialization of the preliminary three-dimensional representation rendering, the SDF value and its displacement parameters of the grid vertices remain unchanged, and only the texture-related parameters are optimized.

[0050] Preferably, step 3 specifically includes:

[0051] Step 3.1: Through the multi-view coupling constraint optimization method, the geometric structure and texture information of the texture grid of the rough three-dimensional mesh representation are optimized and refined respectively;

[0052] Step 3.2: The refined grid extraction is performed on the optimized and refined rough three-dimensional mesh representation to obtain a three-dimensional triangular mesh model with high-quality geometric shape and realistic texture.

[0053] Preferably, step 3.1 specifically includes:

[0054] Step 3.1.1: Based on the initialized deformable tetrahedral mesh, the rough three-dimensional triangular mesh is extracted by the Marching Tetrahedra algorithm, and the vertex normal vectors of the rough three-dimensional mesh are calculated using the nvdiffrast library of Python;

[0055] Step 3.1.2: The vertex normals are converted to texture space coordinates through UV mapping, and the normal vectors are stored in the form of RGB values to complete the optimization of the deformable tetrahedral mesh and generate the surface normal map of the 3D mesh;

[0056] Step 3.1.3: Replace the multi-view target rendering map in Step 1.3 with the generated surface normal map, calculate the optimization gradient through Formulas (5) and (6), and update the parameters of the deformable tetrahedral mesh through backpropagation;

[0057] Step 3.1.4: Repeat Step 3.1.1 - Step 3.1.3 until the optimization converges or reaches the specified number of optimization steps to complete the optimization and refinement of the rough three-dimensional mesh geometry;

[0058] Step 3.1.5: Fix the optimized deformable tetrahedral mesh, extract the rough three-dimensional mesh using the Marching Tetrahedra algorithm, and query the coordinates of each surface sampling point of the mesh;

[0059] Step 3.1.6: Hash-encode the spatial coordinates of the queried sampling points and input them into the MLP to output the color information of the corresponding sampling points;

[0060] Step 3.1.7: Render the rough three-dimensional mesh using the nvdiffrast library to generate multi-view rendering images;

[0061] Step 3.1.8: Calculate the optimization gradient for the multi-view rendering images according to Formulas (5) and (6), and update the hash-encoding network and MLP network layers through backpropagation;

[0062] Step 3.1.9: Repeat Step 3.1.5 - Step 3.1.8 until the optimization converges or reaches the specified number of optimization steps to complete the optimization and refinement of the rough three-dimensional mesh texture.

[0063] Preferably, Step 3.2 specifically includes:

[0064] Step 3.2.1: Based on the optimized deformable tetrahedral mesh, extract the rough three-dimensional mesh using the Marching Tetrahedra algorithm to obtain the mesh vertex coordinates and patch topology;

[0065] Step 3.2.2: Sample each point on the mesh surface to obtain the spatial coordinates of the sampling points, and calculate the color information of the sampling points through the optimized hash-encoding network and MLP network;

[0066] Step 3.2.3: Generate a texture map according to the color information of the sampling points output by the MLP network, and bind the texture map to the corresponding triangular network to obtain a three-dimensional triangular mesh model with high-quality geometric shape and realistic texture.

[0067] The beneficial effects of the present invention are:

[0068] (1) Get rid of dataset dependence: This paper uses a joint optimization strategy by coupling a pre-trained 2D diffusion model with a fine-tuned multi-view generation diffusion model, effectively reducing the dependence on large-scale text-3D datasets, significantly enhancing the versatility and applicability of the technology, and breaking through the limitation of data scarcity on 3D modeling.

[0069] (2) Achieving multi-view geometric consistency: Through the multi-view joint optimization strategy, the present invention successfully solves the core problem of multi-view geometric inconsistency in the current artificial intelligence optimized 3D modeling method. The generated 3D model shows high consistency and accuracy in all view angles, greatly improving the authenticity and reliability of the modeling results.

[0070] (3) Improving modeling efficiency and quality: Based on a step-by-step optimization modeling strategy, the present invention significantly improves the efficiency of 3D modeling, while achieving the generation of high-fidelity 3D models with fine details and rich textures, providing an efficient and accurate solution for high-quality 3D modeling.

[0071] (4) Wide application potential: The present invention has wide applicability and can serve many fields such as architectural modeling, game development, virtual reality and industrial design, and provides intelligent and efficient technical support for complex three-dimensional modeling tasks.

[0072] In summary, the present invention realizes efficient, accurate and universal 3D modeling generation, opens up intelligent and diversified application paths for 3D modeling technology, and has important academic value and broad industrial prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 A schematic diagram of a flow chart of a general 3D intelligent modeling method with multi-view coupling constraints provided by the present invention;

[0074] Figure 2 A visualization diagram of the three-dimensional representation rendering process provided by the present invention;

[0075] Figure 3 A preliminary three-dimensional representation flow chart of multi-view coupled constraint optimization provided by the present invention;

[0076] Figure 4 A schematic diagram of the geometric extraction and texture initialization of the texture grid provided by the present invention;

[0077] Figure 5 A schematic diagram of a preliminary three-dimensional morphology representation generated by the optimized 3D Gaussian Splatting provided by the present invention;

[0078] Figure 6 A schematic diagram of a high-precision three-dimensional triangular mesh model generated by an optimized texture mesh provided by the present invention;

[0079] Figure 7 Schematic diagram of the generated 3D triangular mesh model of the building provided by the present invention. Detailed implementation manners

[0080] Detailed implementation manner 1: In combination with Figure 1 Illustrate this implementation manner. As Figure 1 shown, the steps of a general 3D intelligent modeling method with multi-view coupling constraints described in this implementation manner include:

[0081] S1: Optimize the preliminary 3D representation with multi-view coupling constraints;

[0082] In this implementation manner, 3D Gaussian Splatting is used as a learnable 3D representation example. Combining a pre-trained text-to-image diffusion model and a fine-tuned multi-view generation diffusion model, 3D Gaussian Splatting is optimized to generate a preliminary 3D representation that conforms to the text semantics. Although this implementation manner takes 3D Gaussian Splatting as an example, its method is also applicable to other optimizable 3D representation forms, including the following steps:

[0083] S101: Multi-view differentiable rendering of the 3D representation: Randomly initialize the parameters of the 3D Gaussian Splatting model, and randomly select four horizontal direction views that are mutually orthogonal and have the same elevation angle and camera parameters. Through rasterization rendering technology, generate the corresponding four rendered views;

[0084] The rendering process of each pixel in the rendered image is as follows:

[0085]

[0086]

[0087] In formulas (1) and (2), is the color of each pixel in the rendered image, G(l) is each Gaussian primitive in 3D Gaussian Splatting, and l is the local coordinate centered at μ; is the ordered set of Gaussian primitives overlapping with the pixel along the ray r, and c i is the color information of the learnable 3D representation at the query point;

[0088] The calculation formula for the rendered view is:

[0089]

[0090] In formula (3), θ is the parameter of the 3D Gaussian to be rendered; vi Information for the rendering view; g is a differentiable rendering function, is the rendered image obtained from the 3D representation at view v i ; The rendering process for each pixel is calculated according to formulas (1) and (2); represents the rendered image obtained from the 3D representation at view v i ; The multi-view rendering process is as shown Figure 2 .

[0091] In addition, the learnable 3D representation adopted in this embodiment S1 can be applied to all forms, such as NeRF, 3D Gaussian Splatting, etc. For different 3D representations, the way of parameterizing θ and the form of the rendering function are different, but they all satisfy the calculation forms of formulas (1) and (3). Taking NeRF and 3D Gaussian Splatting as examples: For the NeRF form, the 3D representation θ is parameterized as the learnable neural network parameters (denoted as F θ ), and the transparency information σ i and color information c i in the rendering process are obtained through the neural network F θ (r(i)); For the 3D Gaussian Splatting form, the 3D object is represented as multiple Gaussian primitives, and each Gaussian primitive is parameterized with the following parameters: the center position coordinates , covariance matrix Σ, color parameter , and transparency parameter . Among them, the covariance matrix Σ is parameterized by the scaling factor and the rotation quaternion ; The parameter set of 3D Gaussian Splatting is θ = {μ k , s k , q k , c k , α k}, where k is the serial number of the Gaussian primitive; the color information c i in the rendering process comes from the parameter set, and the transparency information calculation formula is l is the local coordinate centered on μ i .

[0092] S102: Add noise to the multi-view rendering images through diffusion and predict the target rendering images of multiple views in combination with the text input. Based on the forward diffusion noise addition process of the diffusion model, perform step-by-step noise addition processing on the multi-view rendering images generated in S101. The specific diffusion noise addition process for each rendering image is as follows:

[0093]

[0094] In formula (4), is the image generated after adding noise according to the diffusion time step T, where α t and σ t are hyperparameters, and ∈ is the noise randomly sampled from a multi-dimensional normal distribution;

[0095] S103: Input the multi-view rendering map with noise added at the t-th step and the text prompt into the pre-trained text-to-image diffusion model and the fine-tuned multi-view generation diffusion model, and predict the noise views at the (t - 1)-th step for single-view and multi-view respectively. These predicted views are used as the target views of the multi-view rendering map at the (t - 1)-th step to calculate the loss function and perform supervised learning. After simplification, the formula for the gradient of the loss function with respect to the parameters θ of the 3D representation is as follows:

[0096]

[0097]

[0098] In formulas (5) and (6), are four orthogonal views, is the noise rendering map of the corresponding view, y is the input text condition, is the diffusion time step of the diffusion model, ∈ is the random noise added during the diffusion process, ∈ φ is the pre-trained text-to-image diffusion model, ∈ φ parameters are fixed during the optimization process, s is the classification guidance parameter, λ is the regularization parameter, ∈ M represents the fine-tuned multi-view generation diffusion model, ∈ M parameters are fixed during the optimization process, is the pre-trained text-to-image diffusion model with the LoRA module added. The parameters of the LoRA module are alternately updated with the learnable 3D representation model parameters θ during the optimization process, and its optimization calculation formula is:

[0099]

[0100] S104: Repeat S101 - S103 until the iteration converges or reaches the preset number of iterations, and finally generate a preliminary 3D representation that is highly consistent with the text semantics. The iterative optimization process is as Figure 3 shown.

[0101] S2: Coarse texture grid extraction;

[0102] S201: Map the optimized learnable 3D representation to (-1, 1) 3In a cube space, the cube space is divided into 16 3 sub-blocks. For each sub-block, Gaussian basis elements with the center position μ located within the corresponding sub-block are selected. Each sub-block is divided into 8 3 grid cells. The transparency of the center position of each grid cell within all Gaussian basis elements in the sub-block is calculated and accumulated to obtain the transparency of each grid cell. The cube space of (-1, 1) 3 is converted into 256 3 triangle grid cells, and the transparency information of each grid cell is calculated. The specific calculation formula for the grid transparency is as follows:

[0103]

[0104] In formula (8), d(x) is the transparency of the grid cell with x as the center coordinate, B is the set of all Gaussian basis elements within the sub-block where the grid is located, μ i is the center coordinate of the i-th Gaussian basis element, and Σ i is the covariance matrix parameter of the i-th Gaussian basis element.

[0105] S202: Set a threshold β to screen the transparency information, convert the grid cell values into binary form to obtain a rough voxel grid, calculate the SDF value, i.e., the signed distance function value, of each cell in the voxel grid using the distance_transform_edt function in the SciPy library in Python. By initializing DMTet, a deformable tetrahedral grid can be obtained. Combining with the Marching Tetrahedra algorithm, the cube space of (-1, 1) 3 is divided into tetrahedral cells, and SDF values are assigned to the vertices of the tetrahedral cells. The distance from the corresponding vertex to the surface of the modeled three-dimensional object is obtained through the SDF value;

[0106] S203: For the surface texture of the triangle grid, use hash grid to encode the spatial position and combine with a single-layer MLP network for simulation. Among them, the MLP network is a multi-layer perceptron network. The input of the MLP network is the representation of the three-dimensional coordinates of any point in space after being encoded by the hash network position, and the output is the color of the three-dimensional modeled object at the corresponding position in space. Based on the generated triangle grid and the colors of the surface sampling points output by the MLP, under the condition of a given rendering perspective, use the nvdiffrast library in Python to render and generate the initial rendering of the triangle grid;

[0107] S204: According to the multi-view strategy, select different camera views to render the preliminary 3D representation and the triangular mesh after mesh texture initialization respectively. Use the rendered image of the preliminary 3D representation as the target label and calculate the loss with the mesh rendered image;

[0108] S205: Through gradient backpropagation, optimize the MLP network and hash grid encoding parameters to achieve the mesh texture initialization of the preliminary 3D representation rendered image. Among them, during the mesh texture initialization process of the preliminary 3D representation rendered image, the SDF value and its displacement parameters of the mesh vertices remain unchanged, and only the texture-related parameters are optimized. The geometric and texture initialization results of the optimized preliminary 3D representation and the corresponding texture mesh are as Figure 4 shown.

[0109] S3: Fine texture mesh fine-tuning with multi-view coupling constraints;

[0110] S301: Through the multi-view coupling constraint optimization method, optimize and refine the geometric structure and texture information of the texture mesh of the rough 3D mesh representation;

[0111] S30101: Based on the initialized deformable tetrahedral mesh, extract the rough 3D mesh through the Marching Tetrahedra algorithm, and use the nvdiffrast library of Python to calculate the vertex normal vectors of the rough 3D mesh;

[0112] S30102: Convert the vertex normals to texture space coordinates through UV mapping, and store the normal vectors in the form of RGB values to complete the optimization of the deformable tetrahedral mesh and generate the surface normal map of the 3D mesh;

[0113] S30103: Replace the multi-view target rendered image in step 1.2 with the generated surface normal map, calculate the optimization gradient through formulas (5) and (6), and update the parameters of the deformable tetrahedral mesh through backpropagation;

[0114] S30104: Repeat S30101 - S30103 until the optimization converges or reaches the specified number of optimization steps to complete the optimization and refinement of the geometric structure of the rough 3D mesh;

[0115] S30105: Fix the optimized deformable tetrahedral mesh, extract the rough 3D mesh using the Marching Tetrahedra algorithm, and query the coordinates of each surface sampling point of the mesh;

[0116] S30106: Hash-encode the spatial coordinates of the queried sampling points and input them into the MLP to output the color information of the corresponding sampling points;

[0117] S30107: Render the rough 3D mesh using the nvdiffrast library to generate multi-view rendered images;

[0118] S30108: Calculate the optimization gradient for the multi-view rendered images according to Formulas (5) and (6), and update the hash encoding network and the MLP network layer through backpropagation;

[0119] S30109: Repeat S30105 - S30108 until the optimization converges or reaches the specified number of optimization steps to complete the optimization and fine-tuning of the rough 3D mesh texture.

[0120] S302: Extract a refined mesh from the optimized and fine-tuned rough 3D mesh representation to obtain a 3D triangular mesh model with high-quality geometric shapes and realistic textures;

[0121] S30201: Based on the optimized deformable tetrahedral mesh, extract the rough 3D mesh using the Marching Tetrahedra algorithm to obtain the mesh vertex coordinates and the patch topology;

[0122] S30202: Sample points on the mesh surface point by point to obtain the spatial coordinates of the sampled points, and calculate the color information of the sampled points through the optimized hash encoding network and the MLP network;

[0123] S30203: Generate a texture map according to the color information of the sampled points output by the MLP network, and bind the texture map to the corresponding triangular network to obtain a 3D triangular mesh model with high-quality geometric shapes and realistic textures.

[0124] Specific Embodiment 2: Taking 3D Gaussian Splatting as an example, a general 3D intelligent modeling method with multi-view coupling constraints proposed in this embodiment includes the following steps:

[0125] 1. Guide the selection of the optimization model. According to the multi-view coupling constraint optimization formula in Formula (5), this embodiment uses a total of three pre-trained diffusion models to achieve the 3D optimization of multi-view coupling constraints. Specifically:

[0126] (1). For in This embodiment selects the pre-trained text-to-image diffusion model stablediffusion - 2 - 1 - base as the guiding model for approximation;

[0127] (2). For In this embodiment, the pre-trained text-to-image diffusion model StableDiffusion-2-1 with an attached learnable LoRA module is selected as the guiding model for approximation.

[0128] (3). For In this embodiment, the fine-tuned multi-view generation diffusion model MVDream is used for approximation.

[0129] During the optimization process, the network parameters of StableDiffusion-2-1-base, StableDiffusion-2-1, and MVDream remain unchanged, while the parameters of the LoRA module are fine-tuned according to the optimization progress of 3D Gaussian Splatting and the texture grid.

[0130] 2. Optimized selection of camera view parameter settings: In each iteration, the camera view is randomly sampled according to a specific parameter range. The sampling radius of the camera pose is limited within the range of [2.0, 2.5], the field of view (FOV) ranges from [40°, 70°], the azimuth angle coverage ranges from [-180°, 180°], and the depression angle is limited within the range of [-90°, 30°].

[0131] 3. Optimize 3D Gaussian Splatting with multi-view coupling constraints to obtain a preliminary three-dimensional representation. In this embodiment, 3D Gaussian Splatting is optimized through multi-view coupling constraints to generate a preliminary three-dimensional representation. The specific process is as follows:

[0132] (1). Initialize the parameters of 3D Gaussian Splatting: Initialize 1000 3D Gaussian primitives; the transparency of each primitive is set to 0.1, and the color is gray; the positions of the primitives are randomly distributed within a unit sphere with a radius of 0.5.

[0133] (2). 3D Gaussian Splatting optimization settings: The rendering resolution is gradually increased, starting from 128, and adjusted to 256, 512, and 1024 at the 400th, 1200th, and 2000th iteration steps respectively; the rendered background color is randomly selected between white and black; in the first 1500 iteration steps, 3D Gaussian Splatting is cropped and densified every 250 steps, where the gradient threshold is set to 0.01, and primitives with a transparency less than 0.01 or a covariance matrix scale exceeding 0.05 will be removed.

[0134] (3). Multi - perspective Coupling Constraint Optimization Strategy: Optimize 3D Gaussian Splatting based on formula (5) and the pre - trained guided diffusion model. The λ parameter in formula (5) is set to 0.5, and the CFG parameter is set to 7.5, with a total of 4000 optimization iteration steps; in the first 2000 optimization iteration steps, the diffusion step Subsequently anneals to During the optimization process, the Adam optimizer is used, and independent learning rates are set for different parameters. The learning rate of the position parameter decreases non - linearly from 1×10 -3 To 2×10 -5 in the first 1500 iteration steps, and then remains unchanged. The learning rate of the color parameter is set to 0.01, the learning rate of the transparency parameter is set to 0.05, the learning rate of the scale parameter is set to 5×10 -3 and the learning rate of the rotation parameter is set to 1×10 -3 .

[0135] Through the above optimization strategy, multiple preliminary 3D representations finally generated are as Figure 5 shown, demonstrating the high quality and consistency of the rendering results.

[0136] 4. Coarse Texture Mesh Extraction. Extract the coarse texture mesh from the optimized 3D Gaussian Splatting, and the specific steps are as follows:

[0137] (1). Voxel Grid Conversion: Convert the optimized 3D Gaussian Splatting into a coarse voxel grid according to formula (8), and the threshold for voxel conversion is set to 0.2 to mark significant voxel regions.

[0138] (2). Distance Field Calculation and Tetrahedral Mesh Initialization: Calculate the SDF value of the mesh using the distance_transform_edt function in the SciPy library of Python. Initialize the deformable tetrahedral mesh (DMTet) using the calculated SDF to generate a geometric mesh representation of the target 3D object, ensuring the capture of detailed geometric features.

[0139] (3). Mesh Texture Initialization: Initialize the hash grid encoding parameters and the prediction layer of the multi - layer perceptron (MLP). Based on the results of the optimized 3D Gaussian Splatting, pre - initialize the hash grid encoding parameters and the MLP network to provide reasonable initial values for subsequent texture optimization.

[0140] 5. Fine-tuning of the fine-grained texture grid with multi-view coupling constraints. To generate a high-quality 3D texture grid, this embodiment decouples the geometry and texture of the 3D grid and combines multi-view coupling constraints to optimize the geometry and texture respectively. The specific steps are as follows:

[0141] (1). Geometry optimization: First, extract the triangular mesh through the Marching Tetrahedra algorithm, and use the nvdiffrast library in Python to render the surface normal map of the generated mesh. According to formula (5), optimize the DMTet parameters to improve the mesh geometry. The geometry optimization process iterates a total of 15,000 steps. During the optimization process, the λ parameter in formula (5) is set to 0.5, the CFG parameter is set to 100, and the diffusion step Subsequently anneals to The learning rate is set to 5×10 -3 , and the rendering resolution is set to 512.

[0142] (2). Texture optimization: After the geometry optimization is completed, fix the fine-tuned DMTet parameters, and extract the triangular mesh again through Marching Tetrahedra. Sample the mesh surface, predict the colors of the sampled points, and use the nvdiffrast library in Python to render multi-view images. According to formula (5), update the hash encoding network and the MLP network layer through backpropagation to further optimize the texture representation. The texture optimization iterates a total of 3,000 steps. During the optimization process, in the first 1,000 steps, the diffusion step Subsequently anneals to The learning rates of the hash encoding network and the MLP layer are set to 0.05 and 5×10 -3 respectively, the λ parameter in formula (5) is set to 0.1, and the rendering resolution is set to 512.

[0143] After the optimization is completed, extract the mesh through the Marching Tetrahedra algorithm and calculate the UV texture map (i.e., the texture map) to generate a 3D triangular mesh model with high-quality geometry and realistic texture, as Figure 6 and Figure 7 shown.

[0144] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the above-disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments within the spirit and principle of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A general 3D intelligent modeling method with multi-view coupled constraints, characterized in that: The steps of the general 3D intelligent modeling method with multi-view coupling constraints include: Step 1: Optimize the learnable 3D representation through the pre-trained text-to-image diffusion model and the fine-tuned multi-view generative diffusion model to generate a preliminary 3D representation that is fully consistent with the text semantics; Step 2: Based on the generated preliminary 3D representation, extract and convert the rough 3D mesh representation containing geometric structure and texture information according to the transparency prediction information of each point in the preliminary 3D representation and the multi-view rendering of the preliminary 3D representation; Step 3: Use the geometry and texture decoupling optimization method, combined with multi-view coupling constraints, to fine-tune and optimize the geometry and surface texture of the triangular mesh respectively, and generate a three-dimensional triangular mesh model with high-quality geometry and realistic texture.

2. A general 3D intelligent modeling method with multi-view coupled constraints according to claim 1, characterized in that: Step 1 specifically includes: Step 1.1: Randomly initialize the parameters of the learnable 3D representation model, including the parameters that control the geometric structure and texture properties of the model; Step 1.2: Randomly select four mutually orthogonal horizontal viewpoints in the learnable 3D representation with the same elevation angle and camera parameters, render each viewpoint by rasterization rendering method, and generate four corresponding rendering views; Step 1.3: Diffusion and noise are added to the four rendered views, combined with the text prompts, and input into the pre-trained diffusion model to output multi-view target renderings that meet the text conditions, and supervise the learnable 3D representation; Step 1.4: Repeat steps 1.2 and 1.3 until the iteration converges or reaches a preset number of iterations, completing the optimization of the learnable three-dimensional representation and generating a preliminary three-dimensional representation that is completely consistent with the text semantics; The rendering expression for each pixel is: In formulas (1) and (2), is the color of each pixel in the rendered image along the light ray r, T i is the cumulative transmittance of the learnable 3D representation at the query point, σ i is the opacity of the learnable 3D representation at the query point, c i Color information at the query point for a learnable 3D representation; The calculation formula for rendering the view is: In formula (3), θ is the parameter of the learnable three-dimensional representation; v i is the information of the rendering perspective; g is a differentiable rendering function, is a three-dimensional representation at a viewing angle v i The resulting rendered image is shown below.

3. A general 3D intelligent modeling method with multi-view coupling constraints according to claim 2, characterized in that: Step 1.3 specifically includes: Step 1.3.1: Perform gradual diffusion noise addition on the generated rendering view. Input the random single-view rendering image with noise added in step t and the text prompt into the pre-trained text generation picture diffusion model to predict the t-1-step noise view of the single view. Input the multi-view rendering image with noise added in step t and the text prompt into the fine-tuned multi-view generation diffusion model to predict the t-1-step noise view of the multi-view. Step 1.3.2: Take the t-1th step noise view of the single view and the t-1th step noise view of the multi-view as the target view of the four rendered views at the t-1th step, calculate the loss function of the learnable 3D representation model parameter θ according to the target view of the four rendered views at the t-1th step and the pre-trained text generation picture diffusion model with the LoRA module added, and perform supervised learning to update the 3D representation parameter θ; The expression of diffuse noise for each rendering is: In formula (4), for The image generated after adding noise according to the diffusion time step t, α t and σ t is a hyperparameter, ∈ is the noise randomly sampled from a multidimensional normal distribution; The loss function gradient calculation formula of the learnable three-dimensional representation model parameter θ is: In formulas (5) and (6), are four orthogonal perspectives, is the noise rendering of the corresponding viewing angle, y is the input text condition, is the diffusion time step of the diffusion model, ∈ is the random noise added in the diffusion process, ∈ φ Generate image diffusion model for pre-trained text, ∈ φ The parameters of are fixed in the optimization process, s is the classification guidance parameter, λ is the regularization parameter, ∈ M Generate a diffusion model for fine-tuning multi-view, ∈ M The parameters of are fixed during the optimization process. To add the pre-trained text generation image diffusion model of the LoRA module, the LoRA module parameters are updated alternately with the learnable 3D representation model parameters θ during the optimization process; The expression for LoRA module parameter optimization is:

4. A general 3D intelligent modeling method with multi-view coupled constraints according to claim 1, characterized in that: Step 2 specifically includes: Step 2.1: Map the optimized learnable 3D representation to (-1,1) 3 In the cube space, divide the cube space into 16 3 sub-blocks; Step 2.2: Divide each sub-block into 8 3 grid cells, calculate the transparency information of the center position of each grid cell, and then convert (-1,1) 3 The cube space is converted to 256 3 Grid cells, and obtain transparency information of each grid cell; Step 2.3: Set the threshold β to filter the transparency information, convert the grid cell values ​​into binary form, and obtain a rough voxel grid. Use the distance_transform_edt function in the SciPy library in Python to calculate the signed distance function value of each unit in the voxel grid. Use the initialized deformable tetrahedral grid and combine the Marching Tetrahedra algorithm to convert (-1,1) 3 The cube space is divided into tetrahedral units, and the vertices of the tetrahedral units are assigned SDF values, and the distance from the corresponding vertex to the surface of the modeled three-dimensional object is obtained through the SDF value; Step 2.4: Initialize the mesh texture of the triangle mesh based on the multi-view strategy and the optimized preliminary 3D representation to generate a rough 3D mesh representation containing geometric structure and texture information; The calculation formula of transparency information is: d(x)=f θ (x) (8); In formula (8), d(x) is the transparency information of the grid unit with x as the center coordinate, and f θ A function that calculates transparency information at coordinate point x for a three-dimensional representation.

5. A general 3D intelligent modeling method with multi-view coupled constraints according to claim 4, characterized in that: Initializing the mesh texture of the triangular mesh in step 2.4 in combination with the optimized preliminary three-dimensional representation includes: Step 2.4.1: For the surface texture of the triangular mesh, a hash grid is used to encode the spatial position and combined with a single-layer MLP network for simulation, wherein the MLP network is a multi-layer perceptron network, the input of the MLP network is the representation of the three-dimensional coordinates of any point in space after being encoded by the hash network position, and the output is the color of the three-dimensional modeling object at the corresponding position in space. Based on the generated triangular mesh and the color of the MLP output surface sampling point, under the condition of a given rendering perspective, the Python nvdiffrast library is used to render and generate the initial rendering of the triangular mesh; Step 2.4.2: According to the multi-view strategy, different camera perspectives are selected to render the preliminary 3D representation and the triangular mesh after mesh texture initialization respectively, and the rendering of the preliminary 3D representation is used as the target label and the loss is calculated with the mesh rendering; Step 2.4.3: Optimize the MLP network and hash grid encoding parameters through gradient back propagation to achieve mesh texture initialization of the preliminary three-dimensional representation rendering. In the process of mesh texture initialization of the preliminary three-dimensional representation rendering, the SDF value of the mesh vertex and its displacement parameters remain unchanged, and only the texture-related parameters are optimized.

6. A general 3D intelligent modeling method with multi-view coupled constraints according to claim 1, characterized in that: Step 3 specifically includes: Step 3.1: Optimize and fine-tune the geometric structure and texture information of the texture mesh represented by the coarse 3D mesh by using a multi-view coupled constraint optimization method; Step 3.2: Perform fine mesh extraction on the optimized and finely adjusted coarse 3D mesh representation to obtain a 3D triangular mesh model with high-quality geometry and realistic texture.

7. A general 3D intelligent modeling method with multi-view coupled constraints according to claim 6, characterized in that: Step 3.1 specifically includes: Step 3.1.1: Based on the initialized deformable tetrahedral mesh, a rough 3D triangle mesh is extracted by the Marching Tetrahedra algorithm, and the vertex normal vectors of the rough 3D mesh are calculated using the Python nvdiffrast library; Step 3.1.2: Convert vertex normals to texture space coordinates through UV mapping, and store the normal vectors in the form of RGB values ​​to complete the optimization of the deformable tetrahedral mesh and generate the surface normal map of the 3D mesh; Step 3.1.3: Replace the multi-view target rendering image in step 1.3 with the generated surface normal map, calculate the optimization gradient using formula (5) and formula (6), and update the parameters of the deformable tetrahedral mesh through back propagation; Step 3.1.4: Repeat steps 3.1.1 to 3.1.3 until the optimization converges or reaches the specified number of optimization steps, completing the optimization and fine-tuning of the rough three-dimensional mesh geometry; Step 3.1.5: Fix the optimized deformable tetrahedral mesh, use the Marching Tetrahedra algorithm to extract the rough 3D mesh, and perform coordinate query on each surface sampling point of the mesh; Step 3.1.6: Hash-code the spatial coordinates of the sampling points obtained by the query and input them into the MLP, and output the color information of the corresponding sampling points; Step 3.1.7: Use the nvdiffrast library to render the rough 3D mesh and generate multi-view rendering images; Step 3.1.8: Calculate the optimization gradient of the multi-view rendered image according to formula (5) and formula (6), and update the hash coding network and MLP network layer through back propagation; Step 3.1.9: Repeat steps 3.1.5 to 3.1.8 until the optimization converges or reaches the specified number of optimization steps to complete the optimization and fine-tuning of the coarse three-dimensional mesh texture.

8. A general 3D intelligent modeling method with multi-view coupled constraints according to claim 6, characterized in that: Step 3.2 specifically includes: Step 3.2.1: Based on the optimized deformable tetrahedral mesh, extract the rough 3D mesh by Marching Tetrahedra algorithm to obtain the mesh vertex coordinates and facet topology structure; Step 3.2.2: Sample the grid surface point by point to obtain the spatial coordinates of the sampling points. Pass the spatial coordinates of the sampling points through the optimized hash coding network and MLP network to calculate the color information of the sampling points. Step 3.2.3: Generate a texture map based on the color information of the sampling points output by the MLP network, and bind the texture map to the corresponding triangle network to obtain a three-dimensional triangle mesh model with high-quality geometry and realistic texture.

Citation Information

Cited By

  • Three-dimensional data synthesis method, electronic equipment, storage medium and program product

    CN120259590A

  • Three-dimensional data synthesis method, electronic device, storage medium and program product

    CN120259590B

  • 3D wordart generation method based on structure-view angle double-stage diffusion model

    CN121414972A