Three-dimensional scene simulation modeling method in game making process
Through multi-view coupling constraint optimization of three-dimensional representation and texture grid fine-tuning, the problems of inefficiency and inconsistency in multiple perspectives in the existing three-dimensional modeling methods are solved, and efficient and accurate three-dimensional model generation is achieved. The generated models are highly consistent in each perspective and have rich texture details.
Patent Information
- Application Number
- CN202510646036.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-25
AI Technical Summary
The existing three-dimensional modeling methods rely on manual experience and are inefficient, making it difficult to achieve large-scale and rapid generation of diversified models. The multi-view joint optimization method that directly trains 3D generative models and fine-tuned multi-view joint optimization methods are restricted by the scarcity of resources of high-quality and large-scale text-3D datasets, resulting in poor generalization capabilities and easy overfitting of the generation results; methods based on single-view optimization are prone to the problem of geometric inconsistencies in multiple perspectives.
The preliminary three-dimensional representation is optimized by multi-view coupling constraints. By randomly initializing the optimized three-dimensional representation and rendering multiple orthogonal view angles, the loss calculation is performed in combination with the pre-trained diffusion model and the fine-tuned multi-view diffusion model to optimize the three-dimensional representation; the coarse texture grid is extracted and fine-tuned fine-tuned by multi-view coupling constraints to ensure geometric consistency and texture richness.
The efficiency and accuracy of three-dimensional modeling are significantly improved. The generated models are highly consistent from various perspectives and rich in texture details. They reduce their dependence on large-scale data sets and realize the intelligent generation of high-quality three-dimensional models.
Smart Images

Figure CN120374896A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a three-dimensional scene simulation modeling method, belonging to the technical field of game production. Background Art
[0002] With the rapid development of digital technology and the increasing demands of architectural design and industrial manufacturing, three-dimensional modeling technology is gradually becoming an important tool in fields such as game development and virtual reality (VR). In these application scenarios, the quality, complexity, and generation efficiency of models directly affect the effectiveness and cost of the design and production processes. However, existing three-dimensional modeling methods still face many challenges in practical applications.
[0003] Traditional three-dimensional modeling mainly relies on professional design software, such as Maya, 3ds Max, Rhino, and Revit. These tools support the full process of design from geometric shape construction to detailed texture addition by providing designers with rich modeling functions. However, these tools usually require designers to spend a lot of time manually adjusting model details, which not only has low work efficiency but also highly depends on the experience and skill level of designers, making it difficult to achieve large-scale and rapid generation of diverse models. In addition, frequent adjustments and trial-and-error in the design process significantly extend the design cycle, and the innovation and consistency of design results are often difficult to guarantee. How to get rid of the dependence on manual experience, generate high-quality three-dimensional models with both rationality and innovation, and achieve efficient and accurate modeling optimization in multiple fields and scenarios has become the core issue in promoting the intelligent and automated development of three-dimensional modeling.
[0004] In recent years, with the rapid development of machine learning and deep learning technologies, three-dimensional modeling generation methods based on artificial intelligence have gradually emerged. The application of generative models (such as generative adversarial networks GAN, variational autoencoders VAE, and diffusion models) makes it possible to learn the latent sample distribution from massive data. These methods achieve efficient and automated generation of three-dimensional models by analyzing the geometric features and distribution laws of three-dimensional models. Currently, three-dimensional modeling methods based on artificial intelligence mainly include the following two approaches:
[0005] (1) Direct modeling method based on three-dimensional generative models: Use a large-scale text-three-dimensional model dataset to train a generative model and directly generate a three-dimensional model according to text input.
[0006] (2) Three-dimensional modeling method based on single-view optimization: Randomly perform single-view optimization on a learnable three-dimensional representation through a pre-trained two-dimensional diffusion model, and gradually optimize the three-dimensional representation to achieve generation from text to three-dimensional model.
[0007] (3) Fine-tuned multi-view joint optimization 3D modeling method: A multi-view diffusion model is obtained by fine-tuning the pre-trained 2D diffusion model, and its multi-view view output is used to perform multi-view joint optimization on the 3D representation, thereby achieving geometrically consistent 3D model generation. Although these methods have significantly improved the efficiency of 3D modeling, there are still some problems that need to be solved. For example, direct training of 3D generation models and fine-tuning multi-view joint optimization methods are constrained by the lack of high-quality, large-scale text-3D dataset resources, resulting in poor generalization ability and easy overfitting of generated results; while methods based on single-view optimization are prone to multi-view geometric inconsistency problems. Therefore, developing a general, efficient 3D modeling generation method that can ensure multi-view geometric consistency has become an important research direction in this field. Summary of the invention
[0008] The present invention aims to solve the problem that the multi-view joint optimization method for directly training a 3D generative model and fine-tuning is restricted by the lack of high-quality, large-scale text-3D dataset resources, resulting in poor generalization ability and easy overfitting of the generated results; while the method based on single-view optimization is prone to the problem of multi-view geometric inconsistency. Therefore, a 3D scene simulation modeling method in the game production process is proposed.
[0009] The technical solution adopted by the present invention to solve the above problems is: the steps of the present invention include:
[0010] Step 1: Preliminary three-dimensional representation of multi-view coupled constraint optimization;
[0011] Step 2: Extract a rough texture grid;
[0012] Step 3: Fine-tune the texture mesh with multi-view coupling constraints.
[0013] Furthermore, step 1 specifically includes:
[0014] Randomly initialize the optimizable 3D representation and randomly select four rendering perspectives that are mutually orthogonal in the horizontal direction. Render the 3D representation using a differentiable rendering function to obtain orthogonal views at a given perspective.
[0015] Couple the pre-trained diffusion model with a fine-tuned multi-view diffusion model, compute the loss between the rendered orthogonal view and the target view given the text, and optimize the 3D representation via gradient backpropagation;
[0016] This process is repeated iteratively until the optimization converges, and finally a preliminary three-dimensional representation consistent with the text semantics is obtained.
[0017] Furthermore, step 2 specifically includes:
[0018] Based on the preliminary three-dimensional representation optimized in Step 1, according to the transparency prediction information of each point in space, extract and convert it into a rough three-dimensional voxel representation;
[0019] Meanwhile, calculate the signed distance function for each voxel unit to generate a preliminary signed distance field;
[0020] Initialize the deformable tetrahedral mesh through the signed distance field to generate a geometric mesh for 3D modeling;
[0021] In terms of texture generation, use spatial position hash encoding and a single-layer fully connected perception network to perform texture learning on the geometric mesh, and combine the optimized preliminary three-dimensional representation to supervise and initialize the encoding network and MLP, thereby generating a rough three-dimensional mesh representation containing geometric structure and texture information.
[0022] Furthermore, Step 3 specifically includes:
[0023] Through the multi-view coupling constraint optimization method, finely tune the geometric structure and texture information of the texture mesh respectively to ensure that the generated three-dimensional texture mesh has high geometric consistency and rich details;
[0024] After the fine-tuning is completed, use the marching tetrahedra algorithm to extract a high-quality triangular mesh, and convert the texture information on the surface into the UV map of the mesh, and finally generate a fine triangular mesh 3D model with high-fidelity texture.
[0025] The beneficial effects of the present invention are:
[0026] 1. By coupling the pre-trained two-dimensional diffusion model with the fine-tuned multi-view generation diffusion model and adopting a joint optimization strategy, this method effectively reduces the dependence on large-scale text-3D datasets, significantly enhances the generality and applicability of the technology, and breaks through the limitation of data scarcity on 3D modeling;
[0027] 2. Through the multi-view joint optimization strategy, this method successfully solves the core problem of multi-view geometric inconsistency in the current 3D modeling methods optimized by artificial intelligence; the generated 3D model shows high consistency and accuracy in each view, greatly improving the authenticity and reliability of the modeling results;
[0028] 3. Based on the step-by-step optimization modeling strategy, this method significantly improves the efficiency of 3D modeling, and at the same time realizes the generation of high-fidelity 3D models with fine details and rich textures, providing an efficient and accurate solution for high-quality 3D modeling;
[0029] 4. This method provides intelligent and efficient technical support for complex 3D modeling tasks;
[0030] 5. The present invention optimizes the effect of 3D modeling through machine learning and deep learning, improving the efficiency of 3D modeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a flowchart of a general 3D intelligent modeling method with multi-view coupling constraints;
[0032] Figure 2 is a visualization diagram of the 3D representation rendering process;
[0033] Figure 3 is a flowchart of optimizing 3D representation with multi-view coupling constraints;
[0034] Figure 4 is a schematic diagram of geometric extraction and texture initialization of the texture grid;
[0035] Figure 5 is a schematic diagram of the preliminary 3D shape representation generated by optimizing 3D Gaussian Splatting;
[0036] Figure 6 is a schematic diagram of a high-precision 3D triangular mesh model generated by optimizing the texture grid;
[0037] Figure 7 is a schematic diagram of the generated 3D triangular mesh model of the building. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] Embodiment 1: As Figures 1 to 7 shown, a 3D scene simulation modeling method during game production includes the following steps:
[0039] Step 1. Optimize the preliminary 3D representation with multi-view coupling constraints;
[0040] Step 101. Differentiable rendering of the 3D representation from multiple views; First, randomly initialize the parameters of the 3D Gaussian Splatting model, and randomly select four horizontal direction views that are mutually orthogonal and have the same elevation angle and camera parameters; Through rasterization rendering technology, generate the corresponding four rendering views; The rendering process of each pixel in the rendering image is as follows:
[0041]
[0042] where represents the color of each pixel in the rendering image; G(l) represents each Gaussian primitive in 3D Gaussian Splatting, and its definition is as follows: Each Gaussian primitive is parameterized with the following parameters: the center position coordinate the covariance matrix ∑, through the scaling factor and the rotation quaternion Parametric representation); color parameter and transparency parameter Then the parameter set of 3D Gaussian Splatting can be expressed as θ = {μ k , s k , q k , c k , α k}, where k represents the serial number of the Gaussian primitive; in the formula, l represents the local coordinate centered on μ; represents the ordered set of Gaussian primitives overlapping with the pixel along the ray r; the rendering expression of 3D Gaussian Splatting is obtained as follows:
[0043]
[0044] where θ represents the parameters of the 3D Gaussian to be rendered; v i represents the information of the rendering view; g represents the differentiable rendering function, and the rendering process of each pixel is calculated according to formulas (1) and (2); represents the rendering image obtained by the three-dimensional representation under the view v i ;
[0045] Step 102, perform diffusion noise addition on the multi-view rendering images, and combine the text input to predict the target rendering images of multiple views; based on the forward diffusion noise addition process of the diffusion model, gradually add noise to the multi-view rendering images generated in step 101. The specific diffusion noise addition process for each rendering image is as follows:
[0046]
[0047] where, represents the image generated after adding noise according to the diffusion time step t, α t and σ t are hyperparameters, and ∈ is the noise randomly sampled from the multi-dimensional normal distribution;
[0048] Input the multi-view rendering image with noise added at the t-th step and the text prompt into the pre-trained text-to-image diffusion model and the fine-tuned multi-view generation diffusion model, and predict the noise views at the (t - 1)-th step for single-view and multi-view respectively; these predicted views are used as the target views of the multi-view rendering image at the (t - 1)-th step to calculate the loss function and perform supervised learning; after simplification, the gradient calculation formula of the loss function for the parameters θ of the 3D representation is as follows:
[0049]
[0050] where Represents four orthogonal views The corresponding noise rendering map; y represents the input text condition; Represents the diffusion time step of the diffusion model; ∈ is the random noise added during the diffusion process; ∈ φ Represents a pre-trained text-to-image diffusion model, whose parameters are fixed during the optimization process; s is the classification guidance parameter; λ is the regularization parameter; ∈ M Represents a fine-tuned multi-view generation diffusion model, whose parameters are also unchanged during the optimization process; Represents a pre-trained text-to-image diffusion model with a LoRA module added. The LoRA module parameters of this model are alternately updated with the parameters θ of the three-dimensional representation during the optimization process. The optimization calculation formula is:
[0051]
[0052] Repeat the above steps until the iteration converges or reaches the preset number of iterations, and finally generate a preliminary three-dimensional representation that is highly consistent with the text semantics. The iterative optimization process is as Figure 3 shown;
[0053] Step 2, Coarse texture mesh extraction; Initialize a coarse texture mesh form according to the preliminary three-dimensional representation optimized in Step 1 for subsequent refined optimization to generate a high-quality triangular mesh three-dimensional model; Specifically include:
[0054] Step 201, Geometric initialization of the mesh; First, map the optimized 3D Gaussian to the cube space of (-1, 1) 3 and divide this space into 16 3 sub-blocks; For each sub-block, screen out the Gaussian basis elements whose center position μ is within this sub-block. Further, divide each sub-block into 8 3 meshes, calculate the transparency of the center position of each mesh cell in each Gaussian basis element within the sub-block and accumulate it to obtain the transparency of each mesh cell; Finally, convert the entire (-1, 1) 3 into a 256 3 mesh and calculate the transparency information of each mesh cell; The specific calculation formula for the mesh transparency is:
[0055]
[0056] where d(x) represents the transparency of the mesh cell with x as the center coordinate, B represents the set of all Gaussian basis elements within the sub-block where this mesh is located, μ i and ∑ i represent the center coordinate and covariance matrix parameter of the i-th Gaussian basis element respectively;
[0057] By setting a threshold β, screening transparency information, converting the values into binary form (0 or 1), a rough voxel grid is obtained; based on the obtained rough voxel grid, using the distance_transform_edt function in the SciPy library in Python, the signed distance function value of each voxel cell is calculated, and it is used to initialize the deformable tetrahedral grid, and the SDF value is assigned to the vertices of each tetrahedral cell; where DMTet divides (-1, 1) 3 into tetrahedral cells, and the vertices of each cell contain SDF values, which reflect the distance from the vertex to the surface of the modeled three-dimensional object. The displacement and SDF value of each vertex are learnable parameters, and the mesh can be effectively extracted through the Marching Tetrahedra algorithm;
[0058] Step 202, texture initialization of the mesh; for the surface texture of the triangular mesh, the present invention uses a hash grid to encode the spatial position and combines a single-layer multi-layer perceptron to approximately simulate; the input of the MLP network is the representation of the three-dimensional coordinates of any point in space after being encoded by the hash grid position, and the output is the color of the three-dimensional modeling object at this spatial position; based on the triangular mesh generated in step 201 and the color of the surface sampling points output by the MLP, under the condition of a given rendering perspective, the initial rendering of the triangular mesh is generated using the nvdiffrast library in Python;
[0059] Through a multi-view strategy, multiple camera views are selected to render the preliminary three-dimensional representation generated in the first step and the initialized triangular mesh respectively, and the rendering map optimized from the preliminary three-dimensional representation is used as the target label to calculate the loss with the mesh rendering map; through gradient backpropagation, the MLP network for texture prediction and the hash grid encoding parameters are optimized to achieve texture initialization; during this process, the SDF value of the mesh vertices and their displacement parameters remain unchanged, and only the texture-related parameters are optimized;
[0060] Step 3, fine-tuning of the high-precision texture mesh with multi-view coupling constraints; in this step, in the process of texture mesh fine-tuning, the present invention uses a method of decoupling geometric and texture optimization, combined with multi-view coupling constraints, to optimize the geometric shape and surface texture of the triangular mesh respectively, so as to generate a high-precision texture mesh; the specific process includes the following steps:
[0061] Step 301: Optimized fine-tuning of grid geometry; Based on the DMTet initialized in Step 2, extract the triangular mesh through the MarchingTetrahedra algorithm, and use the nvdiffrast library in Python to calculate the vertex normal vectors of the triangular mesh. Subsequently, convert the vertex normals to texture space coordinates through UV mapping, and store the normal vectors in the form of RGB values to generate the surface normal map of the 3D mesh; Replace the rendered image in Step 102 with the generated surface normal map, calculate the optimization gradient according to Formulas (5) and (6), and update the parameters of DMTet through backpropagation. Repeat the above process until the optimization converges or reaches the specified number of optimization steps to achieve the optimization and refinement of the triangular mesh geometry;
[0062] Step 302: Optimized fine-tuning of grid texture. Fix the DMTet optimized in Step 301, extract the triangular mesh using the MarchingTetrahedra algorithm, and then perform coordinate queries on each surface sampling point of the mesh; Further, hash-encode the spatial coordinates of the sampling points and input them into the MLP to output the color information of the point. Then, use the nvdiffrast library to render the mesh to generate multi-view rendered images; Subsequently, calculate the optimization gradient for the multi-view rendered images according to Formulas (5) and (6) as well, and update the hash encoding network and the MLP network layer through backpropagation. Repeat the above process until the optimization converges or reaches the specified number of optimization steps to achieve the optimization and refinement of the triangular mesh texture;
[0063] Step 303: Extraction of refined texture mesh; Use the optimized DMTet to extract the triangular mesh through the Marching Tetrahedra algorithm to obtain the mesh vertex coordinates and patch topology; Subsequently, perform point-by-point sampling on the mesh surface to obtain the spatial coordinates of the sampling points, and pass them through the optimized hash encoding network and MLP network to calculate the color information of the sampling points; Generate a texture map from the color values output by the MLP and bind it to the vertices or patches of the triangular mesh. Finally, obtain a three-dimensional triangular mesh model with high-quality geometric shapes and realistic textures.
[0064] Embodiment
[0065] Taking 3D Gaussian Splatting as an example, a specific embodiment of the general 3D intelligent modeling method with multi-view coupling constraints is provided.
[0066] 1. Guiding the optimization model selection; According to the multi-view coupling constraint optimization formula in Formula (5), the present invention uses a total of three pre-trained diffusion models to achieve 3D optimization with multi-view coupling constraints; Specifically:
[0067] (1). For in Select the pre-trained text-to-image diffusion model stablediffusion-2-1-base as the guiding model for approximation;
[0068] (2). For Select the pre-trained text-to-image diffusion model stablediffusion-2-1 with an attached learnable LoRA module as the guiding model for approximation;
[0069] (3). For Use the fine-tuned multi-view generation diffusion model MVDream for approximation;
[0070] During the optimization process, the network parameters of stablediffusion-2-1-base, stablediffusion-2-1, and MVDream remain unchanged, while the parameters of the LoRA module are fine-tuned according to the optimization process of 3D Gaussian Splatting and the texture grid;
[0071] 2. Optimize the selection of camera view parameters; in each iteration, the camera view is randomly sampled according to a specific parameter range; the sampling radius of the camera pose is limited within the range of [2.0, 2.5], the field of view angle ranges from [40°, 70°], the azimuth angle coverage ranges from [-180°, 180°], and the depression angle is limited within the range of [-90°, 30°];
[0072] 3. Optimize 3D Gaussian Splatting with multi-view coupling constraints to obtain a preliminary three-dimensional representation; the present invention optimizes 3D Gaussian Splatting through multi-view coupling constraints to generate a preliminary three-dimensional representation; the specific process is as follows:
[0073] (1). Initialize the 3D Gaussian Splatting parameters: Initialize 1000 3D Gaussian primitives; the transparency of each primitive is set to 0.1, and the color is gray; the positions of the primitives are randomly distributed within a unit sphere with a radius of 0.5.
[0074] (2). Optimization settings for 3D Gaussian Splatting: The rendering resolution is gradually increased. It starts at 128 and is adjusted to 256, 512, and 1024 at the 400th, 1200th, and 2000th iteration steps respectively; the background color for rendering is randomly selected between white and black; in the first 1500 iteration steps, 3D Gaussian Splatting is cropped and densified every 250 steps, where the gradient threshold is set to 0.01, and primitives with a transparency less than 0.01 or a covariance matrix scale exceeding 0.05 will be removed;
[0075] (3). Optimization strategy for multi-view coupling constraints: Based on Equation (5) and the pre-trained guided diffusion model, 3D Gaussian Splatting is optimized. The λ parameter in Equation (5) is set to 0.5, and the CFG parameter is set to 7.5. A total of 4000 iteration steps are optimized; in the first 2000 iteration steps of the optimization, the diffusion steps Subsequently anneal to During the optimization process, the Adam optimizer is used, and independent learning rates are set for different parameters. The learning rate of the position parameter is non-linearly reduced from 1×10 -3 to 2×10 -5 in the first 1500 iteration steps and then remains unchanged. The learning rate of the color parameter is set to 0.01, the learning rate of the transparency parameter is set to 0.05, the learning rate of the scale parameter is set to 5×10 -3 and the learning rate of the rotation parameter is set to 1×10 -3 ;
[0076] Through the above optimization strategy, multiple preliminary 3D representations finally generated are as Figure 5 shown, demonstrating the high quality and consistency of the rendering results;
[0077] 4. Coarse texture mesh extraction; Coarse texture meshes are extracted from the optimized 3D Gaussian Splatting. The specific steps are as follows:
[0078] (1). Voxel grid conversion: The optimized 3D Gaussian Splatting is converted into a coarse voxel grid according to Equation (8). The threshold for voxel conversion is set to 0.2, which is used to mark significant voxel regions;
[0079] (2). Distance field calculation and tetrahedral mesh initialization: The SDF value of the grid is calculated based on the distance_transform_edt function in the SciPy library of Python. The calculated SDF is used to initialize the deformable tetrahedral mesh (DMTet) to generate a geometric mesh representation of the target 3D object, ensuring the capture of geometric features with details;
[0080] (3). Mesh Texture Initialization: Initialize the hash grid encoding parameters and the prediction layer of the multi-layer perceptron. Based on the results of the optimized 3D Gaussian Splatting, pre-initialize the hash grid encoding parameters and the MLP network to provide reasonable initial values for subsequent texture optimization;
[0081] 5. Fine-tuning of the Fine Texture Grid with Multi-view Coupling Constraints; To generate a high-quality three-dimensional texture grid, the present invention decouples the geometry and texture of the 3D grid and combines multi-view coupling constraints to optimize the geometry and texture respectively. The specific steps are as follows:
[0082] (1). Geometry Optimization: First, extract the triangular mesh through the Marching Tetrahedra algorithm, and use the nvdiffrast library in Python to render the surface normal map of the generated mesh; According to formula (5), optimize the DMTet parameters to improve the mesh geometry; The geometry optimization process iterates a total of 15,000 steps. During the optimization process, the λ parameter in formula (5) is set to 0.5, and the CFG parameter is set to 100. In the first 5,000 iteration steps of the optimization, the diffusion step Subsequently anneal to The learning rate is set to 5×10 -3 , and the rendering resolution is set to 512;
[0083] (2). Texture Optimization: After the geometry optimization is completed, fix the fine-tuned DMTet parameters, and extract the triangular mesh again through Marching Tetrahedra; Sample the mesh surface, predict the color of the sampled points, and use the nvdiffrast library in Python to render multi-view images; According to formula (5), update the hash encoding network and the MLP network layer through backpropagation to further optimize the texture representation; The texture optimization iterates a total of 3,000 steps. During the optimization process, in the first 1,000 steps, the diffusion step Subsequently anneal to The learning rates of the hash encoding network and the MLP layer are set to 0.05 and 5×10 -3 respectively, the λ parameter in formula (5) is set to 0.1, and the rendering resolution is set to 512;
[0084] After the optimization is completed, extract the mesh through the Marching Tetrahedra algorithm, calculate the UV texture map, and generate a three-dimensional triangular mesh model with high-quality geometry and realistic texture, as shown in Figure 6 and Figure 7 shown.
[0085] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art, without departing from the scope of the technical solution of the present invention, may make some changes or modifications using the above-disclosed technical content to form equivalent embodiments of equivalent changes. However, as long as it does not depart from the content of the technical solution of the present invention, and according to the technical essence of the present invention, within the spirit and principle of the present invention, any simple modifications, equivalent replacements, and improvements made to the above embodiments still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for three-dimensional scene simulation modeling in the process of game production, characterized in that The specific steps of the method include: Step 1, optimize the preliminary 3D representation with multi-view coupling constraints; Step 2, extract the rough texture grid; Step 3, fine-tune the fine texture grid with multi-view coupling constraints.
2. A three-dimensional scene simulation modeling method during game production according to claim 1, characterized in that Step 1 specifically includes: Randomly initialize the optimizable 3D representation and randomly select four rendering views that are orthogonal to each other in the horizontal direction. Render the 3D representation through a differentiable rendering function to obtain the orthogonal views under the given views; Couple the pre-trained diffusion model with the fine-tuned multi-view diffusion model, calculate the loss between the rendered orthogonal views and the target views under the given text conditions, and optimize the 3D representation through gradient backpropagation; This process is iterated until the optimization converges, and finally a preliminary 3D representation consistent with the text semantics is obtained.
3. A three-dimensional scene simulation modeling method during game production according to claim 1, characterized in that, Step 2 specifically includes: Based on the preliminary 3D representation optimized in Step 1, extract and convert it into a rough 3D voxel representation according to the transparency prediction information of each point in space; At the same time, calculate the signed distance function for each voxel unit to generate a preliminary signed distance field; Initialize the deformable tetrahedral mesh through the signed distance field to generate the geometric mesh for 3D modeling; In terms of texture generation, use spatial position hash encoding and a single-layer fully connected perception network to perform texture learning on the geometric mesh, and combine the optimized preliminary 3D representation to supervise the learning initialization of the encoding network and the MLP, so as to generate a rough 3D mesh representation containing geometric structure and texture information.
4. A three-dimensional scene simulation and modeling method during game production according to claim 1, characterized in that Step 3 specifically includes: Through the multi-view coupling constraint optimization method, finely tune the geometric structure and texture information of the texture grid respectively to ensure that the generated 3D texture grid has high geometric consistency and rich details; After the fine-tuning is completed, use the marching tetrahedra algorithm to extract a high-quality triangular mesh, and convert the texture information on the surface into the UV map of the mesh, and finally generate a fine triangular mesh 3D model with high-fidelity texture.