3D scene generation method and device guided by structure control information, equipment and medium
By generating multi-view wireframe structure diagrams and depth maps, combined with text-guided image generation models, the problems of high time cost, insufficient style uniformity, and homogenization in large-scale architectural 3D scene modeling are solved, realizing structured and controllable 3D scene generation and meeting personalized design needs.
Patent Information
- Application Number
- CN202511136431.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing technologies suffer from high time costs, heavy data processing burdens, high labor costs, insufficient style uniformity, and scene homogenization in the 3D scene scanning and modeling process of large buildings or complex sites, making it difficult to meet complex or personalized design needs.
By constructing a 3D architectural model of the target building structure, wireframe structural diagrams and depth maps from multiple perspectives are generated. Combined with text prompts in the target 3D scene to guide a pre-trained image generation model, multiple target style images from different perspectives are generated and mapped onto the 3D architectural model as texture information, thus achieving full-process embedding of structural control information.
It achieves structured and controllable generation of 3D scenes, ensuring consistency in scene style, breaking through the geometric constraints of traditional generation models, freely adjusting scene style, avoiding scene homogenization, and meeting complex or personalized design needs.
Smart Images

Figure CN120833439A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a 3D scene generation method and device guided by structure control information, equipment and medium. BACKGROUND
[0002] With the rapid development of computer technology, virtual shooting and digital scene generation technology is gradually becoming a key support in ancient style film and television creation. This kind of technology combines high-precision modeling, real-time rendering and real light simulation capabilities, successfully realizes the digital reconstruction of historical scenes and fictional ancient style world, and provides creators with greater visual expression freedom and production efficiency.
[0003] In the prior art, laser scanning is generally used as a high-precision modeling method, and based on laser scanning and photogrammetry, high-fidelity building data is obtained, and standardized and reusable component resources (such as eaves, dougong, arches, carved window frames, etc.) are constructed to form a "digital building block" asset library. These assets can be quickly combined, copied, scaled, and other operations to quickly build 3D scenes of different eras and styles.
[0004] However, this approach has the following problems: 1. In large buildings or complex sites, complete scanning often takes several days or even weeks, and there are problems such as high time cost and heavy data processing burden. In addition, due to the large amount of point cloud data, there are high requirements for subsequent processing equipment and software performance, and data compression and structure simplification are still needed in the actual production process.
[0005] 2. Digital asset modeling initially relies on a large amount of manual involvement, including component re-topology modeling, UV unfolding, material division, etc., with high production cycle and labor cost. Secondly, the splicing of module assets may lead to insufficient style uniformity, such as mixing Song Dynasty garden elements with Ming Dynasty palace components, which may cause visual confusion without strict specifications. In addition, the digital resource library technology is limited by preset models, and asset reuse may lead to scene homogenization, resulting in scenes in different video works being too similar in vision, lacking innovation and unique style, and being difficult to meet complex or personalized design requirements. SUMMARY
[0006] Therefore, the purpose of the present application is to provide a 3D scene generation method and device guided by structure control information to ensure scene style consistency and reduce scene homogenization.
[0007] In a first aspect, a 3D scene generation method and device guided by structure control information is provided, comprising: constructing a 3D building model of a target building structure based on scene generation requirements; generating a multi-view line frame structure diagram and a depth map based on the 3D building model; The multi-view line frame structure diagram and the depth map are taken as the structure control information, and a pre-trained image generation model is guided to generate multiple target style images in different views in combination with a target 3D scene text prompt word. The multiple target style images in different views generated are taken as texture information and mapped to a 3D building model to obtain a target 3D scene model.
[0008] Optionally, constructing the 3D building model of the target building structure based on the scene generation requirement comprises: determining a target building structure based on the scene generation requirement; obtaining a multi-view RGB image sequence of a target building conforming to the target building structure; inference on the multi-view RGB image sequence by using a VGGT model to obtain point cloud data of a target 3D scene; preprocessing the point cloud data of the target 3D scene to obtain a 3D building model conforming to the target building structure.
[0009] Optionally, preprocessing the point cloud data of the target 3D scene to obtain a 3D building model conforming to the target building structure comprises: correcting structural defects in an initial point cloud model; performing lightweight processing on the corrected point cloud model by using a preset lightweight tool to obtain a 3D building model conforming to the target building structure.
[0010] Optionally, the image generation model comprises a text encoder, a variational autoencoder and a noise prediction network U-Net, and a training process of the image generation model comprises: obtaining a 3D scene image training data set; textually labeling each 3D scene image in the 3D scene image training data set to obtain a text description label; inputting the 3D scene image training data set and the corresponding text description label into an initial image generation model for iterative training to update parameters of the initial 3D scene generation model, thereby obtaining an optimal 3D scene generation model; wherein, a low-rank adapter is added to an attention layer of the noise prediction network U-Net during the training, and the parameters of the initial 3D scene generation model are updated through a low-rank adaptation strategy.
[0011] Optionally, generating multiple target style images in different views in combination with a target 3D scene text prompt word by using a multi-view line frame structure diagram and a depth map as structure control information and guiding a pre-trained image generation model comprises: extracting a text vector representation of the target 3D scene text prompt word based on a text encoder of the image generation model; The feature maps of the line frame structure diagram and the depth diagram are extracted based on a pre-trained ControlNet network model. The feature maps of the line frame structure diagram and the depth diagram, and a text vector representation of a target 3D scene text prompt word are injected into a noise prediction network U-Net of an image generation model, to guide a target style image in a generated latent space of a to-be-de-noised feature image to be generated at random to generate an image; The target style image in the latent space is decoded into a real target style image based on a variational autoencoder.
[0012] Optionally, the network structure of the ControlNet network model is composed of a down-sampling part and a zero convolution layer of the noise prediction network U-Net; and the guiding of the target style image in the generated latent space of the to-be-de-noised feature image to be generated at random includes: The feature maps of the line frame structure diagram and the depth diagram, the to-be-de-noised feature image, and the text vector representation of the target 3D scene text prompt word are fused and de-noised in the down-sampling part of the noise prediction network U-Net; After the fusion de-noising processing for a preset number of time steps, the fused image output by the last layer of the down-sampling part of the noise prediction network U-Net is up-sampled and restored to obtain the target style image in the latent space.
[0013] Optionally, the down-sampling fusion processing process is: The to-be-de-noised feature image of each layer of the down-sampling of the noise prediction network U-Net is fused with the feature maps of the line frame structure diagram and the depth diagram extracted in the same layer of the ControlNet network model to obtain a fused image; The text vector representation of the target 3D scene text prompt word is injected into the fused image; wherein the to-be-de-noised feature image of the initial time step of the first layer of the noise prediction network U-Net is generated at random, and the fused image after each time step of fusion processing is taken as the to-be-de-noised feature image of the next time step.
[0014] In a second aspect, a 3D scene generation device guided by structure control information is provided, comprising: A construction unit configured to construct a 3D building model of a target building structure based on a scene generation requirement; A generation unit configured to generate a line frame structure diagram and a depth diagram in multiple perspectives based on the 3D building model; A guided generation unit configured to guide a pre-trained image generation model to generate a plurality of target style images in different perspectives by taking the line frame structure diagram and the depth diagram in multiple perspectives as structure control information, and combining a target 3D scene text prompt word; A mapping unit configured to map the plurality of target style images in different perspectives as texture information to the 3D building model to obtain a target 3D scene model.
[0015] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is configured to store a computer program. The processor is configured to execute the program stored in the memory to implement the method of the first aspect.
[0016] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method of any one of the first aspect.
[0017] The application provides a structure control information guided 3D scene generation method, device, equipment and medium, constructs a 3D building model of target building structure based on scene generation demand;Based on the 3D building model, a wireframe structure diagram and a depth map of multiple perspectives are generated;The wireframe structure diagram and the depth map of multiple perspectives are used as structure control information, and a pre-trained image generation model is guided to generate multiple target style images under different perspectives in combination with target 3D scene text prompt words;The multiple target style images under different perspectives generated are used as texture information and are mapped to the 3D building model to obtain a target 3D scene model.The application breaks through the geometric constraint limitation of the traditional generation model through the full-process embedding of structure control information, realizes the structural controllability of the 3D scene, and then guarantees the consistency of the scene style according to the structural controllable 3D scene model generated.Meanwhile, the generation is guided by the target 3D scene text prompt words, the scene style and element combination can be freely adjusted according to the creation demand, and the limitation of reality and preset template is broken through.The scene homogenization problem caused by asset reuse in the existing digital asset library technology can be effectively solved.
[0018] In order to make the above objectives, characteristics and advantages of the present application more apparent, clear and easy to understand, the following preferred embodiments are specifically described below with reference to the attached drawings. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation to the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0020] Figure 1 A flow chart of a structure control information guided 3D scene generation method provided by an embodiment of the present application is shown; Figure 2An effect diagram of three-dimensional reconstruction using VGGT is shown; Figure 3 A structural schematic diagram of ControlNet acting on Stable Diffusion is shown; Figure 4 An example diagram of a target style image generated based on structural control information and a text prompt word is shown; Figure 5 A schematic diagram of the overall flow of the 3D ancient style scene generation method is shown; Figure 6 A structural schematic diagram of a 3D scene generation device guided by structural control information is shown; Figure 7 A structural schematic diagram of an electronic device is shown. DETAILED DESCRIPTION
[0021] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0022] Considering that the prior art has the following problems: 1. In large buildings or complex sites, complete scanning often takes several days or even weeks, and there are problems such as high time cost and heavy data processing burden. In addition, due to the large amount of point cloud data, there are high requirements for the performance of subsequent processing equipment and software, and data compression and structure simplification are still needed in the actual production process.
[0023] 2. Digital asset modeling in the early stage relies on a large amount of manual participation, including component re-topology modeling, UV unfolding, material division, etc., and the production cycle and labor cost are high. Secondly, the splicing of module assets may lead to insufficient style uniformity, such as mixing Song Dynasty garden elements with Ming Dynasty palace components, which may cause visual confusion without strict specifications. In addition, the digital resource library technology is limited by preset models, and asset reuse may lead to scene homogenization, resulting in scenes in different video works being too similar in vision, lacking innovation and unique style, and being difficult to meet complex or personalized design requirements.
[0024] Based on this, the embodiment of the present application provides a 3D scene generation method and device guided by structure control information, which combines text and structure control, can be applied not only to 3D scene generation in films and games, but also extended to fields such as virtual restoration of cultural heritage, education and popular science, and tourism promotion, for example, generating virtual restoration models of Dunhuang Mogao Grottoes in different historical periods through text instructions, which has both academic value and artistic expression.
[0025] The following is described through embodiments.
[0026] The embodiment of the present application provides a 3D scene generation method guided by structure control information, as shown in the following steps. Figure 1 The steps are as follows: Step S101: Construct a 3D building model of the target building structure based on the scene generation requirements.
[0027] Step S102: Generate a wireframe structure diagram and a depth map based on the 3D building model.
[0028] In a feasible implementation, the generation process of the wireframe structure diagram and the depth map includes the following steps: Step S102A: Camera setting Configure a wireframe material (such as aiWireframe material in Arnold renderer) for the 3D building model generated in the above step. Then, establish a multi-view camera layout (including front view, rear view, left view, right view, top view and bottom view) around the model, and ensure that all cameras are configured as orthogonal projection mode. Adjust the position and direction of each camera so that it can fully cover all key perspectives of the model.
[0029] Step S102B: Wireframe structure rendering Switch to Arnold renderer and activate aiWireframe material in Arnold rendering settings. Set the output format to PNG and the resolution to 512x512 pixels. According to the pre-set camera sequence, render each view in turn and save the generated wireframe image.
[0030] Step S102C: Depth map rendering Add the Z-depth channel in the AOVs (Arbitrary Output Variables) settings of the Arnold renderer, and select the output format as 8-bit PNG. According to the size of the 3D building model, set the Near Clip Plane and Far Clip Plane reasonably to ensure that the 3D building model is not cropped at all during rendering, and to obtain the best depth value precision. Perform the batch rendering process to generate depth maps for each view, and normalize the exported depth maps to single-channel grayscale images in the range of [0, 1]. This step is crucial for subsequent processing or analysis of the depth information of the model.
[0031] Step S103: Use the multi-view wireframe structure diagram and the depth map as structure control information, and combine the target 3D scene text prompt word to guide the pre-trained image generation model to generate multiple target style images under different views.
[0032] Step S104: Map the generated multiple target style images under different views as texture information to the 3D building model to obtain the target 3D scene model.
[0033] In a feasible implementation, for example, in Maya software, camera parameters can be set and camera projection-based UV mapping can be implemented through a Python script (MayaPython API).
[0034] The entire process can be briefly divided into the following steps: Step S104A: Synchronize camera parameters; Ensure that the camera parameters of the current model are consistent with the rendering of the wireframe structure diagram and the depth map. Use the MayaPython API (such as the cmds module) to set the rotation angle, clipping plane distance, focal length, and other parameters of the camera; to ensure the consistency of the view angle.
[0035] Step S104B: Perform camera projection UV mapping; Use camera projection technology to generate UV layouts (U and V represent the horizontal and vertical directions in the texture space) that adapt to each camera view. The specific operations include: creating a camera projection node, connecting the camera with the projection node, selecting the target mesh of the 3D building model (dividing the 3D building model into a 3D building mesh model) for projection to generate UV coordinates under each camera view.
[0036] This process supports loops and parameterized settings, making it easy to quickly generate multi-view UV mapping for complex models.
[0037] Step S104C: Apply UDIM texture mapping technology; The UDIM (U-Dimension) texture mapping technology is used to realize efficient mapping of multi-view textures. Each camera view corresponds to an independent UDIM area, and the corresponding texture image is named according to the UDIM number. By associating each camera view with the corresponding UDIM number, accurate matching of three-dimensional space structure and two-dimensional image content is realized.
[0038] Step S104D: final rendering and quality optimization; In the Arnold renderer, combined with fine adjustment of light and shadow settings, high-quality ancient style scene 3D model rendering is completed. In this way, not only the three-dimensional structure of the model is accurate, but also the unique ancient style aesthetic characteristics are given, and the works with artistic beauty and technical precision are presented.
[0039] The embodiment of the present application can freely adjust the scene style and element combination according to the creation requirements, such as combining the Tang Dynasty palace with the Xianxia element, or generating the specific dynasty's street scene based on historical research, breaking through the limitations of reality and preset templates, and avoiding the problems of "excessive realism" of laser scanning and "homogenization" of digital resource library.
[0040] On the basis of the above embodiment, the 3D building model of the target building structure is constructed based on the scene generation requirement, and the 3D building model of the target building structure comprises: Step S101A: determining a target building structure based on a scene generation requirement.
[0041] In this step, the scene generation requirement includes building style and scene style, etc., the building style is, for example, Tang Dynasty building style, Song Dynasty building style, etc., and the scene style is, for example, Xianxia scene, etc.
[0042] Step S101B: obtaining a multi-view RGB image sequence of a target building conforming to the target building structure.
[0043] In this step, the multi-view RGB image sequence of the Tang Dynasty style building can be collected by an image collection device, such as a high-precision camera, wherein the multi-view RGB image sequence is obtained by shooting from different angles by the same fixed camera.
[0044] Step S101C: using a VGGT (Visual Geometry Grounded Transformer, visual model based on pure feedforward Transformer architecture) model to infer the multi-view RGB image sequence to obtain point cloud data of a target 3D scene.
[0045] In this step, VGGT is a general 3D vision model based on a pure feedforward Transformer architecture. Core geometric information such as camera parameters, depth maps, point clouds, and 3D point trajectories can be directly inferred. Moreover, all of this is done in one forward inference without any post-processing optimization; compared to traditional 3D reconstruction which relies on optimization methods such as beam adjustment, the calculation is complex, time-consuming, and requires repeated iterations. VGGT directly discards this cumbersome process and uses a pure feedforward design to complete all geometric inference tasks in one forward propagation, greatly improving modeling efficiency.
[0046] After obtaining the multi-view RGB image sequence, first, the RGB image sequence is uniformly preprocessed and aligned.
[0047] Then input it into the VGGT model to obtain the camera intrinsic parameters, extrinsic parameters, depth map, and dense point cloud of the scene, as shown in Figure 2 , and the dense point cloud of the scene is taken as the initial point cloud model of the scene.
[0048] Step S101D: Preprocessing the point cloud data of the target 3D scene to obtain a 3D building model conforming to the target building structure.
[0049] In one possible implementation, preprocessing the initial point cloud model to obtain a 3D building model conforming to the target building structure includes: Step S101D1: Correcting structural defects in the initial point cloud model.
[0050] In the embodiments of the present application, structural defects include topological errors (such as non-manifold geometry, overlapping surfaces), missing parts, or unreasonable structures.
[0051] In one example, for example, for overlapping surfaces or unreasonable structures, the editing tools of the modeling software can be used to adjust the positions of vertices, edges, and surfaces, correct the shape of the model, and make it more conform to the structural characteristics of the real 3D scene.
[0052] For complex structural parts such as corbel arches and carvings, subdivision surface technology can be used to increase the detail accuracy of the model. For example, Autodesk Maya and 3ds Max modeling software both provide subdivision surface options "Subdivision Surfaces", allowing users to refine the model surface by increasing the number of polygons to obtain smoother results.
[0053] Step S101D2: Using a pre-set lightweight tool to perform lightweight processing on the corrected point cloud model to obtain a 3D building model conforming to the target building structure.
[0054] In one example, the model is automatically lightened using the "Mesh>Reduce" function of Autodesk Maya modeling software.
[0055] In this step, the adjusted model is optimized, such as reducing the number of faces and vertices, to reduce the complexity of the model and make it lightweight, while ensuring the clarity of the model's geometric structure, to meet the needs of real-time rendering or further processing.
[0056] Based on the above embodiment, the image generation model can use the Stable Diffusion model, including a text encoder, a variational autoencoder (VEA), and a noise prediction network U-Net.
[0057] The text encoder uses the CLIP model, which can convert input text information into a vector representation (embedding). This embedding vector is used to capture the semantic information of the text and guide the U-Net network during image generation.
[0058] In Stable Diffusion, VAE is mainly used for image encoding and decoding. The encoder part of VAE compresses the original image into a lower-dimensional representation (latent space representation), while the decoder part can reconstruct the image from the low-dimensional representation. In the embodiment of the present application, VAE is used to map images to the latent space of the noise prediction network U-Net and from the latent space to images, rather than directly generating images.
[0059] The noise prediction network U-Net is one of the core components of the Stable Diffusion model, responsible for learning how to gradually remove the noise added to the image, thereby gradually generating the target image. The U-Net architecture contains multiple down-sampling and up-sampling paths, allowing the model to capture features at different scales. During the training phase, the model receives a pair of noisy images and corresponding clean images as input, learning to predict the noise added to the image. In the generation phase, the noise is gradually reduced through the reverse process, and finally a clear image is generated.
[0060] Based on the network structure of the image generation model, the training process of the image generation model includes: Step 1: Obtain a 3D scene image training dataset.
[0061] In this step, images can be obtained through various channels, such as collecting real 3D scene images from historical documents, museum collections, and ancient architecture photography works; and collecting fictional 3D scene images from films, games, and illustrations.
[0062] And the collected 3D scene image is preprocessed, for example, the images that are blurred, repeated, or do not meet the requirements are removed, and the images are uniformly converted and sized (such as 512*512 pixels).
[0063] Second step: text annotation is performed on each 3D scene image in the 3D scene image training data set, and a text description label is obtained.
[0064] In this step, for each image, an accurate and detailed text description is written. The description includes information such as the era, regional style, building structure, and scene atmosphere of the scene.
[0065] In one example, the efficiency and accuracy of annotation can be improved by combining manual annotation with AI-assisted annotation. For example, the open-source Florence-2 model AI can be used to assist in generating basic descriptions.
[0066] Or manually annotate using tools such as Label Studio and add the Ancientry label at the top. Organize the images and corresponding text descriptions into a CSV or JSON format dataset for easy use later.
[0067] Third step: input the 3D scene image training data set and its corresponding text description label into the initial image generation model for iterative training to update the parameters of the initial 3D scene generation model and obtain the optimal 3D scene generation model.
[0068] During training, a low-rank adapter is added to the attention layer of the noise prediction network U-Net to update the parameters of the initial 3D scene generation model through a low-rank adaptation strategy.
[0069] The update formula of the low-rank adaptation strategy LoRA (Low-Rank Adaptation) for the original weight matrix W of the 3D scene generation model is as follows: (1); In the formula, W is the parameter of the initial 3D scene generation model, is the final initial 3D scene generation model parameter affected by the LoRA model. A and B are low-rank matrices, and BA represents the parameters of the LoRA model. r is the parameter of the low-rank matrix, which controls the rank of matrix decomposition; α is the scaling factor, which balances the contribution of the original weight and the LoRA weight. α / r as a whole scaling factor determines the degree of modification of the original weight by the LoRA module.
[0070] When a = r, the scaling factor is 1, and the LoRA weight is directly superimposed on the original weight; increasing a will enhance the influence of the LoRA module, making the model adapt to new data faster, but may cause overfitting; reducing a will weaken the influence of the LoRA module and retain more knowledge of the original model. By reasonably setting the LoRA parameters, the effect close to full-parameter fine-tuning can be achieved while maintaining efficient training.
[0071] Therefore, at the beginning of training, first, the Stable Diffusion model is fine-tuned, the LoRA parameters are configured, the low-rank matrix dimension r = 8, the scaling factor a = 8 for balancing the original weight and the LoRA weight, and the training parameters such as the number of training rounds, the batch size, and the learning rate are configured.
[0072] The constructed 3D scene dataset is divided into a training set and a validation set according to an 8:2 ratio.
[0073] The core components of the Stable Diffusion model are loaded: the image encoder VAE, the text encoder CLIP, and the noise prediction network U-Net, wherein only the attention layer (Q / K / V projection matrix) of the U-Net is added with a low-rank adapter, and other layers are frozen to reduce trainable parameters.
[0074] The images and text description labels in the 3D scene image training dataset are input into the Stable Diffusion model, and the parameters of specific layers of the model are adjusted through the LoRA module, so that the Stable Diffusion model learns the features and styles of the 3D scene.
[0075] Every certain number of steps, the quality of the generated images is evaluated on the validation set. Specifically, the quality of the generated images can be evaluated by the mean square error loss function and the similarity index between the generated images and the real images.
[0076] Then, according to the evaluation results, the parameters of the low-rank matrices A and B are updated in a backward propagation manner, while the remaining weights of the pre-trained model remain unchanged. During the training process, checkpoints should be saved regularly.
[0077] Finally, the effects of different checkpoints are compared, and the models that are underfitting or overfitting are excluded, and the LoRA weight file with the best visual quality is selected. The LoRA weight file is combined with the basic Stable Diffusion model to obtain an image generation model.
[0078] On the basis of the above embodiment, a wireframe structure diagram and a depth map in multiple perspectives are used as structure control information, and a target 3D scene text prompt word is used to guide a pre-trained image generation model to generate multiple target style images in different perspectives, which includes the following steps. Step S103A: A text encoder based on an image generation model extracts a text vector representation of the target 3D scene text prompt.
[0079] Step S103B: A feature map of the line frame structure diagram and the depth map is extracted based on a pre-trained ControlNet network model.
[0080] In an embodiment of the present application, the ControlNet network model is a neural network structure for enhancing Stable Diffusion, which constructs two control branches to guide the Stable Diffusion model. One ControlNet branch is dedicated to processing the line frame diagram to extract the contour structure, and the other ControlNet branch processes the depth map to provide spatial and perspective guidance. Both are fused into each layer of the noise prediction network U-Net through residual addition, strengthening the structural consistency.
[0081] Step S103C: The feature maps of the line frame structure diagram and the depth map, and the text vector representation of the target 3D scene text prompt are injected into the noise prediction network U-Net of the image generation model to guide the generation of the target style image in the random generated denoising feature image generation latent space.
[0082] In an embodiment of the present application, the network structure of the ControlNet network model is composed of the down-sampling part of the noise prediction network U-Net and the zero convolution layer. As shown in Figure 3 The encoding part of the ControlNet network model is consistent with the structure of the down-sampling part of the noise prediction network U-Net, which ensures that the generation ability of the original model is not damaged, and at the same time, new spatial constraint conditions can be flexibly introduced. This design allows ControlNet to efficiently introduce subsequent structural information constraint conditions, thereby achieving fine control of the image generation process while maintaining the stability of the original Stable Diffusion model.
[0083] On this basis, guiding the generation of the target style image in the random generated denoising feature image generation latent space includes: Step S103C1: The feature maps of the line frame structure diagram and the depth map, the denoising feature image, and the text vector representation of the target 3D scene text prompt are fused and denoised in the down-sampling part of the noise prediction network U-Net.
[0084] In a specific embodiment, the down-sampling fusion process is: Step A: The to-be-de-noised feature image of each layer of the down-sampling part of the noise prediction network U-Net is fused with the feature image of the line structure diagram and the depth map extracted from the same layer of the ControlNet network model to obtain a fused image.
[0085] The to-be-de-noised feature image of the initial time step of the first layer of the noise prediction network U-Net is randomly generated, and the fused image after each time step of fusion processing is taken as the to-be-de-noised feature image of the next time step.
[0086] In this step, the two feature image tensors of the line structure diagram and the depth map are fused by element-wise addition, and the fusion formula is as follows: (2); Wherein, is the to-be-de-noised feature image of the noise prediction network U-Net; is the feature image of the line structure diagram; is the feature image of the depth map; weight is the weight, and the influence strength of the ControlNet network model can be controlled by adjusting the value of to ensure that the style of the image comes from the text prompt word, and the structure comes from the structure control diagram of the ControlNet network model.
[0087] Step B: Inject the text vector representation of the target 3D scene text prompt word into the fused image.
[0088] On the basis of structure control, the input target 3D scene text prompt word plays a decisive role in guiding the model to generate 3D scene content with classical charm.
[0089] Step S103C2: After the fusion de-noising processing of the preset number of time steps, the fused image output by the last layer of the down-sampling part of the noise prediction network U-Net is up-sampled and restored to obtain the target style image in the latent space.
[0090] Step S103D: Based on the variational autoencoder, the target style image in the latent space is decoded into a real target style image.
[0091] As shown in Figure 4 , based on the fine-tuned Stable Diffusion model and ControlNet network model, the edge line structure diagram, depth condition diagram and text prompt word are input at the same time, and for the multi-view paired edge line structure diagram-depth map, after multiple inferences, a plurality of style images under different perspectives are obtained to obtain the target style image.
[0092] The core innovation of the application is to realize structure guidance through the double-branch structure (wireframe contour + depth perspective) of the ControlNet network model, and the residual fusion mechanism strictly follows the three-dimensional structure constraint while the text semantic dominates the style generation. Finally, through the camera projection UV mapping technology, the multi-view style images are accurately texture mapped according to the structure coordinates, forming a 3D scene with geometric accuracy and cultural charm. This method breaks through the geometric constraint limitation of traditional generation models through the whole-process embedding of structure information, realizes the structured and controllable generation of 3D scenes, and effectively solves the scene homogenization problem caused by asset reuse in the existing digital asset library technology.
[0093] In order to make the method of the embodiment of the application clearer, taking the generation of an ancient style scene as an example, as shown in the figure, a flowchart for generating the ancient style scene is given. Figure 5
[0094] Based on the same inventive concept, the embodiment of the application provides a 3D scene generation device guided by structure control information, as shown in the figure, the device comprises: Figure 6 The construction unit 601 is configured to construct a 3D building model of a target building structure based on scene generation requirements. The generation unit 602 is configured to generate a wireframe structure graph and a depth graph of multiple views based on the 3D building model. The guided generation unit 603 is configured to use the wireframe structure graph and the depth graph of multiple views as structure control information, and combine a target 3D scene text prompt word to guide a pre-trained image generation model to generate a plurality of target style images under different views. The mapping unit 604 is configured to map the plurality of target style images under different views as texture information to the 3D building model to obtain a target 3D scene model.
[0095] Based on the same technical concept, the embodiment of the application further provides an electronic device, as shown in the figure, comprising a processor 701, a communication interface 702, a memory 703 and a communication bus 704, wherein the processor 701, the communication interface 702 and the memory 703 communicate with each other through the communication bus 704. Figure 7 The memory 703 is configured to store a computer program.
[0096] The processor 701 is configured to execute the program stored in the memory 703 to realize the steps of the 3D scene generation method guided by structure control information.
[0097] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0098] The communication interface is used for communication between the above electronic device and other devices.
[0099] The memory can include a Random Access Memory (RAM) and can also include a Non-Volatile Memory (NVM), for example, at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.
[0100] The processor mentioned above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; can also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0101] The computer program product for generating a 3D scene by guiding structure control information provided by the embodiment of the application includes a computer readable storage medium storing program codes, the program codes include instructions for executing the method described in the foregoing method embodiments, and the specific implementation can be referred to the method embodiments, which will not be described here.
[0102] The apparatus for generating a 3D scene guided by structure control information provided by the embodiments of the present application can be specific hardware on a device or software or firmware installed on the device, etc. The apparatus provided by the embodiments of the present application has the same implementation principle and technical effects as the foregoing method embodiments, and for brief description, the apparatus embodiment part is not mentioned in the foregoing method embodiment. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the foregoing described system, apparatus and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0103] In the embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. The embodiments of the apparatus described above are merely schematic; for example, the division of the units is only a logical function division, and there can be another division manner in actual implementation; for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, and electrical, mechanical or other forms.
[0104] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0105] In addition, each functional unit in the embodiments provided by the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0106] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part of the technical solutions that make essential contributions to the prior art can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0107] It should be noted that similar reference numbers and letters refer to similar items in the following drawings, and therefore, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.
[0108] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present application, and are used to illustrate the technical solutions of the present application, but are not limiting. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some of the technical features, without departing from the technical scope disclosed by the present application. These modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A 3D scene generation method guided by structural control information, characterized in that: The method comprises the following steps: constructing a 3D building model of a target building structure based on a scene generation requirement; generating a multi-view line frame structure diagram and a depth map based on the 3D building model; using the multi-view line frame structure diagram and the depth map as structure control information, combining a target 3D scene text prompt word to guide a pre-trained image generation model to generate a plurality of target style images under different perspectives; mapping the plurality of target style images under different perspectives generated as texture information to the 3D building model to obtain a target 3D scene model.
2. The method of claim 1, wherein, The method of constructing a 3D building model of a target building structure based on a scene generation requirement comprises the following steps: determining a target building structure based on a scene generation requirement; obtaining a multi-view RGB image sequence of a target building conforming to the target building structure; using a VGGT model to infer the multi-view RGB image sequence to obtain point cloud data of a target 3D scene; preprocessing the point cloud data of the target 3D scene to obtain a 3D building model conforming to the target building structure.
3. The method of claim 2, wherein, The method of preprocessing the point cloud data of the target 3D scene to obtain a 3D building model conforming to the target building structure comprises the following steps: correcting structural defects in the initial point cloud model; using a pre-set lightweight tool to perform lightweight processing on the corrected point cloud model to obtain a 3D building model conforming to the target building structure.
4. The method of claim 1, wherein, The image generation model comprises a text encoder, a variational autoencoder and a noise prediction network U-Net, and the training process of the image generation model comprises the following steps: obtaining a 3D scene image training data set; textually labeling each 3D scene image in the 3D scene image training data set to obtain a text description label; inputting the 3D scene image training data set and the corresponding text description label into an initial image generation model for iterative training to update the parameters of the initial 3D scene generation model and obtain an optimal 3D scene generation model; wherein, a low-rank adapter is added to the attention layer of the noise prediction network U-Net during training, and the parameters of the initial 3D scene generation model are updated through a low-rank adaptation strategy.
5. The method of claim 4, wherein, The method of using the multi-view line frame structure diagram and the depth map as structure control information, combining a target 3D scene text prompt word to guide a pre-trained image generation model to generate a plurality of target style images under different perspectives comprises the following steps: extracting a text vector representation of the target 3D scene text prompt word based on the text encoder of the image generation model; extracting feature maps of the line frame structure diagram and the depth map based on a pre-trained ControlNet network model; injecting the feature maps of the line frame structure diagram and the depth map and the text vector representation of the target 3D scene text prompt word into the noise prediction network U-Net of the image generation model to guide a target style image in a randomly generated de-noising feature image generation latent space; decoding the target style image in the latent space into a real target style image based on the variational autoencoder.
6. The method of claim 5, wherein, The network structure of the ControlNet network model is composed of a down-sampling part of the noise prediction network U-Net and a zero convolution layer; the target style image in the latent space is generated by guiding the randomly generated to-be-de-noised feature image to generate the target style image in the latent space, including: The feature maps of the line frame structure and the depth map, the to-be-de-noised feature image, and the text vector representation of the target 3D scene text prompt word are fused and de-noised in the down-sampling part of the noise prediction network U-Net; After the fusion de-noising processing of the preset time steps, the fused image output by the last layer of the down-sampling part of the noise prediction network U-Net is up-sampled and restored to obtain the target style image in the latent space.
7. The method of claim 6, wherein, The down-sampling fusion processing process is: The to-be-de-noised feature image of each layer of the down-sampling part of the noise prediction network U-Net is fused with the feature maps of the line frame structure and the depth map extracted from the same layer in the ControlNet network model to obtain a fused image; The text vector representation of the target 3D scene text prompt word is injected into the fused image; wherein the to-be-de-noised feature image of the initial time step of the first layer of the noise prediction network U-Net is randomly generated, and the fused image after each time step fusion processing is used as the to-be-de-noised feature image of the next time step.
8. A 3D scene generation apparatus of structure control information guidance, characterized by, Including: A construction unit configured to construct a 3D building model of a target building structure based on a scene generation requirement; A generation unit configured to generate a line frame structure and a depth map in multiple perspectives based on the 3D building model; A guided generation unit configured to guide a pre-trained image generation model to generate target style images in multiple perspectives by taking the line frame structure and the depth map in multiple perspectives as structure control information and combining a target 3D scene text prompt word; A mapping unit configured to map the generated target style images in multiple perspectives as texture information to the 3D building model to obtain a target 3D scene model.
9. An electronic device, comprising: A processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store a computer program; The processor is used to execute the program stored in the memory to implement the method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-7.
Citation Information
Patent Citations
Text-guided neural radiation field building scene stylization method
CN117541732A
Dynamic Gaussian scene reconstruction method based on depth regularization
CN119991974A
Single building three-dimensional reconstruction method based on point cloud semantic segmentation and structure fitting
WO2024077812A1
Cited By
Interactive multi-parameter fabric texture generation method and system
CN121353287A
Multi-model collaborative two-dimensional Gaussian splash three-dimensional reconstruction method
CN121392157A
A two-dimensional gaussian splashing three-dimensional reconstruction method based on multi-model cooperation
CN121392157B