Indoor three-dimensional scene controllable generation model training and application method, equipment and medium
By using neural radiation fields in the adversarial generation network framework and using Internet image data without labeling information to train the generation model, the problem of difficulty in generating high-reality indoor three-dimensional scenes in the prior art is solved, and efficient and controllable indoor three-dimensional scene generation is achieved.
Patent Information
- Application Number
- CN202510152150.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The prior art is difficult to effectively generate high-reality indoor three-dimensional scenes, especially in massive Internet image data without labeling information, and it is impossible to generate indoor three-dimensional scenes controllable according to user needs.
Adopting an adversarial generation network framework based on neural radiation field, an indoor three-dimensional scene controllable generation model supporting multiple conditional control is obtained through indoor scene image data training with unlabeled information collected from the Internet to generate indoor three-dimensional scenes.
It realizes efficient and controllable generation of high-realistic indoor three-dimensional scenes without any additional labeling information, reducing the requirements for training data quality and scale, and improving the generation efficiency and quality.
Smart Images

Figure CN119625462B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of controllable generation of indoor three-dimensional scenes, and particularly to a method, device, and medium for training and applying a controllable generation model of indoor three-dimensional scenes. Background Art
[0002] In recent years, people's demand for three-dimensional digital content (which refers to three-dimensional digital models created by three-dimensional modeling software for design and subsequent processing work) has been increasing. Generating a large number of high-quality three-dimensional digital content is the key to many applications, such as video games, movie production, virtual reality, and augmented reality. As the indoor space is the main place for human life, the generation of indoor three-dimensional scene data is an indispensable technology in the generation of three-dimensional digital content. With the development of industries such as interior design, smart home, indoor robots, and games, the need for high-fidelity indoor three-dimensional scene data is also increasing. For example, most of the latest methods in research fields related to artificial intelligence algorithms and embodied intelligence in indoor scenes are based on data-driven deep learning algorithms. More indoor three-dimensional scene data can significantly improve the performance of deep learning algorithms, especially a large amount of high-fidelity indoor three-dimensional scene data, which is crucial for enhancing the understanding of indoor spaces in the real world by intelligent agents such as indoor robots.
[0003] There are two traditional ways to generate indoor three-dimensional scenes. One is for professional modelers and interior designers to manually design, model, and construct highly complex indoor three-dimensional scenes containing a large number of furniture and decorative items. The other is to use scanning devices to scan real indoor scenes and use three-dimensional reconstruction algorithms to perform three-dimensional reconstruction on the scanned single real indoor scene to obtain the indoor three-dimensional scene of the real indoor scene. However, these two ways of generating indoor three-dimensional scenes require professional modeling design software or scanning devices, as well as certain professional knowledge, with a relatively high technical threshold, and are time-consuming and laborious. The cost of collecting a large amount of rich and diverse indoor three-dimensional scene data is extremely high.
[0004] In recent years, a large number of researchers have attempted to use artificial intelligence technology to automatically generate indoor three-dimensional scene data. Some researchers use deep learning algorithms to learn the data distribution of indoor three-dimensional scenes from a small existing indoor three-dimensional scene dataset (about tens of thousands of indoor three-dimensional scenes), so as to automatically generate new indoor three-dimensional scenes. However, this method is limited by the data quality and scale of the existing indoor three-dimensional scene dataset, with limited performance and generalization ability, and it is difficult to generate high-fidelity indoor three-dimensional scenes similar to those in the real world. Although high-fidelity indoor three-dimensional scene datasets (in the tens of thousands) are very scarce, two-dimensional image data (in the hundreds of millions) containing a variety of indoor scenes in the real world are very rich and easy to obtain. Therefore, some researchers have tried to learn information about the three-dimensional world from a large number of two-dimensional image data from different indoor scenes and generate indoor three-dimensional scenes. However, this method has high requirements for the image dataset, requires the shooting perspectives or their distributions of the images in the known image dataset, or requires multi-perspective image data of indoor scenes, and cannot be applied to the massive Internet image data without any annotation information, nor can it generate indoor three-dimensional scenes controllably according to user needs. Summary of the Invention
[0005] The purpose of this application is to provide a method, device, and medium for training and applying a controllable indoor three-dimensional scene generation model, which can be used to train a controllable indoor three-dimensional scene generation model that supports multiple condition controls based on a large amount of indoor scene-related image data collected directly from the Internet without any annotation information, and further generate indoor three-dimensional scenes.
[0006] To achieve the above purpose, this application provides the following solutions.
[0007] In the first aspect, this application provides a method for training a controllable indoor three-dimensional scene generation model, and the method for training the controllable indoor three-dimensional scene generation model includes the following steps.
[0008] Obtain training control conditions and multiple indoor scene images without annotation information, and use the indoor scene images without annotation information as training images; the training control conditions include training scene images and training room wall boundaries.
[0009] Respectively use the training scene image and multiple training images as inputs, and use the camera parameter predictor to determine the first control camera parameters corresponding to the training scene image and the first training camera parameters corresponding to each training image.
[0010] Use the training control conditions and noise that meets a preset distribution as inputs, and use the three-dimensional scene generator to generate the first training neural radiance field of the indoor three-dimensional scene.
[0011] Render a first control synthetic image based on the first training neural radiance field and the first control camera parameters; render a second control synthetic image based on the first training neural radiance field and the second control camera parameters; render a first training synthetic image based on the first training neural radiance field and all the first training camera parameters; the first control synthetic image, the second control synthetic image, and the first training synthetic image all include a color image and a depth image.
[0012] Calculate the depth image of the training room wall boundary based on the training room wall boundary and the second control camera parameters.
[0013] Use a two-dimensional image discriminator to discriminate between the training images and the first training synthetic image to obtain a discrimination result, and calculate a discrimination loss according to the discrimination result.
[0014] Calculate a scene image conditional control consistency loss according to the similarity between the training scene image and the first control synthetic image.
[0015] Calculate a room wall boundary conditional control consistency loss according to the difference between the depth image of the training room wall boundary and the depth image of the second control synthetic image.
[0016] Update the three-dimensional scene generator and the two-dimensional image discriminator according to the discrimination loss, the scene image conditional control consistency loss, and the room wall boundary conditional control consistency loss to obtain an updated generator and an updated discriminator.
[0017] Determine whether the iteration termination condition is reached.
[0018] If so, stop the iteration and use the updated generator as the indoor three-dimensional scene controllable generation model.
[0019] If not, continue the iteration, use the updated generator as the three-dimensional scene generator for the next iteration, use the updated discriminator as the two-dimensional image discriminator for the next iteration, and return to the step of "obtaining the training control conditions and multiple unlabeled indoor scene images".
[0020] In a second aspect, the present application provides a method for applying an indoor three-dimensional scene controllable generation model, and the method for applying the indoor three-dimensional scene controllable generation model includes the following steps.
[0021] Obtain the control conditions input by the user; the control conditions are scene images or room wall boundaries.
[0022] Using the control conditions and noise that meets the preset distribution as inputs, a neural radiance field of an indoor three-dimensional scene is generated by using a controllable generation model of an indoor three-dimensional scene; the controllable generation model of the indoor three-dimensional scene is trained by using the above-mentioned training method of the controllable generation model of the indoor three-dimensional scene.
[0023] Based on the neural radiance field, an indoor three-dimensional scene is generated.
[0024] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned training method of the controllable generation model of the indoor three-dimensional scene or the above-mentioned application method of the controllable generation model of the indoor three-dimensional scene.
[0025] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned training method of the controllable generation model of the indoor three-dimensional scene or the above-mentioned application method of the controllable generation model of the indoor three-dimensional scene.
[0026] According to the specific embodiments provided by the present application, the present application has the following technical effects.
[0027] The present application provides a method, device, and medium for training and applying a controllable generation model of an indoor three-dimensional scene. The first training neural radiance field of the indoor three-dimensional scene is generated by using a three-dimensional scene generator with control conditions and noise that meets the preset distribution as inputs. Based on the first training neural radiance field, the first control composite image, the second control composite image, and the first training composite image are rendered. Further, the discriminant loss, the scene image conditional control consistency loss, and the room wall boundary conditional control consistency loss are calculated to update the three-dimensional scene generator and the two-dimensional image discriminator, obtaining the updated generator and the updated discriminator, and continuously iterating until the iteration termination condition is reached, obtaining a controllable generation model of the indoor three-dimensional scene that supports multiple condition controls. By designing an adversarial generation network framework, the present application can train a controllable generation model of the indoor three-dimensional scene that supports multiple condition controls based on a large amount of image data related to indoor scenes without any annotation information directly collected from the Internet, further generating an indoor three-dimensional scene. The requirements for the image dataset are not high, avoiding the problem of affecting the model accuracy due to the small content of the image dataset, and can generate indoor three-dimensional scenes with high realism. Description of the Drawings
[0028] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0029] Figure 1 Schematic diagram of the generation process of the existing HouseCrafter provided in Embodiment 1 of the present application.
[0030] Figure 2 Schematic diagram of the adversarial generation network framework for controllable generation of room-level indoor three-dimensional scenes provided in Embodiment 1 of the present application.
[0031] Figure 3 Schematic diagram of the controllable generation of residential-level indoor three-dimensional scenes provided in Embodiment 2 of the present application.
[0032] Figure 4 Schematic diagram of the structure of a computer device provided in Embodiment 3 of the present application. Detailed implementation manners
[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0034] Embodiment 1.
[0035] The existing method of learning information about the three-dimensional world from a large amount of two-dimensional image data from different indoor scenes and generating indoor three-dimensional scenes can be: HouseCrafter, which can generate residential-level indoor three-dimensional scenes. The generation process of HouseCrafter is as Figure 1As shown, it is necessary to provide the floor plans of all rooms and the layout diagrams of the furniture in the rooms as inputs. According to the given scene layout diagrams (i.e., the floor plans of all rooms and the layout diagrams of the furniture in the rooms), first sample a sequence of camera viewpoints in the scene, and then generate the color images and depth images for each camera viewpoint one by one in an autoregressive manner. Each time a new color image and depth image for a camera viewpoint are generated, the color images and depth images for the adjacent (i.e., neighboring) camera viewpoints that have already been generated are used as references. Finally, the color images and depth images for all camera viewpoints that are generated are fused through the Truncated Signed Distance Function (TSDF) technology and an indoor three-dimensional scene represented by a mesh is reconstructed.
[0036] The core technology of HouseCrafter is to generate the color images and depth images for a new camera viewpoint guided by the scene layout diagram and the color images and depth images for the adjacent camera viewpoints that have already been generated. To this end, HouseCrafter designs a diffusion model guided by the scene layout diagram and the reference images (i.e., the color images and depth images for the adjacent camera viewpoints that have already been generated), and uses the cross-attention mechanism to achieve conditional control of the scene layout diagram and the reference images, supporting the simultaneous generation of color images and depth images. Therefore, HouseCrafter needs to train this diffusion model on approximately 2000 scenes in the artificially designed indoor scene dataset 3D-FRONT. Specifically, the scene layout diagram is obtained from the three-dimensional meshes of the furniture objects and the house walls in the scene. Camera viewpoints are sampled in each scene and a total of 2 million images are rendered for the training of the diffusion model.
[0037] HouseCrafter has certain drawbacks: (1) This method has relatively strict requirements for the training data. It requires the scene layout diagrams of all residential-level indoor three-dimensional scenes in the training data and the sequences of color images and depth images under a dense camera viewpoint sequence in the scene. It is difficult to obtain large-scale indoor three-dimensional scene data that meet the conditions, which will affect the generation quality and generalization ability of the diffusion model; (2) This method uses an autoregressive manner to generate the color images and depth images for different camera viewpoints, and cannot guarantee the consistency of the content in the color images and depth images for different camera viewpoints (for example, the appearance and shape of the same object in adjacent images are inconsistent), resulting in a poor quality of the indoor three-dimensional scene reconstructed by fusion (such as Figure 1The reconstructed three-dimensional scene represented by the generated grid shown on the right is relatively blurred and unclean; (3) The generation efficiency of this method is low. It is necessary to autoregressively generate color images and depth images from one camera view after another, and fuse and reconstruct the color images and depth images from all camera views to obtain the final indoor three-dimensional scene, resulting in a large computational overhead and time consumption.
[0038] Obviously, the above method has high requirements for the image dataset. It is necessary to know the shooting perspectives or their distributions of the images in the image dataset, or multi-view image data of the indoor scene, and it cannot be applied to the massive Internet image data without any annotation information. Moreover, the generation quality and efficiency of the above method are both low.
[0039] To solve the above problems, this embodiment designs an indoor three-dimensional scene generation method based on neural radiance fields. It does not require indoor three-dimensional scene data for training and can directly learn the information of the indoor three-dimensional scene from the massive indoor scene-related image data collected from the Internet without any additional annotation information. It can generate a large number of diverse room-level indoor three-dimensional scenes, thus having low requirements for training data and not affecting the generation quality and generalization ability of the model due to the difficulty of obtaining large-scale qualified indoor three-dimensional scene data. It can generate high-fidelity indoor three-dimensional scenes. At the same time, by generating a neural radiance field, this method can directly generate an indoor three-dimensional scene subsequently, without the need to generate images from multiple camera views one by one and then perform fusion reconstruction. It can ensure the consistency of the content in the color images and depth images from different camera views, and the computational overhead and time consumption are both small.
[0040] This embodiment designs a high-fidelity room-level indoor three-dimensional scene generation scheme driven by Internet image sets. To use Internet images (such as photos) of indoor scenes as training data to generate indoor three-dimensional scenes, this embodiment adopts an adversarial generation network framework for three-dimensional perception based on neural radiance fields. This adversarial generation network framework includes a three-dimensional scene generator based on neural radiance fields, a differentiable volume renderer, and a two-dimensional image discriminator. For the neural radiance field of the indoor three-dimensional scene generated by the three-dimensional scene generator, that is, using the neural radiance field to represent the indoor three-dimensional scene, the volume renderer is used to render the generated neural radiance field into a two-dimensional synthetic image from a specified camera view, and then the two-dimensional image discriminator is used. By using the discriminative loss function, it learns the distribution of the indoor scene from the training images (i.e., Internet images), provides an optimization gradient direction for the rendered two-dimensional synthetic image, and realizes the backpropagation of the gradient through the volume renderer, and backpropagates the optimization gradient back to the three-dimensional scene generator, so as to realize using Internet images as training data to update the three-dimensional scene generator based on neural radiance fields.
[0041] To ensure that the two-dimensional image discriminator can correctly learn the distribution of indoor scenes in the training images, it is necessary to ensure that the distribution of camera viewpoints used when rendering synthetic images from the generated neural radiance field is consistent with the distribution of camera viewpoints in the training images, so as to prevent the three-dimensional scene generator from learning a biased indoor three-dimensional scene. However, since the images from the Internet have no additional annotation information and the camera viewpoints are unknown, in order to apply the adversarial generation network framework for three-dimensional perception based on neural radiance fields to Internet images, this embodiment adds an online-trained camera parameter predictor, which is trained on the generated neural radiance field and applied to the training images to predict the camera parameters of the training images, thereby guiding the sampling and rendering of camera viewpoints in the generated neural radiance field, enabling the three-dimensional scene generator to learn the correct indoor three-dimensional scene.
[0042] The training phase and generation phase of the entire adversarial generation network framework are as Figure 2 shown. The adversarial generation network framework contains three trainable modules: a three-dimensional scene generator, a two-dimensional image discriminator, and a camera parameter predictor. The adversarial generation network framework also contains two camera samplers: a real camera sampler and a random camera sampler. The adversarial generation network framework also contains a voxel renderer.
[0043] In each iteration of the training phase, the three-dimensional scene generator inputs a latent code (i.e., Gaussian noise) randomly sampled from a specified distribution (such as a Gaussian distribution) and outputs a neural radiance field of a room-level three-dimensional scene. Then, according to the camera parameters sampled by the real camera sampler, the voxel renderer is used to render the generated neural radiance field to obtain a synthetic image. The synthetic image is input to the two-dimensional image discriminator, and the two-dimensional image discriminator updates the three-dimensional scene generator and the two-dimensional image discriminator through a discriminative loss function by forcing the distribution of the generated synthetic image to be closer to the distribution of the training images. After the training is completed, a trained three-dimensional scene generator can be obtained. At this time, the distribution of the synthetic images rendered in the neural radiance field generated by the trained three-dimensional scene generator is close to the distribution of the training images, that is, the synthetic images rendered in the generated neural radiance field are realistic enough.
[0044] Among them, the three-dimensional scene generator can adopt the generator of StyleGAN2, including multiple convolutional layers connected in sequence.
[0045] The indoor three-dimensional scene is represented using a neural radiance field. The neural radiance field used in this embodiment adopts a three-plane design, including three orthogonal feature planes and a multi-layer perceptron. The input of the neural radiance field is the coordinate position in the three-dimensional space. Bilinear interpolation is performed in the three orthogonal feature planes according to the coordinate position to obtain three features of the coordinate position. The three features are added together, and the added feature is input to the multi-layer perceptron, and finally the density value and color value of the coordinate position in the neural radiance field are output.
[0046] The role of the real camera sampler is to avoid sampling invalid camera parameters during the training process. For example, sampling camera parameters inside furniture is unreasonable. Therefore, it is preferable to select camera parameters at positions where there are no furniture objects (i.e., positions with lower density values). The input of the real camera sampler is the neural radiance field of the indoor three-dimensional scene and a set of camera parameters from the training images One training image corresponds to a set of camera parameters. For the th training image, the camera parameters are denoted as which selects a set of camera parameters from multiple sets of camera parameters, that is, the output is a set of camera parameters suitable for the indoor three-dimensional scene. Specifically, first calculate the density values of the positions of the set of camera parameters in the neural radiance field Then use the softmin function to normalize the density values into the sampling probability of the camera According to the probability sample a set of camera parameters from the set of camera parameters for subsequent rendering and synthesizing images in the neural radiance field. This sampling strategy reduces the possibility of sampling invalid camera parameters and also retains the randomness required for camera parameter sampling. Among them, the camera parameters include the field of view angle, position, and rotation angle of the camera.
[0047] For the generated neural radiance field, a voxel renderer can be used to render a synthesized image from a specified camera view. The synthesized image includes a color image and a depth image.
[0048] The camera parameters of the training images are predicted by a camera parameter predictor, which includes multiple convolutional layers connected in sequence. The input of the camera parameter predictor is the training image, and the output is the camera parameters corresponding to the training image. To train the camera parameter predictor, in this embodiment, input images required for training and the corresponding true camera parameters are generated from the generated neural radiance field. Specifically, multiple groups of camera parameters are randomly sampled by a random camera sampler, and combined with the neural radiance field, synthetic images rendered under each group of camera parameters are obtained. The synthetic images and their corresponding camera parameters are used as the training data of the camera parameter predictor, and a prediction loss function is used to calculate the loss to update the camera parameter predictor.
[0049] The prediction loss function is as follows.
[0050] 。
[0051] Among them, is the prediction loss; is the value of the th element in the camera parameters predicted by the camera parameter predictor; is the value of the th element in the true camera parameters; is the value of the th element in the true camera parameters.
[0052] The two-dimensional image discriminator can adopt the discriminator of StyleGAN2, which includes multiple convolutional layers connected in sequence.
[0053] The discriminant loss function is the loss function used in the existing generative adversarial network, which will not be elaborated here. Moreover, the process of updating the three-dimensional scene generator and the two-dimensional image discriminator is also the update method used in the existing generative adversarial network, which will not be elaborated here.
[0054] In the generation stage, the three-dimensional scene generator can generate different indoor three-dimensional scenes according to different given latent codes. Thanks to the neural radiance field representation of the indoor three-dimensional scene, the indoor three-dimensional scene represented by a grid can be reconstructed from the neural radiance field, and color images and depth images at any viewing angle in the indoor three-dimensional scene can be obtained through a voxel renderer.
[0055] The above method can only realize the generation of indoor three-dimensional scenes, but cannot realize the controllable generation of indoor three-dimensional scenes according to user needs. To achieve the purpose of controllable generation, this embodiment designs a method for controllable generation of indoor three-dimensional scenes based on neural radiance fields, introduces control conditions input by users, supports users to provide control conditions such as scene images (such as photos), scene text descriptions, or room wall boundaries, and can controllably generate a large number of diverse room-level indoor three-dimensional scenes.
[0056] This embodiment designs a controllable generation scheme for indoor three-dimensional scenes. To meet the user's customized control requirements for indoor three-dimensional scene generation, this embodiment supports using control conditions such as user-given scene images, scene text descriptions, or room wall boundaries as inputs to generate corresponding room-level indoor three-dimensional scenes. In order to enable the three-dimensional scene generator to support different types of control conditions as inputs and generate indoor three-dimensional scenes that meet the control conditions, the following will introduce different control conditions respectively from the signal injection of the control conditions and the constraint loss function corresponding to the control conditions.
[0057] When the user gives a scene image as a control condition, a scene image encoder (which includes multiple sequentially connected convolutional layers) is used to extract the two-dimensional feature map of the scene image, a camera parameter predictor is used to estimate the camera parameters of the scene image, and the three-dimensional space is discretized into a voxelized three-dimensional space according to the resolution of the three planes of the neural radiance field. For each voxel block in the voxelized three-dimensional space, it is projected onto the two-dimensional space of the two-dimensional feature map according to the estimated camera parameters, and the pixel in the two-dimensional feature map that is closest to it is indexed according to the nearest neighbor interpolation principle, and the feature at the pixel position in the two-dimensional feature map is assigned to the voxel block to obtain a three-dimensional feature map. To facilitate injecting the scene image features (i.e., the three-dimensional feature map) into the neural radiance field expressed by the three planes, a max-pooling operation is used to pool the three-dimensional feature map along three orthogonal coordinate axis directions into the feature maps of three orthogonal planes, and the feature maps of the three orthogonal planes are injected into the neural radiance field through signal modulation. Specifically, the feature maps of the three orthogonal planes are used as the input of each convolutional layer of the three-dimensional scene generator, multiplied by the original input of the convolutional layer first, and then input into the convolutional layer. To promote the three-dimensional scene generator to generate a scene that is consistent with the input scene image, during the update process of the three-dimensional scene generator, an image consistency loss function is added, requiring that the two-dimensional image rendered from the generated indoor three-dimensional scene under the corresponding camera parameters is consistent with the input scene image in the perceptual domain, so as to constrain the generated indoor three-dimensional scene to meet the input scene image conditions. To make the training of the network more robust and stable, it is only required that the two images are consistent in the perceptual domain without requiring the two images to be pixel-by-pixel consistent in the two-dimensional image space. The specific approach is to require that the image features obtained after the two images pass through the same fixed pre-trained image feature extraction network (such as VGG) are consistent. The specific calculation process is as follows: Use the pre-trained image feature extraction network to extract features from the given scene image and the synthetic image obtained by rendering the neural radiance field under the camera parameters of the given scene image to obtain the first feature map and the second feature map , are the height, width, and number of channels of the feature map respectively. Calculate the first feature map and the second feature map the distance, that is, calculate the image consistency loss function .
[0058] When the user gives a scene text description as a control condition, first use the existing text-to-image generation network to convert the scene text description into a scene image, and then the above scene image control method can be followed.
[0059] When the user gives the room wall boundary, first generate a three-dimensional room based on the room wall boundary, set the pixel value of the pixel points inside the three-dimensional room to 1, and set the pixel value of the pixel points outside the three-dimensional room to 0 to obtain a three-dimensional room mask. Then project the three-dimensional room mask onto three orthogonal planes along three orthogonal coordinate axes respectively to obtain the room wall boundary masks of the three orthogonal planes. Next, use a room wall boundary encoder (which includes multiple sequentially connected convolutional layers) to extract the feature maps of the three orthogonal planes respectively, and inject the feature maps of the three orthogonal planes into the neural radiance field by means of signal modulation. Specifically, use the feature maps of the three orthogonal planes as the input of each convolutional layer of the three-dimensional scene generator, multiply them with the original input of the convolutional layer first, and then input them into the convolutional layer. In order to make the generated indoor three-dimensional scene meet the given room wall boundary, during the update process of the three-dimensional scene generator, add a boundary consistency loss function, use a voxel renderer to render the depth image of the indoor three-dimensional scene, and require the depth value at each pixel to be equal to the depth value of the room wall boundary, so as to constrain the generated indoor three-dimensional scene to meet the input room wall boundary conditions. The specific calculation process is as follows: First, generate a three-dimensional room wall based on the room wall boundary, calculate the depth image of the room wall boundary under random camera parameters , and then calculate the depth image of the synthesized image obtained by rendering the neural radiance field under random camera parameters , are the height and width of the depth image respectively, calculate the distance between the depth image of the room wall boundary and the depth image of the synthesized image, that is, calculate the boundary consistency loss function .
[0060] After introducing the control conditions, in each iteration of the training phase, first, the camera parameter predictor is used to determine the first control camera parameters corresponding to the training scene image and the first training camera parameters corresponding to each training image. The input of the 3D scene generator is the latent code (i.e., Gaussian noise) randomly sampled from a specified distribution (such as Gaussian distribution) and the training control conditions (including the training scene image and the training room wall boundaries), and the output is the first training neural radiance field of the room-level 3D scene. The voxel renderer is used to render the generated first training neural radiance field according to the first control camera parameters to obtain the first control synthetic image, render the generated first training neural radiance field according to the randomly set second control camera parameters to obtain the second control synthetic image, and render the generated first training neural radiance field according to the camera parameters sampled from all training camera parameters by the real camera sampler to obtain the first training synthetic image. The synthetic image includes both the color image and the depth image. The training image and the first training synthetic image are input to the 2D image discriminator, and the 2D image discriminator updates the 3D scene generator and the 2D image discriminator by forcing the distribution of the generated first training synthetic image to be closer to the distribution of the training image and calculating the discriminant loss through the discriminant loss function. According to the similarity between the training scene image and the first control synthetic image, the scene image conditional control consistency loss is calculated to update the 3D scene generator. Based on the training room wall boundaries and the second control camera parameters, the depth image of the training room wall boundaries is calculated. According to the difference between the depth image of the training room wall boundaries and the depth image of the second control synthetic image, the room wall boundary conditional control consistency loss is calculated to update the 3D scene generator. After the training is completed, a trained 3D scene generator can be obtained. At this time, the distribution of the synthetic images rendered in the neural radiance field generated by the trained 3D scene generator is close to the distribution of the training images, that is, the synthetic images rendered in the generated neural radiance field are realistic enough and satisfy the given room wall boundaries and scene image control conditions.
[0061] At this time, the input of the 3D scene generator is the latent code and the given room wall boundaries and scene image control conditions, and the output is three feature planes of the neural radiance field of the indoor 3D scene. The latent code and the control conditions are modulated into each convolutional layer, that is, the feature maps of the three orthogonal planes of the control conditions are used as the input of each convolutional layer of the 3D scene generator. They are first multiplied with the original input of the convolutional layer and then input into the convolutional layer, and finally, three feature planes of the neural radiance field of the indoor 3D scene under the given room wall boundaries and scene image control conditions are output.
[0062] In the generation stage, the 3D scene generator can generate different indoor 3D scenes according to different given latent codes and control conditions input by the user. Thanks to the neural radiance field representation of the indoor 3D scene, the indoor 3D scene represented by a mesh can be reconstructed from the neural radiance field, and color images and depth images at any camera view in the indoor 3D scene can be obtained through a voxel renderer.
[0063] This embodiment provides a training method for a controllable generation model of an indoor 3D scene based on a neural radiance field. According to the room wall boundaries, scene images, or scene text descriptions given by the user as control conditions, a large number of diverse room-level indoor 3D scenes can be controllably generated. The training method for the controllable generation model of the indoor 3D scene includes the following steps.
[0064] S101, obtain training control conditions and multiple unannotated indoor scene images, and use the unannotated indoor scene images as training images; the training control conditions include training scene images and training room wall boundaries.
[0065] The unannotated indoor scene images can be indoor scene images obtained from the Internet, which do not have any annotation information.
[0066] S102, respectively use the training scene image and multiple training images as inputs, and use the camera parameter predictor to determine the first control camera parameters corresponding to the training scene image and the first training camera parameters corresponding to each training image.
[0067] Before respectively using the training scene image and multiple training images as inputs and using the camera parameter predictor to determine the first control camera parameters corresponding to the training scene image and the first training camera parameters corresponding to each training image, the training method for the controllable generation model of the indoor 3D scene further includes: using the training control conditions and noise satisfying a preset distribution as inputs, and using the 3D scene generator of the current iteration to generate a second training neural radiance field of the indoor 3D scene; using a random camera sampler to randomly sample multiple second training camera parameters; for each second training camera parameter, based on the second training neural radiance field and the second training camera parameter, render a second training synthetic image; use the second training synthetic image as an input and the second training camera parameter as a label to train the camera parameter predictor of the previous iteration to obtain the camera parameter predictor of the current iteration.
[0068] S103, use the training control conditions and noise satisfying a preset distribution as inputs, and use the 3D scene generator to generate a first training neural radiance field of the indoor 3D scene.
[0069] Among them, using the training control conditions and noise that meets the preset distribution as inputs, a first training neural radiance field of the indoor three-dimensional scene is generated by using a three-dimensional scene generator, which specifically includes: using the training scene image as an input, and extracting a two-dimensional feature map by using a scene image encoder; processing the two-dimensional feature map based on the first control camera parameters to obtain a three-dimensional feature map; performing max pooling on the three-dimensional feature map along three orthogonal coordinate axes respectively to obtain scene image conditional control feature maps of three orthogonal planes; generating a three-dimensional room based on the training room wall boundaries, setting the pixel values of the pixels located inside the three-dimensional room to 1, and setting the pixel values of the pixels located outside the three-dimensional room to 0 to obtain a three-dimensional room mask; projecting the three-dimensional room mask along three orthogonal coordinate axes respectively onto three orthogonal planes to obtain room wall boundary mask of three orthogonal planes; using a room wall boundary encoder to extract features from the room wall boundary masks of three orthogonal planes respectively to obtain room wall boundary conditional control feature maps of three orthogonal planes; using the noise that meets the preset distribution, the scene image conditional control feature maps of three orthogonal planes and the room wall boundary conditional control feature maps of three orthogonal planes as inputs, and generating a first training neural radiance field of the indoor three-dimensional scene by using a three-dimensional scene generator. Specifically, the noise and the two feature maps are modulated into each convolutional layer of the three-dimensional scene generator, that is, the two feature maps are used as the inputs of each convolutional layer of the three-dimensional scene generator, first multiplied with the original input of the convolutional layer, and then input into the convolutional layer, and finally outputting the first training neural radiance field of the indoor three-dimensional scene under the training control conditions.
[0070] S104, based on the first training neural radiance field and the first control camera parameters, rendering to obtain a first control synthetic image; based on the first training neural radiance field and the second control camera parameters, rendering to obtain a second control synthetic image; based on the first training neural radiance field and all the first training camera parameters, rendering to obtain a first training synthetic image; the first control synthetic image, the second control synthetic image and the first training synthetic image all include a color image and a depth image.
[0071] Among them, based on the first training neural radiance field and all the first training camera parameters, a first training synthetic image is rendered, which specifically includes: using the first training neural radiance field and all the first training camera parameters as inputs, determining appropriate first training camera parameters by means of a real camera sampler; rendering the first training neural radiance field based on the appropriate first training camera parameters to obtain the density value and color value of each first visible position point in the indoor three-dimensional scene, where the first visible position point is a position point visible in the indoor three-dimensional scene under the appropriate first training camera parameters; using the density value and color value of each first visible position point in the indoor three-dimensional scene as inputs and rendering with a volume renderer to obtain the first training synthetic image.
[0072] S105. Calculate a depth image of the training room wall boundary based on the training room wall boundary and the second control camera parameters.
[0073] Among them, calculating a depth image of the training room wall boundary based on the training room wall boundary and the second control camera parameters specifically includes: generating a three-dimensional room wall based on the training room wall boundary, calculating the depth of the camera from each second visible position point of the three-dimensional room wall under the second control camera parameters, and forming the depth image of the training room wall boundary, where the second visible position point is a position point visible in the three-dimensional room wall under the second control camera parameters.
[0074] S106. Use a two-dimensional image discriminator to discriminate between the training image and the first training synthetic image to obtain a discrimination result, and calculate a discrimination loss according to the discrimination result.
[0075] S107. Calculate a scene image conditional control consistency loss according to the similarity between the training scene image and the first control synthetic image.
[0076] Among them, calculating a scene image conditional control consistency loss according to the similarity between the training scene image and the first control synthetic image specifically includes: using a pre-trained image feature extraction network to extract features from the training scene image and the first control synthetic image respectively to obtain a first feature map corresponding to the training scene image and a second feature map corresponding to the first control synthetic image; using the first feature map and the second feature map as inputs and calculating the scene image conditional control consistency loss with an image consistency loss function.
[0077] Among them, the image consistency loss function is as follows.
[0078] .
[0079] Among them, is the scene image conditional control consistency loss; is the number of pixel points of the first feature map and the second feature map in the height direction, = 1, 2,... ; is the number of pixel points of the first feature map and the second feature map in the width direction, = 1, 2,... ; is the pixel value of the first feature map at the pixel point ( ); is the pixel value of the second feature map at the pixel point ( ).
[0080] S108. Calculate the control consistency loss of the room wall boundary condition according to the difference between the depth image of the training room wall boundary and the depth image of the second control synthetic image.
[0081] Among them, calculating the control consistency loss of the room wall boundary condition according to the difference between the depth image of the training room wall boundary and the depth image of the second control synthetic image specifically includes: using the depth image of the training room wall boundary and the depth image of the second control synthetic image as inputs, and calculating the control consistency loss of the room wall boundary condition by using the boundary consistency loss function.
[0082] Among them, the boundary consistency loss function is as follows.
[0083] .
[0084] Among them, is the control consistency loss of the room wall boundary condition; is the number of pixel points of the depth image in the height direction, = 1, 2,... ; is the number of pixel points of the depth image in the width direction, = 1, 2,... ; is the pixel value of the depth image of the training room wall boundary at the pixel point ( ); is the pixel value of the depth image of the second control synthetic image at the pixel point ( ).
[0085] S109. Update the three-dimensional scene generator and the two-dimensional image discriminator according to the discriminant loss, the control consistency loss of the scene image condition, and the control consistency loss of the room wall boundary condition, and obtain the updated generator and the updated discriminator.
[0086] S110, Determine whether the iteration termination condition is reached.
[0087] The iteration termination condition can be reaching the maximum number of iterations.
[0088] S111, If so, stop the iteration and use the updated generator as the controllable generation model for the indoor three-dimensional scene.
[0089] S112, If not, continue the iteration, use the updated generator as the three-dimensional scene generator for the next iteration, use the updated discriminator as the two-dimensional image discriminator for the next iteration, and return to the step of "obtaining the control conditions for training and multiple unlabeled indoor scene images".
[0090] This embodiment further provides a method for applying a controllable generation model for an indoor three-dimensional scene. The method for applying the controllable generation model for an indoor three-dimensional scene includes the following steps.
[0091] S201, Obtain the control conditions input by the user; the control conditions are scene images or room wall boundaries.
[0092] S202, Use the control conditions and noise that satisfies a preset distribution as inputs, and use the controllable generation model for the indoor three-dimensional scene to generate the neural radiance field of the indoor three-dimensional scene; the controllable generation model for the indoor three-dimensional scene is trained by using the above-mentioned training method for the controllable generation model for the indoor three-dimensional scene.
[0093] S203, Generate an indoor three-dimensional scene based on the neural radiance field.
[0094] In this embodiment, the control conditions further include scene text descriptions. At this time, using the control conditions and noise that satisfies a preset distribution as inputs, and using the controllable generation model for the indoor three-dimensional scene to generate the neural radiance field of the indoor three-dimensional scene specifically includes: using a text-to-image generation network to convert the scene text description into a converted scene image; using the converted scene image and noise that satisfies a preset distribution as inputs, and using the controllable generation model for the indoor three-dimensional scene to generate the neural radiance field of the indoor three-dimensional scene.
[0095] After generating the neural radiance field of the indoor three-dimensional scene, the method for applying the controllable generation model for the indoor three-dimensional scene in this embodiment further includes: rendering to obtain a synthetic image of the indoor three-dimensional scene under the preset camera parameters based on the neural radiance field and the preset camera parameters.
[0096] Embodiment 2.
[0097] A method for generating room-level indoor three-dimensional scenes controllable based on room wall boundaries. In this embodiment, the generation of residential-level indoor three-dimensional scenes can be achieved. In this embodiment, an independent room-level three-dimensional scene generator is trained for different room types respectively to generate indoor three-dimensional scenes of different room types. When the user gives a residential floor plan as a control condition, according to the room types and room wall boundaries of each room in the residential floor plan, the corresponding room-level three-dimensional scene generators are used respectively to generate room-level indoor three-dimensional scenes that meet the room wall boundaries, and together they form the final residential-level indoor three-dimensional scene, as Figure 3 shown.
[0098] The main problem to be solved by this application is how to learn and understand the geometric structure, appearance texture, and spatial layout of indoor three-dimensional scenes from indoor scene image data from the Internet without any additional annotations, and automatically, efficiently, and controllably generate a large number of rich and diverse high-fidelity room-level or residential-level indoor three-dimensional scenes. To solve this problem, the content studied in this application is the controllable generation of high-fidelity room-level indoor three-dimensional scenes based on neural radiance fields driven by Internet image sets. Specifically, a three-dimensional perception controllable adversarial generation network framework that can be trained on indoor scene image data from the Internet without any additional annotations is constructed, which can learn and understand the geometric structure, appearance texture, and spatial layout of indoor three-dimensional scenes, support the user to provide control conditions such as scene images, scene text descriptions, room wall boundaries, or residential floor plans as inputs, and automatically, efficiently, and controllably generate a large number of rich and diverse high-fidelity room-level or residential-level indoor three-dimensional scenes that meet the given control conditions.
[0099] This application is verified on the indoor scene image dataset LSUN collected from the Internet, and can realize learning and understanding the geometric structure, appearance texture, and spatial layout of indoor three-dimensional scenes from millions of indoor scene image data from the Internet without any additional annotations, and automatically, efficiently, and controllably generate a large number of rich and diverse high-fidelity room-level or residential-level indoor three-dimensional scenes. Thanks to the fact that this application directly learns and trains on a large amount of unrestricted indoor scene image data from the Internet, this application has extremely low requirements for training data, and the generated indoor three-dimensional scenes are more consistent with the real world in terms of geometric structure, appearance texture, and spatial layout, which is convenient for the use of industries such as virtual reality and augmented reality, interior design, games, and embodied intelligence, and improves the effectiveness and generalization ability of research fields such as data-driven indoor scene understanding and analysis.
[0100] Embodiment 3.
[0101] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 4As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the above-mentioned indoor three-dimensional scene controllable generation model training method or the above-mentioned indoor three-dimensional scene controllable generation model application method.
[0102] Those skilled in the art can understand that Figure 4 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0103] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the above-mentioned indoor three-dimensional scene controllable generation model training method or the above-mentioned indoor three-dimensional scene controllable generation model application method.
[0104] Embodiment 4.
[0105] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, it implements the above-mentioned indoor three-dimensional scene controllable generation model training method or the above-mentioned indoor three-dimensional scene controllable generation model application method.
[0106] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0107] In this text, specific examples are used to illustrate the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for training a controllable generative model for an indoor three-dimensional scene, characterized in that: The indoor three-dimensional scene controllable generation model training method comprises: Acquire training control conditions and a plurality of indoor scene images without labeled information, and use the indoor scene images without labeled information as training images; the training control conditions include training scene images and training room wall boundaries; Taking the training scene image and the plurality of training images as input respectively, using a camera parameter predictor to determine a first control camera parameter corresponding to the training scene image and a first training camera parameter corresponding to each training image; Using a training control condition and noise satisfying a preset distribution as input, a first training neural radiation field of an indoor three-dimensional scene is generated by a three-dimensional scene generator; Based on the first training neural radiation field and the first control camera parameters, a first control synthetic image is rendered; based on the first training neural radiation field and the second control camera parameters, a second control synthetic image is rendered; based on the first training neural radiation field and all the first training camera parameters, a first training synthetic image is rendered; the first control synthetic image, the second control synthetic image and the first training synthetic image all include a color image and a depth image; Based on the training room wall boundary and the second control camera parameters, a depth image of the training room wall boundary is calculated; Using a two-dimensional image discriminator to discriminate the training image and the first training synthetic image to obtain a discrimination result, and calculating a discrimination loss according to the discrimination result; Calculating the scene image conditional control consistency loss according to the similarity between the training scene image and the first control synthetic image; The room wall boundary condition control consistency loss is calculated based on the difference between the depth image of the training room wall boundary and the depth image of the second control synthetic image; According to the discrimination loss, the scene image condition control consistency loss and the room wall boundary condition control consistency loss, the three-dimensional scene generator and the two-dimensional image discriminator are updated to obtain an updated generator and an updated discriminator; Determine whether the iteration termination condition is reached; If yes, stop the iteration and use the updated generator as the controllable generation model of the indoor 3D scene; If not, continue iterating, using the updated generator as the 3D scene generator of the next iteration, using the updated discriminator as the 2D image discriminator of the next iteration, and return to the step of "obtaining control conditions for training and multiple indoor scene images without labeled information".
2. The indoor three-dimensional scene controllable generation model training method according to claim 1 is characterized in that: Before using the training scene image and the plurality of training images as inputs respectively and using a camera parameter predictor to determine the first control camera parameter corresponding to the training scene image and the first training camera parameter corresponding to each training image, the indoor three-dimensional scene controllable generation model training method further includes: Using the training control condition and the noise satisfying the preset distribution as input, the current iterative 3D scene generator is used to generate a second training neural radiation field of the indoor 3D scene; Using a random camera sampler to randomly sample and obtain a plurality of second training camera parameters; For each second training camera parameter, rendering a second training synthetic image based on the second training neural radiation field and the second training camera parameter; The second training synthetic image is used as input and the second training camera parameters are used as labels to train the camera parameter predictor of the previous iteration to obtain the camera parameter predictor of the current iteration.
3. The indoor three-dimensional scene controllable generation model training method according to claim 1 is characterized in that: Taking the training control condition and the noise satisfying the preset distribution as input, a first training neural radiation field of an indoor three-dimensional scene is generated by using a three-dimensional scene generator, specifically including: Taking the training scene image as input, a two-dimensional feature map is extracted using a scene image encoder; Processing the two-dimensional feature map based on the first control camera parameter to obtain a three-dimensional feature map; The three-dimensional feature map is subjected to maximum pooling along three orthogonal coordinate axis directions to obtain the scene image conditional control feature maps of three orthogonal planes; Generate a three-dimensional room based on the wall boundary of the training room, set the pixel value of the pixel point inside the three-dimensional room to 1, and set the pixel value of the pixel point outside the three-dimensional room to 0, to obtain a three-dimensional room mask; Project the three-dimensional room mask onto three orthogonal planes along three orthogonal coordinate axis directions respectively to obtain the room wall boundary mask of the three orthogonal planes; The room wall boundary encoder is used to extract features of the room wall boundary masks of three orthogonal planes respectively, and the room wall boundary condition control feature maps of the three orthogonal planes are obtained; A first training neural radiation field for an indoor three-dimensional scene is generated using a three-dimensional scene generator, with noise that satisfies a preset distribution, scene image condition control feature maps of three orthogonal planes, and room wall boundary condition control feature maps of three orthogonal planes as input.
4. The indoor three-dimensional scene controllable generation model training method according to claim 1 is characterized in that: Based on the first training neural radiation field and all the first training camera parameters, rendering to obtain a first training synthetic image specifically includes: Taking the first training neural radiation field and all the first training camera parameters as input, determining appropriate first training camera parameters using a real camera sampler; Rendering the first training neural radiation field based on the appropriate first training camera parameters to obtain a density value and a color value of each first visible position point in the indoor three-dimensional scene; the first visible position point is a position point in the indoor three-dimensional scene that is visible under the appropriate first training camera parameters; The density value and color value of each first visible position point in the indoor three-dimensional scene are used as input, and a voxel renderer is used to render to obtain a first training synthetic image.
5. The indoor three-dimensional scene controllable generation model training method according to claim 1 is characterized in that: Based on the training room wall boundary and the second control camera parameters, a depth image of the training room wall boundary is calculated, specifically including: A three-dimensional room wall is generated based on the training room wall boundary, and the depth of the camera from each second visible position point on the three-dimensional room wall under the second control camera parameters is calculated to form a depth image of the training room wall boundary; wherein the second visible position point is a position point on the three-dimensional room wall that is visible under the second control camera parameters.
6. The indoor three-dimensional scene controllable generation model training method according to claim 1 is characterized in that: According to the similarity between the training scene image and the first control synthetic image, the scene image condition control consistency loss is calculated, specifically including: Using a pre-trained image feature extraction network, feature extraction is performed on the training scene image and the first control synthetic image to obtain a first feature map corresponding to the training scene image and a second feature map corresponding to the first control synthetic image; Taking the first feature map and the second feature map as input, the image consistency loss function is used to calculate the scene image condition control consistency loss; Among them, the image consistency loss function is: ; in, Controlling consistency loss for scene image conditions; is the number of pixels in the height direction of the first feature map and the second feature map, =1, 2, ..., ; is the number of pixels in the width direction of the first feature map and the second feature map, =1, 2, ..., ; is the first feature map at pixel point ( )’s pixel value; is the second feature map at pixel point ( )’s pixel value; According to the difference between the depth image of the training room wall boundary and the depth image of the second control synthetic image, the room wall boundary condition control consistency loss is calculated, specifically including: Taking the depth image of the training room wall boundary and the depth image of the second control synthetic image as input, the boundary consistency loss function is used to calculate the control consistency loss of the room wall boundary condition; Among them, the boundary consistency loss function is: ; in, Control consistency loss for room wall boundary conditions; is the number of pixels in the depth image in the height direction, =1, 2, ..., ; is the number of pixels in the width direction of the depth image, =1, 2, ..., ; The depth image of the training room wall boundary at pixel point ( )’s pixel value; The depth image of the second control synthesized image is at pixel point ( )’s pixel value.
7. An application method for a controllable generation model of an indoor three-dimensional scene, characterized in that: The indoor three-dimensional scene controllable generation model application method comprises: Obtaining a control condition input by a user; the control condition is a scene image or a room wall boundary; Taking control conditions and noise satisfying a preset distribution as input, a controllable generative model of an indoor three-dimensional scene is used to generate a neural radiation field of the indoor three-dimensional scene; the controllable generative model of the indoor three-dimensional scene is trained using the controllable generative model training method for indoor three-dimensional scenes described in any one of claims 1 to 6; Generate indoor 3D scenes based on neural radiation fields.
8. The indoor three-dimensional scene controllable generation model application method according to claim 7, characterized in that: The control condition also includes a scene text description. At this time, the control condition and the noise that meets the preset distribution are used as input, and the neural radiation field of the indoor three-dimensional scene is generated by using the controllable generation model of the indoor three-dimensional scene, which specifically includes: Use a text-to-image generation network to convert scene text descriptions into transformed scene images; Taking the transformed scene image and noise satisfying the preset distribution as input, the neural radiation field of the indoor 3D scene is generated using the controllable generation model of the indoor 3D scene.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the indoor three-dimensional scene controllable generation model training method described in any one of claims 1 to 6 or the indoor three-dimensional scene controllable generation model application method described in any one of claims 7 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for training a controllable generation model of an indoor three-dimensional scene described in any one of claims 1-6 or the method for applying a controllable generation model of an indoor three-dimensional scene described in any one of claims 7-8 is implemented.
Citation Information
Patent Citations
Iterative three-dimensional neural radiation field reconstruction method based on visual prompt
CN118470207A
Radiance Fields for Three-Dimensional Reconstruction and Novel View Synthesis in Large-Scale Environments
US20240420413A1