Three-dimensional indoor scene generation method based on improved diffusion model
By combining deep learning and convolutional networks, three-dimensional indoor scenes are generated, and the problem of generating high-quality and diverse three-dimensional scenes in the existing technology is solved, and more efficient and precise scene generation is achieved.
Patent Information
- Application Number
- CN202510338940.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to generate high-quality and diverse three-dimensional indoor scenarios, especially in scenarios with complex semantic and geometric relationships, and the computational efficiency is low.
A three-dimensional indoor scene generation method based on an improved diffusion model is adopted to build a convolutional network through deep learning strategies, combining local features and global features, a room layout feature extraction module and a relationship matrix prediction module are built to support unconditional generation of three-dimensional scenes.
It realizes more flexible and accurate three-dimensional indoor scene generation, improves the accuracy of scene understanding and generation, and reduces computing complexity and resource consumption.
Smart Images

Figure CN120198593A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of high-quality three-dimensional mesh generation, and particularly to a three-dimensional indoor scene generation method based on an improved diffusion model. Background Art
[0002] Synthesizing realistic, semantically meaningful, and diverse three-dimensional indoor scenes has been a long-standing challenge in the field of computer graphics. Traditional scene modeling and synthesis methods typically treat this task as an optimization problem, guiding the optimization process through predefined scene priors, such as layout guidelines, frequency distributions of object classes, affordance maps of human-object interactions, or scene arrangement instances. Such methods usually start from an initial scene and iteratively optimize the configuration multiple times. However, designing explicit rules is not only time-consuming and requires a large amount of manual intervention by experienced artists, but also the scene optimization process is often cumbersome and computationally inefficient. In addition, the predefined design rules can only express simple and intuitive scene layout patterns and cannot cover all possible scene arrangements.
[0003] To address this issue, in recent years, some methods have attempted to leverage deep generative models to learn the prior knowledge of scenes from large-scale datasets, thereby capturing more complex scene arrangement patterns and achieving diverse scene synthesis. Among these methods, the generative adversarial network (GAN)-based solutions implicitly fit the scene distribution through adversarial training. Although these methods can generate high-quality results, they often face the problem of mode collapse. In contrast, the recently proposed autoregressive models sequentially predict the next object based on known objects. Although this method has made progress in some aspects, the characteristics of the sequential prediction process fail to effectively utilize the relative attributes between objects and may accumulate prediction errors.
[0004] Inspired by the successful application of diffusion generative models in image synthesis and shape generation, diffusion models have demonstrated excellent visual quality in generative tasks, especially in multiple fields of two-dimensional image synthesis, including applications such as image inpainting, super-resolution, text-to-image generation, and video generation. However, different from the applications in the two-dimensional field, diffusion models in the three-dimensional field have not received sufficient attention, especially in three-dimensional scene synthesis. Most existing studies focus on single-object generation tasks. Compared with single objects, the synthesis of three-dimensional scenes involves more complex semantic and geometric relationships, and the spatial scope of the scene is also much broader. Summary of the Invention
[0005] The object of the present invention is to overcome the shortcomings and deficiencies of the prior art, and propose a three-dimensional indoor scene generation method based on an improved diffusion model. This method uses a deep learning strategy to construct a convolutional network, combining local features and global features. This method does not require the user to provide an initial scene, supports unconditional generation of three-dimensional scenes, and can further achieve more flexible and accurate downstream applications.
[0006] To achieve the above object, the technical solution provided by the present invention is: a three-dimensional indoor scene generation method based on an improved diffusion model. The improved diffusion model is obtained by adding modules to the original diffusion model, including: a room layout feature extraction module, which is used to extract the features of the binary image of the room layout of the scene through convolutional layers with different depths and combine them to obtain the room layout contour features, so as to generate a more accurate edge detection result and ensure that all objects are correctly mapped into the room layout, where the binary image of the room layout of the scene represents the room contour of each scene; a relationship matrix prediction module, which is used to extract the scene relationship matrix and incorporate the spatial constraints and interaction relationships between objects into the training process of the model, thereby improving the accuracy of scene understanding and generation, where the scene relationship matrix represents the position interaction relationship between objects in the scene.
[0007] The specific implementation of this three-dimensional indoor scene generation method includes the following steps:
[0008] 1) Obtain a three-dimensional indoor scene dataset and perform a structured expression operation;
[0009] 2) Augment the structured three-dimensional indoor scene data to obtain an augmented dataset;
[0010] 3) Input the augmented dataset into the improved diffusion model for training, and after multiple training iterations until the loss is minimized and stable, finally obtain a trained model;
[0011] 4) Randomly sample from the standard normal distribution to obtain a structured expression of a three-dimensional indoor scene, and then obtain the room layout contour features and the scene relationship matrix through the room layout feature extraction module and the relationship matrix prediction module to form feature variables. Input the feature variables and the structured expression of the three-dimensional indoor scene randomly sampled from the standard normal distribution into the trained model, and output the generated three-dimensional indoor scene.
[0012] Furthermore, in step 1), a three-dimensional indoor scene dataset is obtained, including the 3D-FRONT dataset and the 3D-FUTURE dataset, and then a structured expression operation is performed on the three-dimensional indoor scene dataset: First, each scene x is represented as an unordered set of objects. Each scene is in a world frame centered at the origin and consists of no more than N objects. Therefore, a fully connected scene graph with N graph nodes is used to represent such a scene, where each node represents an object o, and o i represents the i-th object, and the object o i is defined by its category c i , the axis-aligned three-dimensional bounding box size s i , the position l i , the rotation angle θ around the vertical axis i and the shape encoding f i . Therefore, the feature of the object o i is the concatenation of all attributes, that is where represents the real number space, and D refers to the dimension of the object features. Therefore, the scene x can be represented as a set of N objects o, that is
[0013] Furthermore, in step 2), the structured three-dimensional scene data is enhanced, including adding room layout contour features and a scene relationship matrix. First, the room layout of each scene is orthogonally projected by a camera to generate a binary image representation of the room layout of the scene, where the interior area of the room is set to pure white and the exterior area of the room is set to pure black to clearly distinguish the spatial range. Then, the binary image of the room layout of the scene passes through the room layout feature extraction module to obtain the room layout contour features, and the room layout contour features are introduced as additional features of the scene. Secondly, in order to more finely depict the spatial relationship between objects in the scene, a scene relationship matrix is used for modeling, and at the same time, the scene relationship matrix is introduced as an additional feature of the scene. This matrix encodes the relative position relationship between objects in the scene, considering the front, back, left, right, up and adjacent relationships to ensure the rationality of the spatial layout.
[0014] Furthermore, in step 3), the improved diffusion model uses a U-shaped neural network U-Net as the basic structure, including a room layout feature extraction module, a relationship matrix prediction module, an encoder, an intermediate layer, and a decoder. The specific situation is as follows:
[0015] The room layout feature extraction module consists of a seven-layer structure. The first six layers are convolutional layers, and the last layer is a fully connected layer. The convolutional layers sequentially use convolutional kernels with sizes of 3×3, 3×3, 3×3, 1×1, 1×1, and 1×1, with a convolutional stride of 1 for each. The ReLU activation function is used for non-linear transformation. After each convolutional layer, the spatial dimension of the feature map is gradually compressed. A max pooling layer with a size of 2×2 and a stride of 2 is connected after each convolutional layer to reduce the size of the feature map. Finally, multi-scale features are fused in the fully connected layer to obtain a 64-dimensional low-dimensional feature vector.
[0016] The relationship matrix prediction module consists of two graph convolutional layers. The graph convolutional layer uses a graph convolution operation with a convolutional kernel size of 1×1, and the stride is 1 for each. The ReLU activation function is used for non-linear transformation after the first graph convolutional layer.
[0017] The encoder consists of a four-layer structure. Each layer includes three operations: a residual module, a time embedding, and a downsampling. Each residual module uses two convolutional operations, combined with a residual connection. The activation function for each convolution is SiLU.
[0018] The intermediate layer contains a series of convolutional blocks for processing features and an attention mechanism module. Each layer includes a residual module, a cross-attention, and a time embedding operation.
[0019] The decoding layer is a four-layer structure. Each layer includes three operations: an upsampling, a feature concatenation, and a residual connection. The SiLU activation function is used. Between the encoder and the decoder, features are shared through cross-layer connections.
[0020] Furthermore, in step 3), the steps for training the improved diffusion model are as follows:
[0021] Given a structured representation of a three-dimensional scene x as the initial input x0, the binary image of the room layout of the scene is passed through the room layout feature extraction module to obtain the room layout contour feature L. At the same time, x0 is passed through the relationship matrix prediction module to obtain the scene relationship matrix R. Then, the room layout contour feature L and the scene relationship matrix R are concatenated to obtain a new feature variable con. Then, Gaussian noise is gradually added to x0 according to the predefined linearly increasing noise variances β1, β2, …, β t , …, β T , where β1 represents the noise variance at the first time step, β t represents the noise variance at the t-th time step, and β T represents the noise variance at the T-th time step, and β1 < β2 < … < β t < … < β T . Therefore, the forward diffusion process is expressed as:
[0022]
[0023] In the formula, q(x 1:T |x0) represents the forward diffusion process from the initial input x0 after adding Gaussian noise and converting to x1, x2, …, x t , …, x T , and q(x t |x t-1 ) represents the forward diffusion process from x t-1 to x t . Here, T represents the total number of time steps, t represents a specific time step, and x t is the three-dimensional scene structured representation at the t-th step. The forward diffusion process defined at time t is:
[0024]
[0025] In the formula, is a Gaussian distribution, is the mathematical expectation, β t I is the variance, is the standard normal distribution; at each step of the forward diffusion, the model gradually adds noise to x0 until the scene structured representation x T obtained after T steps of adding noise operations is pure noise;
[0026] Next is the reverse diffusion process. Taking as the initial input, using an improved diffusion model with parameters θ and a structure of U-Net to predict its true distribution, the formula for the reverse diffusion process is as follows:
[0027]
[0028] In the formula, p θ (x t-1 |x t , con) represents the reverse diffusion process from x t to x t-1 . is a Gaussian distribution, where μ θ (x t , t, con) is the mathematical expectation, and ∑ θ (x t , t) is the variance. Here, x t-1 is the scene structured representation at the (t - 1)-th step, and Σ θ (x t , t) is a fixed value. con is the feature variable obtained by splicing the room layout contour feature L and the scene relationship matrix R. According to Bayes' theorem, μ θ (x t , t, con) can be obtained through the following formula:
[0029]
[0030] where α t = 1 - β t 、 ε θ (x t , t, con) can be predicted by improving the diffusion model. Similarly, repeating the reverse diffusion process for T steps can generate a new three-dimensional scene structured representation x0'. This three-dimensional scene structured representation x0' is also defined by its category c', the size s' of the axis-aligned three-dimensional bounding box, the position l', the rotation angle θ' around the vertical axis, and the shape encoding f'. Index the furniture categories in the three-dimensional indoor scene dataset and insert them into the position defined by the size s', the position l', and the rotation angle θ' around the vertical axis of the three-dimensional bounding box to complete the generation of the new three-dimensional scene;
[0031] Among them, for the loss function of the model design, the relationship matrix prediction module uses the cross-entropy loss between the output and the ground truth as the loss function, and the room layout feature extraction module and the original diffusion model use the mean square error between the output and the ground truth as the loss function. The Adam optimization method is used to optimize the loss function. After multiple training iterations until the loss is minimized and stable, the trained improved diffusion model is finally obtained.
[0032] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0033] 1. Model the three-dimensional scene as a set of latent variables, which represent the objects in the scene, their layout, and other relevant geometric and semantic information. Through this transformation, the originally complex three-dimensional geometric and semantic information is mapped into a relatively simplified two-dimensional latent space, enabling the model to learn and optimize in a lower-dimensional space. This not only reduces the dimension of the data but also effectively reduces the computational complexity and resource consumption during the training process.
[0034] 2. The present invention uses the binary image of the room layout of the scene as the training input of the room layout feature extraction module, and outputs the room layout contour feature as an additional feature of the scene. This room layout feature extraction module realizes the global edge detection task based on multi-scale features. By extracting features from convolutional layers at different depths and combining these features, more accurate edge detection results can be generated, which can ensure that the objects in the three-dimensional indoor scene are generated within the room layout.
[0035] 3. By generating a relationship matrix for each scene and using it as an additional feature of the scene, the present invention can incorporate the spatial constraints and interaction relationships between objects into the model training process, thereby improving the accuracy of scene understanding and generation.
[0036] 4. The method of the present invention has a wide range of application spaces in computer vision tasks, is simple to operate, highly adaptable, and has broad application prospects. Description of the Drawings
[0037] Figure 1 It is a schematic diagram of the logic flow of the present invention.
[0038] Figure 2 It is a schematic diagram of the structure of the improved diffusion model. Detailed Embodiments
[0039] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0040] As Figure 1 and Figure 2 shown, this embodiment discloses a three-dimensional indoor scene generation method based on an improved diffusion model, including the following steps:
[0041] Step 1: Obtain a three-dimensional indoor scene dataset and perform structured expression operations on the three-dimensional scene data, specifically as follows:
[0042] Obtain the 3D-FRONT dataset and the 3D-FUTURE dataset. The 3D-FRONT dataset is a large-scale comprehensive dataset containing 18,968 indoor scenes filled with 3D furniture models with professional design layouts and high-quality textures. This dataset is derived from professionally created layout designs and contains 13,151 furniture objects with high-quality textures. The 3D-FUTURE dataset contains 20,240 photos captured in 5,000 different scenes, as well as 9,992 photos related to unique industrial 3D CAD furniture shapes developed by professional designers, with high-resolution information textures.
[0043] The indoor scene is in a world coordinate system with the origin at the center of the floor, and each scene x is a combination of no more than N objects where o represents an object, and each scene is represented by a fully connected scene graph with N graph nodes, where each node represents an object o, o i represents the i-th object, and the object o i is encoded by its class semantics where represents the real number space, C refers to the total number of furniture categories, the axis-aligned three-dimensional bounding box size position shape encoding where F is the dimension of the shape encoding, and the rotation angle θ around the vertical axis iIt is defined as follows. Since the number of objects in different scenarios is different, we define an additional empty object and place it in the scenario to make the number of objects in different scenarios fixed. In summary, object o i is characterized by the concatenation of all attributes, that is where D refers to the sum of the dimensions of the concatenated attributes. Therefore, scene x can be represented as a set of N objects, that is
[0044] Step 2: Enhance the structured three-dimensional scene data to obtain an enhanced dataset, which is specifically as follows:
[0045] 2.1 Binary image of the room layout of the scene
[0046] First, use the camera to perform an orthographic projection on the scene. The room size is set to 3.1m, the camera target is the origin of the scene, the camera position is set 4 unit distances above the scene, and the projection matrix of the camera uses an orthographic projection matrix. The boundary of the projection matrix is set to the room size. Therefore, a binary image of the room layout of the corresponding scene can be obtained for each scene, where the inside of the room is set to pure white and the outside of the room is set to pure black.
[0047] 2.2 Relative relationship matrix of room objects
[0048] To better describe the relative relationships between objects in the scene, each scene uses a scene relationship matrix R to represent the spatial relationships between objects. This scene relationship matrix is an N×N matrix, where N represents the number of objects in the scene. The element R in the scene relationship matrix i,j represents the relative relationship between the i-th object and the j-th object. If there are front, back, left, right, up, and adjacent relative relationships between the first object and the second object, then the values of R 1,2 and R 2,1 in the scene relationship matrix are both 1, indicating that there is a relative relationship between these two objects; if there is no relative relationship, the corresponding element value is 0. Currently, this scene relationship matrix only considers the front, back, left, right, up, and adjacent relationships between objects within the scene.
[0049] Step 3: Construct an improved diffusion model and train it, which is specifically as follows:
[0050] 3.1 Forward noise addition process
[0051] The forward noise addition process is an image generation method based on the diffusion process. It simulates all scene graphs through a Markov chain based on discrete time. Each scene is represented as Where N refers to the number of objects, i.e., furniture, in the scene, and D refers to the dimension of object features, as described in step 1. Given a structured representation of a three-dimensional scene x as the initial input x0, the binary image of the room layout of the scene is passed through the room layout feature extraction module to obtain the room layout contour feature L. At the same time, x0 is passed through the relational matrix prediction module to obtain the scene relational matrix R. Then, the room layout contour feature L and the scene relational matrix R are concatenated to obtain a new feature variable con. Then, Gaussian noise is gradually added to x0 according to the predefined linearly increasing noise variances β1, β2, …, β t , …, β T , where β1 represents the noise variance at the first time step, β t represents the noise variance at the t-th time step, and β T represents the noise variance at the T-th time step, and β1 < β2 < … < β t < … < β T . Therefore, the forward diffusion process can be expressed as:
[0052]
[0053] In the formula, q(x 1:T |x0) represents the forward diffusion process of converting from the initial input x0 after adding Gaussian noise to x1, x2, …, x t , …, x T . q(x t |x t-1 ) represents the forward diffusion process of converting from x t-1 to x t . Here, T represents the total number of time steps, t represents a specific time step, and x t is the structured representation of the three-dimensional scene at the t-th step. The forward diffusion process defined at time t is:
[0054]
[0055] In the formula, is the Gaussian distribution, is the mathematical expectation, β t I is the variance. In particular, is the standard normal distribution. At each step of the forward diffusion, the model gradually adds noise to x0 until the structured representation x T of the scene that is almost pure noise is obtained after the noise addition operation at the T-th step.
[0056] 3.2. Reverse Denoising Process
[0057] The reverse denoising process is parameterized as a Markov chain of learnable reverse Gaussian transitions, that is, by reversing the forward noise addition process in step 3.1, we get p(x t-1 |xt )'s distribution, by continuously denoising to restore the original input x0 from Given a noisy scene graph from a standard multivariate Gaussian distribution as the initial input, an improved diffusion model with parameters θ and structure U-Net is used to predict its true distribution. The formula for the reverse diffusion process is as follows:
[0058]
[0059] In the formula, p θ (x t-1 |x t , con) represents the reverse diffusion process from x t to x t-1 . is a Gaussian distribution, where μ θ (x t , t, con) is the mathematical expectation, and ∑ θ (x t , t) is the variance. Here, x t-1 is the scene structured expression at the (t - 1)-th step, and ∑ θ (x t , t) is a fixed value. con is the feature variable obtained by concatenating the room layout contour feature L and the scene relationship matrix R. According to Bayes' theorem, μ θ (x t , t, con) can be obtained through the following formula:
[0060]
[0061] In the formula, α t = 1 - β t . ε θ (x t , t, con) can be predicted by the improved diffusion model. Similarly, repeating the reverse diffusion process for T steps can generate a new three-dimensional scene structured expression x0'. Similarly, this three-dimensional scene structured expression x0' is defined by its category c', the size s' of the axis-aligned three-dimensional bounding box, the position l', the rotation angle θ' around the vertical axis, and the shape encoding f'. Indexing the furniture categories in the three-dimensional indoor scene dataset and inserting them into the position defined by the size s', position l', and rotation angle θ' around the vertical axis of the three-dimensional bounding box can complete the generation of the new three-dimensional scene.
[0062] The improved diffusion model uses the U-shaped neural network U-Net as the basic structure, including a room layout feature extraction module, a relationship matrix prediction module, an encoder, an intermediate layer, and a decoder. The specific situation is as follows:
[0063] The room layout feature extraction module consists of seven layers. The first six layers are convolutional layers, and the last layer is a fully connected layer. The convolutional layers sequentially use convolutional kernels with sizes of 3×3, 3×3, 3×3, 1×1, 1×1, and 1×1, and the convolutional stride is 1 for each. The ReLU activation function is used for non-linear transformation. After each convolutional layer, the spatial dimension of the feature map is gradually compressed. A max pooling layer with a size of 2×2 and a stride of 2 is connected after each convolutional layer to reduce the size of the feature map. Finally, the multi-scale features are fused in the fully connected layer to obtain a 64-dimensional low-dimensional feature vector.
[0064] The relationship matrix prediction module consists of two graph convolutional layers. The graph convolutional layers sequentially use GCN convolutional operations with a convolutional kernel size of 1×1 and a stride of 1 for each. The ReLU activation function is used for non-linear transformation after the first convolutional layer. After each convolutional operation, the node features will propagate according to the graph structure to learn more rich relationship features between nodes. The first convolutional layer maps the input feature dimension from 65 to 256, and the second convolutional layer further maps the hidden features to 21, and outputs the relationship prediction result between node pairs through the Sigmoid activation function. This network effectively captures the local structural information between nodes through two GCN convolutional layers and generates the relationship prediction probability between each pair of nodes.
[0065] The encoder consists of 4 layers. Each layer includes three operations: a residual module, a time embedding, and downsampling. Each residual module uses two convolutional operations, combined with a residual connection. The activation function for each convolution is the Sigmoid LinearUnit, the size of each convolutional kernel is 1×1, the stride is 1, and the number of convolutional kernels is 512.
[0066] The middle layer contains a series of convolutional blocks for processing features and an attention mechanism module. Each layer includes a residual module, cross-attention, and time embedding operations. The residual module is similar to the residual modules in the encoder part but is used to extract deep features. And self-attention mechanism and cross-attention mechanism are added to this layer to further enhance the dependence relationship between features. Similarly, the time embedding is passed to these blocks to enhance the time correlation.
[0067] The decoding layer also has a 4-layer structure. Each layer includes three operations: upsampling, feature concatenation, and residual connection. The SiLU activation function is used. Between the encoder and the decoder, features are shared through cross-layer connections. Each decoder layer first upsamples its input, and then concatenates the upsampled features with the feature map from the corresponding encoder layer. The concatenated feature map will be used as the input of the current layer of the decoder for further feature learning through adaptive convolution.
[0068] The dataset is divided into a training set, a validation set, and a test set in a ratio of 7:2:1 to train the model. The validation set is used to evaluate the model in real time and calculate evaluation metrics, and the test set is used to test the performance of the trained network. The device used has an Intel Xeon Silver 4216 processor and an NVIDIA 3090 graphics card. For the scene generation task, training is first carried out with a batch size of 128 and a learning rate of 0.0002. Every 20,000 epochs, the learning rate will decrease once. Each time it is adjusted, the learning rate is multiplied by a decay factor of 0.5. A total of 150,000 epochs of training are carried out, and the entire training process takes seven days. In the training process, the relationship matrix prediction module uses the cross-entropy loss between the output and the ground truth as the loss function, and the room layout feature extraction module and the diffusion model use the mean square error between the output and the ground truth as the loss function. The Adam method is used to optimize the loss function, and finally the improved diffusion model after training is obtained.
[0069] 4. Randomly sample from the standard normal distribution to obtain a three-dimensional structured expression of the indoor scene. Then, through the room layout feature extraction module and the relationship matrix prediction module, the room layout contour features and the scene relationship matrix are obtained to form feature variables. The feature variables and the three-dimensional structured expression of the indoor scene randomly sampled from the standard normal distribution are input into the trained model to output the generated three-dimensional indoor scene.
[0070] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A three-dimensional indoor scene generation method based on an improved diffusion model, characterized in that: The improved diffusion model adds modules to the original diffusion model, including: a room layout feature extraction module, which is used to extract the features of the binary image of the room layout of the scene through convolution layers of different depths and combine them to obtain the room layout contour features, thereby generating more accurate edge detection results to ensure that all objects are correctly mapped to the room layout, where the binary image of the room layout of the scene represents the room contour of each scene; a relationship matrix prediction module, which is used to extract the scene relationship matrix and incorporate the spatial constraints and interaction relationships between objects into the model training process, thereby improving the accuracy of scene understanding and generation, where the scene relationship matrix represents the position interaction relationship between objects in the scene; The specific implementation of the three-dimensional indoor scene generation method includes the following steps: 1) Obtain a 3D indoor scene dataset and perform structured expression operations; 2) Enhance the structured 3D indoor scene data to obtain an enhanced data set; 3) Input the enhanced data set into the improved diffusion model for training. After multiple training iterations, the loss is minimized and stable, and finally a trained model is obtained. 4) Randomly sample from the standard normal distribution to obtain a structured expression of the three-dimensional indoor scene, and then obtain the room layout contour features and the scene relationship matrix to form feature variables through the room layout feature extraction module and the relationship matrix prediction module. The feature variables and the structured expression of the three-dimensional indoor scene obtained by random sampling from the standard normal distribution are input into the trained model together to output the generated three-dimensional indoor scene.
2. The method for generating a three-dimensional indoor scene based on an improved diffusion model according to claim 1, characterized in that: In step 1), a 3D indoor scene dataset is obtained, including the 3D-FRONT dataset and the 3D-FUTURE dataset, and then a structured expression operation is performed on the 3D indoor scene dataset: First, each scene x is represented as an unordered set of objects. Each scene is in a world frame centered on the origin and is composed of no more than N objects. Therefore, a fully connected scene graph with N graph nodes is used to represent such a scene, where each node represents an object o, o i Represents the i-th object, object o i By its category c i , axis-aligned 3D bounding box size s i 、Location i , the rotation angle θ around the vertical axis i And the shape code f i To define, so the object o i The characteristic of is the concatenation of all attributes, namely in represents the real number space, and D refers to the dimension of the object features, so the scene x can be represented as a set of N objects o, that is, 3. The method for generating a three-dimensional indoor scene based on an improved diffusion model according to claim 2, characterized in that: In step 2), the structured expression of the three-dimensional scene data is enhanced, including adding room layout outline features and scene relationship matrices. First, the room layout of each scene is orthogonally projected by a camera to generate a binary image representation of the room layout of the scene, wherein the interior area of the room is set to pure white and the exterior area of the room is set to pure black to clearly distinguish the spatial range. The binary image of the room layout of the scene is then passed through a room layout feature extraction module to obtain the room layout outline features, and the room layout outline features are introduced as additional features of the scene. Secondly, in order to more finely characterize the spatial relationship between objects in the scene, the scene relationship matrix is used for modeling, and the scene relationship matrix is introduced as an additional feature of the scene. The matrix encodes the relative position relationship between objects in the scene, taking into account the front, back, left, right, top and adjacent relationships to ensure the rationality of the spatial layout.
4. The method for generating a three-dimensional indoor scene based on an improved diffusion model according to claim 3, characterized in that: In step 3), the improved diffusion model uses a U-type neural network U-Net as the basic structure, including a room layout feature extraction module, a relationship matrix prediction module, an encoder, an intermediate layer, and a decoder. The specific situation is as follows: The room layout feature extraction module includes a seven-layer structure, the first six layers are convolutional layers, and the last layer is a fully connected layer. The convolutional layers use convolution kernels of sizes 3×3, 3×3, 3×3, 1×1, 1×1, and 1×1 in sequence, and the convolution step size is 1. The ReLU activation function is used for nonlinear transformation. After each convolution layer, the spatial dimension of the feature map is gradually compressed. Each convolutional layer is followed by a maximum pooling layer with a pooling size of 2×2 and a step size of 2 to reduce the size of the feature map. Finally, the multi-scale features are fused in the fully connected layer to obtain a 64-dimensional low-dimensional feature vector. The relationship matrix prediction module consists of two graph convolution layers, which use a graph convolution operation with a convolution kernel size of 1×1 and a step size of 1. The ReLU activation function is used for the nonlinear transformation after the first layer of graph convolution. The encoder consists of a four-layer structure, each layer includes three operations: residual module, time embedding and downsampling, where each residual module uses two convolution operations, with residual connection, and the activation function of each convolution is SiLU; The intermediate layer contains a series of convolution blocks and attention mechanism modules for processing features, and each layer includes a residual module, a cross attention and a time embedding operation; The decoding layer has a four-layer structure, each layer includes three operations: upsampling, feature concatenation and residual connection, and adopts SiLU activation function. Between the encoder and the decoder, the features are shared through cross-layer connections.
5. The method for generating a three-dimensional indoor scene based on an improved diffusion model according to claim 4, characterized in that: In step 3), the steps for training the improved diffusion model are as follows: Given a structured representation of a three-dimensional scene x as the initial input x0, the binary image of the room layout of the scene is passed through the room layout feature extraction module to obtain the room layout outline feature L, and x0 is passed through the relationship matrix prediction module to obtain the scene relationship matrix R. The room layout outline feature L and the scene relationship matrix R are then concatenated to obtain a new feature variable con, and then Gaussian noise is gradually added to x0, and the noise variance β1, β2, … β is linearly increased according to the pre-defined t ,…,β T , where β1 represents the noise variance at the first time step, β t represents the noise variance at the tth time step, β T represents the noise variance at the Tth time step, and β1<β2<…<β t <…<β T , so the forward diffusion process is expressed as: In the formula, q(x 1:T |x0) represents the conversion from the initial input x0 to x1, x2, …, x t ,…,x T The forward diffusion process, q(x t |x t-1 ) represents the t-1 Convert to x t The forward diffusion process, where T represents the total number of time steps, t represents the specific time step, and x t It is the structured expression of the three-dimensional scene at the tth step. The forward diffusion process defined at time t is: In the formula, is a Gaussian distribution, is the mathematical expectation, β t I is the variance, is a standard normal distribution; at each step of the forward diffusion, the model gradually adds noise to x0 until the scene structured expression x of pure noise is obtained after T steps of adding noise. T ; This is followed by a reverse diffusion process, As the initial input, the improved diffusion model is used, where the parameter is θ and the structure is U-Net to predict its true distribution. The formula of the reverse diffusion process is as follows: In the formula, p θ (x t-1 |x t ,con) represents the t Convert to x t-1 The reverse diffusion process, is a Gaussian distribution, where μ θ (x t ,t,con) is the mathematical expectation, ∑ θ (x t ,t) is the variance, where x t-1 is the structural expression of the scene at step t-1, ∑ θ (x t ,t) is a fixed value, con is the feature variable obtained by concatenating the room layout outline feature L and the scene relationship matrix R. According to the Bayesian theorem μ θ (x t ,t,con) can be obtained by the following formula: In the formula, α t =1-β t , ε θ (x t ,t,con) can be predicted by improving the diffusion model. Similarly, the reverse diffusion process can be repeated by T steps to generate a new three-dimensional scene structured expression x0′. Similarly, the three-dimensional scene structured expression x0′ is defined by its category c', axis-aligned three-dimensional bounding box size s', position l', rotation angle θ' around the vertical axis and shape code f'. The furniture is indexed in the three-dimensional indoor scene dataset and then inserted into the position defined by the size s', position l' and rotation angle θ' of the three-dimensional bounding box to complete the generation of the new three-dimensional scene. Among them, for the loss function of the model design, the relationship matrix prediction module uses the cross entropy loss between the output and the true value as the loss function, the room layout feature extraction module and the original diffusion model use the mean square error between the output and the true value as the loss function, and the Adam optimization method is used to optimize the loss function. After multiple training iterations until the loss is minimal and stable, a trained improved diffusion model is finally obtained.
Citation Information
Cited By
Indoor three-dimensional scene layout optimization method based on wall features
CN120912833A
Three-dimensional indoor scene layout generation method and system based on scene graph control
CN121564225A