Open-domain Indoor Scene Hierarchy Generation Method and System Based on Large Language Model
Through a hierarchical scene structure and fine-grained relative position inference network, the problem of object overlap and out of bounds by LLM when generating indoor scene layout is solved, and more reasonable and feasible scene generation is achieved, supporting open domain and interactive design.
Patent Information
- Application Number
- CN202510019290.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-07
AI Technical Summary
In the prior art, large language model (LLM) lacks spatial reasoning ability when generating indoor scene layouts, resulting in serious overlap and out-of-boundary phenomena of objects. It is difficult for existing methods to convert text descriptions into digital layouts to meet dense and complex spatial relationships at the same time, resulting in inconsistent or unreasonable generation results.
A hierarchical scene structure is adopted, scene nodes are defined through a three-level hierarchy, and relationships between objects are represented using simple text phrases. Fine-grained relative positions are inferred in combination with pre-trained visual semantic models, and partition layout optimization strategies are designed to optimize object positions locally and globally to generate physically feasible scene layouts.
It improves the rationality and physical feasibility of the scene layout, and the generated scenes are more in line with user needs, with better generalization capabilities and interactive design support.
Smart Images

Figure CN119416330B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of indoor scene synthesis, and particularly relates to an open-domain indoor scene hierarchical generation method and system based on a large language model. Background Art
[0002] Indoor scene design requires comprehensive consideration of space division, function arrangement, and aesthetic creativity to determine the selection and placement of objects, thereby forming a scene layout. Its goal is to automatically generate reasonable, realistic, and diverse three-dimensional indoor scenes, especially considering arbitrary user requirements. However, due to the complexity of indoor scenes, most of them are limited to the scope of training data and cannot be generalized to arbitrary conditions. Recently, some pioneering work has utilized the powerful generalization ability of pre-trained large language models (LLMs) to solve open-domain scene synthesis tasks, where the LLM is responsible for interpreting any text requirement as a detailed scene configuration. The challenge of this method lies in obtaining a reasonable and physically feasible scene layout from the LLM output.
[0003] The inventors found that the prior art has the following technical defects: on the one hand, the LLM can utilize original knowledge and example sets to directly output a numerical layout. However, due to the lack of spatial reasoning ability of the LLM, it is unable to understand the spatial relationships of the numerical layout, resulting in serious problems such as object overlap and out-of-bounds. On the other hand, compared with directly generating a numerical layout, having the LLM generate a text description of the scene spatial relationship can obtain more reliable answers. But this requires a method to convert the text description into a digital layout while maintaining the generalization of the entire pipeline. Existing methods pre-define text phrases and numerical rules for several types of spatial relationships to obtain the layout. However, dense spatial relationships often cannot be satisfied simultaneously, resulting in inconsistent configurations output by the LLM and generated results, while rough relationships are difficult to represent complex spatial positions, leading to unreasonable object placement. Summary of the Invention
[0004] Aiming at the problems existing in the above prior art, the present invention provides an open-domain indoor scene hierarchical generation method and system based on a large language model, which ensures the physical feasibility of the scene while significantly improving the layout rationality.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] In a first aspect, the present invention provides an open-domain indoor scene hierarchical generation method based on a large language model, including:
[0007] Define the scene structure as a three - level hierarchical structure. The first level is the root node representing the entire scene. The second level is the internal nodes where each node represents a rectangular functional area. The third level is the leaf nodes representing the objects belonging to the corresponding area. Use simple text phrases to represent the relationships between objects. Construct a prompt according to the user's needs and the definition of the scene structure as input to guide the pre - trained large language model to divide the functional areas and output structured text to describe the hierarchical scene representation, including the size of the objects, text descriptions, and the attributes of the coarse - grained relative positions between objects.
[0008] Train a fine - grained relative position inference network to infer the fine - grained relative positions between objects with spatial relationships. Based on the hierarchical scene structure and with the help of a pre - trained visual - semantic large model, it can infer reasonable relative positions in an open - domain setting.
[0009] Design a divide - and - conquer layout optimization strategy to optimize the scene layout from the hierarchical scene representation with fine - grained relative positions. It first performs local optimization within each functional area and then global optimization to organize the areas into a physically feasible scene layout. According to the text descriptions of the objects and the pictures of the objects in the dataset, use the pre - trained CLIP to calculate the cosine similarity, retrieve the corresponding 3D object models, and then scale and place the object models according to the scene layout to generate a complete scene.
[0010] As an alternative implementation, in the three - level hierarchical structure, the nodes are connected to two types of edges, namely, the parent - child relationship representing the hierarchical structure and the pairwise relationship between objects to represent their spatial relationships. Specifically, to reduce redundancy, an anchor object is set for each functional area, and only pairwise relationships between the anchor object and other objects belonging to the same functional area are allowed.
[0011] Furthermore, each node in the hierarchical scene structure contains the definition of attributes. Assuming an axis - aligned rectangular floor plan of the scene, the root node includes the size attribute and the text description of the scene , that is, , where is a two - dimensional vector representing the length and width of the scene; the internal node includes the size attribute , text description , central position and orientation attribute , that is, , where, is a two - dimensional vector representing the length and width of the functional area, is a two - dimensional coordinate, is a binary value representing the horizontal or vertical direction; the leaf node including text descriptions , category labels , corresponding 3D models and the size of the object-oriented bounding box , central position , orientation , that is , where is a 3D vector representing the length, width, and height of the object, is a 3D coordinate; in addition, paired spatial relationships store rough text descriptions and fine-grained relative position coordinates and relative orientations , that is .
[0012] As an alternative implementation, the construction prompt is specifically:
[0013] First, assign a role and task description to the LLM, including a brief definition of the node meanings and their connection hierarchy. Second, give a description of the preferred data format and predefined constraints, including the types of functional areas, possible anchor objects, and spatial relationships. Finally, show a simple scenario example to the LLM in the preferred format and specific user requirements.
[0014] As an alternative implementation, the process of constructing a fine-grained relative position inference network includes: constructing an input graph , where represents the set of object nodes, represents the set of directed edges from node to within the same functional area. Take the object description contained in each vertex from the input graph , object size information, and the rough text description of the spatial relationship contained in each edge information.
[0015] Furthermore, use linear embedding encoding for , and use a pre-trained CLIP text encoder for encoding and . During training, use the real data of relative position coordinates to enrich each edge information and encode it using linear embedding, where is a binary indicator of alignment between two objects. Note that the relative position coordinate information is only used during the training process, that is, it is not required during the inference process. After encoding, the node embeddings And edge embedding is as follows:
[0016] 。
[0017] Furthermore, incorporate and into each node and edge of the input graph and conceptualize it as a context graph , and use a variational graph neural network for encoding and decoding. The encoding process can be expressed in the following way:
[0018] ,
[0019] where and represent the node embeddings during the k-th round of information passing, represents the edge embeddings during the k-th round of information passing; is the edge embedding update function, which means updating the edge embedding information with the node embeddings of the k-th round; is the node embedding update function, which means updating the node embedding information with the neighbor nodes in the k-th round; represents the set of neighbor nodes connected to node ; represents the operation of taking the mean; after encoding, the edge embeddings are parameterized as a Gaussian distribution; the decoder takes the updated context graph as input and randomly samples from the Gaussian distribution to obtain the features of the relative positions; finally, a separate MLP is used to decode the relative position information and output , where are the decoded relative position coordinates, is the decoded relative direction, is the predicted alignment binary indicator.
[0020] When training the fine-grained relative position inference network, the objective function is:
[0021] ,
[0022] where is the Kullback-Leibler divergence between the Gaussian distribution and the posterior distribution of the edge feature components, is the L1 loss on the relative positions, and are the cross-entropy losses of the discrete relative direction angles and the alignment binary indicators.
[0023] As an alternative implementation, the divide-and-conquer layout optimization includes two processes: local optimization and global optimization.
[0024] Furthermore, the local optimization is specifically as follows: for each functional area, perform local optimization to minimize the relative position of the object and the relative position output by the fine-grained relative position inference module, and impose constraints to avoid object overlap and out-of-bounds, that is
[0025] ,
[0026] where and represent the positions of the object and the corresponding anchor object within the functional area, including the central position coordinates and direction. represents calculating the relative position between two objects, is the relative position between the object predicted by the network and the corresponding anchor object. represents the set of objects within the area. Constrain the overlap between the oriented bounding boxes of any two objects to be as small as possible, aiming to avoid the object bounding box being outside the area boundary, is the size of the area boundary, which is generated by the pre-trained LLM.
[0027] Furthermore, the global optimization is specifically as follows: organize the areas of local optimization to form a scene. Each functional area takes the direction of its anchor object as its own direction. Based on observations in daily life, it is required that the functional areas be placed against the wall, far from each other, with the direction pointing inside the scene, while avoiding object overlap and out-of-bounds. The optimization function is as follows:
[0028] ,
[0029] where represents the area placement, including the central position coordinates and direction. is the distance between the back of the area and the scene boundary. is the distance between the bounding boxes of two domains. is the set of areas in the scene. and are the same as those in the local optimization.
[0030] Furthermore, calculate the cosine similarity between the vector obtained from the object picture through the CLIP Image Encoder and the vector obtained from the text description through the CLIPText Encoder.
[0031] In a second aspect, the present invention provides an open-domain indoor scene hierarchical generation system based on a large language model, including:
[0032] The hierarchical scene generation module is configured to: define the scene structure as a three - level hierarchical structure, where the first level is the root node representing the entire scene, the second level is the internal node where each node represents a rectangular functional area, and the third level is the leaf node representing the objects within the corresponding area; construct a prompt according to the user requirements and the definition of the scene structure to guide the pre - trained large language model to divide the functional areas and output structured text to describe the hierarchical scene representation, including the size of the objects, text descriptions, and the attributes of the coarse - grained relative positions between the objects.
[0033] The fine - grained relative position reasoning module is configured to: infer the fine - grained relative positions between objects with spatial relationships; based on the hierarchical scene structure and with the help of a pre - trained visual semantic large model, it can infer reasonable relative positions in an open - domain setting.
[0034] The divide - and - conquer layout optimization module is configured to: optimize the scene layout from the hierarchical scene representation with fine - grained relative positions; it first performs local optimization within each functional area and then global optimization to organize the areas into a physically feasible scene layout; according to the text descriptions of the objects and the pictures of the objects in the dataset, it uses the pre - trained CLIP to calculate the cosine similarity, retrieves the corresponding 3D object models, and then scales and places the object models according to the scene layout to generate a complete scene.
[0035] The present invention has achieved the following technical effects:
[0036] The present invention proposes a new open - domain indoor scene hierarchical generation method based on a large language model. It uses a hierarchical scene structure as an intermediate representation to guide the refinement of spatial relationships and the optimization of layout positions, ensuring the physical feasibility of the scene while improving the layout rationality. It mainly includes three modules: the hierarchical scene generation module, the fine - grained relative position reasoning module, and the divide - and - conquer layout optimization module. First, the hierarchical scene generation module generates a hierarchical scene structure with text descriptions by prompting the pre - trained large language model. This structure has three levels, with the entire scene as the root node, functional areas as internal nodes, and objects as leaf nodes. It uses simple text phrases to represent the spatial relationships of the objects, and the spatial relationships only exist within the areas. This avoids the contradictory placement caused by complex and dense spatial relationships. Subsequently, to avoid the problem of inaccurate object placement due to rough spatial relationships, the fine - grained relative position reasoning module is used to further infer the fine - grained relative positions between objects with text - based spatial relationships that are difficult to describe with text phrases. Based on the hierarchical structure and with the help of a pre - trained visual semantic large model, this module can infer reasonable relative positions in an open - domain setting. Subsequently, the present invention designs a divide - and - conquer optimization to optimize each functional area separately and then arrange them to form the entire scene, which can effectively generate a physically feasible scene layout.
[0037] It is worth noting that using the hierarchical scene representation module has twofold advantages. First, the hierarchical scene structure provides a rough basis for object arrangement, alleviates the contradictory placements with dense relationships, and enhances the generalization ability of the network to infer fine-grained placements. Second, it naturally supports a divide-and-conquer optimization to more effectively generate scene layouts that match the descriptions generated by the LLM. Extensive comparative experiments, ablation studies, and in-depth analyses were conducted qualitatively and quantitatively. The experimental results show that the present invention can generate more reasonable and physically feasible scenes, which are more in line with user requirements and LLM arrangements. In addition, the present invention can perform scene generation and interactive design in the open domain, with generalization. Brief Description of the Drawings
[0038] Figure 1 Schematic diagram of an open-domain indoor scene hierarchy generation method based on a large language model in Embodiment 1 of the present invention;
[0039] Figure 2 Schematic diagram of the hierarchical scene structure in Embodiment 1 of the present invention;
[0040] Figure 3 Schematic diagram of eight possible alignment situations in Embodiment 1 of the present invention;
[0041] Figure 4 Qualitative comparison chart of scene generation results of different methods in Embodiment 1 of the present invention; among them, (a) is the scene generation result of ATISS, (b) is the scene generation result of DiffuScene, (c) is the scene generation result of LayoutGPT, (d) is the scene generation result of HOLODECK, and (e) is the scene generation result of the present invention;
[0042] Figure 5 Schematic diagram of the qualitative evaluation of the ablation study in Embodiment 1 of the present invention; among them, (a) is the input hierarchical scene structure diagram, (b) is the output result of the basic method, (c) is the output result after adding the fine-grained relative position inference network, and (d) is the output result after adding both the fine-grained relative position inference network and the divide-and-conquer optimization strategy; the arrows in (a) represent the relationships between nodes, and the arrows in (b), (c), and (d) represent the directions of each piece of furniture;
[0043] Figure 6 Different result displays of the open-domain indoor scene generation task in Embodiment 1 of the present invention; (a) is the scene of generating different layouts under the same requirement; (b) is the scene of generating different types according to the requirements of a specific user;
[0044] Figure 7Shows the different results of the open-domain indoor scene editing task in Embodiment 1 of the present invention; among them, (a) is the living room of an artist, (b) is the dining room of a family of four, (c) is the bedroom of an actress, and (d) is an arcade;
[0045] Figure 8 Schematic diagram of an open-domain indoor scene hierarchical generation framework based on a large language model in Embodiment 2 of the present invention. Detailed implementation manners
[0046] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0047] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0048] The present invention will be further described below in conjunction with embodiments.
[0049] Embodiment 1
[0050] As Figure 1 shown, the present invention provides an open-domain indoor scene hierarchical generation method based on a large language model. The method includes three stages. First is the hierarchical scene generation stage, which prompts the pre-trained LLM to generate a hierarchical scene structure with text descriptions. The hierarchical scene structure has three levels. The entire scene serves as the root node, the functional area serves as the internal node, and the object serves as the leaf node. The nodes are connected by two types of edges, namely the parent-child relationship representing the hierarchical structure and the pairwise relationship between objects to represent their spatial relationship. Specifically, to reduce redundancy, an anchor object is set for each functional area, and only pairwise relationships between the anchor object and other objects belonging to the same functional area are allowed. As Figure 2As shown. In this structure, simple text phrases are used to represent the spatial relationships of objects, and the spatial relationships only exist within the regions. This avoids the contradictory placement caused by complex and dense spatial relationships. Secondly is the fine-grained relative position reasoning stage. To avoid the problem of inaccurate object placement caused by rough spatial relationships, the present invention trains a fine-grained relative position reasoning network to further infer the fine-grained relative positions between objects with text spatial relationships that are difficult to describe with text phrases. Based on the hierarchical structure and with the help of a pre-trained large vision-language model, this network can infer reasonable relative positions in an open-vocabulary setting. Thirdly is the divide-and-conquer layout optimization stage. The present invention develops a divide-and-conquer optimization to optimize each functional region separately and then arrange them to form the entire scene to effectively solve the physically feasible scene layout.
[0051] 1. Hierarchical scene generation stage
[0052] Construct a prompt according to the user requirements and the definition of the scene structure to guide the pre-trained large language model to divide the functional regions. The pre-trained LLM takes the constructed prompt as input and outputs structured text to describe the hierarchical scene representation, including the size of the objects, the text description, and the attributes of the coarse-grained relative positions between the objects. The key challenge is to generate reasonable and information-rich spatial relationships to specify the scene layout.
[0053] Although existing work defines dense relationships between objects to describe the layout, due to the insufficient spatial reasoning ability of the LLM, the more detailed the description, the more errors or self-contradictory results they produce. Therefore, the present invention requires the LLM to generate a hierarchical room structure to describe the objects, only generate the spatial relationships between the objects belonging to the same region, and roughly specify their arrangement.
[0054] The present invention constructs the input prompt with three components. First is the description of the role and task assigned to the LLM, including a brief definition of the hierarchical structure with node meanings and their connections. Second is the description of the preferred data format and predefined constraints, including the types of functional regions, possible anchor objects, and spatial relationships. Finally, show a simple scene example to the LLM in the preferred format and specific user requirements. In this stage, the LLM generates the text descriptions and size attributes of the functional regions and objects, as well as the text descriptions of the spatial relationships.
[0055] Each node in the hierarchical scene structure contains the definition of attributes. Assuming an axis-aligned rectangular floor plan of the scene, the root node includes the size attribute and the text description of the scene i.e., where, is a two-dimensional vector representing the length and width of the scene; the internal node including size attributes , text description , central position and orientation attributes , that is , where is a two-dimensional vector representing the length and width of the functional area, is a two-dimensional coordinate, is a binary value representing the horizontal or vertical direction; leaf nodes include text description , category label , corresponding 3D model and the size of the object-oriented bounding box , central position , orientation , that is , where is a three-dimensional vector representing the length, width and height of the object, is a three-dimensional coordinate; in addition, paired spatial relationships store rough text description and fine-grained relative position coordinates and relative orientation , that is .
[0056] 2. Fine-grained relative position inference stage
[0057] To avoid the problem of inaccurate object placement caused by rough spatial relationships, the present invention proposes a hierarchical perception-based fine-grained relative position inference network to infer the fine-grained relative positions between related objects. The relative positions within each functional area exhibit more compact and generalizable priors, enabling the training of the network to infer positions in various scenarios. As Figure 1 shown, given the hierarchy generated by the LLM, construct the input graph , where represents the set of object nodes, represents the set of directed edges from node to within the same functional area. Although the input includes all objects in the scene, the functional areas are isolated from each other.
[0058] Furthermore, extract information from the hierarchical scene structure generated by the LLM, where each vertex contains the description of the object , the size of the object information and each edge contains rough text description of the spatial relationship information. Furthermore, for Using linear embedding encoding, the pre-trained CLIP text encoder is used for and encoding. During training, the ground-truth data of relative position coordinates is used to enrich each edge information and encode it using linear embedding, where is the binary indicator of the alignment between two objects. Note that the relative position coordinate information is only used during the training process, i.e., it is not required during the inference process.
[0059] These information are incorporated into each node and edge of the input graph and conceptualized as a context graph where the node embedding and the edge embedding are:
[0060] .
[0061] Since there are only text space relationships between the anchor object and other objects, for the edges without corresponding text space relationships, all-zero vectors are used as text embeddings.
[0062] For the context graph with node embedding and edge embedding , a variational graph neural network is used for encoding and decoding. The encoding process can be expressed in the following way:
[0063] ,
[0064] where and represent the node embeddings at the k-th round of information passing, represents the edge embeddings at the k-th round of information passing; is the edge embedding update function, which means updating the edge embedding information with the node embeddings at the k-th round; is the node embedding update function, which means updating the node embedding information with the neighbor nodes in the k-th round; represents the set of neighbor nodes connected to the node ; represents the operation of taking the mean; after encoding, the edge embeddings are parameterized as a Gaussian distribution; the decoder takes the updated context graph as input and randomly samples from the Gaussian distribution to obtain the features of the relative positions; finally, a separate MLP is used to decode the relative position information and output where is the decoded relative position coordinates, is the decoded relative direction, is the predicted alignment binary indicator.
[0065] In the post - processing stage, based on the predicted correct the predicted relative positions . Specifically, as Figure 3 shown, there are a total of eight possible alignments between any two objects. If the indicator classifies two objects as aligned, the alignment solution closest to the network - predicted position will be selected from the eight possible cases.
[0066] During training, freeze the CLIP text encoder and update all other network layers. The objective function is:
[0067] ,
[0068] where is the Kullback - Liebler divergence between the Gaussian distribution and the posterior distribution of the edge feature components. is the L1 loss on the relative positions. and is the cross - entropy loss of the discrete relative direction angles and the binary alignment indicators.
[0069] 3. Divide - and - Conquer Layout Optimization Stage
[0070] Given a hierarchical scenario of the relative positions between related objects, the present invention designs a divide - and - conquer optimization to solve the final layout. First, local optimization is performed on each functional area, and then global optimization organizes the areas into a scenario. This optimization is more effective than simple global optimization in generating a reasonable and physically feasible layout, or iteratively optimizing the positions of each object.
[0071] Local optimization. For each functional area, use local optimization to solve the positions of the objects within the functional area. Local optimization is formulated to minimize the relative placement of the objects and the positions inferred by the network, with constraints to avoid object overlap and outside the boundaries, i.e.,
[0072] ,
[0073] where and represent the positions of the object and the corresponding anchor object within the functional area, including the central position coordinates and orientation. represents calculating the relative position between two objects, is the relative position between the object predicted by the network and the corresponding anchor object. represents the set of objects within the area. constrain the overlap between the oriented bounding boxes of any two objects to be as small as possible, aiming to avoid the object bounding boxes being outside the area boundaries, is the size of the regional boundary, which is generated by the pre-trained LLM.
[0074] Global optimization. These regions are organized using global optimization to form a scene. Each functional region takes the orientation of its anchor object as its own orientation. Based on observations in daily life, an optimization strategy is developed to place the functional regions on the walls, far from each other, with their orientations pointing inside the scene, while avoiding object overlap and going out of bounds. The optimization function is as follows:
[0075] ,
[0076] where, represents the regional placement, including the central position coordinates and orientation. is the distance between the back of the region and the scene boundary. is the distance between the bounding boxes of two domains. is the set of regions in the scene. and are the same as those in local optimization.
[0077] After optimization, the coordinate system is transformed to obtain the object positions and orientations in the scene frame. Finally, according to The scoring retrieves 3D object models from the Objaverse and 3D-Front datasets, that is, calculates the cosine similarity between the vector obtained from the object image through the CLIP ImageEncoder and the vector obtained from the text embedding through the CLIP Text Encoder. Then, the object models are scaled and placed according to the scene layout, thereby generating a complete scene.
[0078] To achieve the indoor scene generation task in the open domain, existing research based on pre-trained large language models can be divided into two categories, including methods that directly output numerical layouts and methods that indirectly output descriptions of spatial relationships.
[0079] It should be understood that the method of directly outputting numerical layouts requires the LLM to be able to directly output the final result using the original knowledge and example sets. Currently, many studies focus on the 2D layouts of controllable scene image synthesis. For example, UI Grammar introduces a grammar tree to guide the LLM to generate UI layouts. LayoutGPT first tried to use the retrieved scene layouts and required the LLM to output the numerical bounding boxes containing each object in the format of CSS. However, due to the lack of spatial reasoning ability, the LLM cannot handle the complex relationships and changes in 3D scenes, resulting in serious object overlap and out-of-bounds problems.
[0080] It should be understood that the method of indirectly outputting spatial relationship descriptions aims to use an LLM to generate text scene descriptions and convert them into 3D scenes. For example, Aladdin samples from an abstract description and generates a set of three-dimensional texture assets, and manually organizes them to build a scene. HOLODECK and AnyHome require the LLM to describe object relationships by selecting from a set of predefined atomic relationships, which are interpreted as fixed relative positions between objects and refined using a rule-based optimization algorithm. However, it is very difficult to define a set of compact and informative atomic relationships. Dense and detailed object relationships can provide precise spatial arrangements, but usually lead to self-contradictions in the LLM output, while sparse and coarse-grained relationship sets lead to coherent arrangements but cannot capture the different spatial positions between objects.
[0081] To verify the effectiveness of the algorithm, the present invention conducted comparative experiments and ablation studies on the 3D-Front dataset. According to the settings of LayoutGPT, scenes with irregular floors were also filtered out from the dataset. Finally, the sizes of the training sets for the bedroom and living room were 3397 and 690, respectively, while the corresponding test sets were 60 and 53.
[0082] The present invention evaluates the generated scenes from two perspectives. One is the physical feasibility of the scene, which is estimated by the ratio of the overlapping of oriented bounding boxes, i.e., the overlap rate, and the ratio outside the boundary, i.e., the out-of-bounds rate. The other is the rationality of the scene layout. For this, the inventors selected some common object pairs, namely bed-nightstand, table-chair, coffee table-sofa, and measured the average KL-divergence between the relative position distribution of the real scenes in the test set and the generated scenes.
[0083] For all LLM-assisted methods, GPT-4 was used for evaluation. The Adam optimizer was set to train the fine-grained relative position inference network for 500 epochs, with a batch size of 4 and a learning rate of 1e-4. The divide-and-conquer optimization was implemented using the GUROBI solver.
[0084] 1. Performance Comparison
[0085] The present invention compared two types of state-of-the-art methods, including those that train deep neural networks from scratch, namely ATISS and DiffuScene, and LLM-assisted indoor scene synthesis methods, namely LayoutGPT and HOLODECK.
[0086] As Figure 4As shown, the deep learning methods ATISS and DiffuScene generate results with reasonably placed objects. However, the network cannot guarantee the physical feasibility of the scene layout and sometimes results in object overlap and out-of-bounds. LayoutGPT uses context learning to infer digital layouts based on demonstrated examples and generates many incorrect orientations and positions. HOLODECK produces relatively good results in terms of physical feasibility, but some objects are not placed in the optimal positions specified by the LLM. In contrast, the present invention is capable of generating a more reasonable and feasible scene layout.
[0087] Table 1 Quantitative comparison of scene generation results of different methods
[0088] 。
[0089] Table 1 validates the observations of the visual results (in the table, ↑ indicates that the higher the value of the evaluation metric, the better the generated result; ↓ indicates that the lower the value of the evaluation metric, the better the generated result). Among all the methods, the present invention achieves the best results in terms of physical feasibility (overlap rate and out-of-bounds rate) and reasonable relative positions (KL Div.). It is worth noting that data-driven methods are good at the relative positions of objects, and the LLM-assisted optimization method, namely HOLODECK, obtains more feasible results, while the present invention takes advantage of both and achieves the best results in all metrics.
[0090] Furthermore, the present invention provides an additional quantitative comparison with HOLODECK. Both use the LLM to generate text scene descriptions and then solve the scene layout. The difference is that HOLODECK requires dense and detailed spatial relationships, while the present invention uses a hierarchical scene structure with sparse relationships and a neural network to infer fine-grained relative positions. Table 2 reports the semantic alignment results between the descriptions generated by the LLM and the generated scenes. Specifically, #Rel. calculates the percentage of relative positions that match the spatial relationships generated by the LLM; #Obj. calculates the percentage of objects in the generated scene that match the objects specified by the LLM. Obviously, the results of the present invention are better, which means the advantage of using a hierarchical scene representation.
[0091] Table 2 Semantic alignment results between the descriptions generated by the LLM and the generated scenes
[0092] 。
[0093] 2. User survey
[0094] The survey evaluated the quality of scenarios generated by different methods. The survey showed the renderings and top views of 25 generated scenarios to 30 participants. The participants were asked to rate the scenarios on a 5-point scale in terms of scenario effectiveness, physical feasibility, and layout rationality.
[0095] Table 3 shows that the results of the present invention obtained the highest scores in terms of scenario effectiveness, physical feasibility, and layout rationality. This is because the present invention uses a hierarchical scenario structure to prompt the LLM to generate, which can effectively generate feasible scenarios aligned with the LLM arrangement. Otherwise, although HOLODECK can prevent objects from overlapping or going out of bounds, its results may disrupt the LLM arrangement and may affect human activities, such as Figure 4 shown in the first row case in
[0096] Table 3 Average scores of scenarios generated by different models in the user survey in terms of scenario effectiveness, physical feasibility, and layout rationality
[0097] 。
[0098] 3. Ablation Study
[0099] Table 4 Quantitative evaluation of the ablation study to verify the key designs of the present invention: fine-grained relative position inference network and divide-and-conquer layout optimization strategy
[0100] 。
[0101] The ablation study verified the key designs of the present invention, namely the fine-grained relative position inference network and the divide-and-conquer layout optimization strategy. Specifically, corresponding stages were removed based on the method of the present invention. Among them, when removing the fine-grained relative position inference network, relative position coordinates of text space relations were predefined to replace the network's prediction; when removing the divide-and-conquer layout optimization strategy, the LLM was directly required to generate the coordinates of the anchor object, and the relative positions of other objects were converted to global coordinates without optimization refinement. The results are shown in Table 4 and Figure 5 as shown. The results show that the fine-grained relative position inference network is of great significance in capturing reasonable relative positions between objects, and the divide-and-conquer optimization ensures the physical feasibility of the generated scene layout.
[0102] 4. Open-Vocabulary Scene Synthesis
[0103] To show the generalization of the present invention, Figure 6The results of open-domain scene generation are shown. In this part, the constraints restricting object categories are removed from the LLM prompts. The results show that the LLM can generate a reasonable hierarchical scene structure for any requirements. In addition, the fine-grained relative position inference network and the divide-and-conquer layout optimization strategy can produce reasonable and realistic scenes with various hierarchical descriptions.
[0104] 5. Interactive Scene Editing
[0105] The present invention also supports user-friendly language-guided interactive scenes. Specifically, the current state of the scene and the editing instructions are described as the input of the LLM, and an additional constraint is added in the divide-and-conquer optimization to keep the positions of the unchanged objects as much as possible. As Figure 7 shown, the LLM can modify the current scene by adding and deleting objects. In addition, through the hierarchical scene structure and the method of the present invention, the edited scene can show the smallest changes from the original scene while meeting the user's expectations and the LLM output results.
[0106] Embodiment 2
[0107] As Figure 8 shown, this embodiment provides an open-domain indoor scene hierarchical generation system based on a large language model, including:
[0108] A hierarchical scene generation module, configured to: define the scene structure as a three-level hierarchical structure, where the first layer is the root node representing the entire scene, the second layer is the internal node where each node represents a rectangular functional area, and the third layer is the leaf node representing the objects within the corresponding area; construct prompts according to the user's needs and the definition of the scene structure, guide the pre-trained large language model to divide the functional areas, and output structured text to describe the hierarchical scene representation, including the size of the objects, text descriptions, and the attributes of the coarse-grained relative positions between the objects;
[0109] A fine-grained relative position inference module, configured to: infer the fine-grained relative positions between objects with spatial relationships; based on the hierarchical scene structure, with the help of a pre-trained visual semantic large model, it can infer reasonable relative positions in the open-domain setting;
[0110] A divide-and-conquer layout optimization module, configured to: optimize the scene layout from the hierarchical scene representation with fine-grained relative positions; it first performs local optimization within each functional area, and then global optimization to organize the areas into a physically feasible scene layout; according to the text descriptions of the objects and the pictures of the objects in the dataset, use the pre-trained CLIP to calculate the cosine similarity, retrieve the corresponding 3D object models, and then scale and place the object models according to the scene layout to generate a complete scene.
[0111] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0112] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An open-domain indoor scene hierarchical generation method based on large language models, characterized in that, Including: Define the scene structure as a three - level hierarchical structure. The first layer is the root node representing the entire scene. The second layer is the internal node where each node represents a rectangular functional area. The third layer is the leaf node representing the objects belonging to the corresponding area. Use simple text phrases to represent the relationships between objects. Construct prompts according to user requirements and the definition of the scene structure to guide the pre - trained large language model to divide functional areas and output structured text to describe the hierarchical scene representation, including the size of objects, text descriptions, and the attributes of the coarse - grained relative positions between objects. Train a fine - grained relative position inference network to infer the fine - grained relative positions between objects with spatial relationships. Based on the hierarchical scene structure and with the help of a pre - trained visual - semantic large model, it can infer reasonable relative positions in an open - domain setting. Design a divide - and - conquer layout optimization strategy to optimize the scene layout from the hierarchical scene representation with fine - grained relative positions. It first performs local optimization within each functional area and then global optimization to organize the areas into a physically feasible scene layout. According to the text descriptions of objects and the pictures of objects in the dataset, use the pre - trained CLIP to calculate the cosine similarity to retrieve the corresponding 3D object models, and then scale and place the object models according to the scene layout to generate a complete scene. Each node in the hierarchical scene structure contains the definition of attributes. Assuming an axis-aligned rectangular floor plan of the scene, the root node includes the size attribute and the text description of the scene , that is , where is a two-dimensional vector representing the length and width of the scene; the internal node includes the size attribute , the text description , the central position and the orientation attribute , that is , where is a two-dimensional vector representing the length and width of the functional area, is a two-dimensional coordinate, is a binary value representing the horizontal or vertical direction; the leaf node includes the text description , the category label , the corresponding 3D model and the size of the object-oriented bounding box , the central position , the orientation , that is , where is a three-dimensional vector representing the length, width and height of the object, is a three-dimensional coordinate; in addition, the paired spatial relationship stores the rough text description and the fine-grained relative position coordinates and the relative orientation , that is ; The specific construction of the prompt is as follows: First, assign a role and task description to the LLM, including a brief definition of the node meanings and their connected hierarchical structures. Second, give a description of the preferred data format and predefined constraints, including the types of functional areas, possible anchor objects, and spatial relationships. Finally, show a simple scene example to the LLM in the preferred format and specific user requirements. The process of constructing a fine-grained relative position inference network includes: constructing an input graph according to the hierarchical scene structure , where represents a set of object nodes, represents a set of directed edges from node to within the same functional area; taking the object description contained in each vertex from the input graph , object size information, and the rough text description of the spatial relationship contained in each edge information.
2. The open-domain indoor scene hierarchical generation method based on a large language model according to claim 1, wherein It also includes: for Use linear embedding encoding for and Use a pre-trained CLIP text encoder for encoding; during training, use the ground truth data of relative position coordinates to enrich each edge information and encode it using linear embedding, where is a binary indicator of the alignment between two objects; the relative position coordinate information is only used during the training process, that is, it is not required during the inference process; after encoding, the node embedding and the edge embedding are: 。 3. The open-domain indoor scene hierarchical generation method based on a large language model according to claim 2, characterized in that, Also including: Include and into each node and edge of the input graph and conceptualize it as a context graph , and use a variational graph neural network for encoding and decoding; the encoding process is expressed in the following way: , Among them, and represent the node embeddings during the k-th round of message passing, represents the edge embeddings during the k-th round of message passing; is the edge embedding update function, which means updating the edge embedding information with the node embeddings in the k-th round; is the node embedding update function, which means updating the node embedding information with the neighbor nodes in the k-th round; represents the set of neighbor nodes connected to the node ; represents the operation of taking the mean; after encoding, the edge embeddings are parameterized as a Gaussian distribution; the decoder takes the updated context graph as input and randomly samples from the Gaussian distribution to obtain the features of the relative positions; finally, a separate MLP is used to decode the relative position information and output , where is the decoded relative position coordinate, is the decoded relative direction, is the predicted alignment binary indicator.
4. An open-domain indoor scene hierarchical generation method based on a large language model according to claim 1, wherein, When training the fine - grained relative position inference network, the objective function is: , Among them, is the Kullback-Leibler divergence between the Gaussian distribution and the posterior distribution of the edge feature components, is the L1 loss on the relative position, and is the cross-entropy loss of the discrete relative direction angle and the binary indicator of alignment.
5. A method for generating an open-domain indoor scene hierarchy based on a large language model according to claim 1, characterized in that, The divide - and - conquer layout optimization includes two processes: local optimization and global optimization.
6. The open-domain indoor scene hierarchical generation method based on a large language model according to claim 5, wherein, The specific local optimization is: For each functional area, perform local optimization to minimize the relative positions of objects and the relative positions output by the fine - grained relative position inference module, with constraints to avoid object overlap and out - of - bounds, that is , Among them, and represent the positions of the object and the corresponding anchor object within the functional area, including the central position coordinates and the direction; represents calculating the relative position between two objects, is the relative position between the object predicted by the network and the corresponding anchor object; represents the set of objects within the area; constrains the overlap between the oriented bounding boxes of any two objects to be as small as possible, aims to avoid the object bounding box being outside the area boundary, is the size of the area boundary and is generated by the pre-trained LLM; The specific global optimization is: Organize the locally optimized areas to form a scene, and each functional area takes the direction of its anchor object as its own direction. Based on observations in daily life, require the functional areas to be placed against the wall, away from each other, with the direction pointing into the scene, while avoiding object overlap and out - of - bounds. The optimization function is as follows: , Among them, represents region placement, including the center position coordinates and direction; is the distance between the back of the region and the scene boundary; is the distance between the bounding boxes of two domains; is the set of regions in the scene; and is the same as that in the local optimization.
7. The generation system of an open-domain indoor scene hierarchy generation method based on a large language model according to any one of claims 1-6, characterized in that, Including: A hierarchical scene generation module, configured to: Define the scene structure as a three - level hierarchical structure, where the first layer is the root node representing the entire scene, the second layer is the internal node where each node represents a rectangular functional area, and the third layer is the leaf node representing the objects belonging to the corresponding area. Construct prompts according to user requirements and the definition of the scene structure to guide the pre - trained large language model to divide functional areas and output structured text to describe the hierarchical scene representation, including the size of objects, text descriptions, and the attributes of the coarse - grained relative positions between objects. The fine-grained relative position reasoning module is configured to: infer the fine-grained relative positions between objects with spatial relationships; based on the hierarchical scene structure and with the help of a pre-trained large vision-language model, it can infer reasonable relative positions in an open-domain setting; The divide-and-conquer layout optimization module is configured to: optimize the scene layout from the hierarchical scene representation with fine-grained relative positions; It first performs local optimization within each functional area and then global optimization to organize the areas into a physically feasible scene layout; according to the text descriptions of the objects and the pictures of the objects in the dataset, it uses the pre-trained CLIP to calculate the cosine similarity to retrieve the corresponding 3D object models, and then scales and places the object models according to the scene layout to generate a complete scene.
Citation Information
Patent Citations
Scene map generation method
CN111462282A
Three-dimensional scene generation method based on pre-training language model and related components
CN117475089A