Indoor scene construction method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202510598798.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2045-05-09
AI Technical Summary
[0004]本发明提供一种室内场景构建方法、装置、电子设备及存储介质,用以解决现有技术中基于生成模型的方法、基于大语言模型的方法,以及程序化内容生成方法在场景构建中的缺陷
[0016] The present invention provides an indoor scene construction method, apparatus, electronic device, and storage medium that acquires multimodal input, extracts scene description from the multimodal input, and then generates an initial layout 3D scene based on the scene description to obtain an initial indoor scene. The initial indoor scene is then vertically decoupled to obtain a multi-layer scene structure. The multi-layer scene structure sequentially includes a floor layer, a furniture layer, and a stacking layer. The stacking layer describes objects stacked on top of other objects. The multi-layer scene structure is horizontally decoupled to obtain a scene graph. Nodes in the scene graph represent objects, and edges in the scene graph represent the positional semantic relationships between objects. The scene graph is then optimized to obtain the target indoor scene. On the one hand, a coarse-to-fine generation method is first adopted to vertically decouple the scene into a multi-level structure of floors, furniture, and stacked elements. Then, the scene structure is horizontally decoupled to obtain a scene graph. The process is optimized step by step from the macro-level furniture layout to the precise arrangement of micro-level objects (such as tableware and books), improving the placement accuracy and physical rationality of fine-grained objects. On the other hand, the scene graph is optimized to further improve the construction of indoor scenes, enhancing the accuracy and adaptability of indoor scene construction. Thus, the layered generation architecture and feedback optimization mechanism work together to achieve refined control of the scene.
Smart Images

Figure CN120765829B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of scene construction technology, and in particular to an indoor scene construction method, apparatus, electronic device and storage medium. Background Technology
[0002] With the rapid development of virtual reality, interior design, and embodied intelligence, generating high-precision, realistic 3D interior scenes has become a key technology. Realistic 3D scenes not only enhance the immersive experience of virtual environments but also provide interior designers with intuitive visualization tools to help users quickly validate design solutions. Simultaneously, they provide high-quality training data for embodied intelligence systems, enabling them to perform precise interactions and planning in complex environments. Furthermore, the automated generation of physically accurate and richly detailed scenes is also of great significance in film production, game development, and other fields.
[0003] Existing 3D indoor scene generation technologies mainly revolve around generative models, large language models (LLMs), and procedural content generation. Generative model-based methods directly synthesize scene layouts using generative adversarial networks or diffusion models. However, limited by the generative model's ability to model fine-grained spatial relationships, it struggles to precisely control the placement of small objects (such as tableware and books), resulting in scenes lacking detail, realism, and functionality. LLM-based methods generate 3D scenes through text prompts, but due to the limited spatial reasoning capabilities of language models, the generated layouts often suffer from object misalignment, orientation deviations, or collisions, failing to guarantee physical plausibility. Procedural content generation methods rely on predefined rules to generate deterministic scenes. While they can guarantee stability and collision-free layouts, their coarse-grained rule system struggles to support fine-grained spatial reasoning, exhibiting shortcomings in the precise control of small object positions. Summary of the Invention
[0004] This invention provides an indoor scene construction method, apparatus, electronic device, and storage medium to address the shortcomings of existing generative model-based methods, large language model-based methods, and procedural content generation methods in scene construction.
[0005] This invention provides a method for constructing a three-dimensional indoor scene, comprising the following steps: Acquire multimodal input and extract scene description from the multimodal input; Based on the scene description, an initial layout 3D scene is generated to obtain the initial indoor scene; The initial indoor scene is vertically decoupled to obtain a multi-layer scene structure; the multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence; the stacking layer is used to reflect objects stacked on top of other objects. The multi-layer scene structure is horizontally decoupled to obtain a scene graph; the nodes in the scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects; The scene image is optimized to obtain the target indoor scene.
[0006] According to a method for constructing a three-dimensional indoor scene provided by the present invention, optimizing the scene map to obtain a target indoor scene includes: The scene graph is subjected to single-attribute verification, and the first indoor scene is determined based on the single-attribute verification result; the single-attribute verification is used to verify whether objects in each level of the scene graph belong to the same parent node; Perform intra-layer collision detection and / or inter-layer collision detection on the first indoor scene to obtain the second indoor scene; Based on the second indoor scene, the target indoor scene is determined.
[0007] According to a method for constructing a three-dimensional indoor scene provided by the present invention, determining the target indoor scene based on a second indoor scene includes: Determine the target position of the center of gravity projection of the target object in the second indoor scene. Based on the positional relationship between the target position of the center of gravity projection and the support surface of the target object, adjust the posture of the target object. Based on the adjusted second indoor scene, determine the target indoor scene.
[0008] According to a method for constructing a three-dimensional indoor scene provided by the present invention, adjusting the posture of the target object based on the positional relationship between the target position of the center of gravity projection and the support surface of the target object includes: If the target position of the center of gravity projection exceeds the support surface, adjust the projection position of the target object to be within the support surface.
[0009] According to a method for constructing a three-dimensional indoor scene provided by the present invention, determining the target indoor scene based on an adjusted second indoor scene includes: Based on the alignment optimization model, the adjusted second indoor scene is aligned and optimized to obtain the target indoor scene. The alignment optimization model is trained based on sample multimodal input and the corresponding label indoor scene of the sample multimodal input.
[0010] According to a method for constructing a three-dimensional indoor scene provided by the present invention, the training step of the alignment optimization model includes: Obtain the initial alignment optimization model, the sample multimodal input, and the label position and label orientation of the object in the label indoor scene corresponding to the sample multimodal input; Extract the predicted scene description from the multimodal input of the sample, and generate a three-dimensional scene of the initial layout based on the predicted scene description to obtain the predicted initial indoor scene. The predicted initial indoor scene is vertically decoupled to obtain a predicted multi-layer scene structure; the predicted multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence; The predicted multi-layer scene structure is decoupled at the scene level to obtain a predicted scene graph; the nodes in the predicted scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects. Based on the initial alignment optimization model, the predicted scene map is optimized to obtain the predicted position and predicted direction of the object in the predicted target indoor scene; Based on the predicted position and the label position, as well as the predicted direction and the label direction, a target loss is determined, and the initial alignment optimization model is iterated based on the target loss to obtain the alignment optimization model.
[0011] According to a three-dimensional indoor scene construction method provided by the present invention, the stacking level is a sub-level of the furniture level, and the furniture level is a sub-level of the floor level; The objects in each sub-level belong to the coordinate system of the parent level corresponding to each sub-level.
[0012] The present invention also provides a three-dimensional indoor scene construction device, comprising the following units: The acquisition unit is used to acquire multimodal input and extract scene description from the multimodal input; An initial scene generation unit is used to generate an initial layout 3D scene based on the scene description to obtain an initial indoor scene; A scene vertical decoupling unit is used to perform scene vertical decoupling on the initial indoor scene to obtain a multi-layer scene structure; the multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence; the stacking layer is used to reflect objects stacked on top of other objects. A scene horizontal decoupling unit is used to perform scene horizontal decoupling on the multi-layer scene structure to obtain a scene graph; the nodes in the scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects; An optimization unit is used to optimize the scene map to obtain the target indoor scene.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the three-dimensional indoor scene construction method as described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the three-dimensional indoor scene construction method as described above.
[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the three-dimensional indoor scene construction method as described above.
[0016] The present invention provides an indoor scene construction method, apparatus, electronic device, and storage medium that acquires multimodal input, extracts scene description from the multimodal input, and then generates an initial layout 3D scene based on the scene description to obtain an initial indoor scene. The initial indoor scene is then vertically decoupled to obtain a multi-layer scene structure. The multi-layer scene structure sequentially includes a floor layer, a furniture layer, and a stacking layer. The stacking layer describes objects stacked on top of other objects. The multi-layer scene structure is horizontally decoupled to obtain a scene graph. Nodes in the scene graph represent objects, and edges in the scene graph represent the positional semantic relationships between objects. The scene graph is then optimized to obtain the target indoor scene. On the one hand, a coarse-to-fine generation method is first adopted to vertically decouple the scene into a multi-level structure of floors, furniture, and stacked elements. Then, the scene structure is horizontally decoupled to obtain a scene graph. The process is optimized step by step from the macro-level furniture layout to the precise arrangement of micro-level objects (such as tableware and books), improving the placement accuracy and physical rationality of fine-grained objects. On the other hand, the scene graph is optimized to further improve the construction of indoor scenes, enhancing the accuracy and adaptability of indoor scene construction. Thus, the layered generation architecture and feedback optimization mechanism work together to achieve refined control of the scene. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is one of the flowcharts illustrating the three-dimensional indoor scene construction method provided by the present invention.
[0019] Figure 2 This is the second flowchart illustrating the three-dimensional indoor scene construction method provided by the present invention.
[0020] Figure 3 This is a schematic diagram of the structure of the three-dimensional indoor scene construction device provided by the present invention.
[0021] Figure 4This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] The terms "first," "second," etc., used in this invention are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and that the objects distinguished by "first," "second," etc., are generally of the same class.
[0024] Figure 1 This is one of the flowcharts illustrating the three-dimensional indoor scene construction method provided by the present invention. Figure 2 This is the second flowchart illustrating the three-dimensional indoor scene construction method provided by the present invention, as shown below. Figure 1 and Figure 2 As shown, the method includes steps 110, 120, 130, 140 and 150.
[0025] Step 110: Obtain multimodal input and extract the scene description from the multimodal input.
[0026] Specifically, multimodal input can be acquired. Here, multimodal input may include text, images, and sketches, etc., but the embodiments of the present invention do not specifically limit this.
[0027] It should be noted that before each scene generation task begins, the scene description in the user's multimodal input is received and extracted using a large language model. The scene description includes structured scene descriptions such as room type, furniture list, and spatial relationships. This embodiment of the invention does not impose specific limitations on this.
[0028] Wherein, the example input is: “Please generate a dining room with a dining table.Place two chairs on either side of the table. On both sides of the table,place a set of two stacked plates. To the right of each stacked plate, placea fork and a knife in sequence.” Output scene description: { "room_type": " dining room ", "objects": ["dining table 1", " dining chair 2", "plate 4", "fork 2","knife 2"], "spatial_relations": "The dining table is placed in the dining room.", "Two dining chairs are positioned on opposite sides of the diningtable, one in the front and one in the back.", "Two plates are placed in front of the chair at the front side of thedining table, with one plate stacked on top of the other.", "Two plates are placed in front of the chair at the back side of thedining table, with one plate stacked on top of the other.", "A fork is placed to the right of the lower plate in front of the chair at the front side of the table.", " A knife is positioned to the right of the fork at the front side of the table.", "A fork is placed to the right of the lower plate at the back side of the table.", "A knife is positioned to the right of the fork at the back side of the table." } .
[0029] For another example, the multimodal input is: "Create a room with a sofa against the wall. The sofa has two adjacent cushions on the left and one cushion on the right. In front of the sofa is a coffee table with a cup at its center, a book to the left of the cup, and another book on the right side of the table.", the large language model prompts "You are an experienced room designer. Please help me analyze the given information and give output in the specified format: "1. room type:… 2. object list:… 3. spatial relation:…"".
[0030] Step 120: based on the scene description, generate a three-dimensional scene of the initial layout, to obtain an initial indoor scene.
[0031] Specifically, after obtaining the scene description, a 3D scene with an initial layout can be generated based on the scene description to obtain the initial indoor scene. That is, coarse-grained room generation is performed based on the scene description to construct a 3D scene framework (initial indoor scene) containing the initial layout of walls, doors, windows and furniture.
[0032] Step 130: Perform vertical decoupling on the initial indoor scene to obtain a multi-layer scene structure; the multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence; the stacking layer is used to reflect objects stacked on top of other objects.
[0033] Specifically, after obtaining the initial indoor scene, the initial indoor scene can be vertically decoupled to obtain a multi-layer scene structure, which includes a floor layer, a furniture layer, and a stacking layer. Here, the floor layer includes large furniture (dining table, dining chair), the furniture layer includes tabletop objects (plate, fork, knife), and the stacking layer describes objects (plate) stacked on top of other objects.
[0034] Vertical scene decoupling refers to separating and processing a 3D scene according to a hierarchical structure (i.e., from top to bottom or from bottom to top). Each layer is responsible for different functions or data abstractions.
[0035] It should be noted that the stacking level is a child level of the furniture level, and the furniture level is a child level of the floor level. In other words, the floor level is the parent level of the furniture level, and the furniture level is the parent level of the stacking level.
[0036] In this context, objects in each sub-level belong to the coordinate system of their corresponding parent level. For example, a fork belongs only to the dining table level.
[0037] Here, the furniture hierarchy is refined to the precise placement of tabletop objects (such as tableware and books), and the positions of child-level objects are dynamically adjusted based on the coordinate system of the parent hierarchy. Furthermore, as needed, other stacked layers can be added to the furniture hierarchy, and spatial constraints are passed between layers through parent-child relationship chains to ensure that child-level objects do not escape the spatial boundaries of the parent hierarchy.
[0038] Step 140: Decouple the multi-layer scene structure horizontally to obtain a scene graph; the nodes in the scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects.
[0039] Specifically, after obtaining the multi-layer scene structure, the multi-layer scene structure can be decoupled at the scene level to obtain a scene graph. In the scene graph, the nodes represent objects, and the edges represent the positional semantic relationships between objects. Here, the positional semantic relationships between objects refer to the semantic associations that exist between different objects in spatial location.
[0040] Scene horizontal decoupling refers to the separation and independent processing of different functional modules or objects in a 3D scene during construction, arranged horizontally (i.e., at the same level). These modules or objects are at the same level in the hierarchical structure, but are independent of each other in terms of function, data, or logic.
[0041] Here, a large language model can be used to decouple the multi-layered scene structure at the scene level, resulting in a scene graph. Example: At the furniture level, the above multi-layered scene structure is input into the large language model, and the resulting scene graph output by the large language model shows the positional semantic relationships between objects (i.e., transforming the scene description into positional semantic relationships in the scene graph): plate-0 | diningtable,on | diningtable,front plate-1 | diningtable,on | diningtable,back fork-0 | diningtable,on | plate-0,right fork-1 | diningtable,on | plate-1,right kinfe-0 | diningtable,on | fork-0,right kinfe-1 | diningtable,on | fork-1,right In addition, the scene graph can be recursively parsed based on a greedy algorithm, and objects can be placed in order of priority (e.g., place the plate first, and then recursively place the fork and the knife).
[0042] Step 150: Optimize the scene map to obtain the target indoor scene.
[0043] Specifically, after obtaining the scene graph, it can be optimized to obtain the target indoor scene. For example, single-attribute verification, intra-level collision detection, and inter-level collision detection can be performed on the scene graph. This embodiment of the invention does not specifically limit these methods.
[0044] The target indoor scene refers to the indoor scene ultimately constructed by the user.
[0045] Here, the scene graph is optimized, which can resolve layout conflicts and improve physical rationality through iterative adjustments.
[0046] The method provided in this invention involves acquiring multimodal input, extracting scene descriptions from the multimodal input, generating an initial 3D scene based on the scene descriptions to obtain an initial indoor scene, and performing vertical decoupling on the initial indoor scene to obtain a multi-layer scene structure. The multi-layer scene structure sequentially includes a floor layer, a furniture layer, and a stacking layer. The stacking layer describes objects stacked on top of other objects. The multi-layer scene structure is then horizontally decoupled to obtain a scene graph. Nodes in the scene graph represent objects, and edges in the scene graph represent the positional semantic relationships between objects. The scene graph is then optimized to obtain the target indoor scene. On the one hand, a coarse-to-fine generation method is first adopted to vertically decouple the scene into a multi-level structure of floors, furniture, and stacked elements. Then, the scene structure is horizontally decoupled to obtain a scene graph. The process is optimized step by step from the macro-level furniture layout to the precise arrangement of micro-level objects (such as tableware and books), improving the placement accuracy and physical rationality of fine-grained objects. On the other hand, the scene graph is optimized to further improve the construction of indoor scenes, enhancing the accuracy and adaptability of indoor scene construction. Thus, the layered generation architecture and feedback optimization mechanism work together to achieve refined control of the scene.
[0047] Based on the above embodiments, step 150 includes: Step 151: Perform single-attribute verification on the scene graph, and determine the first indoor scene based on the single-attribute verification result; the single-attribute verification is used to verify whether objects in each level of the scene graph belong to the same parent node. Step 152: Perform intra-layer collision detection and / or inter-layer collision detection on the first indoor scene to obtain the second indoor scene; Step 153: Determine the target indoor scene based on the second indoor scene.
[0048] Specifically, a single-attribute verification can be performed on the scene graph, and the first indoor scene can be determined based on the single-attribute verification result. The single-attribute verification verifies whether objects in each level of the scene graph belong to the same parent node. For example, if a fork belongs to both the dining table level and the dining chair level, a conflict is triggered, and it is forcibly assigned to the dining table level.
[0049] After obtaining the first indoor scene, collision detection within a layer and / or collision detection between layers can be performed on the first indoor scene to obtain the second indoor scene. Here, collision detection within a layer, collision detection between layers, or both can be performed on the first indoor scene, etc., and this embodiment of the invention does not specifically limit the specific methods used.
[0050] Here, the bounding box (BBOX) can be used to detect whether the fork and knife on the coffee table overlap. If they overlap, the position of the knife is translated along the X / Y axis.
[0051] Inter-level detection can be performed by checking whether the projection of the stacked plates overlaps with the fork through axial projection (Proj_z). If they overlap, the position of the fork is moved.
[0052] After obtaining the second indoor scene, the target indoor scene can be determined based on the second indoor scene.
[0053] Based on the above embodiments, step 153 includes: Step 1531: Determine the target position of the center of gravity projection of the target object in the second indoor scene; adjust the posture of the target object based on the positional relationship between the target position of the center of gravity projection and the support surface of the target object; and determine the target indoor scene based on the adjusted second indoor scene.
[0054] Specifically, the target position of the center of gravity projection of the target object in the second indoor scene can be determined, and the posture of the target object can be adjusted based on the positional relationship between the target position of the center of gravity projection and the supporting surface of the target object to ensure that the center of gravity projection of the object is located within the polygon of the supporting surface, thus preventing the object from tilting or floating.
[0055] Finally, based on the adjusted second indoor scene, the target indoor scene is determined.
[0056] Based on the above embodiments, step 1531, which involves adjusting the posture of the target object based on the positional relationship between the target position of the center of gravity projection and the support surface of the target object, includes: Step 1531-1: If the target position of the center of gravity projection exceeds the support surface, adjust the projection position of the target object to be within the support surface.
[0057] Specifically, when the target position of the center of gravity projection exceeds the support surface, the projection position of the target object is adjusted to be within the support surface, thereby preventing the object from tilting or floating.
[0058] The method provided in this invention adjusts the projection position of the target object to within the support surface when the target position of the center of gravity projection exceeds the support surface, thereby preventing the object from tilting or suspending.
[0059] Based on the above embodiments, step 1531, which involves determining the target indoor scene based on the adjusted second indoor scene, includes: Based on the alignment optimization model, the adjusted second indoor scene is aligned and optimized to obtain the target indoor scene. The alignment optimization model is trained based on sample multimodal input and the corresponding label indoor scene of the sample multimodal input.
[0060] Specifically, the adjusted second indoor scene can be aligned and optimized based on the alignment optimization model to obtain the target indoor scene.
[0061] Here, the training steps for the alignment optimization model are as follows: First, obtain the initial alignment optimization model, sample multimodal input, and the label indoor scene corresponding to the sample multimodal input.
[0062] Then, the predicted scene description is extracted from the multimodal input of the sample, and based on the predicted scene description, the 3D scene of the initial layout is generated to obtain the predicted initial indoor scene.
[0063] Vertical decoupling is performed on the initial indoor scene to obtain a multi-layered scene structure, which includes a floor layer, a furniture layer, and a stacking layer.
[0064] The predicted multi-layer scene structure is decoupled at the scene level to obtain the predicted scene graph, where nodes in the predicted scene graph represent objects and edges in the scene graph represent the positional semantic relationships between objects.
[0065] Based on the initial alignment optimization model, the predicted scene map is optimized to obtain the predicted target indoor scene.
[0066] Based on the difference between the predicted indoor scene and the labeled indoor scene, the target loss is determined, and the parameters of the initial alignment optimization model are iterated based on the target loss. The initial alignment optimization model after parameter iteration is used as the alignment optimization model.
[0067] Understandably, the greater the difference between the predicted indoor scene and the labeled indoor scene, the greater the target loss; conversely, the smaller the difference between the predicted indoor scene and the labeled indoor scene, the smaller the target loss.
[0068] Based on the above embodiments, the training steps of the alignment optimization model include: Step 10: Obtain the initial alignment optimization model, the sample multimodal input, and the label position and label orientation of the object in the label indoor scene corresponding to the sample multimodal input; Step 20: Extract the predicted scene description from the multimodal input of the sample, and generate a 3D scene of the initial layout based on the predicted scene description to obtain the predicted initial indoor scene. Step 30: Perform vertical decoupling on the predicted initial indoor scene to obtain a predicted multi-layer scene structure; the predicted multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence. Step 40: Decouple the predicted multi-layer scene structure at the scene level to obtain a predicted scene graph; the nodes in the predicted scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects. Step 50: Based on the initial alignment optimization model, optimize the predicted scene map to obtain the predicted position and predicted direction of the object in the predicted target indoor scene; Step 60: Based on the predicted position and the label position, as well as the predicted direction and the label direction, determine the target loss, and perform parameter iteration on the initial alignment optimization model based on the target loss to obtain the alignment optimization model.
[0069] Specifically, first, the initial alignment optimization model, sample multimodal input, and the label position and label orientation of the objects in the label indoor scene corresponding to the sample multimodal input are obtained.
[0070] Here, the parameters of the initial alignment optimization model can be randomly generated or preset, and the embodiments of the present invention do not specifically limit this.
[0071] Then, the predicted scene description is extracted from the multimodal input of the sample, and based on the predicted scene description, the 3D scene of the initial layout is generated to obtain the predicted initial indoor scene.
[0072] Furthermore, the initial indoor scene is vertically decoupled to obtain a predicted multi-layer scene structure; the predicted multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence.
[0073] The predicted multi-layer scene structure is decoupled at the scene level to obtain the predicted scene graph; in the predicted scene graph, the nodes represent objects, and the edges in the scene graph represent the positional semantic relationships between objects.
[0074] After obtaining the predicted scene map, the predicted scene map can be optimized based on the initial alignment optimization model to obtain the predicted position and predicted direction of objects in the predicted indoor scene.
[0075] Finally, the target loss can be determined based on the predicted position and label position, as well as the predicted direction and label direction. The parameters of the initial alignment optimization model can then be iterated based on the target loss to obtain the alignment optimization model.
[0076] It is understandable that the greater the difference between the predicted position and the label position, the greater the target loss; the smaller the difference between the predicted position and the label position, the smaller the target loss.
[0077] It is understandable that the greater the difference between the predicted direction and the label direction, the greater the target loss; the smaller the difference between the predicted direction and the label direction, the smaller the target loss.
[0078] Here, the formula for the target loss is as follows: in, This represents the target loss, also known as the alignment loss function. Represents the predicted objects in the target indoor scene. The predicted location and predicted direction, Represents objects in an indoor scene. The label position and label orientation.
[0079] Here, target loss can be used to adjust the position of the items through gradient descent.
[0080] The method provided in this invention ensures semantic consistency between the generated scene and user instructions (text / image / sketch) based on the alignment loss function (target loss).
[0081] Based on any of the above embodiments, this invention proposes a refined 3D indoor scene construction method based on hierarchical layout generation, aiming to solve the problems of insufficient precision in fine-grained object placement, lack of physical rationality, and poor input adaptability in existing technologies. This method achieves refined scene control through a hierarchical generation architecture and a feedback optimization mechanism: First, a coarse-to-fine generation method is adopted, vertically decoupling the scene into multi-level structures such as floors and furniture. Combined with scene graph reasoning of horizontal spatial relationships using a large language model, the precise arrangement of objects (such as tableware and books) is optimized level by level. Second, ownership conflicts between levels are eliminated through single-ownership verification, collision detection and physical stability constraints are used to correct object overlap and positional anomalies, and the semantic consistency between the generated scene and user instructions (text / image / sketch) is ensured based on an input alignment loss function.
[0082] This invention proposes a hierarchical layout generation method that solves the problem of fine-grained control in 3D indoor scene generation through progressive optimization from coarse to fine. This invention is the first to combine refined scene generation with feedback-driven optimization, decomposing complex scenes into multiple levels and refining the layout layer by layer. In the coarse-grained stage, the overall room structure, including the initial positions of walls, doors, windows, and furniture, can be generated based on multimodal inputs (text, images, sketches). In the fine-grained stage, precise object layout is achieved through vertical and horizontal decoupling: vertically, layers are created along the Z-axis (e.g., floor layer → furniture layer → stacking layer), ensuring that objects belong to only a single parent node (e.g., a cup is only placed on the table); horizontally, a scene graph is constructed using a Large Language Model (LLM), transforming spatial relationships (e.g., "a teaspoon is placed to the left of the teacup") into structured constraints, and a greedy algorithm recursively optimizes the positions of sub-objects.
[0083] To address the issues of physical plausibility and input alignment in layout generation, this invention designs a feedback-driven optimization module: First, it eliminates ownership conflicts between levels through single-ownership verification (e.g., preventing a cup from simultaneously belonging to both a table and a chair); second, it solves the problem of object overlap by combining collision detection algorithms; further, it introduces physical stability constraints, adjusting objects based on the relationship between the object's center of gravity projection and the supporting surface to prevent objects from floating; finally, it minimizes the positional and directional deviations between the generated scene and user commands through an input alignment loss function. This invention supports full-scale generation from macroscopic furniture to microscopic decorations, significantly improving the layout accuracy of fine-grained objects (such as tableware and office supplies) while ensuring global structural consistency. It overcomes the limitations of existing methods in terms of missing details, frequent collisions, and physical inconsistencies, providing a highly controllable and high-fidelity 3D scene generation solution for applications such as virtual reality and interior design.
[0084] Experiments show that this method significantly outperforms mainstream generative models in terms of object boundary crossing rate (16.2%) and orientation accuracy (89.2%). User surveys show that the overall preference score (8.5 / 10.0) is higher than the comparison method. It effectively solves the problem of co-optimization between global planning and local details, and can be widely applied in the fields of virtual reality environment construction, smart home design and embodied intelligence training, promoting the practical application of high-precision 3D scene generation technology.
[0085] The three-dimensional indoor scene construction device provided by the present invention is described below. The three-dimensional indoor scene construction device described below can be referred to in correspondence with the three-dimensional indoor scene construction method described above.
[0086] Based on any of the above embodiments, the present invention provides a three-dimensional indoor scene construction device. Figure 3 This is a structural schematic diagram of the three-dimensional indoor scene construction device provided by the present invention, as shown below. Figure 3 As shown, the device includes: The acquisition unit 310 is used to acquire multimodal input and extract scene description from the multimodal input; The initial scene generation unit 320 is used to generate a three-dimensional scene of the initial layout based on the scene description, so as to obtain an initial indoor scene; The scene vertical decoupling unit 330 is used to perform scene vertical decoupling on the initial indoor scene to obtain a multi-layer scene structure; the multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence; the stacking layer is used to reflect objects stacked on top of other objects. The scene horizontal decoupling unit 340 is used to perform scene horizontal decoupling on the multi-layer scene structure to obtain a scene graph; the nodes in the scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects; The optimization unit 350 is used to optimize the scene map to obtain the target indoor scene.
[0087] The apparatus provided in this invention acquires multimodal input, extracts scene description from the multimodal input, and then generates an initial layout 3D scene based on the scene description to obtain an initial indoor scene. The initial indoor scene is then vertically decoupled to obtain a multi-layer scene structure. The multi-layer scene structure sequentially includes a floor layer, a furniture layer, and a stacking layer. The stacking layer describes objects stacked on top of other objects. The multi-layer scene structure is horizontally decoupled to obtain a scene graph. Nodes in the scene graph represent objects, and edges in the scene graph represent the positional semantic relationships between objects. The scene graph is then optimized to obtain the target indoor scene. On the one hand, a coarse-to-fine generation method is first adopted to vertically decouple the scene into a multi-level structure of floors, furniture, and stacked elements. Then, the scene structure is horizontally decoupled to obtain a scene graph. The process is optimized step by step from the macro-level furniture layout to the precise arrangement of micro-level objects (such as tableware and books), improving the placement accuracy and physical rationality of fine-grained objects. On the other hand, the scene graph is optimized to further improve the construction of indoor scenes, enhancing the accuracy and adaptability of indoor scene construction. Thus, the layered generation architecture and feedback optimization mechanism work together to achieve refined control of the scene.
[0088] Based on any of the above embodiments, the optimization unit 350 specifically includes: A single-attribute verification unit is used to perform single-attribute verification on the scene graph and determine the first indoor scene based on the single-attribute verification result; the single-attribute verification is used to verify whether objects in each level of the scene graph belong to the same parent node. The collision detection unit is used to perform intra-layer collision detection and / or inter-layer collision detection on the first indoor scene to obtain the second indoor scene. A target indoor scene determination unit is used to determine the target indoor scene based on the second indoor scene.
[0089] Based on any of the above embodiments, determining the target indoor scene unit specifically includes: A determination subunit is used to determine the target position of the center of gravity projection of the target object in the second indoor scene, adjust the posture of the target object based on the positional relationship between the target position of the center of gravity projection and the support surface of the target object, and determine the target indoor scene based on the adjusted second indoor scene.
[0090] Based on any of the above embodiments, the determining subunit is specifically used for: If the target position of the center of gravity projection exceeds the support surface, adjust the projection position of the target object to be within the support surface.
[0091] Based on any of the above embodiments, the determining subunit is specifically used for: Based on the alignment optimization model, the adjusted second indoor scene is aligned and optimized to obtain the target indoor scene. The alignment optimization model is trained based on sample multimodal input and the corresponding label indoor scene of the sample multimodal input.
[0092] Based on any of the above embodiments, a training unit is further included, wherein the training unit is specifically used for: Obtain the initial alignment optimization model, the sample multimodal input, and the label position and label orientation of the object in the label indoor scene corresponding to the sample multimodal input; Extract the predicted scene description from the multimodal input of the sample, and generate a three-dimensional scene of the initial layout based on the predicted scene description to obtain the predicted initial indoor scene. The predicted initial indoor scene is vertically decoupled to obtain a predicted multi-layer scene structure; the predicted multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence; The predicted multi-layer scene structure is decoupled at the scene level to obtain a predicted scene graph; the nodes in the predicted scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects. Based on the initial alignment optimization model, the predicted scene map is optimized to obtain the predicted position and predicted direction of the object in the predicted target indoor scene; Based on the predicted position and the label position, as well as the predicted direction and the label direction, a target loss is determined, and the initial alignment optimization model is iterated based on the target loss to obtain the alignment optimization model.
[0093] Based on any of the above embodiments, the stacking level is a sub-level of the furniture level, and the furniture level is a sub-level of the floor level; The objects in each sub-level belong to the coordinate system of the parent level corresponding to each sub-level.
[0094] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a 3D indoor scene construction method. This method includes: acquiring multimodal input and extracting scene descriptions from the multimodal input; generating an initial layout 3D scene based on the scene descriptions to obtain an initial indoor scene; performing vertical scene decoupling on the initial indoor scene to obtain a multi-layer scene structure; the multi-layer scene structure sequentially includes a floor layer, a furniture layer, and a stacking layer; the stacking layer reflects objects stacked on top of other objects; performing horizontal scene decoupling on the multi-layer scene structure to obtain a scene graph; nodes in the scene graph represent objects, and edges in the scene graph represent positional semantic relationships between objects; optimizing the scene graph to obtain a target indoor scene.
[0095] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0096] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the three-dimensional indoor scene construction method provided by the above methods. The method includes: acquiring multimodal input and extracting scene description from the multimodal input; generating a three-dimensional scene with an initial layout based on the scene description to obtain an initial indoor scene; performing vertical scene decoupling on the initial indoor scene to obtain a multi-layer scene structure; the multi-layer scene structure sequentially includes a floor layer, a furniture layer, and a stacking layer; the stacking layer is used to reflect objects stacked on other objects; performing horizontal scene decoupling on the multi-layer scene structure to obtain a scene graph; the nodes in the scene graph represent objects, and the edges in the scene graph represent positional semantic relationships between objects; and optimizing the scene graph to obtain a target indoor scene.
[0097] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for constructing a three-dimensional indoor scene provided by the methods described above. This method includes: acquiring multimodal input and extracting a scene description from the multimodal input; generating a three-dimensional scene with an initial layout based on the scene description to obtain an initial indoor scene; performing vertical scene decoupling on the initial indoor scene to obtain a multi-layer scene structure; the multi-layer scene structure sequentially includes a floor layer, a furniture layer, and a stacking layer; the stacking layer reflects objects stacked on top of other objects; performing horizontal scene decoupling on the multi-layer scene structure to obtain a scene graph; nodes in the scene graph represent objects, and edges in the scene graph represent positional semantic relationships between objects; and optimizing the scene graph to obtain a target indoor scene.
[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0099] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a three-dimensional indoor scene, characterized in that, include: Acquire multimodal input and extract scene description from the multimodal input; Based on the scene description, an initial layout 3D scene is generated to obtain the initial indoor scene; The initial indoor scene is vertically decoupled to obtain a multi-layer scene structure; the multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence. The stacking hierarchy is used to reflect objects stacked on top of other objects; The multi-layer scene structure is horizontally decoupled to obtain a scene graph; the nodes in the scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects; The scene image is optimized to obtain the target indoor scene.
2. The method for constructing a three-dimensional indoor scene according to claim 1, characterized in that, The optimization of the scene map to obtain the target indoor scene includes: The scene graph is subjected to single-attribute verification, and the first indoor scene is determined based on the single-attribute verification result; the single-attribute verification is used to verify whether an object in each level of the scene graph belongs to only one parent node; Perform intra-layer collision detection and / or inter-layer collision detection on the first indoor scene to obtain the second indoor scene; Based on the second indoor scene, the target indoor scene is determined.
3. The method for constructing a three-dimensional indoor scene according to claim 2, characterized in that, Determining the target indoor scene based on the second indoor scene includes: Determine the target position of the center of gravity projection of the target object in the second indoor scene. Based on the positional relationship between the target position of the center of gravity projection and the support surface of the target object, adjust the posture of the target object. Based on the adjusted second indoor scene, determine the target indoor scene.
4. The method for constructing a three-dimensional indoor scene according to claim 3, characterized in that, The adjustment of the target object's posture based on the positional relationship between the target position projected by the center of gravity and the supporting surface of the target object includes: If the target position of the center of gravity projection exceeds the support surface, adjust the projection position of the target object to be within the support surface.
5. The method for constructing a three-dimensional indoor scene according to claim 3, characterized in that, Determining the target indoor scene based on the adjusted second indoor scene includes: Based on the alignment optimization model, the adjusted second indoor scene is aligned and optimized to obtain the target indoor scene. The alignment optimization model is trained based on sample multimodal input and the corresponding label indoor scene of the sample multimodal input.
6. The method for constructing a three-dimensional indoor scene according to claim 5, characterized in that, The training steps of the alignment optimization model include: Obtain the initial alignment optimization model, the sample multimodal input, and the label position and label orientation of the object in the label indoor scene corresponding to the sample multimodal input; Extract the predicted scene description from the multimodal input of the sample, and generate a three-dimensional scene of the initial layout based on the predicted scene description to obtain the predicted initial indoor scene. The predicted initial indoor scene is vertically decoupled to obtain a predicted multi-layer scene structure; the predicted multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence; The predicted multi-layer scene structure is decoupled at the scene level to obtain a predicted scene graph; the nodes in the predicted scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects. Based on the initial alignment optimization model, the predicted scene map is optimized to obtain the predicted position and predicted direction of the object in the predicted target indoor scene; Based on the predicted position and the label position, as well as the predicted direction and the label direction, a target loss is determined, and the initial alignment optimization model is iterated based on the target loss to obtain the alignment optimization model.
7. The method for constructing a three-dimensional indoor scene according to any one of claims 1 to 6, characterized in that, The stacking level is a sub-level of the furniture level, and the furniture level is a sub-level of the floor level; The objects in each sub-level belong to the coordinate system of the parent level corresponding to each sub-level.
8. A three-dimensional indoor scene construction device, characterized in that, include: The acquisition unit is used to acquire multimodal input and extract scene description from the multimodal input; An initial scene generation unit is used to generate an initial layout 3D scene based on the scene description to obtain an initial indoor scene; A scene vertical decoupling unit is used to perform scene vertical decoupling on the initial indoor scene to obtain a multi-layer scene structure; the multi-layer scene structure includes a floor layer, a furniture layer, and a stacking layer in sequence; the stacking layer is used to reflect objects stacked on top of other objects. A scene horizontal decoupling unit is used to perform scene horizontal decoupling on the multi-layer scene structure to obtain a scene graph; the nodes in the scene graph represent objects, and the edges in the scene graph represent the positional semantic relationships between objects; An optimization unit is used to optimize the scene map to obtain the target indoor scene.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the three-dimensional indoor scene construction method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional indoor scene construction method as described in any one of claims 1 to 7.